Thinking Machines
Posted 2w ago

Research Engineer, Infrastructure, Kernels

Thinking Machines
San Francisco, California, United States
$350k-$475k/yrHybridFull Time
Responsibilities
  • designing kernels
  • profiling performance
  • scaling infrastructure
Requirements
  • Bachelor's or equivalent
  • Strong GPU/kernel engineering (CUDA/CuTe/Triton)
  • Experience with PyTorch/JAX
  • Profiling and optimizing compute-intensive workloads
  • Collaborative research orientation
Technical tools mentioned
CUDACuTeTritonPyTorchJAXXLATVM

Job description

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We’re looking for an infrastructure research engineer to design, optimize, and maintain the compute foundations that power large-scale language model training. You will develop high-performance ML kernels (e.g., CUDA, CuTe, Triton), enable efficient low-precision arithmetic, and improve the distributed compute stack that makes training large models possible.

This role is perfect for an engineer who enjoys working close to the metal and across the research boundary. You’ll collaborate with researchers and systems architects to bridge algorithmic design with hardware efficiency. You’ll prototype new kernel implementations, profile performance across hardware generations, and help define the numerical and parallelism strategies that determine how we scale next-generation AI systems.

Note: This is an "evergreen role" that we keep open on an on-going basis to express interest. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. Still, we encourage you to apply. We continuously review applications and reach out to applicants as new opportunities open. You are welcome to reapply if you get more experience, but please avoid applying more than once every 6 months. You may also find that we put up postings for singular roles for separate, project or team specific needs. In those cases, you're welcome to apply directly in addition to an evergreen role.

What You’ll Do

  • Design and implement custom ML kernels (e.g., CUDA, CuTe, Triton) for core LLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU and accelerator architectures.

  • Design and think through compute primitives to reduce memory bandwidth bottlenecks and improve kernel compute efficiency.

  • Collaborate with research teams to align kernel-level optimizations with model architecture and algorithmic goals.

  • Develop and maintain a library of reusable kernels and performance benchmarks that serve as the foundation for internal model training.

  • Contribute to infrastructure stability and scalability, ensuring reproducibility, consistency across precision formats, and high utilization of compute resources.

  • Document and share insights through internal talks, technical papers, or open-source contributions to strengthen the broader ML systems community.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases

  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

  • Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.

  • Demonstrated ability to analyze, profile, and optimize compute-intensive workloads.

Preferred qualifications — we encourage you to apply if you meet some but not all of these:

  • Experience training or supporting large-scale language models with tens of billions of parameters or more.

  • Track record of improving research productivity through infrastructure design or process improvements.

  • Experience developing or tuning kernels for deep learning frameworks such as PyTorch, JAX, or custom accelerators.

  • Familiarity with tensor parallelism, pipeline parallelism, or distributed data processing frameworks.

  • Experience implementing low-precision formats (FP8, INT8, block floating point) or contributing to related compiler stacks (e.g., XLA, TVM).

  • Contributions to open-source GPU, ML systems, or compiler optimization projects.

  • Prior research or engineering experience in numerical optimization, communication-efficient training, or scalable AI infrastructure.

Logistics

  • Location: This role is based in San Francisco, California. 

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

About Thinking Machines

Building AI systems to extend human will and judgment.

Year founded
2025
Employees
100
Organization type
Private
Latest investment
Raised $2.00B Seed (2025) — led by Andreessen Horowitz
Headquarters
US

Similar jobs

Research Engineer roles near San Francisco, California
1d
Save
Mark Applied
Hide
Research Engineer, Preference Data
San Francisco, California, United States
$250k-$400k/yr OnsiteFull Time
Vizcom
Vizcom: AI-powered tools for industrial designers to visualize concepts instantly.
Experience building training-data or large-scale data pipelines; experimental mindset. Model training, labeling, human feedback, evaluation operations, and privacy or contractual data constraints are preferred.
LeetCode
1d
Save
Mark Applied
Hide
Research Engineer – Benchmarking
San Francisco or New York City or London
$130k-$500k/yr OnsiteFull Time
Mercor
Mercor: Connecting expert human intelligence with frontier AI model development.
Applied AI research, model evaluation or benchmarking, strong coding and ML experience, data structures and algorithms, backend systems, APIs, SQL or NoSQL, cloud platforms, and model behavior analysis.
SQL, NoSQL, NeurIPS, ICML, ACL
4d
Save
Mark Applied
Hide
Research Engineer, Lab Automation
Menlo Park or San Francisco
$200k-$250k/yr OnsiteFull Time
Periodic Labs
Periodic Labs: Builds autonomous laboratories for AI-driven scientific discovery.
PhD or equivalent research experience in materials science, chemistry, chemical engineering, or related field; materials lab hardware expertise; Python proficiency; and ability to translate scientific workflows into automation requirements.
Python, Electronic Lab Notebooks, LIMS
4d
Save
Mark Applied
Hide
Research Engineer - New Grad (2027)
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$140k-$200k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
Recent MSc or PhD graduate in machine learning, computer vision, autonomy, robotics, or related field; experience with Python, PyTorch, computer vision, robotics, and distributed model training.
Python, PyTorch
4d
Save
Mark Applied
Hide
Research Engineer, LangSmith Engine
New York City or San Francisco
OnsiteFull Time
LangChain
LangChain: Tools for building and deploying production-ready AI agents.
4+ YOERequires 4+ years in ML/AI research, a relevant master's or Ph.D., LLM and AI agent experience, benchmark and experiment design, and strong software engineering skills.
LangSmith, LangChain, LangGraph, Deep Agents, LLMs, AI agents, GPU infrastructure, SFT, RLHF, RLAIF
4d
Save
Mark Applied
Hide
Research Engineer, Synthetic Data
San Francisco or Singapore
$150k-$250k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
2+ YOERequires 2–4 years in software, ML engineering, or AI research; Python, Linux, Docker, synthetic data pipelines, evaluation frameworks, structured datasets, and independent project ownership.
Python, Linux, Docker
5d
Save
Mark Applied
Hide
Lead Research Engineer, Search & Retrieval
New York City or Frisco or Toronto or Ann Arbor or Eagan or San Francisco or Los Angeles or Irvine or McLean or Washington
$137k-$255k/yr HybridFull Time
Thomson Reuters
Thomson ReutersNASDAQ: TRI: Provides professional software, data, and news services globally.
7+ YOEBachelor's or master's in computer science, engineering, or related field; 7+ years building production software; search and retrieval expertise; Python, AWS, OpenSearch or Vespa, distributed systems, and technical leadership.
OpenSearch, Vespa, Elasticsearch, Solr, Lucene, Python, AWS, Kafka, RAG, A/B tests
5d
Save
Mark Applied
Hide
New College Grad - AI Innovation Research Engineer
San Jose, California, United States
$113k-$242k/yr OnsiteFull Time
Micron Technology
Micron TechnologyNASDAQ: MU: Designs and manufactures semiconductor memory and data storage solutions.
Recent Master's or PhD in a technical field; AI research or project experience; understanding of generative AI, LLMs, agents, machine learning, or analytics; Python programming; strong analytical and communication skills.
Generative AI, Large Language Models (LLMs), Python, Microsoft Copilot, Azure AI, OpenAI, Anthropic, Google AI