Thinking Machines
Posted 2w ago

Research Engineer, Infrastructure, Training Systems

Thinking Machines
San Francisco, California, United States
$350k-$475k/yrOnsiteFull Time
Responsibilities
  • designing systems
  • optimizing throughput
  • building frameworks
Requirements
  • Bachelor's or equivalent
  • Strong systems engineering
  • Experience with distributed training and deep learning frameworks (PyTorch, JAX)
  • Contributions to ML infra preferred
Technical tools mentioned
PyTorchJAXXLAMegatron-LMDeepSpeed

Job description

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We’re looking for an infrastructure research engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Thinking Machines fast and reliable to ensure our research teams can focus on science, not system bottlenecks.

This role is ideal for someone who blends deep systems and performance expertise with a curiosity for machine learning at scale. You’ll take ownership of the training stack end to end, ensuring every GPU cycle drives scientific progress.

Note: This is an "evergreen role" that we keep open on an on-going basis to express interest. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. Still, we encourage you to apply. We continuously review applications and reach out to applicants as new opportunities open. You are welcome to reapply if you get more experience, but please avoid applying more than once every 6 months. You may also find that we put up postings for singular roles for separate, project or team specific needs. In those cases, you're welcome to apply directly in addition to an evergreen role.

What You’ll Do

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes for large-scale training workloads.

  • Develop high-performance optimizations to maximize throughput and efficiency.

  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.

  • Establish standards for reliability, maintainability, and security, ensuring systems are robust under rapid iteration.

  • Collaborate with researchers and engineers to build scalable infrastructure.

  • Publish and share learnings through internal documentation, open-source libraries, or technical reports that advance the field of scalable AI infrastructure.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases

  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Preferred qualifications — we encourage you to apply if you meet some but not all of these:

  • Past experience working on distributed training for the world’s largest models to make them stable, reliable, and performant.

  • Track record of improving research productivity through infrastructure design or process improvements.

  • Contributions to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.

Logistics

  • Location: This role is based in San Francisco, California. 

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

About Thinking Machines

Building AI systems to extend human will and judgment.

Year founded
2025
Employees
100
Organization type
Private
Latest investment
Raised $2.00B Seed (2025) — led by Andreessen Horowitz
Headquarters
US

Similar jobs

Research Engineer roles near San Francisco, California
1d
Save
Mark Applied
Hide
Research Engineer, Preference Data
San Francisco, California, United States
$250k-$400k/yr OnsiteFull Time
Vizcom
Vizcom: AI-powered tools for industrial designers to visualize concepts instantly.
Experience building training-data or large-scale data pipelines; experimental mindset. Model training, labeling, human feedback, evaluation operations, and privacy or contractual data constraints are preferred.
LeetCode
1d
Save
Mark Applied
Hide
Research Engineer – Benchmarking
San Francisco or New York City or London
$130k-$500k/yr OnsiteFull Time
Mercor
Mercor: Connecting expert human intelligence with frontier AI model development.
Applied AI research, model evaluation or benchmarking, strong coding and ML experience, data structures and algorithms, backend systems, APIs, SQL or NoSQL, cloud platforms, and model behavior analysis.
SQL, NoSQL, NeurIPS, ICML, ACL
1d
Save
Mark Applied
Hide
Robotics Research Engineer - Robot Simulation and Evaluation
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOEPhD in computer science, robotics, or related field; 8+ years in robotics, simulation, or robot learning; software design expertise; Python and deep learning software stack proficiency.
NVIDIA Omniverse, Python, PyTorch, JAX, PhysX, Isaac Gym, Isaac Lab, CUDA, Warp, ROS
4d
Save
Mark Applied
Hide
Research Engineer, Lab Automation
Menlo Park or San Francisco
$200k-$250k/yr OnsiteFull Time
Periodic Labs
Periodic Labs: Builds autonomous laboratories for AI-driven scientific discovery.
PhD or equivalent research experience in materials science, chemistry, chemical engineering, or related field; materials lab hardware expertise; Python proficiency; and ability to translate scientific workflows into automation requirements.
Python, Electronic Lab Notebooks, LIMS
4d
Save
Mark Applied
Hide
Research Engineer - New Grad (2027)
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$140k-$200k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
Recent MSc or PhD graduate in machine learning, computer vision, autonomy, robotics, or related field; experience with Python, PyTorch, computer vision, robotics, and distributed model training.
Python, PyTorch
5d
Save
Mark Applied
Hide
Research Engineer, LangSmith Engine
New York City or San Francisco
OnsiteFull Time
LangChain
LangChain: Tools for building and deploying production-ready AI agents.
4+ YOERequires 4+ years in ML/AI research, a relevant master's or Ph.D., LLM and AI agent experience, benchmark and experiment design, and strong software engineering skills.
LangSmith, LangChain, LangGraph, Deep Agents, LLMs, AI agents, GPU infrastructure, SFT, RLHF, RLAIF
5d
Save
Mark Applied
Hide
Research Engineer, Synthetic Data
San Francisco or Singapore
$150k-$250k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
2+ YOERequires 2–4 years in software, ML engineering, or AI research; Python, Linux, Docker, synthetic data pipelines, evaluation frameworks, structured datasets, and independent project ownership.
Python, Linux, Docker
5d
Save
Mark Applied
Hide
Lead Research Engineer, Search & Retrieval
New York City or Frisco or Toronto or Ann Arbor or Eagan or San Francisco or Los Angeles or Irvine or McLean or Washington
$137k-$255k/yr HybridFull Time
Thomson Reuters
Thomson ReutersNASDAQ: TRI: Provides professional software, data, and news services globally.
7+ YOEBachelor's or master's in computer science, engineering, or related field; 7+ years building production software; search and retrieval expertise; Python, AWS, OpenSearch or Vespa, distributed systems, and technical leadership.
OpenSearch, Vespa, Elasticsearch, Solr, Lucene, Python, AWS, Kafka, RAG, A/B tests