Thinking Machines Lab
Posted 2mo ago

Reliability Engineer, Supercomputing

Thinking Machines Lab
San Francisco, California, United States
$350k-$475k/yrOnsiteFull Time
Responsibilities
  • investigating issues
  • owning drivers
  • automating monitoring
Requirements
  • Ensure reliability of GPU supercomputing fleet across hardware
  • Firmware, and OS
  • Debug kernel/driver/hardware issues
  • Engage vendors
  • Automate monitoring and runroot-cause analysis
Technical tools mentioned
PythonRustKubernetesSlurmLinuxBMCiDRACIPMIRedfishDCGMNVLinkNVSwitchLinux kernel

Job description

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We're hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and operating system. You will track the long tail of hardware issues: We are conducting frontier research in AI and a single bad NIC, HBM or a kernel driver edge case can compromise an experiment. Your job is to diagnose these issues, track their root cause down to the hardware, and resolve them internally or directly with vendors so that our researchers can run at scale and with confidence.

Note: This is an "evergreen role" that we keep open on an on-going basis to express interest. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. Still, we encourage you to apply. We continuously review applications and reach out to applicants as new opportunities open. You are welcome to reapply if you get more experience, but please avoid applying more than once every 6 months. You may also find that we put up postings for singular roles for separate, project or team specific needs. In those cases, you're welcome to apply directly in addition to an evergreen role.

What You’ll Do

  • Investigate, reproduce, and remediate issues across large GPU clusters.

  • Own the drivers, kernel surface, and diagnostics that span hardware, firmware, and OS.

  • Automate the monitoring of fleet reliability and analyze error rates to validate whether a fix or firmware change measurably reduced failures rather than shifting them around.

  • Drive the firmware lifecycle: tracking, qualification, staged rollout, and regression analysis.

  • Engage vendors directly — GPUs, server OEMs, NIC vendors, and storage vendors — to get real fixes rather than ticket numbers. Manage RMA flows when hardware needs to come out.

  • Monitor and improve GPU hardware health signals and turn them into actionable reliability improvements.

  • Write clear postmortems and vendor cases that move issues forward.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.

  • Proficiency in at least one backend language (we use Python or Rust).

  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).

  • Comfort operating across the stack and owning projects end-to-end.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Preferred qualifications — we encourage you to apply if you meet some but not all of these:

  • Fluency with Linux systems and debugging tools.

  • Proven statistical rigor in analyzing reliability.

  • A track record of debugging a problem from application symptom to the root cause in hardware.

  • Comfort reading vendor errata, firmware release notes, and kernel changelogs.

  • Experience engaging hardware vendors directly — not just through escalation portals.

  • Linux kernel literacy: the scheduler, memory management, IRQ paths, and the driver model.

  • Out-of-band management experience: BMC / iDRAC / IPMI / Redfish.

  • Depth in GPU hardware health: Xid error taxonomy, NVLink, NVSwitch, fabric manager, and DCGM.

  • Proficiency in at least one backend language (we use Python and Rust).

  • Significant ownership of the hardware reliability function at scale.

  • Strong writing skills for vendor cases and postmortems.

  • An instinct for telling apart a flaky machine, a flaky workload, and a flaky test.

Logistics

  • Location: This role is based in San Francisco, California. 

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

About Thinking Machines Lab

Private AI research and product building customizable multimodal systems for researchers and the wider public.

Similar jobs

Reliability Engineer roles near San Francisco, California
1d
Save
Mark Applied
Hide
RELIABILITY ENGINEER
Carolina or Caguas or Boston or San Francisco or United States or Puerto Rico or Dominican Republic or Mexico or Germany or Canada or South America
FieldFull Time
Mentor Technical Group
Mentor Technical Group: Technical, engineering, and compliance solutions for life sciences.
Bachelor's degree in engineering required, mechanical engineering preferred. Reliability, vibration analysis, thermography, or project management certifications are preferred or advantageous.
Computer Maintenance Management System (CMMS)
1d
Save
Mark Applied
Hide
Reliability Engineer
Littleton or Sunnyvale
$79k-$147k/yr OnsiteFull Time
Lockheed Martin
Lockheed MartinNYSE: LMT: Global security and aerospace driving innovative defense solutions.
Experience with space, missile, or related systems; systems engineering knowledge; modeling or statistical analysis skills; U.S. citizenship; and ability to obtain or maintain a Top Secret clearance.
HALT/HASS, Design of Experiments (DOE), MIL-HDBK-217, MIL-STD-1629, MIL-STD-1543, MIL-HDBK-338, MIL-HDBK-472
2d
Save
Mark Applied
Hide
Reliability Engineer 3 (Observability Specialist)
Brookfield or Atlanta or Hopkins or Cupertino or Charlotte or New York City or Chicago or Gresham or Englewood or Cincinnati or Irving or Earth City
$98k-$116k/yr HybridFull Time
U.S. Bank
U.S. BankNYSE: USB: Provider of personal, business, and institutional financial services.
5+ YOEBachelor's degree or equivalent experience and 5–7 years in IT service management, production support, risk analysis, product/project management, or application development; observability, SRE, cloud, distributed systems, and Kubernetes expertise preferred.
Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, OpenTelemetry, Kubernetes
3d
Save
Mark Applied
Hide
Senior Reliability Engineer, MicroLED
Fremont, California, United States
$159k-$230k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
4+ YOEBachelor's degree or equivalent practical experience; 4 years in optoelectronic device reliability; 3 years in qualification testing and analysis; 2 years with databases and programming; JMP or Reliasoft experience.
JMP, Reliasoft, databases
3d
Save
Mark Applied
Hide
Senior Engineer, Reliability Engineering
San Jose or Rio Robles
$110k-$151k/yr OnsiteFull Time
Analog Devices
Analog DevicesNASDAQ: ADI: Designs and manufactures semiconductors for signal processing and power management.
2+ YOEBachelor’s or master’s degree in electrical engineering, device physics, or related discipline; 2+ years in semiconductor reliability; knowledge of device physics, JEDEC, IPC, reliability testing, Weibull models, and JMP.
Failure Modes and Effects Analysis (FMEA), Failure Mechanisms Matrix (FMM), 8D, JEDEC, IPC, HTOL, HAST, JMP, Weibull
3d
Save
Mark Applied
Hide
Reliability Engineer
Pittsburg, California, United States
$95k-$115k/yr OnsiteFull Time
Douglas Products
Douglas Products: Privately held U.S. specialty chemical manufacturer and marketer serving agricultural and structural pest-control professionals.
5+ YOEBachelor’s degree in an engineering field and 5+ years in reliability, maintenance, or process engineering in industrial or manufacturing settings. Requires CMMS, mechanical systems, maintenance, RCA, and regulatory knowledge.
IBM Maximo, Microsoft Office, Microsoft Excel, Intelex, API, ASME, P&IDs, CMMS
3d
Save
Mark Applied
Hide
Senior Reliability Engineer
Mountain View, California, United States
$175k-$250k/yr OnsiteFull Time
Rhoda AI
Rhoda AI: Private robotics startup building robot foundation models for autonomous industrial tasks in manufacturing, logistics, automotive, and ecommerce.
5+ YOEBachelor's or Master's in engineering; 5+ years in reliability engineering; experience with mission profiles, accelerated testing, FMEA/FMECA, Weibull, MTBF/MTTF, electromechanical systems, and cross-functional communication.
HALT, HASS, FMEA, FMECA, Weibull analysis, MTBF, MTTF, Python, Arrhenius, Coffin-Manson
3d
Save
Mark Applied
Hide
Sr. Reliability Engineer
Mountain View or San Jose or California
$210k-$265k/yr OnsiteFull Time
1AI Energy
1AI Energy: Private renewable-energy delivering rapidly deployable solar-and-storage power systems to data centers and energy-intensive industries.
8+ YOEBS or MS in engineering and 8+ years of reliability experience in renewable energy, power electronics, industrial power systems, or electrification; expertise in testing, statistical reliability, failure analysis, and root cause methods.
ReliaSoft Weibull++, JMP, Minitab, MATLAB, Python, Altium, power analyzers, oscilloscopes, thermal chambers, vibration systems, environmental chambers, DAQ, DFMEA, PFMEA, DRBFM, Weibull analysis, MTBF, MTTF, Reliability Block Diagrams (RBD), Fault Tree Analysis (FTA), Physics of Failure (PoF), 8D, Fishbone, 5-Why, SEM, EDS, X-ray, CT scanning, DVP&R, UL 1741, UL 62109, UL 9540, UL 9540A, IEEE 1547, IEC 62477, IEC 60068, IEC 62093