Thinking Machines
Posted 1mo ago

Reliability Engineer, Supercomputing

Thinking Machines
San Francisco, California, United States
$350k-$475k/yrOnsiteFull Time
Responsibilities
  • investigating issues
  • owning drivers
  • automating monitoring
Requirements
  • Ensure reliability of GPU supercomputing fleet across hardware
  • Firmware, and OS
  • Debug kernel/driver/hardware issues
  • Engage vendors
  • Automate monitoring and runroot-cause analysis
Technical tools mentioned
PythonRustKubernetesSlurmLinuxBMCiDRACIPMIRedfishDCGMNVLinkNVSwitchLinux kernel

Job description

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We're hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and operating system. You will track the long tail of hardware issues: We are conducting frontier research in AI and a single bad NIC, HBM or a kernel driver edge case can compromise an experiment. Your job is to diagnose these issues, track their root cause down to the hardware, and resolve them internally or directly with vendors so that our researchers can run at scale and with confidence.

Note: This is an "evergreen role" that we keep open on an on-going basis to express interest. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. Still, we encourage you to apply. We continuously review applications and reach out to applicants as new opportunities open. You are welcome to reapply if you get more experience, but please avoid applying more than once every 6 months. You may also find that we put up postings for singular roles for separate, project or team specific needs. In those cases, you're welcome to apply directly in addition to an evergreen role.

What You’ll Do

  • Investigate, reproduce, and remediate issues across large GPU clusters.

  • Own the drivers, kernel surface, and diagnostics that span hardware, firmware, and OS.

  • Automate the monitoring of fleet reliability and analyze error rates to validate whether a fix or firmware change measurably reduced failures rather than shifting them around.

  • Drive the firmware lifecycle: tracking, qualification, staged rollout, and regression analysis.

  • Engage vendors directly — GPUs, server OEMs, NIC vendors, and storage vendors — to get real fixes rather than ticket numbers. Manage RMA flows when hardware needs to come out.

  • Monitor and improve GPU hardware health signals and turn them into actionable reliability improvements.

  • Write clear postmortems and vendor cases that move issues forward.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.

  • Proficiency in at least one backend language (we use Python or Rust).

  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).

  • Comfort operating across the stack and owning projects end-to-end.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Preferred qualifications — we encourage you to apply if you meet some but not all of these:

  • Fluency with Linux systems and debugging tools.

  • Proven statistical rigor in analyzing reliability.

  • A track record of debugging a problem from application symptom to the root cause in hardware.

  • Comfort reading vendor errata, firmware release notes, and kernel changelogs.

  • Experience engaging hardware vendors directly — not just through escalation portals.

  • Linux kernel literacy: the scheduler, memory management, IRQ paths, and the driver model.

  • Out-of-band management experience: BMC / iDRAC / IPMI / Redfish.

  • Depth in GPU hardware health: Xid error taxonomy, NVLink, NVSwitch, fabric manager, and DCGM.

  • Proficiency in at least one backend language (we use Python and Rust).

  • Significant ownership of the hardware reliability function at scale.

  • Strong writing skills for vendor cases and postmortems.

  • An instinct for telling apart a flaky machine, a flaky workload, and a flaky test.

Logistics

  • Location: This role is based in San Francisco, California. 

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

About Thinking Machines

Building AI systems to extend human will and judgment.

Year founded
2025
Employees
100
Organization type
Private
Latest investment
Raised $2.00B Seed (2025) — led by Andreessen Horowitz
Headquarters
US

Similar jobs

Reliability Engineer roles near San Francisco, California
2d
Save
Mark Applied
Hide
Reliability Engineer 4 (Observability Specialist )
Chicago or Atlanta or Cupertino or Gresham or Denver or Charlotte or Brookfield or Irving or Hopkins or Earth City
$124k-$146k/yr HybridFull Time
U.S. Bank
U.S. BankNYSE: USB: Provider of personal, business, and institutional financial services.
6+ YOEBachelor's degree or equivalent experience and 6–8 years in reliability, SRE, IT service management, production support, application development, or related work; expertise in observability and stakeholder leadership.
Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, OpenTelemetry, Kubernetes
3d
Save
Mark Applied
Hide
Senior Reliability Engineer
Alameda, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
4+ YOEAssociate's degree, 4+ years in an FDA/ISO-regulated environment, failure analysis, hardware/software integration, reliability investigations, embedded systems, electrical design, and project management experience.
embedded systems
3d
Save
Mark Applied
Hide
Senior Reliability Engineer
Alameda, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
4+ YOEAssociate's degree, 4+ years in an FDA/ISO-regulated environment, and experience with failure analysis, hardware/software integration, reliability investigations, embedded systems, electrical design, and project management.
embedded systems
2w
Save
Mark Applied
Hide
Staff Reliability Engineer
Menlo Park, California, United States
$151k-$178k/yr OnsiteFull Time
Mainspring Energy
Mainspring Energy: Manufactures fuel-flexible linear generators for onsite power generation.
8+ YOEBachelor's degree in electrical or mechanical engineering and 8+ years of reliability testing experience, or 5+ years with a related master's degree. Requires FMEA, failure analysis, instrumentation, and data analysis expertise.
FMEA, HALT, Python, Matlab, Google Docs, Weibull ++, JMP, CAPA, RCA, 8D, DAQ, X-ray, CT, UL, NFPA, IEC
2w
Save
Mark Applied
Hide
Reliability Engineer
Beaver Brook or Santa Clara
$122k-$232k/yr OnsiteFull Time
Intel
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
4+ YOEBS/MS/PhD in EE/ME or 4+ years experience; authoring reliability specs, RAS, FMEA, statistical reliability (Weibull,FIT), fleet telemetry and thermal/power redundancy experience.
Python, SQL
2w
Save
Mark Applied
Hide
Reliability Engineer (Hardware)
Mountain View, California, United States
$142k-$200k/yr HybridFull Time
Lightmatter
Lightmatter: Develops photonic processors and interconnects for AI data centers
5+ YOEBachelor's in engineering,5+ years reliability experience in semiconductor/photonics,knowledge of reliability test methods and failure analysis tools,statistical reliability modeling and strong communication skills.
SEM, TEM, FIB, EMMI, Minitab, JMP
2w
Save
Mark Applied
Hide
Reliability Engineering Technical Leader
San Jose, California, United States
$163k-$205k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
8+ YOEBachelor's in Engineering with 12+ years or Master's with 8+ years; deep hardware reliability and PCBA knowledge; expertise in RAS, risk management, data-driven reliability, and executive influence; proven leadership and mentoring skills.
2w
Save
Mark Applied
Hide
Senior Reliability Engineer, Labs
San Francisco or Oakland or United States
$139k-$205k/yr OnsiteFull Time
DoorDash
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
5+ YOEFive years of reliability validation or hardware testing experience in robotics or automated vehicles; bachelor's or higher in engineering; experience with test equipment, CAD, shop tools, Python, and hardware-software validation.
Python, CAD, DAQ, FRACAS, FMEA, FTA, HIL