Fluidstack
Posted 4w ago

Production Engineering Lead, Compute

Fluidstack
San Francisco, California, United States
$260k-$361k/yrOnsiteFull Time
Responsibilities
  • leading teams
  • operating fleet
  • building automation
Requirements
  • Experience leading SRE or production engineering teams running large compute fleets
  • Proven reliability improvements
  • Automated remediation
  • Hiring and growing engineers
  • Familiarity with GPU/HPC, Kubernetes or Slurm is a plus
Technical tools mentioned
KubernetesSlurm

Job description

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.


We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate

  • Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

  • Velocity. We drive everything forward as fast as possible.

  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

The Data Center Operations Team

Examples of key problems the team is working on

  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 100 GW.

  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.

  • Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.

Role Scope

  • Lead the compute production engineering team keeping tens of thousands of GPUs serving customers.

  • Own fleet availability for compute: define the SLOs, build the tooling, and move the number.

  • Build automation for the node lifecycle, provisioning, health checks, remediation, return to service, without human touch.

  • Set the on-call and escalation model that keeps response sharp without burning the team out.

What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You've led SRE or production engineering teams running large fleets.

  • You've moved an availability number and can explain exactly how.

  • You've shipped automated remediation that retired a runbook.

  • You hire and grow strong engineers, and they say so.

  • Bonus: GPU or HPC fleets. Kubernetes or Slurm. Hardware failure analytics. Customer-facing reliability.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email [email protected] with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.

About Fluidstack

Provides high-performance cloud GPU infrastructure for AI development.

Year founded
2017
Employees
145
Organization type
Private
Latest investment
Raised $450.00M Series B (2026) — led by Situational Awareness
Headquarters
US

Similar jobs

Production Engineer roles near San Francisco, California
2d
Save
Mark Applied
Hide
Sr. Staff Production Engineer
San Jose or Bellevue
$144k-$205k/yr HybridFull Time
Zscaler
ZscalerNASDAQ: ZS: Provides cloud-native cybersecurity solutions through zero trust architecture.
8+ YOERequires 8+ years managing reliability, scalability, and availability for large-scale production services; Python, Go, or C/C++; networking, Linux/FreeBSD, distributed architecture, incident management, ITIL, and 24/7 on-call experience.
AWS, Azure, GCP, Python, Go, C/C++, Prometheus, Grafana, OpenTelemetry, Linux, FreeBSD, ITIL, Ansible, Terraform, BGP, GRE, IPSec, HAProxy, DNS
1w
Save
Mark Applied
Hide
Applied Machine Learning Production Engineer Graduate (AML-Production Engineer) - 2027 Start
San Jose, California, United States
OnsiteInternship
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Completing or recently completed a bachelor's or master's degree in software development, computer science, computer engineering, or a related discipline; programming, algorithms, and distributed systems knowledge required.
Python, C++, Go, Linux, Unix, TCP/IP, HTTP, Docker, Kubernetes, TensorFlow, PyTorch
1w
Save
Mark Applied
Hide
Production Engineer
Fremont, California, United States
$90k-$110k/yr OnsiteFull Time
Air Liquide
Air LiquideEuronext Paris: AI: Supplies industrial gases, medical gases, and related technologies.
Bachelor's degree in Chemical Engineering; process or production engineering knowledge; English communication; industrial process, safety, KPI, automation, and improvement experience preferred.
Northwest Analytics (NWA), PLC, Process Automation, Six Sigma, 5S, Lean Manufacturing, LOPA, HAZOP
2w
Save
Mark Applied
Hide
Senior Production Engineer, Compute
Sunnyvale, California, United States
$170k-$205k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
5+ YOE5+ years in compute SRE or Linux system engineering; deep Linux kernel and virtualization experience (KVM/QEMU/Xen/VMware); Go/C/Rust proficiency; system-level debugging (kdump,kexec); IaC and CI/CD for bare-metal/cloud.
KVM, QEMU, Xen, VMware, Go, C, Rust, kdump, kexec
2w
Save
Mark Applied
Hide
Production Engineer, Network
Menlo Park, California, United States
$154k-$217k/yr OnsiteFull Time
Meta
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
5+ YOEBachelor's in CS or equivalent, 5+ years with network device configurations, 5+ years coding (Python, Go, C++), experience automating operations, and strong TCP/IPv4/6 and routing protocol knowledge.
Python, Go, C++
3w
Save
Mark Applied
Hide
Production Engineer
Fremont, California, United States
$30-$34/hr OnsiteFull Time
Pacific Coast Building Products
Pacific Coast Building Products: Manufacturing and distributing building products and specialty contracting services.
Advanced MS Office skills, fluent Spanish, strong communication, valid CA driver’s license, ability to coordinate vendors and analyze operations.
Microsoft Word, Microsoft Excel, Microsoft Outlook
3w
Save
Mark Applied
Hide
Senior Software Engineer, Production Engineering (Cloud & On-Prem) - W&B
Livingston or New York or Sunnyvale or Bellevue or San Francisco
$139k-$185k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Specialized cloud infrastructure provider for high-performance AI workloads.
Deep experience operating production infrastructure across cloud and on-prem, expertise with IaC, Kubernetes, CI/CD, observability, scripting, and on-call participation; technical leadership and mentorship expected.
AWS, GCP, Azure, Terraform, CloudFormation, CDK, Ansible, Kubernetes, Go, Python, Bash, GitHub Actions, Prometheus, Grafana, Datadog, PostgreSQL, MySQL, ClickHouse, PagerDuty, Backstage
1mo
Save
Mark Applied
Hide
Production Engineer (Contract)
San Francisco or South San Francisco
$50-$60/hr OnsiteContract
VantageScore
VantageScore: A credit modeling that develops predictive and inclusive consumer credit scoring models to expand access to financial products.
1+ YOEBachelor's degree, 1–2 years software development or production support experience, proficiency in Python, databases, Git/CI-CD, cloud platforms, and strong problem-solving and communication skills.
Python, SAS, PostgreSQL, MySQL, MongoDB, Git, CI/CD, AWS, GCP, Azure, Lambda, EC2, DynamoDB, S3, API Gateway, Terraform, Docker, Kubernetes, GraphQL, RESTful