The Role
Training robot foundation models is expensive and iteration speed is everything: the faster our robotics engineers can launch a job, get results, and try the next idea, the faster the whole company moves. We need a founding engineer to build the training infrastructure that removes that gate, making it trivial to launch a training job, and making sure that job runs as efficiently as possible wherever it runs.
Concretely, that means building a system that lets robotics engineers launch training jobs with a single, simple interface, without needing to think about which cloud provider, which cluster, or which hardware is underneath it. You'll own the abstraction that decides where a job actually runs, whether that's optimizing for cost, availability, or performance across providers, and you'll be responsible for making sure GPUs aren't sitting idle and jobs aren't silently running slower than they should be.
This is a foundational role. The infrastructure you build will be the thing every robotics engineer touches every day, so the decisions you make about reliability, usability, and cost-efficiency will directly shape how fast this company can iterate on its core technology.
What You'll Do
Build a streamlined, self-serve way for robotics engineers to launch training jobs, abstracting away the underlying cloud provider or cluster
Design and build the orchestration layer that schedules and manages training jobs across multiple cloud providers
Continuously monitor and optimize training jobs for cost, GPU utilization, and throughput
Identify and eliminate bottlenecks in data loading, checkpointing, and distributed training that waste compute
Build tooling to make it easy to compare performance and cost across providers and hardware types
Set up monitoring and alerting so failed or underperforming jobs are caught quickly, not discovered days later
Work closely with the robotics engineering team to understand their workflows and remove friction wherever it shows up
What We're Looking For
Must-haves:
Strong systems engineering fundamentals, with experience building infrastructure that other engineers or robotics engineers depend on daily
Experience with distributed training and large-scale ML workloads
Proficiency with PyTorch
Experience working across multiple cloud providers (e.g. AWS, GCP, Azure) and reasoning about tradeoffs in cost and performance
Comfort building and operating orchestration or scheduling systems
A track record of shipping infrastructure that had to actually hold up under real, daily use
Ability to work in-person in Seattle on weekdays
Nice-to-haves:
Experience with distributed training frameworks (e.g. PyTorch Distributed, DeepSpeed, Ray)
Experience with Kubernetes or other container orchestration systems
Experience optimizing GPU utilization, data loading pipelines, or checkpointing at scale
Background in cloud cost optimization or FinOps for ML workloads
Experience building internal developer tools or platforms for robotics engineering teams
We care much more about how you think and what you've built than a specific degree or years-of-experience number. If you've done work that maps to this but doesn't check every box above, we'd still like to hear from you.