The Role
Every robot we ship goes out with edge compute attached to it, and that compute has to work reliably in someone else's factory, not in our lab. We need a founding software engineer to build the system that makes that possible: software that can be dropped onto a heterogeneous fleet of edge devices, updated and controlled remotely, and trusted to keep running when nobody from our team is standing next to it.
This isn't a narrow infra role. You'll own the layer that sits between our models and the physical robot, which means you're directly responsible for how fast and how reliably a policy's decisions turn into robot motion. Latency here isn't an abstract metric, it's the difference between a robot that feels responsive and one that doesn't work at all. You'll also be the person who can log into a robot on a factory floor from anywhere, see what's actually happening, and fix it, because when something breaks in production, it needs to get fixed without anyone getting on a plane.
You'll be one of the first engineers on this system, so the architecture decisions you make now, how we deploy, how we monitor, how we recover from failure, will shape how the whole company scales its hardware footprint.
The system doesn't end at the robot. Factory workers, not robotics engineers, are the ones using it day to day, so you'll also build the customer-facing apps they interact with directly. That means getting close to how people on the floor actually use the product, listening to their feedback, and using it to keep smoothing out the experience until running the system feels like second nature for factory workers anywhere in the world.
What You'll Do
Design and build a deployment system that runs across a heterogeneous set of edge compute, from different robots and different hardware generations
Build remote management tooling: pushing software updates, controlling and monitoring edge devices, and troubleshooting issues without physical access to the robot
Build robust data recording pipelines so we can capture what a robot saw and did, on-device and in the field
Profile and optimize the full control loop to identify bottlenecks and minimize latency between the controller and the policy
Make hard tradeoffs to keep the system lightweight enough to run well on constrained, low-power edge hardware
Instrument the system so failures are diagnosable remotely, and design for graceful recovery when they happen
Work closely with the robotics and ML team to understand how their systems actually run on hardware, not just in theory
Build and iterate on the customer-facing apps factory workers use to operate and monitor the system, folding in their feedback to make the experience smooth and intuitive
What We're Looking For
Must-haves:
Strong systems engineering fundamentals: you can reason about latency, concurrency, resource constraints, and failure modes, not just write application code
Experience building and operating distributed or fleet-style systems (remote deployment, monitoring, OTA updates, or similar)
Comfort working close to hardware and across heterogeneous environments, where "it works on my machine" isn't good enough
A track record of shipping systems that had to actually run reliably in the real world, not just pass tests
Ability to work in-person in Seattle on weekdays
Nice-to-haves:
Experience with embedded Linux, edge compute.
Experience building remote device management or fleet infrastructure (IoT, robotics, or similar)
Experience with real-time or low-latency control systems
Background in performance profiling and optimization on resource-constrained hardware
Experience with observability and remote debugging tooling for systems you can't physically access
Familiarity with PyTorch and optimizing model inference using ONNX, TensorRT, quantization, and JIT
We care much more about how you think and what you've built than a specific degree or years-of-experience number. If you've done work that maps to this but doesn't check every box above, we'd still like to hear from you.