Spellbrush: Develops anime-themed video games using proprietary generative AI technology.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yrHybridFull Time
Waymo: Autonomous driving technology for ride-hailing and logistics.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
MakerMaker: Small San Francis-based team building autonomous ML systems
6+ YOESenior ML engineer with 6+ years building production-grade ML systems; strong Python; distributed systems experience; familiar with Ray, Kubernetes, and experimentation infrastructure.
Sciforium: Building multimodal AI models and high-performance model serving infrastructure.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
San Francisco or Los Angeles or Denver or Austin or Chicago or New York City or Canada or Seattle or Santa Barbara or San Diego or Toronto
$152k-$228k/yrRemoteFull Time
Invoca: AI platform for conversation intelligence and revenue execution.
5+ YOE5+ years of ML engineering experience; advanced Python, PyTorch, and deep learning; production NLP model deployment; SLM/LLM fine-tuning; inference infrastructure; production APIs; MLOps and model monitoring.
Staff Research Engineer, Scientific Computing and ML/Physics Infrastructure
Cambridge or London or San Francisco
$224k-$294k/yrOnsiteFull Time
Lila Sciences: Develops an AI platform for autonomous scientific research and discovery.
Strong software engineering in Python, experience with ML/scientific computing, distributed systems, GPU performance, PyTorch/JAX/CUDA, Linux and containers, orchestration systems, and working with research teams.
Petah Tikva or Austin or New York City or Los Angeles or San Francisco or London or Berlin or Singapore or Tel Aviv
HybridFull Time
Digital TurbineNASDAQ: APPS: Platform for mobile application distribution and advertising monetization.
8+ YOE8+ years in infrastructure, platform, or back-end engineering; distributed systems, AWS or GCP, programming, Kubernetes, data or ML infrastructure, infrastructure-as-code, CI/CD, GitOps, observability, and production operations experience.
San Francisco or Minneapolis or Washington, D.C. or United States
$120k-$215k/yrRemoteFull Time
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
4+ YOE2+ MgmtBachelor's degree or 4+ years equivalent, 4+ years Python, 4+ years cloud infrastructure (AWS/Azure/GCP), 4+ years AI/ML infrastructure experience, 2+ years team lead, 1+ year LLM experience.
Python, AWS, Azure, GCP, Large Language Models (LLMs), GitHub, GitHub Actions, Docker, Terraform, CI/CD
OpenAI: Develops artificial intelligence models and generative AI software services.
7+ YOE7+ years of professional software engineering, experience with large-scale distributed systems or ML infrastructure, built ML workflows and data pipelines, low-latency, reliable systems, observability, and cross-functional collaboration.
Epsilon Health: Provides AI-powered radiology diagnostic services and medical imaging software.
5+ YOE5+ years building production ML infrastructure and data pipelines, strong Python and PyTorch/JAX skills, distributed training and cloud experience, familiarity with data pipeline technologies and containerization.
San Francisco or New York or Los Angeles or Seattle
$200k-$345k/yrHybridFull Time
Whatnot: Social marketplace for buying and selling via live streams
4+ YOE4+ years building ML systems,3+ years software engineering,1+ year Python,experience with databases,monitoring,cloud services and production ML deployments.
Sesame: Designing wearable computers with lifelike voice-driven AI agents.
3+ YOEStrong systems thinker with reliability engineering experience; 3+ years in infrastructure, platform, or ML systems; Kubernetes production experience; strong communication.
Denver or San Francisco or New York City or San Jose or Scottsdale
$160k-$240k/yrHybridFull Time
Gusto: Cloud-based payroll and HR software for small businesses.
4+ YOERequires 4+ years of software engineering experience in Python, Ruby, or Java; ML infrastructure and platform-service experience; cloud-platform experience; and familiarity with AI frameworks and AI-assisted development tools.
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yrOnsiteFull Time
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
Watney Robotics: Develops autonomous robotic systems for data center infrastructure buildout.
Experience building ML platforms and large-scale distributed training, fluency in Python and Rust or C/C++, familiarity with JAX and PyTorch, knowledge of distributed backbones (FSDP, DeepSpeed, Megatron, Ray Train), GPU/TPU optimization, and cluster orchestration.
JAX, PyTorch, FSDP, DeepSpeed, Megatron, Ray Train, vLLM, Triton, Ray Serve, Python, Rust, C/C++