496 ml infrastructure engineer jobs at 200 companies in San Bruno, CA

1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yr OnsiteFull Time
xAI
xAI: Artificial intelligence research and development.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
Python, C++, Rust, JAX, PyTorch, NVIDIA drivers, CUDA, Linux, Slurm, Puppet, Ansible
2mo
Save
Mark Applied
Hide
Staff ML Infrastructure Engineer
Sunnyvale, California, United States
$189k-$291k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Designing, building, and selling the world's best vehicles.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Python, C++, PyTorch, TensorFlow, Bazel, Buck, Blaze, CMake, Docker, Kubernetes
3d
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Mateo, California, United States
OnsiteFull Time
Clera
Clera: AI-powered talent agent matching candidates to startup roles.
5+ YOE5+ years building production ML inference or model-serving systems. Requires scalable distributed systems, Docker, Kubernetes, cloud experience, observability tooling, and proficiency in Python, Go, Rust, C++, or Java.
TensorFlow Serving, TorchServe, Triton, KServe, Docker, Kubernetes, Prometheus, Grafana, ELK, AWS, GCP, Azure, Python, Go, Rust, C++, Java, Neo4j, Amazon Neptune
3mo
Save
Mark Applied
Hide
Founding ML infrastructure Engineer
San Francisco or United States
$200k-$350k/yr RemoteFull Time
uRun
uRun: AI infrastructure helping model labs, builders, and research teams run real-time interactive video and stateful inference.
Experience designing and operating large-scale distributed infrastructure; Kubernetes/Slurm; multi-cloud GPU; reliability and scheduling; startup mindset.
Kubernetes, Slurm, Scheduling, TensorRT-LLM, NCCL, InfiniBand, RoCE, CuTe, Triton, TileLang
2mo
Save
Mark Applied
Hide
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
OnsiteFull Time
Spellbrush
Spellbrush: Generative AI and anime game studio making anime illustrations and video games for artists and players.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
SLURM, Slinky, K8s, Warewulf, MAAS, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, LDAP, dmesg, HGX, VLAN
23h
Save
Mark Applied
Hide
Senior ML Infra Engineer
New York City or San Francisco
$200k-$275k/yr HybridFull Time
General Legal
General Legal: AI-native law firm serving growing companies with flat-fee commercial, corporate, and employment legal services.
5+ YOERequires 5+ years in software, machine learning, or infrastructure engineering; Python proficiency; production ML infrastructure, cloud, distributed systems, containers, independent system design, and a Computer Science degree or equivalent.
Python
2mo
Save
Mark Applied
Hide
Software Engineer, ML Infrastructure
Mountain View, California, United States
$160k-$241k/yr OnsiteFull Time
Nuro
Nuro: Private U.S. autonomous-driving technology developing AI systems for automakers and mobility providers.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Terraform, Pulumi, Crossplane, Kubernetes, KubeRay, Ray, Slurm, Volcano, Apache Spark, Apache Beam, Feast, Hopsworks, Redis, Lustre, Ceph, NVMe, AWS, GCP, Azure, Kubeflow, CNCF
3w
Save
Mark Applied
Hide
Senior Software Engineer, ML Infrastructure
Cupertino or San Francisco or Houston
$209k-$235k/yr OnsiteFull Time
Gridmatic
Gridmatic: AI-powered energy supplying and optimizing clean electricity for commercial and industrial customers.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
GCP, AWS, Azure, Kubernetes, GKE, Terraform, Flyte, Temporal, Airflow, Python, Go, C++, Java, Rust, Grafana, Google Cloud Monitoring
1mo
Save
Mark Applied
Hide
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
TPU, GPU, JAX, PyTorch, TensorFlow, Ray, Pathways
1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Private San Francis neurotechnology startup developing brain-computer interfaces that restore communication for people with severe disabilities.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
3mo
Save
Mark Applied
Hide
ML & Cloud Infrastructure Engineer
Belmont, California, United States
OnsiteFull Time
Gritt Robotics
Gritt Robotics: Private AI-powered construction robotics automating labor-intensive infrastructure work for construction crews.
4+ YOE4+ years deploying high-performance ML pipelines; proficient in Python; comfortable with C++/Go; PyTorch; cloud platforms (AWS, GCP, Azure); Docker/Kubernetes/Airflow; Parquet/HDF5/TFRecord; US work authorization.
Python, C++, Go, PyTorch, Parquet, HDF5, TFRecord, AWS, GCP, Azure, Docker, Kubernetes, Airflow
3w
Save
Mark Applied
Hide
ML Engineer, Foundation Model Infrastructure
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology and robotaxi service provider.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
Python, C++, Flume, JAX, PyTorch, TensorFlow, Spark, Borg, Kubeflow
3mo
Save
Mark Applied
Hide
Staff Software Engineer, ML Infrastructure
San Francisco, California, United States
$220k-$260k/yr HybridFull Time
Voxel
Voxel: Industrial AI providing a site-intelligence platform that helps enterprise safety and operations teams reduce workplace risk.
7+ YOE7+ years software systems, ML infra experience, Python, PyTorch, ML tooling, strong communication.
Python, PyTorch, AWS, TensorRT, ONNX, Weights & Biases, MLflow, ClearML
2mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Privately held container application platform helping developers build, share, and run applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
1mo
Save
Mark Applied
Hide
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Global platform for user-created immersive digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Kubernetes, Terraform, AWS, GCP, Docker, CI/CD
3mo
Save
Mark Applied
Hide
Autonomy Engineer - ML & DL Infrastructure
San Mateo, California, United States
$170k-$278k/yr HybridFull Time
Skydio
Skydio: American autonomous-drone manufacturer serving public safety, government, utility, and enterprise customers.
Experience with data engineering, cloud ML platforms, ML Ops, large-scale data pipelines, model training/deployment, and security/compliance in ML infrastructure; strong collaboration.
Cloud platforms, Containerization, ML Ops, Databases, Data pipelines
3mo
Save
Mark Applied
Hide
ML ENGINEER (GENERAL)
San Francisco, California, United States
OnsiteFull Time
MakerMaker
MakerMaker: Private U.S. AI research building autonomous research agents for recursive self-improvement.
6+ YOESenior ML engineer with 6+ years building production-grade ML systems; strong Python; distributed systems experience; familiar with Ray, Kubernetes, and experimentation infrastructure.
Python, PyTorch, JAX, Ray, Slurm, Kubernetes
3w
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yr OnsiteFull Time
Sciforium
Sciforium: AI infrastructure building multimodal models and high-efficiency serving software for developers and teams.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Ansible, SaltStack, Git, Python, Bash, Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI, enroot, pyxis, Docker, containerd, NVIDIA Container Toolkit, CUDA, cuDNN, NCCL, Fabric Manager, ROCm, RCCL, DKMS, GPUDirect RDMA, GPUDirect Storage, MOFED, DOCA, PyTorch, JAX, DCGM exporter, Prometheus, Grafana, PXE, MaaS, Packer, Foreman, Terraform, Lustre, GPFS, Weka, vLLM, Triton Inference Server, TensorRT-LLM, Nsight Systems, Nsight Compute, rocprof, perf, eBPF, EMR
22h
Save
Mark Applied
Hide
Staff ML Engineer
Pune or Bangalore or Santa Clara
HybridFull Time
Cohesity
Cohesity: AI-powered data security and management solutions provider.
10+ YOERequires 10+ years software engineering, 8+ years distributed backend systems, 3+ years LLM/RAG or ML infrastructure, production Kubernetes experience, technical leadership, and a bachelor's degree in a technical field.
Python, Java, Go, JavaScript, LLMs, Prompt Engineering, Retrieval-Augmented Generation (RAG), Kubernetes, OpenShift, EKS, GKE, Docker, Helm, Terraform, AWS, gRPC, REST APIs, SQL, NoSQL, PostgreSQL, Redis, Elasticsearch, CI/CD, DevOps, Git, Node.js, Jenkins
2w
Save
Mark Applied
Hide
Senior ML Engineer
San Francisco or Los Angeles or Denver or Austin or Chicago or New York City or Canada or Seattle or Santa Barbara or San Diego or Toronto
$152k-$228k/yr RemoteFull Time
Invoca
Invoca: Privately held SaaS platform helping marketing, commerce, and contact-center teams convert customer conversations into revenue.
5+ YOE5+ years of ML engineering experience; advanced Python, PyTorch, and deep learning; production NLP model deployment; SLM/LLM fine-tuning; inference infrastructure; production APIs; MLOps and model monitoring.
Python, PyTorch, HuggingFace Transformers, spaCy, Triton Inference Server, Baseten, Kubernetes, LoRA, QLoRA, PEFT, vLLM, TGI, SageMaker, Vertex AI, Braintrust, MLflow, RLHF