249 ai ml infrastructure engineer jobs at 105 companies in Tiburon, CA

3w
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yr OnsiteFull Time
Sciforium
Sciforium: Building multimodal AI models and high-performance model serving infrastructure.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Ansible, SaltStack, Git, Python, Bash, Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI, enroot, pyxis, Docker, containerd, NVIDIA Container Toolkit, CUDA, cuDNN, NCCL, Fabric Manager, ROCm, RCCL, DKMS, GPUDirect RDMA, GPUDirect Storage, MOFED, DOCA, PyTorch, JAX, DCGM exporter, Prometheus, Grafana, PXE, MaaS, Packer, Foreman, Terraform, Lustre, GPFS, Weka, vLLM, Triton Inference Server, TensorRT-LLM, Nsight Systems, Nsight Compute, rocprof, perf, eBPF, EMR
1w
Save
Mark Applied
Hide
Senior AI/ML Engineer
Lisbon or Atlanta or London or San Francisco or Santiago or Sydney or Tokyo or Toronto or Australia or Canada or United States
HybridFull Time
PagerDuty
PagerDutyNYSE: PD: Platform for real-time incident response and digital operations automation.
5+ YOE5+ years of software engineering experience building production distributed systems and AI systems, with expertise in LLMs, agents, retrieval, cloud infrastructure, containers, Kubernetes, reliability, and evaluation.
LLM, Kubernetes, AWS, GCP, Azure, LangChain, LlamaIndex, Kafka, Airflow, Spark
1w
Save
Mark Applied
Hide
Senior Staff AI/ML Engineer
Santa Clara, California, United States
$131k-$261k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
8+ YOEBachelor's degree and 8+ years, master's and 6+ years, or Ph.D. and 4+ years of related experience. Requires ML infrastructure, cloud, MLOps, AI governance, distributed systems, and leadership expertise.
AWS, Azure, Google Cloud, Docker, Kubernetes, Airflow, Terraform, Azure ML SDK, Azure Data Factory, Databricks, Spark, SQL, NoSQL, CI/CD
2w
Save
Mark Applied
Hide
Lead AI/ML Engineer - Remote
San Francisco, California, United States
$146k-$250k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
12+ YOEBachelor's degree or 4+ years equivalent experience; 12+ years software, data science, or analytics experience; 3+ years AI/ML; 4+ years Python, cloud infrastructure, and technical leadership.
Python, AWS, Azure, GCP, Google Cloud Platform, FHIR, HL7, HIPAA, GitHub, GitHub Actions, Docker, Terraform, CI/CD, Large Language Models (LLMs)
3d
Save
Mark Applied
Hide
Senior AI/ML Platform Engineer
Seattle or San Francisco or San Jose
$152k-$265k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOE5+ years in software engineering, backend infrastructure, data systems, or ML infrastructure; distributed systems and scalable design expertise; imaging model experience; API, pipeline, cloud, and production engineering experience; programming proficiency.
Python, TypeScript, C++, Go, Kafka, Spark, Flink, LLMs, MLOps
3mo
Save
Mark Applied
Hide
Senior AI/ML Platform Engineer
San Mateo, California, United States
$148k-$247k/yr HybridFull Time
Guidewire
GuidewireNYSE: GWRE: Provides a software platform for property and casualty insurers.
10+ YOE10+ years software engineering; 5+ years ML platforms/infrastructure; distributed systems; Python/Go/Java; Docker/Kubernetes; MLOps tools; cloud experience; knowledge of ML models.
Python, Go, Java, Docker, Kubernetes, MLflow, Kubeflow, SageMaker, Vertex AI, Databricks, AWS, GCP, Azure
4w
Save
Mark Applied
Hide
Senior AI / ML Engineer
San Francisco or Melbourne
OnsiteFull Time
Artificial Analysis
Artificial Analysis: Independent AI benchmarking and performance analysis platform.
3+ YOE3+ years software engineering experience, Python and pandas proficiency, OpenAI API experience, cloud infrastructure familiarity, data visualization and communication skills; Bachelor’s or Master’s preferred.
Python, pandas, OpenAI, PyTorch
1w
Save
Mark Applied
Hide
Senior Staff AI/ML Engineer
Santa Clara, California, United States
$131k-$261k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
8+ YOEBachelor's degree in a related field with 8+ years, master's with 6+ years, or Ph.D. with 4+ years. Requires ML infrastructure, cloud, MLOps, distributed systems, CI/CD, and technical leadership experience.
Azure ML SDK, Azure Data Factory, Databricks, Spark, SQL, NoSQL, AWS, Azure, Google Cloud, Docker, Kubernetes, Airflow, Terraform, CI/CD, RAG, MCP, Kanban, Scrum
2mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
2mo
Save
Mark Applied
Hide
Senior Applied AI and AI Infrastructure Engineer - Chip Design and DFX
Santa Clara, California, United States
$200k-$380k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEAdvanced degree or equivalent experience in engineering with 6+–12+ yrs experience in AI infrastructure, applied ML/GenAI; experience with agents, SQL/ETL/data modeling, cloud (AWS/Azure/GCP), Python and C++; strong communication.
SQL, ETL, AWS, Azure, GCP, Python, C++
2mo
Save
Mark Applied
Hide
Senior Applied AI and AI Infrastructure Engineer - Chip Design and DFX
Santa Clara, California, United States
$200k-$380k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOEBSEE/MSEE/PhD with 12+/10+/6+ years experience in AI infrastructure, applied ML and Gen AI. Experience with SQL, ETL, data modeling, cloud (AWS/Azure/GCP), Python, C++, distributed systems, and mentoring.
SQL, ETL, Python, C++, AWS, Azure, GCP
1mo
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yr OnsiteFull Time
Kodiak Robotics
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
PyTorch, PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL, WebDataset, MosaicML Streaming, MDS, Nsight, PyTorch Profiler, Python, C++, CUDA, Triton, NVLink, InfiniBand
2w
Save
Mark Applied
Hide
Software Engineer III, AI/ML, AI and Infrastructure
Sunnyvale, California, United States
$147k-$210k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
2+ YOEBachelor's degree or equivalent practical experience, 2+ years programming in Python or C++, and experience with ML infrastructure and a specialized ML area.
Python, C++, Machine Learning (ML), Machine Learning Infrastructure, Vertex AI, TPUs
1mo
Save
Mark Applied
Hide
Lead AI Infrastructure Engineer, Reinforcement Learning
Santa Clara, California, United States
$179k-$306k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Design and operate distributed RL training infrastructure at scale; strong systems experience in ML platforms; proficiency with PyTorch/JAX, NCCL/MPI-style distributed training, C++/Python performance tuning; Bachelor's degree required.
PyTorch, JAX, NCCL, MPI, C++, Python
3mo
Save
Mark Applied
Hide
Software Engineer, AI/ML (Infrastructure & Platform)
New York City or Tempe or San Francisco
$200k-$235k/yr HybridFull Time
Wealth.com
Wealth.com: Digital estate planning platform for wealth management professionals.
Build AI infrastructure and platform software; strong software engineering, distributed systems, production AI experience.
Python, TypeScript, C#, APIs, Distributed Systems, LLM/AI Systems, Tool/Skill Abstractions, LangChain, Data Pipelines
2w
Save
Mark Applied
Hide
Recommendation Architecture AI/ML Infrastructure Engineer Intern (Data-Arch-TikTok Live) - 2027 Summer
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteInternship
TikTok
TikTok: Global short-form video hosting and social media platform.
Pursuing a bachelor's or master's degree in AI, software development, computer science, computer engineering, or related field; strong C++ or Python skills; data structures, algorithms, systems, PyTorch or TensorFlow, and LLM knowledge.
C++, Python, PyTorch, TensorFlow, CUDA, Triton, FSDP, ZeRO, FlashAttention
1mo
Save
Mark Applied
Hide
Senior MLOps & AI Infrastructure Engineer
San Jose, California, United States
$149k-$216k/yr OnsiteFull Time
Altera
Altera: Manufacturer of field-programmable gate arrays and programmable logic devices.
10+ YOEBachelor's degree, 10+ years ML engineering/MLOps experience, strong Python, cloud ML platforms (AWS/GCP/Azure), Docker/Kubernetes, CI/CD, ML frameworks (PyTorch/TensorFlow/JAX), experience with MLflow/W&B and HPC schedulers.
PyTorch, TensorFlow, JAX, Hugging Face, scikit-learn, XGBoost, MLflow, Kubeflow, Airflow, Weights & Biases, DVC, Feast, AWS SageMaker, GCP Vertex AI, Azure ML, Terraform, CloudFormation, Docker, Kubernetes, Slurm, LSF, Python, Bash, Go, SQL, Prometheus, Grafana, ELK Stack, Evidently AI, Arize
3mo
Save
Mark Applied
Hide
Infrastructure Engineer
New York or San Francisco or United States
$165k-$200k/yr HybridFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
Kubernetes production experience; IaC (Terraform/Helm); cloud (AWS/GCP); Python/Node.js; CI/CD (GitHub Actions/Spacelift); security and ML/AI infrastructure familiarity.
Kubernetes, Terraform, Helm, Python, Node.js, GitHub Actions, Spacelift, AWS, GCP, PyTorch, TensorFlow, Bash
2mo
Save
Mark Applied
Hide
Founding Engineer, AI Infra
San Francisco, California, United States
HybridFull Time
Phenix Space
Phenix Space: Builds robotic systems for on-orbit satellite assembly and upgrades.
5+ YOE5+ years building or operating ML infrastructure; deep GPU and distributed training knowledge; experience with PyTorch/DeepSpeed/Megatron/Ray, inference stacks (vLLM, TGI, Triton), Python and C++/Rust/Go, Kubernetes and IaC, and observability tooling.
FlashAttention, CUDA, Triton, PyTorch, DeepSpeed, Megatron, Ray, vLLM, SGLang, TGI, Python, C++, Rust, Go, Kubernetes, Terraform, Pulumi, Prometheus, Grafana, OpenTelemetry, Llama 3, Qwen, DeepSeek
1mo
Save
Mark Applied
Hide
Lead AI/ML Engineer (GenAI & Agentic Systems)-iCloud
San Francisco, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Technical leadership applying production LLM systems, retrieval-assisted generation (RAG), agentic workflows, skills-based automation, evaluation and orchestration frameworks to improve cloud infrastructure efficiency and engineering productivity.
LLM, Retrieval-assisted generation (RAG), iCloud, Apple Intelligence, Private Cloud Compute (PCC)