249 ai ml infrastructure engineer jobs at 105 companies in Tiburon, CA
3w
Save
Mark Applied
Hide
3w
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yrOnsiteFull Time
Sciforium: Building multimodal AI models and high-performance model serving infrastructure.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Lisbon or Atlanta or London or San Francisco or Santiago or Sydney or Tokyo or Toronto or Australia or Canada or United States
HybridFull Time
PagerDutyNYSE: PD: Platform for real-time incident response and digital operations automation.
5+ YOE5+ years of software engineering experience building production distributed systems and AI systems, with expertise in LLMs, agents, retrieval, cloud infrastructure, containers, Kubernetes, reliability, and evaluation.
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
8+ YOEBachelor's degree and 8+ years, master's and 6+ years, or Ph.D. and 4+ years of related experience. Requires ML infrastructure, cloud, MLOps, AI governance, distributed systems, and leadership expertise.
AWS, Azure, Google Cloud, Docker, Kubernetes, Airflow, Terraform, Azure ML SDK, Azure Data Factory, Databricks, Spark, SQL, NoSQL, CI/CD
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
12+ YOEBachelor's degree or 4+ years equivalent experience; 12+ years software, data science, or analytics experience; 3+ years AI/ML; 4+ years Python, cloud infrastructure, and technical leadership.
Python, AWS, Azure, GCP, Google Cloud Platform, FHIR, HL7, HIPAA, GitHub, GitHub Actions, Docker, Terraform, CI/CD, Large Language Models (LLMs)
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOE5+ years in software engineering, backend infrastructure, data systems, or ML infrastructure; distributed systems and scalable design expertise; imaging model experience; API, pipeline, cloud, and production engineering experience; programming proficiency.
GuidewireNYSE: GWRE: Provides a software platform for property and casualty insurers.
10+ YOE10+ years software engineering; 5+ years ML platforms/infrastructure; distributed systems; Python/Go/Java; Docker/Kubernetes; MLOps tools; cloud experience; knowledge of ML models.
Artificial Analysis: Independent AI benchmarking and performance analysis platform.
3+ YOE3+ years software engineering experience, Python and pandas proficiency, OpenAI API experience, cloud infrastructure familiarity, data visualization and communication skills; Bachelor’s or Master’s preferred.
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
8+ YOEBachelor's degree in a related field with 8+ years, master's with 6+ years, or Ph.D. with 4+ years. Requires ML infrastructure, cloud, MLOps, distributed systems, CI/CD, and technical leadership experience.
Azure ML SDK, Azure Data Factory, Databricks, Spark, SQL, NoSQL, AWS, Azure, Google Cloud, Docker, Kubernetes, Airflow, Terraform, CI/CD, RAG, MCP, Kanban, Scrum
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Senior Applied AI and AI Infrastructure Engineer - Chip Design and DFX
Santa Clara, California, United States
$200k-$380k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEAdvanced degree or equivalent experience in engineering with 6+–12+ yrs experience in AI infrastructure, applied ML/GenAI; experience with agents, SQL/ETL/data modeling, cloud (AWS/Azure/GCP), Python and C++; strong communication.
Senior Applied AI and AI Infrastructure Engineer - Chip Design and DFX
Santa Clara, California, United States
$200k-$380k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOEBSEE/MSEE/PhD with 12+/10+/6+ years experience in AI infrastructure, applied ML and Gen AI. Experience with SQL, ETL, data modeling, cloud (AWS/Azure/GCP), Python, C++, distributed systems, and mentoring.
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yrOnsiteFull Time
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
2+ YOEBachelor's degree or equivalent practical experience, 2+ years programming in Python or C++, and experience with ML infrastructure and a specialized ML area.
Lead AI Infrastructure Engineer, Reinforcement Learning
Santa Clara, California, United States
$179k-$306k/yrHybridFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Design and operate distributed RL training infrastructure at scale; strong systems experience in ML platforms; proficiency with PyTorch/JAX, NCCL/MPI-style distributed training, C++/Python performance tuning; Bachelor's degree required.
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteInternship
TikTok: Global short-form video hosting and social media platform.
Pursuing a bachelor's or master's degree in AI, software development, computer science, computer engineering, or related field; strong C++ or Python skills; data structures, algorithms, systems, PyTorch or TensorFlow, and LLM knowledge.
Altera: Manufacturer of field-programmable gate arrays and programmable logic devices.
10+ YOEBachelor's degree, 10+ years ML engineering/MLOps experience, strong Python, cloud ML platforms (AWS/GCP/Azure), Docker/Kubernetes, CI/CD, ML frameworks (PyTorch/TensorFlow/JAX), experience with MLflow/W&B and HPC schedulers.
Phenix Space: Builds robotic systems for on-orbit satellite assembly and upgrades.
5+ YOE5+ years building or operating ML infrastructure; deep GPU and distributed training knowledge; experience with PyTorch/DeepSpeed/Megatron/Ray, inference stacks (vLLM, TGI, Triton), Python and C++/Rust/Go, Kubernetes and IaC, and observability tooling.