270 ai ml infrastructure engineer jobs at 110 companies in Half Moon Bay, CA
3w
Save
Mark Applied
Hide
3w
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yrOnsiteFull Time
Sciforium: AI infrastructure building multimodal models and high-efficiency serving software for developers and teams.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Lisbon or Atlanta or London or San Francisco or Santiago or Sydney or Tokyo or Toronto or Australia or Canada or United States
HybridFull Time
PagerDutyNYSE: PD: Public software providing AI-powered digital operations and incident management software to business teams.
5+ YOE5+ years of software engineering experience building production distributed systems and AI systems, with expertise in LLMs, agents, retrieval, cloud infrastructure, containers, Kubernetes, reliability, and evaluation.
AbbottNYSE: ABT: Global healthcare technology focused on life-changing medical innovations.
8+ YOEBachelor's degree and 8+ years, master's and 6+ years, or Ph.D. and 4+ years of related experience. Requires ML infrastructure, cloud, MLOps, AI governance, distributed systems, and leadership expertise.
AWS, Azure, Google Cloud, Docker, Kubernetes, Airflow, Terraform, Azure ML SDK, Azure Data Factory, Databricks, Spark, SQL, NoSQL, CI/CD
Optum Insight: Healthcare analytics and technology-services business serving payers, providers, governments and life sciences companies.
12+ YOEBachelor's degree or 4+ years equivalent experience; 12+ years software, data science, or analytics experience; 3+ years AI/ML; 4+ years Python, cloud infrastructure, and technical leadership.
Python, AWS, Azure, GCP, Google Cloud Platform, FHIR, HL7, HIPAA, GitHub, GitHub Actions, Docker, Terraform, CI/CD, Large Language Models (LLMs)
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
5+ YOE5+ years in software engineering, backend infrastructure, data systems, or ML infrastructure; distributed systems and scalable design expertise; imaging model experience; API, pipeline, cloud, and production engineering experience; programming proficiency.
Guidewire SoftwareNYSE: GWRE: Cloud platform provider for property and casualty insurance carriers.
10+ YOE10+ years software engineering; 5+ years ML platforms/infrastructure; distributed systems; Python/Go/Java; Docker/Kubernetes; MLOps tools; cloud experience; knowledge of ML models.
Artificial Analysis: Independent AI benchmarking and analysis helping developers, researchers, and businesses choose AI technologies.
3+ YOE3+ years software engineering experience, Python and pandas proficiency, OpenAI API experience, cloud infrastructure familiarity, data visualization and communication skills; Bachelor’s or Master’s preferred.
AbbottNYSE: ABT: Global healthcare technology focused on life-changing medical innovations.
8+ YOEBachelor's degree in a related field with 8+ years, master's with 6+ years, or Ph.D. with 4+ years. Requires ML infrastructure, cloud, MLOps, distributed systems, CI/CD, and technical leadership experience.
Azure ML SDK, Azure Data Factory, Databricks, Spark, SQL, NoSQL, AWS, Azure, Google Cloud, Docker, Kubernetes, Airflow, Terraform, CI/CD, RAG, MCP, Kanban, Scrum
Docker: Privately held container application platform helping developers build, share, and run applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yrOnsiteFull Time
Kodiak RoboticsNasdaq: KDK: Public autonomous-vehicle technology serving commercial trucking, industrial trucking, defense, and public-sector customers.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
YouTube: Google-owned video-sharing and content-distribution platform serving viewers, creators, and advertisers.
8+ YOEBachelor's degree or equivalent,8+ years software development,5+ years product launches,experience with large-scale distributed systems and ML infrastructure,EMR not mentioned,TensorFlow/PyTorch proficiency preferred.
Lead AI Infrastructure Engineer, Reinforcement Learning
Santa Clara, California, United States
$179k-$306k/yrHybridFull Time
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Design and operate distributed RL training infrastructure at scale; strong systems experience in ML platforms; proficiency with PyTorch/JAX, NCCL/MPI-style distributed training, C++/Python performance tuning; Bachelor's degree required.
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteInternship
TikTok: Short-form mobile video and social media platform.
Pursuing a bachelor's or master's degree in AI, software development, computer science, computer engineering, or related field; strong C++ or Python skills; data structures, algorithms, systems, PyTorch or TensorFlow, and LLM knowledge.
Goaly: Cox Enterprises' AI accelerator and early-stage venture studio backing startups with capital and technical expertise.
5+ YOE5+ years building or operating ML infrastructure; deep GPU and distributed training knowledge; experience with PyTorch/DeepSpeed/Megatron/Ray, inference stacks (vLLM, TGI, Triton), Python and C++/Rust/Go, Kubernetes and IaC, and observability tooling.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOEBachelor's degree or equivalent experience, 8+ years in infrastructure development, scalable cloud infrastructure expertise, AI/ML and data analytics background, and strong object-oriented programming skills in Java or Go.
Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure)
New York City or San Francisco or San Jose or Cambridge or McLean or Plano
$209k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: A technology-driven bank providing diverse financial services.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years developing AI/ML technologies; 6 years programming with Python, Go, Scala, or Java. Leadership and cloud AI experience preferred.
Cargomatic: Private logistics technology connecting shippers with trucking capacity for local and regional freight.
3+ YOE3+ years software engineering with focus on AI/ML or automation; hands-on AI-powered applications; experience with agentic AI concepts, LLMs, and modern AI frameworks; cloud infrastructure (AWS); backend (Node.js); frontend (React).