419 ai infrastructure engineer jobs at 204 companies in Ross, CA
6d
Save
Mark Applied
Hide
6d
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yrOnsiteFull Time
Sciforium: Building multimodal AI models and high-performance model serving infrastructure.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Together AI: Cloud platform for training and deploying artificial intelligence models.
5+ YOE5+ years in AI infrastructure or related roles; BS in CS or equivalent; knowledge of Ansible, Terraform, Kubernetes; programming/scripting; monitoring/observability; cloud services; collaborative work
Paradigm: Venture capital firm focused on crypto and frontier technologies.
Experienced engineer with infra, security, and AI model experience; familiarity with durable control planes, runtimes, secrets, observability, and self-hosted deployment.
Luma AI: Develops multimodal AI for video generation and creative production.
Deep Linux and distributed systems expertise, experience operating GPU/accelerator clusters, Kubernetes fluency, debugging across hardware/kernel/runtime/orchestration, coding and automation skills, technical leadership and hiring experience.
Woodbridge or New York City or Atlanta or Boston or Chicago or Dallas or Delaware or Denver or Garden City or Cayman Islands or Greenwich or Houston or Los Angeles or Miami or Naples or Nevada or Palm Beach or San Diego or San Francisco or Seattle or Stuart or Washington
$160k-$200k/yrHybridFull Time
Bessemer Trust: Wealth management and family office services for affluent clients.
7+ YOE7+ years in DevOps/platform/cloud infrastructure engineering; strong AWS (IAM, CloudFormation, Lambda, API Gateway, VPC, CloudWatch, SSM, Secrets Manager, ECR); AWS CDK/CloudFormation/Terraform; CI/CD (Bitbucket/GitHub Actions, OIDC); container and datastore operations; security fundamentals.
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
TencentHKEX: 0700: Provides integrated internet services, digital entertainment, and cloud technology.
Master's or PhD in related field, expertise in GPGPU/AI accelerator architectures, low-level operator development (CUDA,Triton), distributed systems knowledge, and experience optimizing large-scale accelerator clusters and DL frameworks.
Scale AI: Provides data and infrastructure for training artificial intelligence models.
4+ YOE4+ years building high-performance systems software; deep Linux internals, containerization/virtualization, systems programming (Go/Rust/C/C++); strong debugging and API/SDK design skills.
New York City or Seattle or San Francisco or London
$140k-$274k/yrHybridFull Time
Writer: Platform for building and deploying enterprise generative AI agents.
5+ YOE5+ years infrastructure/DevOps experience, production Kubernetes, Helm, Terraform/Pulumi, major cloud (AWS preferred), Python or Go, observability stacks (Prometheus, Grafana, ELK), and daily AI-assisted workflows.
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yrOnsiteFull Time
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years infrastructure engineering experience, formal training/certification, deep cloud and network knowledge, scripting and automation experience, familiarity with security/segmentation and AI-assisted engineering, strong problem-solving and mentoring skills.
Cisco, Juniper, Arista, Juniper Mist, Cisco/Viptela, Fortinet, BGP, OSPF, MPLS, EVPN/VXLAN, SD-WAN, Wi Fi 6E, Wi Fi 7, 5G, Palo Alto, Zscaler, IPsec, TLS, ZTNA, Python, Ansible, Terraform, Git, GitHub Actions, Jenkins, ThousandEyes, Splunk, Grafana, Wireshark
Sirius Technology: AI-powered retention platform for subscription-based businesses.
1+ YOE1+ year in infrastructure/DevOps/platform engineering or solutions architecture with customer-facing deployment experience; deep experience with AWS, Terraform, container orchestration, and cloud networking (VPC, IAM, DNS); strong operational and stakeholder communication skills.
Phenix Space: Builds robotic systems for on-orbit satellite assembly and upgrades.
5+ YOE5+ years building or operating ML infrastructure; deep GPU and distributed training knowledge; experience with PyTorch/DeepSpeed/Megatron/Ray, inference stacks (vLLM, TGI, Triton), Python and C++/Rust/Go, Kubernetes and IaC, and observability tooling.
Miami or New York or San Francisco or Mexico City or United States
RemoteFull Time
Félix Pago: Remittance platform for money transfers via WhatsApp messaging.
8+ YOE8+ years software engineering experience with infrastructure and systems development; 2+ years building production LLM/AI applications; strong Python, system architecture, containerization (Docker/Kubernetes), CI/CD, observability, RAG and agentic systems experience.