419 ai infrastructure engineer jobs at 204 companies in Ross, CA

6d
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yr OnsiteFull Time
Sciforium
Sciforium: Building multimodal AI models and high-performance model serving infrastructure.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Ansible, SaltStack, Git, Python, Bash, Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI, enroot, pyxis, Docker, containerd, NVIDIA Container Toolkit, CUDA, cuDNN, NCCL, Fabric Manager, ROCm, RCCL, DKMS, GPUDirect RDMA, GPUDirect Storage, MOFED, DOCA, PyTorch, JAX, DCGM exporter, Prometheus, Grafana, PXE, MaaS, Packer, Foreman, Terraform, Lustre, GPFS, Weka, vLLM, Triton Inference Server, TensorRT-LLM, Nsight Systems, Nsight Compute, rocprof, perf, eBPF, EMR
2mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$190k-$270k/yr OnsiteFull Time
Together AI
Together AI: Cloud platform for training and deploying artificial intelligence models.
5+ YOE5+ years in AI infrastructure or related roles; BS in CS or equivalent; knowledge of Ansible, Terraform, Kubernetes; programming/scripting; monitoring/observability; cloud services; collaborative work
Ansible, Terraform, Kubernetes
4w
Save
Mark Applied
Hide
AI Infrastructure Engineer
Fremont, California, United States
OnsiteFull Time
AMAX
AMAXTaiwan Stock Exchange: 6933: Designs and manufactures GPU-accelerated AI and HPC computing infrastructure.
Experience with on-prem/datacenter operations, Infrastructure-as-Code, networking (VLANs/routing/firewalls), container orchestration, scripting, Git; comfortable with hands-on hardware tasks.
Terraform, Terragrunt, Ansible, Vault, Boundary, Keycloak, Prometheus, Grafana, Alertmanager, Docker, Kubernetes, Git, Jira, Confluence
1mo
Save
Mark Applied
Hide
Cloud Infrastructure and AI Efficiency Engineer
San Francisco, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Engineering role focused on cloud infrastructure and AI efficiency; specific experience and certifications not provided in the description.
1w
Save
Mark Applied
Hide
Infrastructure Engineer, Applied AI
San Francisco, California, United States
$250k-$400k/yr OnsiteFull Time
Paradigm
Paradigm: Venture capital firm focused on crypto and frontier technologies.
Experienced engineer with infra, security, and AI model experience; familiarity with durable control planes, runtimes, secrets, observability, and self-hosted deployment.
Reth, Foundry, EVMBench, OpenAI, Centaur, Kubernetes
2w
Save
Mark Applied
Hide
Staff AI Infrastructure Engineer
Redwood City, California, United States
HybridFull Time
Luma AI
Luma AI: Develops multimodal AI for video generation and creative production.
Deep Linux and distributed systems expertise, experience operating GPU/accelerator clusters, Kubernetes fluency, debugging across hardware/kernel/runtime/orchestration, coding and automation skills, technical leadership and hiring experience.
Linux, Kubernetes
2mo
Save
Mark Applied
Hide
Gen AI Infrastructure Engineer
Woodbridge or New York City or Atlanta or Boston or Chicago or Dallas or Delaware or Denver or Garden City or Cayman Islands or Greenwich or Houston or Los Angeles or Miami or Naples or Nevada or Palm Beach or San Diego or San Francisco or Seattle or Stuart or Washington
$160k-$200k/yr HybridFull Time
Bessemer Trust
Bessemer Trust: Wealth management and family office services for affluent clients.
7+ YOE7+ years in DevOps/platform/cloud infrastructure engineering; strong AWS (IAM, CloudFormation, Lambda, API Gateway, VPC, CloudWatch, SSM, Secrets Manager, ECR); AWS CDK/CloudFormation/Terraform; CI/CD (Bitbucket/GitHub Actions, OIDC); container and datastore operations; security fundamentals.
AWS Bedrock, AgentCore, Lambda, API Gateway, AWS CDK, CloudFormation, Terraform, Bitbucket Pipelines, GitHub Actions, OIDC, IAM, VPC, CloudWatch, SSM, Secrets Manager, ECR, Neo4j, Neptune, Redis, Milvus, VectorDB
1mo
Save
Mark Applied
Hide
Infrastructure Engineer
Menlo Park, California, United States
OnsiteFull Time
Shakudo
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
Kubernetes, Rust, CI/CD, DGX, GPU, LLM, ETL
1w
Save
Mark Applied
Hide
Sr. Cloud AI Infrastructure Engineer
Palo Alto, California, United States
$145k-$273k/yr OnsiteFull Time
Tencent
TencentHKEX: 0700: Provides integrated internet services, digital entertainment, and cloud technology.
Master's or PhD in related field, expertise in GPGPU/AI accelerator architectures, low-level operator development (CUDA,Triton), distributed systems knowledge, and experience optimizing large-scale accelerator clusters and DL frameworks.
CUDA, Triton, PyTorch, TensorFlow
3w
Save
Mark Applied
Hide
AI Infrastructure Engineer, Sandbox Platform
San Francisco or Seattle or New York City
$180k-$225k/yr OnsiteFull Time
Scale AI
Scale AI: Provides data and infrastructure for training artificial intelligence models.
4+ YOE4+ years building high-performance systems software; deep Linux internals, containerization/virtualization, systems programming (Go/Rust/C/C++); strong debugging and API/SDK design skills.
Docker, Firecracker, gVisor, QEMU, Kata Containers, Go, Rust, C/C++, Kubernetes, OpenHands, Agent2Agent, MCP, CRIU
1w
Save
Mark Applied
Hide
Infrastructure engineer
New York City or Seattle or San Francisco or London
$140k-$274k/yr HybridFull Time
Writer
Writer: Platform for building and deploying enterprise generative AI agents.
5+ YOE5+ years infrastructure/DevOps experience, production Kubernetes, Helm, Terraform/Pulumi, major cloud (AWS preferred), Python or Go, observability stacks (Prometheus, Grafana, ELK), and daily AI-assisted workflows.
Python, Go, AWS, GCP, Azure, Kubernetes, Helm, Terraform, Pulumi, Claude Code, Droid, Codex, Prometheus, Grafana, ELK
1mo
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yr OnsiteFull Time
Kodiak Robotics
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
PyTorch, PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL, WebDataset, MosaicML Streaming, MDS, Nsight, PyTorch Profiler, Python, C++, CUDA, Triton, NVLink, InfiniBand
1mo
Save
Mark Applied
Hide
Lead Infrastructure Engineer-Network Engineer
San Francisco or Seattle
$143k-$185k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years infrastructure engineering experience, formal training/certification, deep cloud and network knowledge, scripting and automation experience, familiarity with security/segmentation and AI-assisted engineering, strong problem-solving and mentoring skills.
Cisco, Juniper, Arista, Juniper Mist, Cisco/Viptela, Fortinet, BGP, OSPF, MPLS, EVPN/VXLAN, SD-WAN, Wi Fi 6E, Wi Fi 7, 5G, Palo Alto, Zscaler, IPsec, TLS, ZTNA, Python, Ansible, Terraform, Git, GitHub Actions, Jenkins, ThousandEyes, Splunk, Grafana, Wireshark
2mo
Save
Mark Applied
Hide
Forward Deployed Infrastructure Engineer
San Francisco, California, United States
OnsiteFull Time
Sirius Technology: AI-powered retention platform for subscription-based businesses.
1+ YOE1+ year in infrastructure/DevOps/platform engineering or solutions architecture with customer-facing deployment experience; deep experience with AWS, Terraform, container orchestration, and cloud networking (VPC, IAM, DNS); strong operational and stakeholder communication skills.
AWS, Terraform, VPC, IAM, DNS, container orchestration, CRM, AI/ML, LLM
2mo
Save
Mark Applied
Hide
AI Engineer, Agent Infrastructure
San Francisco, California, United States
OnsiteFull Time
Zed
Zed: AI-native neobank providing premium credit services to young professionals.
Experience shipping production LLM/agent systems; strong backend/infrastructure; familiarity with workflows, observability, and production readiness.
Python, Go, REST APIs, Docker, Kubernetes, LLMs, Observability, Monitoring
2mo
Save
Mark Applied
Hide
Founding Engineer, AI Infra
San Francisco, California, United States
HybridFull Time
Phenix Space
Phenix Space: Builds robotic systems for on-orbit satellite assembly and upgrades.
5+ YOE5+ years building or operating ML infrastructure; deep GPU and distributed training knowledge; experience with PyTorch/DeepSpeed/Megatron/Ray, inference stacks (vLLM, TGI, Triton), Python and C++/Rust/Go, Kubernetes and IaC, and observability tooling.
FlashAttention, CUDA, Triton, PyTorch, DeepSpeed, Megatron, Ray, vLLM, SGLang, TGI, Python, C++, Rust, Go, Kubernetes, Terraform, Pulumi, Prometheus, Grafana, OpenTelemetry, Llama 3, Qwen, DeepSeek
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
New York or San Francisco or United States
$165k-$200k/yr HybridFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
Kubernetes production experience; IaC (Terraform/Helm); cloud (AWS/GCP); Python/Node.js; CI/CD (GitHub Actions/Spacelift); security and ML/AI infrastructure familiarity.
Kubernetes, Terraform, Helm, Python, Node.js, GitHub Actions, Spacelift, AWS, GCP, PyTorch, TensorFlow, Bash
2mo
Save
Mark Applied
Hide
AI Engineer
San Francisco, California, United States
OnsiteFull Time
Emanate
Emanate: A technology building AI-powered revenue infrastructure.
Backend/AI engineer role focusing on AI infrastructure, LLM integration, data pipelines, and scalable autonomous workflows.
1mo
Save
Mark Applied
Hide
Staff AI Engineer
Miami or New York or San Francisco or Mexico City or United States
RemoteFull Time
Félix Pago
Félix Pago: Remittance platform for money transfers via WhatsApp messaging.
8+ YOE8+ years software engineering experience with infrastructure and systems development; 2+ years building production LLM/AI applications; strong Python, system architecture, containerization (Docker/Kubernetes), CI/CD, observability, RAG and agentic systems experience.
Python, Docker, Kubernetes, CI/CD, LLM APIs, LLMOps, RAG, vector databases, Open CLAW, Hermes, WhatsApp
2mo
Save
Mark Applied
Hide
Founding Infrastructure Engineer
San Francisco, California, United States
$130k-$200k/yr OnsiteFull Time
Virio
Virio: A fast-growing technology startup focused on scalable infrastructure for AI-driven workloads.
Build and scale distributed infrastructure, own core services, design for reliability and performance, and collaborate with product and AI teams.