109 infrastructure reliability engineer jobs at 82 companies in Fairfield, CA

PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
3w
Save
Mark Applied
Hide
Customer Reliability Engineer - Infrastructure
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Apache Airflow, AWS, Azure, CI/CD, GCP, Infrastructure as Code (IaC), Kubernetes, Linux, Python
2mo
Save
Mark Applied
Hide
Staff Infrastructure Reliability Engineer - Database & Storage
Seattle or San Francisco or Detroit or United States
$180k-$279k/yr HybridFull Time
Rocket Companies
Rocket CompaniesNYSE: RKT: Provides digital mortgage, real estate, and personal finance services.
7+ YOE7+ years AWS/cloud infra; 5+ years PostgreSQL/AWS services; Linux admin/scripting; mentoring; infrastructure as code and security; AI code generation tools; on-call readiness.
AWS, PostgreSQL, Aurora/RDS, S3, ElastiCache, OpenSearch, DynamoDB, Linux, Python, Infrastructure as Code, Security practices, AI code generation tools
3d
Save
Mark Applied
Hide
Reliability Engineer, R&D
Austin or New York City or San Francisco or Seattle
$203k-$232k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience in reliability engineering for infrastructure or complex hardware, building availability/RAM models, leading cross-discipline FMEAs, and mining field failure data.
2mo
Save
Mark Applied
Hide
Research Engineer, RL Infrastructure and Reliability (Knowledge Work)
San Francisco, California, United States
$350k-$850k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
5+ YOEExperienced Python engineer with ML/distributed systems background, on-call/incidence response, SRE mindset, able to read research code, ensure reliable environments.
Python, Distributed systems tooling, Observability stacks, SRE tooling
2w
Save
Mark Applied
Hide
Infrastructure Engineer
San Francisco, California, United States
$200k-$350k/yr OnsiteFull Time
Console
Console: Automates IT support and internal operations using AI agents.
5+ YOE5+ years infrastructure/platform/backend experience; hands-on AWS and Kubernetes; experience with Pulumi or Terraform; comfortable in TypeScript/Node/React codebases; production reliability, observability, and enterprise deployment experience.
AWS, Kubernetes, Pulumi, Terraform, TypeScript, Node, React, Slack, Microsoft Teams
1w
Save
Mark Applied
Hide
Software Engineer, Infrastructure & Reliability
San Francisco, California, United States
HybridFull Time
CrewAI
CrewAI: Platform for orchestrating collaborative multi-agent AI systems.
Experience building and operating production SaaS infrastructure: cloud, containers, CI/CD, observability, secrets, databases, and automation using Python/Ruby/Go/Bash.
AWS, Docker, CI/CD, GitHub Actions, ECS, ECR, Kubernetes, Helm, PostgreSQL, Redis, Celery, FastAPI, Rails, Sentry, OpenTelemetry, Python, Ruby, Go, Bash, Terraform
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - AI Infrastructure
San Francisco or United States
RemoteFull Time
Andromeda Cluster
Andromeda Cluster: AI compute orchestration platform for GPU clusters.
Senior SRE with GPU infra, distributed training, and networking expertise.
NVIDIA GPUs, InfiniBand, RoCE, NVLink, NCCL, CUDA, PyTorch, DeepSpeed, Megatron, FSDP, Linux, Kubernetes, Slurm, Terraform, Helm, Ansible, DCGM, nvidia-smi
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Builds advanced multimodal AI models and model optimization infrastructure.
Bachelor's degree or equivalent experience; distributed systems, cloud infrastructure or SRE; reliability tooling; incident response; strong cross-team communication.
Kubernetes, Docker, Cloud Platforms, Monitoring Tools, Automation
2w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
San Francisco, California, United States
$200k-$250k/yr OnsiteFull Time
Stuut
Stuut: Automates business accounts receivable and collections through AI agents.
7+ YOE7+ years in SRE/infrastructure or backend engineering. Experience with AWS, Kubernetes/EKS, Docker, observability, SLOs/SLIs, Python or TypeScript, CI/CD, and production-grade distributed systems.
Python, TypeScript, AWS, Kubernetes, EKS, Docker, FastAPI, Vue.js, PostgreSQL (RDS), CI/CD
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States or San Francisco or New York City
$101k-$199k/yr OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
1+ YOEMaster's or Bachelor's in CS/IT (or equivalent experience), 1+ years managing physical infrastructure, on-call experience, experience with large-scale cloud/distributed systems preferred, and ability to pass Microsoft security screening.
Azure, InfiniBand, GPUs
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
Ray, Kubernetes, Prometheus, Grafana, Go, Python, AWS, Azure, GCP, Linux kernel
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yr RemoteFull Time
Bedrock Ocean Exploration
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Python, Go, Bash, Terraform, AWS, Docker, Kubernetes, Prometheus, Grafana, ROS 2, ROS, Jetson, Linux, IAM
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
San Francisco, California, United States
$150k-$300k/yr OnsiteFull Time
Reducto
Reducto: AI platform extracting structured data from unstructured documents
5+ YOE5+ years building production infrastructure; proficient in Python; strong cloud, Kubernetes, networking, storage, and automation; focus on reliability.
Python, Kubernetes, Cloud platforms, Networking, Storage, Automation
2mo
Save
Mark Applied
Hide
Founding ML infrastructure Engineer
San Francisco or United States
$200k-$350k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Experience designing and operating large-scale distributed infrastructure; Kubernetes/Slurm; multi-cloud GPU; reliability and scheduling; startup mindset.
Kubernetes, Slurm, Scheduling, TensorRT-LLM, NCCL, InfiniBand, RoCE, CuTe, Triton, TileLang
1w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
1mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
3w
Save
Mark Applied
Hide
Staff Security Reliability Engineer
San Francisco, California, United States
$293k-$385k/yr OnsiteFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
10+ YOE10+ years operating mission-critical on-prem/hybrid infrastructure; strong SRE discipline, IaC and config management (Terraform, Chef, Ansible); identity and Azure/Microsoft Entra experience; observability and incident response.
Terraform, Chef, Ansible, Microsoft Entra, Azure, FleetDM, Azure Virtual Desktop
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Jose or Seattle or San Francisco
$159k-$302k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOEBachelor's or equivalent, 5+ years SRE/infrastructure/backend experience; Kubernetes, Docker, Terraform, AWS, Postgres/Redis, observability, incident response, CI/CD, bash, Node.js/TypeScript experience; on-call participation.
Kubernetes, Docker, bash, CircleCI, Node.js, TypeScript, Postgres, Redis, AWS Aurora (Postgres-compatible), Terraform, AWS

Explore Jobs

Expand Your Job Search