144 sre engineer jobs at 110 companies in Fairfield, CA

PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
3w
Save
Mark Applied
Hide
Sr SRE/Dev Ops Engineer
San Francisco or United States
$170k-$175k/yr OnsiteFull Time
Madison Reed
Madison Reed: Direct-to-consumer hair color products and salon services provider.
5+ YOE5+ years in SRE/DevOps/platform engineering with strong cloud (preferably AWS), IaC (Terraform/CloudFormation), CI/CD, observability, on-call experience, and scripting (Python/Bash/Go).
AWS, Terraform, CloudFormation, Python, Bash, Go, Datadog, New Relic, Grafana, Prometheus, OpenTelemetry, Splunk
3mo
Save
Mark Applied
Hide
Senior SRE
San Francisco, California, United States
$135k-$159k/yr OnsiteFull Time
LiveRamp
LiveRampNew York Stock Exchange: RAMP: Provides a platform for secure data collaboration and identity.
5+ YOE5+ years in SRE/DevOps, IaC with Terraform, CI/CD (Jenkins/CircleCI), Kubernetes and cloud (GCP/AWS), Python or Go, observability, and ability to mentor others.
Terraform, Kubernetes, Jenkins, CircleCI, GCP, AWS, Python, Go, CI/CD, Observability
2mo
Save
Mark Applied
Hide
SRE/Infrastructure Engineer
San Francisco, California, United States
$200k-$350k/yr OnsiteFull Time
E2B
E2B: Open-source cloud infrastructure for running autonomous AI agents.
5+ YOE5+ years producing production cloud infrastructure; strong Terraform and Kubernetes; multi-cloud experience; BYOC deployments; in-person in San Francisco.
Terraform, Kubernetes, Nomad, Google Cloud, Amazon Web Services, Azure, Cloudflare, Go, YAML
1mo
Save
Mark Applied
Hide
Platform Engineer (SRE) - AI Control Plane
San Francisco, California, United States
OnsiteFull Time
Speakeasy
Speakeasy: Automates API SDK and documentation generation for developers.
Platform Engineer (SRE) to own reliability, design deployments, and participate in on-call; strong systems and software engineering.
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Builds advanced multimodal AI models and model optimization infrastructure.
Bachelor's degree or equivalent experience; distributed systems, cloud infrastructure or SRE; reliability tooling; incident response; strong cross-team communication.
Kubernetes, Docker, Cloud Platforms, Monitoring Tools, Automation
13h
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
1mo
Save
Mark Applied
Hide
Senior SRE Engineer - San Francisco
San Francisco, California, United States
HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
AWS, GCP, Azure, Kubernetes, Go, Python, Java, Cursor, GPT models, Gemini, Claude
3mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Palo Alto or San Francisco
$170k-$230k/yr HybridFull Time
Mithril
Mithril: Orchestrates global compute capacity specifically for AI workloads.
3+ YOE3+ years in SRE/Production Engineering; Kubernetes; cloud (AWS/GCP/Azure); Python or Go; Linux; RCA; strong communicator.
Kubernetes, Terraform, Pulumi, Python, Go, Linux, Prometheus, Grafana, OpenTelemetry
3mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Madrid or Lisbon or San Francisco
€55k-€68k/yr RemoteFull Time
Air Apps
Air Apps: Develops AI-powered productivity and utility mobile applications.
4+ YOE4+ years in SRE/DevOps/System Eng; cloud platforms (AWS/Azure/GCP); observability tools; IaC; containers; Linux; incident management; scripting; security; on-call.
Prometheus, Grafana, Datadog, ELK, Terraform, CloudFormation, Pulumi, Docker, Kubernetes, Helm, Python, Go, Bash
1mo
Save
Mark Applied
Hide
Senior Software Engineer - SRE
United States or Carson City or San Francisco or Seattle or New York
$160k-$180k/yr HybridFull Time
Socure
Socure: Provide AI-driven identity verification and fraud prevention software.
Proven experience building, running, and scaling production systems with deep AWS, Terraform, Kubernetes/EKS, Go or Python, CI/CD, GitHub/ArgoCD, and observability (Datadog, SLIs/SLOs).
AWS, Terraform, Kubernetes, Amazon EKS, Go, Python, GitHub, GitHub Actions, ArgoCD, Datadog
2w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
3mo
Save
Mark Applied
Hide
Observability Lead - Cloud SRE & Network Reliability
Fremont, California, United States
$114k-$253k/yr HybridFull Time
Lam Research
Lam ResearchNASDAQ: LRCX: Designs and manufactures wafer fabrication equipment for the semiconductor industry.
12+ YOE6+ MgmtSenior SRE leader with 12+ years infrastructure/SRE/DevOps experience and 6+ years leading teams; deep multi-cloud networking, observability, DR/BCP, backup/restore, automation (Ansible/Terraform/Python), Kubernetes, and incident management experience.
Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, Google Cloud Operations, Splunk, Ansible, Terraform, Python, Kubernetes, AKS, EKS, GKE, ServiceNow
3mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Burlingame, California, United States
$170k-$197k/yr OnsiteFull Time
Xona Space Systems
Xona Space Systems: Operates low Earth orbit satellites for high-precision navigation services.
4+ YOE4+ years cloud operations in AWS/GCP/Azure; Kubernetes (EKS); IaC (Terraform, Ansible, Helm); automation and CI/CD; observability; data systems; Python/C++; Linux internals.
AWS, GCP, Azure, Kubernetes (EKS), Terraform, Ansible, Helm, Python, C++, Prometheus, Grafana, ELK
1mo
Save
Mark Applied
Hide
Senior Infrastructure Engineer
San Francisco, California, United States
$180k-$215k/yr OnsiteFull Time
Casca
Casca: AI-powered loan origination and processing platform for financial institutions.
5+ YOE5+ years in infrastructure, platform engineering, or SRE; strong cloud (AWS); IaC; CI/CD; programming; observability; containerization.
AWS, Terraform, CDK, Pulumi, GitHub Actions, BuildKite, Docker, Cloudflare, Python, Go, TypeScript, Rust
3w
Save
Mark Applied
Hide
Platform Engineer I
Pleasanton, California, United States
$32-$41/hr HybridFull Time
Blackhawk Network
Blackhawk Network: Provider of global branded payment and gift card solutions.
Bachelor's in CS/Engineering or equivalent; experience in platform/DevOps/SRE or similar; strong Linux, AWS, Git; scripting with Python/Bash; production support and incident management; experience with AI-assisted engineering tools.
AWS, Git, Python, Bash, Kubernetes, Docker, Jenkins, Splunk, New Relic, Prometheus, Grafana, OpenTelemetry, ServiceNow, Terraform, CloudFormation, GitHub Copilot, Cursor, Claude
1mo
Save
Mark Applied
Hide
Senior Platform Engineer
Pleasanton or California or United States
HybridFull Time
STN
STN: Provides high-performance GPU infrastructure, cloud, and managed IT services.
6+ YOE6+ years in platform/SRE/cloud engineering, deep Kubernetes expertise, Go and/or Python programming, experience operating GPU/AI infrastructure, bachelor's degree or equivalent experience.
Kubernetes, Slurm, Run:ai, Go, Python, NVIDIA GPU Operator, MIG, MPS, NCCL, KubeRay, Istio, Linkerd, CRDs
2mo
Save
Mark Applied
Hide
Senior Production Engineer, Oeprational Excellence
San Francisco or Sunnyvale
$172k-$209k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
5+ YOE5+ years in Production Engineering or SRE; GPU workloads; Linux; IaC; Kubernetes; Go/Python; strong communication.
Prometheus, Grafana, OpenTelemetry, Terraform, Ansible, Kubernetes, AWS, GCP, Linux
1mo
Save
Mark Applied
Hide
Staff Platform Engineer
San Francisco or New York City or Colorado or California or Washington
$220k-$331k/yr OnsiteFull Time
Amplitude
AmplitudeNasdaq: AMPL: Develops digital analytics software for tracking customer product behavior.
8+ YOE8+ years in software/DevOps/SRE, Bachelor\u000degree in Computer Engineering (required), deep Kubernetes and cloud experience, Terraform/IaC and programming (Golang or Python), strong cross-team leadership and communication.
Kubernetes, EKS, GKE, AKS, Terraform, Helm, Kustomize, Argo CD, Argo Workflows, Argo Rollouts, GitHub Actions, Datadog, Amplitude, Golang, Python, EC2, IAM, VPC, ALB, S3, Backstage, Envoy, GitOps, LLM
3w
Save
Mark Applied
Hide
Platform Engineer I
Pleasanton, California, United States
$41/hr HybridFull Time
Blackhawk Network
Blackhawk Network: Provider of branded payment and gift card solutions.
Bachelor's in CS/Engineering or equivalent experience; platform/DevOps/SRE experience; strong Linux, AWS, Git; scripting in Python/Bash; experience with major incident management and AI-assisted development tools.
AWS, Kubernetes, Git, Python, Bash, GitHub Copilot, Cursor, Claude, Docker, Jenkins, Splunk, New Relic, Prometheus, Grafana, OpenTelemetry, ServiceNow, Terraform, CloudFormation, Linux
1w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD