337 sre jobs at 187 companies in San Francisco, CA

3mo
Save
Mark Applied
Hide
SRE
San Francisco or New York or Remote
$135k-$285k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
Kubernetes expert; scalable infra experience; strong observability; IaC and GitOps; runbooks and incident response; ML knowledge not required.
Kubernetes, EKS, GKE, Terraform, Helm, Flux CD, ArgoCD, Prometheus, Loki, ELK, Grafana, incident.io
2w
Save
Mark Applied
Hide
SRE Engineer - Cloud Environment
Santa Clara, California, United States
$88k-$112k/yr OnsiteFull Time
Bambu Lab
Bambu Lab: Manufacturer of desktop 3D printers and 3D printing accessories.
3+ YOE3+ years SRE/DevOps experience, proficiency with Linux, TCP/IP networking, Kubernetes, cloud (AWS/GCP), IaC (Terraform), scripting (Python/Go/Shell), and ability to support on-call rotation.
Kubernetes, Terraform, Python, Go, Shell, AWS, GCP, Linux, TCP/IP
23h
Save
Mark Applied
Hide
Senior Director – Observability | SRE
San Francisco or Dallas
OnsiteFull Time
Gap Inc.
Gap Inc.NYSE: GAP: Global specialty retailer of apparel and accessories.
10+ YOEStrategic technology leader with 10+ years driving operational transformation, expertise in ITIL, SRE, architecture, observability, infrastructure, cloud operations, automation, and service management, plus large-team leadership experience.
ITIL, SRE, Live Sight Insights
3mo
Save
Mark Applied
Hide
Engineering Lead – Platform & SRE
Santa Clara, California, United States
$175k-$215k/yr HybridFull Time
Kerrigan Robotics
Kerrigan Robotics: AI-powered orchestration platform for factory automation systems.
5+ YOE5+ years SRE/DevOps experience; deep Kubernetes knowledge; expertise with Pulumi/Terraform, Helm; proficiency in Go, Linux networking, observability (Prometheus, Grafana), and CI/CD (GitHub Actions).
Pulumi, Kubernetes, AWS, GCP, Azure, GitHub Actions, Prometheus, Grafana, OIDC, CNIs, Terraform, Helm, Go, Linux
1mo
Save
Mark Applied
Hide
Sr. SRE Platform Architect
San Jose or Austin
HybridFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
10+ YOE10+ years production SRE/platform engineering or infra-architecture (including ≥3 years architect-level). Hands-on GPU/AI compute, multi-region observability, Kubernetes and cluster platforms, data-center operations, DDD and plugin framework experience; BS/MS CS.
NVIDIA, DCGM, MIG, vGPU, NVLink, NVSwitch, XID, NCCL, InfiniBand, RoCE, Lustre, NetApp, Pure, DDN, VAST, NVMe-oF, Kubernetes, GPU Operator, Slurm, Volcano, Kueue, Ray, KubeRay, ZTP, BMC, IPMI, Redfish, GitOps
2mo
Save
Mark Applied
Hide
Evaluation Reliability SRE
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Work on Siri and Apple Intelligence to build AI-driven assistant capabilities with strong focus on privacy and cross-platform impact.
4w
Save
Mark Applied
Hide
AI-First SRE/DevOps Engineer
San Jose, California, United States
$120k-$160k/yr HybridFull Time
Axiad
Axiad: Identity security platform for passwordless authentication and credential management.
5+ YOE5–8 years SRE/DevOps experience with production Kubernetes, infrastructure-as-code, GitOps, CI/CD, cloud provider experience, observability, Docker, AI-First tooling; Go or Python preferred.
Kubernetes, Docker, GitOps, Go, Python, Claude Code, Cursor, Windsurf
2mo
Save
Mark Applied
Hide
SRE/Dev Ops Engineer (Hybrid, Sunnyvale)
Sunnyvale, California, United States
$120k-$180k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
8+ YOE8+ years in DevOps/SRE or platform engineering; production Kubernetes; CI/CD with GitHub Actions/Jenkins/Tekton; IaC (Terraform/Pulumi); GitOps (ArgoCD/Flux); Observability (Prometheus/Grafana); multi-cloud or multi-region experience; able to work in Sunnyvale office 2+ days.
Kubernetes, GitHub Actions, Jenkins, Tekton, Terraform, Pulumi, Crossplane, ArgoCD, Flux, Prometheus, Grafana, Jaeger, OpenTelemetry, Temporal, Argo Workflows, Istio, Linkerd, Go
1w
Save
Mark Applied
Hide
Senior/Lead SRE Platform Services Engineer Technical Leader
San Jose, California, United States
$94k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ YOERequires 8–12+ years in SRE, platform engineering, infrastructure, distributed systems, or cloud operations; software engineering, infrastructure-as-code, cloud, Kubernetes, AWS, CI/CD, security, and technical leadership experience.
Python, Go, Ruby, Kubernetes, AWS, CI/CD, FedRAMP High, DoD IL5, GovCloud, FIPS
1mo
Save
Mark Applied
Hide
Job Posting Title AI/ DevOps Engineer
San Jose, California, United States
$174k-$331k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
6+ YOE6–10 years SRE/infrastructure experience; strong datastore, Kubernetes, cloud (AWS/Azure/GCP), observability and incident response skills; interest in AI/ML ops; automation and reliability focus.
Aerospike, FoundationDB, Postgres, CosmosDB, DynamoDB, Kubernetes, AWS, Azure, GCP, Prometheus, Grafana, OpenTelemetry, Copilot, Claude Code, Codex
3w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
2mo
Save
Mark Applied
Hide
Platform Engineer (SRE) - AI Control Plane
San Francisco, California, United States
OnsiteFull Time
Speakeasy
Speakeasy: Automates API SDK and documentation generation for developers.
Platform Engineer (SRE) to own reliability, design deployments, and participate in on-call; strong systems and software engineering.
2mo
Save
Mark Applied
Hide
Head - SRE
Bengaluru or Las Vegas or San Francisco
HybridFull Time
Skillz
SkillzNYSE: SKLZ: Operates a platform for competitive multiplayer mobile gaming.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
EKS, EC2, VPC, IAM, Cost Explorer, Savings Plans, Kubernetes, Istio, Datadog, Prometheus, Jaeger, X-Ray, ArgoCD, GitHub Actions, Go, Python, Java
2mo
Save
Mark Applied
Hide
Senior SRE Engineer - San Francisco
San Francisco, California, United States
HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
AWS, GCP, Azure, Kubernetes, Go, Python, Java, Cursor, GPT models, Gemini, Claude
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE)
Palo Alto, California, United States
$175k-$229k/yr HybridFull Time
Instrumental
Instrumental: AI-powered software for electronics manufacturing quality and optimization.
5+ YOE5+ years DevOps/SRE experience on public cloud (AWS preferred); expertise in Linux, shell, containers, Kubernetes, terraform, monitoring/logging/APM; strong automation, KPI measurement, and security awareness; U.S. citizenship required for access-controlled work.
AWS, Linux, shell, containerization, Kubernetes, terraform, APM
1mo
Save
Mark Applied
Hide
Machine Learning Ops Engineer, Global SRE
San Jose, California, United States
$245k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS or equivalent; expertise in Linux, networking, storage; programming in Python, Go, C, C++, or Java; troubleshooting and production operations experience; SRE of ML systems preferred.
Linux, Python, Go, C, C++, Java
2mo
Save
Mark Applied
Hide
SRE/Devops Engineer- Sunnyvale, CA, the US
Sunnyvale, California, United States
OnsiteFull Time
Kody
Kody: An agentic commerce platform providing integrated in-person payment solutions.
Senior SRE with deep AWS and GitHub experience, strong monitoring/logging and scripting skills, incident leadership, and absolute fluency in Mandarin and English.
AWS, GitHub
5d
Save
Mark Applied
Hide
Senior Software Engineer, Infrastructure (SRE & Security Focused)
San Francisco, California, United States
$200k-$260k/yr OnsiteFull Time
Nimble
Nimble: Building autonomous robotic systems for supply chain fulfillment.
3+ YOEBachelor's, master's, PhD, or equivalent experience; 3+ years in software, infrastructure, security, or SRE; programming proficiency and experience with cloud infrastructure, Kubernetes, distributed systems, security, and IaC.
Rust, Go, C++, Python, TypeScript, Kubernetes, Terraform, AWS CDK, Pulumi, CloudFormation, AWS, GCP, Azure, OpenTelemetry, Prometheus, VictoriaMetrics, Grafana, Datadog, Microsoft 401k
2mo
Save
Mark Applied
Hide
Senior Software Engineer - SRE
United States or Carson City or San Francisco or Seattle or New York
$160k-$180k/yr HybridFull Time
Socure
Socure: Provide AI-driven identity verification and fraud prevention software.
Proven experience building, running, and scaling production systems with deep AWS, Terraform, Kubernetes/EKS, Go or Python, CI/CD, GitHub/ArgoCD, and observability (Datadog, SLIs/SLOs).
AWS, Terraform, Kubernetes, Amazon EKS, Go, Python, GitHub, GitHub Actions, ArgoCD, Datadog
2w
Save
Mark Applied
Hide
Forward Deployed Engineer - SRE
North America or San Francisco
HybridFull Time
Andromeda Cluster
Andromeda Cluster: AI compute orchestration platform for GPU clusters.
Hands-on experience operating GPU clusters, fabric and driver troubleshooting, Kubernetes and Slurm experience, strong systems-level debugging and incident response, proficiency in Python/Go/Bash and IaC tooling.
Slurm, Kubernetes, NCCL, InfiniBand, RoCE, NVLink, CUDA toolkit, NVIDIA drivers, Linux, Python, Go, Bash, Terraform, Helm, Ansible, DCGM, nvidia-smi, VAST, WEKA, Lustre, GPFS