132 sre engineer jobs at 91 companies in Napa, CA

1w
Save
Mark Applied
Hide
ASE Compute - Senior SRE Software Engineer
San Francisco, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Design, engineer, and operate large-scale infrastructure and systems that power global services and ensure high availability.
2mo
Save
Mark Applied
Hide
SRE/Infrastructure Engineer
San Francisco, California, United States
$200k-$350k/yr OnsiteFull Time
E2B
E2B: Open-source cloud infrastructure for running autonomous AI agents.
5+ YOE5+ years producing production cloud infrastructure; strong Terraform and Kubernetes; multi-cloud experience; BYOC deployments; in-person in San Francisco.
Terraform, Kubernetes, Nomad, Google Cloud, Amazon Web Services, Azure, Cloudflare, Go, YAML
2mo
Save
Mark Applied
Hide
Platform Engineer (SRE) - AI Control Plane
San Francisco, California, United States
OnsiteFull Time
Speakeasy
Speakeasy: Automates API SDK and documentation generation for developers.
Platform Engineer (SRE) to own reliability, design deployments, and participate in on-call; strong systems and software engineering.
2w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
1mo
Save
Mark Applied
Hide
Senior SRE Engineer - San Francisco
San Francisco, California, United States
HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
AWS, GCP, Azure, Kubernetes, Go, Python, Java, Cursor, GPT models, Gemini, Claude
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
2mo
Save
Mark Applied
Hide
Senior Software Engineer - SRE
United States or Carson City or San Francisco or Seattle or New York
$160k-$180k/yr HybridFull Time
Socure
Socure: Provide AI-driven identity verification and fraud prevention software.
Proven experience building, running, and scaling production systems with deep AWS, Terraform, Kubernetes/EKS, Go or Python, CI/CD, GitHub/ArgoCD, and observability (Datadog, SLIs/SLOs).
AWS, Terraform, Kubernetes, Amazon EKS, Go, Python, GitHub, GitHub Actions, ArgoCD, Datadog
1w
Save
Mark Applied
Hide
Forward Deployed Engineer - SRE
North America or San Francisco
HybridFull Time
Andromeda Cluster
Andromeda Cluster: AI compute orchestration platform for GPU clusters.
Hands-on experience operating GPU clusters, fabric and driver troubleshooting, Kubernetes and Slurm experience, strong systems-level debugging and incident response, proficiency in Python/Go/Bash and IaC tooling.
Slurm, Kubernetes, NCCL, InfiniBand, RoCE, NVLink, CUDA toolkit, NVIDIA drivers, Linux, Python, Go, Bash, Terraform, Helm, Ansible, DCGM, nvidia-smi, VAST, WEKA, Lustre, GPFS
1w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Tinker, Kubernetes, LoRA, CI/CD
2mo
Save
Mark Applied
Hide
Head - SRE
Bengaluru or Las Vegas or San Francisco
HybridFull Time
Skillz
SkillzNYSE: SKLZ: Operates a platform for competitive multiplayer mobile gaming.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
EKS, EC2, VPC, IAM, Cost Explorer, Savings Plans, Kubernetes, Istio, Datadog, Prometheus, Jaeger, X-Ray, ArgoCD, GitHub Actions, Go, Python, Java
6d
Save
Mark Applied
Hide
Staff Site Reliability Engineer (SRE) (Hybrid)
San Francisco or San Jose or New York City or Milpitas or Mountain View or Holmdel or Goleta or Redwood City or Fremont or Sunnyvale or Brooklyn or Palo Alto
$187k-$268k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
6+ YOERequires 8+ years with a bachelor's, 6+ with a master's, or 3+ with a PhD; 6+ years in SRE or infrastructure engineering, 5+ years operating Kubernetes, cloud, CI/CD, and Python or Go.
Kubernetes, AWS, GCP, Python, Go, Terraform, MLOps, CI/CD
1w
Save
Mark Applied
Hide
Principal Engineer
San Francisco or Charlotte
$159k-$305k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
7+ YOERequires 7+ years of engineering and software, platform, SRE, or distributed systems experience, with expertise in resilient scalable platforms, cloud technologies, Kubernetes/OpenShift, observability, automation, and technical leadership.
Kubernetes, OpenShift, GitHub Copilot, AIOps, Zelle, ACH, RTP, FedNow, CI/CD
1mo
Save
Mark Applied
Hide
Platform Engineer I
Pleasanton, California, United States
$32-$41/hr HybridFull Time
Blackhawk Network
Blackhawk Network: Provider of global branded payment and gift card solutions.
Bachelor's in CS/Engineering or equivalent; experience in platform/DevOps/SRE or similar; strong Linux, AWS, Git; scripting with Python/Bash; production support and incident management; experience with AI-assisted engineering tools.
AWS, Git, Python, Bash, Kubernetes, Docker, Jenkins, Splunk, New Relic, Prometheus, Grafana, OpenTelemetry, ServiceNow, Terraform, CloudFormation, GitHub Copilot, Cursor, Claude
1mo
Save
Mark Applied
Hide
Senior Platform Engineer
Pleasanton or California or United States
HybridFull Time
STN
STN: Provides high-performance GPU infrastructure, cloud, and managed IT services.
6+ YOE6+ years in platform/SRE/cloud engineering, deep Kubernetes expertise, Go and/or Python programming, experience operating GPU/AI infrastructure, bachelor's degree or equivalent experience.
Kubernetes, Slurm, Run:ai, Go, Python, NVIDIA GPU Operator, MIG, MPS, NCCL, KubeRay, Istio, Linkerd, CRDs
2mo
Save
Mark Applied
Hide
Staff Platform Engineer - Americas
Austin or Los Angeles or Portland or Atlanta or Salt Lake City or Boston or Seattle or New York or San Francisco or Denver or Chicago
$232k-$323k/yr RemoteFull Time
Ashby: Provides all-in-one recruiting software for high-growth businesses.
Experienced infrastructure/platform engineer comfortable coding, scaling systems, SRE practices, SQL, Kubernetes, observability, on-call, and building developer-facing platform tools.
TypeScript, Node.js, React, Apollo GraphQL, Postgres, Redis, Datadog, Sentry, AWS, Kubernetes, SQL
3mo
Save
Mark Applied
Hide
Senior Production Engineer, Oeprational Excellence
San Francisco or Sunnyvale
$172k-$209k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
5+ YOE5+ years in Production Engineering or SRE; GPU workloads; Linux; IaC; Kubernetes; Go/Python; strong communication.
Prometheus, Grafana, OpenTelemetry, Terraform, Ansible, Kubernetes, AWS, GCP, Linux
1mo
Save
Mark Applied
Hide
Staff Platform Engineer
San Francisco or New York City or Colorado or California or Washington
$220k-$331k/yr OnsiteFull Time
Amplitude
AmplitudeNasdaq: AMPL: Develops digital analytics software for tracking customer product behavior.
8+ YOE8+ years in software/DevOps/SRE, Bachelor\u000degree in Computer Engineering (required), deep Kubernetes and cloud experience, Terraform/IaC and programming (Golang or Python), strong cross-team leadership and communication.
Kubernetes, EKS, GKE, AKS, Terraform, Helm, Kustomize, Argo CD, Argo Workflows, Argo Rollouts, GitHub Actions, Datadog, Amplitude, Golang, Python, EC2, IAM, VPC, ALB, S3, Backstage, Envoy, GitOps, LLM
5d
Save
Mark Applied
Hide
IT Systems Engineer - Internal Platforms & SRE
San Francisco or San Jose
$206k-$275k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
Experience with system design, scalable cloud infrastructure, configuration management, programming in Python or Go, distributed systems, automation, documentation, and cross-functional collaboration.
AWS, GCP, Azure, Chef, Ansible, Terraform, GitHub Actions, Python, Go
1mo
Save
Mark Applied
Hide
Platform Engineer I
Pleasanton, California, United States
$41/hr HybridFull Time
Blackhawk Network
Blackhawk Network: Provides branded payment products and corporate incentive solutions globally.
Bachelor's degree or equivalent experience in CS/Engineering, experience in platform/SRE/DevOps or similar, strong Linux and AWS skills, Git, Python/Bash scripting, production incident management, and experience with AI-assisted development tools.
AWS, Linux, Git, Python, Bash, GitHub Copilot, Cursor, Claude, Kubernetes, Docker, Jenkins, Splunk, New Relic, Prometheus, Grafana, OpenTelemetry, ServiceNow, Terraform, CloudFormation, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD