121 sre engineer jobs at 83 companies in Cotati, CA

5d
Save
Mark Applied
Hide
ASE Compute - Senior SRE Software Engineer
San Francisco, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Design, engineer, and operate large-scale infrastructure and systems that power global services and ensure high availability.
2mo
Save
Mark Applied
Hide
SRE/Infrastructure Engineer
San Francisco, California, United States
$200k-$350k/yr OnsiteFull Time
E2B
E2B: Open-source cloud infrastructure for running autonomous AI agents.
5+ YOE5+ years producing production cloud infrastructure; strong Terraform and Kubernetes; multi-cloud experience; BYOC deployments; in-person in San Francisco.
Terraform, Kubernetes, Nomad, Google Cloud, Amazon Web Services, Azure, Cloudflare, Go, YAML
2mo
Save
Mark Applied
Hide
Platform Engineer (SRE) - AI Control Plane
San Francisco, California, United States
OnsiteFull Time
Speakeasy
Speakeasy: Automates API SDK and documentation generation for developers.
Platform Engineer (SRE) to own reliability, design deployments, and participate in on-call; strong systems and software engineering.
2w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
1mo
Save
Mark Applied
Hide
Senior SRE Engineer - San Francisco
San Francisco, California, United States
HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
AWS, GCP, Azure, Kubernetes, Go, Python, Java, Cursor, GPT models, Gemini, Claude
2mo
Save
Mark Applied
Hide
Observability Lead - Cloud SRE & Network Reliability (193698)
Fremont or San Francisco or Oakland
$114k-$253k/yr HybridFull Time
Lam Research
Lam ResearchNASDAQ: LRCX: Manufacturing equipment used to fabricate advanced semiconductor microchips.
12+ YOE6+ MgmtBS/MS/PhD or equivalent, 12+ years in infrastructure/SRE/DevOps/network engineering, 6+ years leading SRE/observability teams; multi-cloud networking, DR/BCP, observability platforms, IaC, automation, Python/Go experience.
Azure, AWS, GCP, Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, Google Cloud Operations, Splunk, Ansible, Terraform, Python, Go, Kubernetes, AKS, EKS, GKE, ServiceNow
2mo
Save
Mark Applied
Hide
Senior Software Engineer - SRE
United States or Carson City or San Francisco or Seattle or New York
$160k-$180k/yr HybridFull Time
Socure
Socure: Provide AI-driven identity verification and fraud prevention software.
Proven experience building, running, and scaling production systems with deep AWS, Terraform, Kubernetes/EKS, Go or Python, CI/CD, GitHub/ArgoCD, and observability (Datadog, SLIs/SLOs).
AWS, Terraform, Kubernetes, Amazon EKS, Go, Python, GitHub, GitHub Actions, ArgoCD, Datadog
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
1w
Save
Mark Applied
Hide
Forward Deployed Engineer - SRE
North America or San Francisco
HybridFull Time
Andromeda Cluster
Andromeda Cluster: AI compute orchestration platform for GPU clusters.
Hands-on experience operating GPU clusters, fabric and driver troubleshooting, Kubernetes and Slurm experience, strong systems-level debugging and incident response, proficiency in Python/Go/Bash and IaC tooling.
Slurm, Kubernetes, NCCL, InfiniBand, RoCE, NVLink, CUDA toolkit, NVIDIA drivers, Linux, Python, Go, Bash, Terraform, Helm, Ansible, DCGM, nvidia-smi, VAST, WEKA, Lustre, GPFS
4d
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Tinker, Kubernetes, LoRA, CI/CD
2mo
Save
Mark Applied
Hide
Head - SRE
Bengaluru or Las Vegas or San Francisco
HybridFull Time
Skillz
SkillzNYSE: SKLZ: Operates a platform for competitive multiplayer mobile gaming.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
EKS, EC2, VPC, IAM, Cost Explorer, Savings Plans, Kubernetes, Istio, Datadog, Prometheus, Jaeger, X-Ray, ArgoCD, GitHub Actions, Go, Python, Java
2d
Save
Mark Applied
Hide
Staff Site Reliability Engineer (SRE) (Hybrid)
San Francisco or San Jose or New York City or Milpitas or Mountain View or Holmdel or Goleta or Redwood City or Fremont or Sunnyvale or Brooklyn or Palo Alto
$187k-$268k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
6+ YOERequires 8+ years with a bachelor's, 6+ with a master's, or 3+ with a PhD; 6+ years in SRE or infrastructure engineering, 5+ years operating Kubernetes, cloud, CI/CD, and Python or Go.
Kubernetes, AWS, GCP, Python, Go, Terraform, MLOps, CI/CD
3d
Save
Mark Applied
Hide
Principal Engineer
San Francisco or Charlotte
$159k-$305k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
7+ YOERequires 7+ years of engineering and software, platform, SRE, or distributed systems experience, with expertise in resilient scalable platforms, cloud technologies, Kubernetes/OpenShift, observability, automation, and technical leadership.
Kubernetes, OpenShift, GitHub Copilot, AIOps, Zelle, ACH, RTP, FedNow, CI/CD
2mo
Save
Mark Applied
Hide
Staff Platform Engineer - Americas
Austin or Los Angeles or Portland or Atlanta or Salt Lake City or Boston or Seattle or New York or San Francisco or Denver or Chicago
$232k-$323k/yr RemoteFull Time
Ashby: Provides all-in-one recruiting software for high-growth businesses.
Experienced infrastructure/platform engineer comfortable coding, scaling systems, SRE practices, SQL, Kubernetes, observability, on-call, and building developer-facing platform tools.
TypeScript, Node.js, React, Apollo GraphQL, Postgres, Redis, Datadog, Sentry, AWS, Kubernetes, SQL
3mo
Save
Mark Applied
Hide
Senior Production Engineer, Oeprational Excellence
San Francisco or Sunnyvale
$172k-$209k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
5+ YOE5+ years in Production Engineering or SRE; GPU workloads; Linux; IaC; Kubernetes; Go/Python; strong communication.
Prometheus, Grafana, OpenTelemetry, Terraform, Ansible, Kubernetes, AWS, GCP, Linux
1mo
Save
Mark Applied
Hide
Staff Platform Engineer
San Francisco or New York City or Colorado or California or Washington
$220k-$331k/yr OnsiteFull Time
Amplitude
AmplitudeNasdaq: AMPL: Develops digital analytics software for tracking customer product behavior.
8+ YOE8+ years in software/DevOps/SRE, Bachelor\u000degree in Computer Engineering (required), deep Kubernetes and cloud experience, Terraform/IaC and programming (Golang or Python), strong cross-team leadership and communication.
Kubernetes, EKS, GKE, AKS, Terraform, Helm, Kustomize, Argo CD, Argo Workflows, Argo Rollouts, GitHub Actions, Datadog, Amplitude, Golang, Python, EC2, IAM, VPC, ALB, S3, Backstage, Envoy, GitOps, LLM
1d
Save
Mark Applied
Hide
IT Systems Engineer - Internal Platforms & SRE
San Francisco or San Jose
$206k-$275k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
Experience with system design, scalable cloud infrastructure, configuration management, programming in Python or Go, distributed systems, automation, documentation, and cross-functional collaboration.
AWS, GCP, Azure, Chef, Ansible, Terraform, GitHub Actions, Python, Go
3w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
2mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
San Francisco, California, United States
$200k-$250k/yr OnsiteFull Time
Stuut
Stuut: Automates business accounts receivable and collections through AI agents.
7+ YOE7+ years in SRE/infrastructure or backend engineering. Experience with AWS, Kubernetes/EKS, Docker, observability, SLOs/SLIs, Python or TypeScript, CI/CD, and production-grade distributed systems.
Python, TypeScript, AWS, Kubernetes, EKS, Docker, FastAPI, Vue.js, PostgreSQL (RDS), CI/CD