143 sre engineer jobs at 106 companies in Vallejo, CA
2w
Save
Mark Applied
Hide
2w
SRE Engineer (Full Time; Multiple Openings)
Belmont, California, United States
HybridFull Time
RingCentralNYSE: RNG: Sells cloud-based business phone and video conferencing software.
2+ YOEMaintain 24x7 production availability, implement automation/orchestration, partner with development, perform root cause analysis; required experience with cloud, containers, scripting, and monitoring.
E2B: Open-source cloud infrastructure for running autonomous AI agents.
5+ YOE5+ years producing production cloud infrastructure; strong Terraform and Kubernetes; multi-cloud experience; BYOC deployments; in-person in San Francisco.
Terraform, Kubernetes, Nomad, Google Cloud, Amazon Web Services, Azure, Cloudflare, Go, YAML
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Plaud: Develops AI-powered voice recorders and automated transcription software.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
Instrumental: AI-powered software for electronics manufacturing quality and optimization.
5+ YOE5+ years DevOps/SRE experience on public cloud (AWS preferred); expertise in Linux, shell, containers, Kubernetes, terraform, monitoring/logging/APM; strong automation, KPI measurement, and security awareness; U.S. citizenship required for access-controlled work.
GEICO: Provides vehicle and property insurance services to consumers.
8+ YOE8+ years in software or site reliability engineering; 5+ years in SRE/DevOps; strong Python; Golang preferred; AWS/Azure/GCP experience; CI/CD and IaC; observability and incident response; security tooling familiarity.
Andromeda Cluster: AI compute orchestration platform for GPU clusters.
Hands-on experience operating GPU clusters, fabric and driver troubleshooting, Kubernetes and Slurm experience, strong systems-level debugging and incident response, proficiency in Python/Go/Bash and IaC tooling.
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
United States or Carson City or San Francisco or Seattle or New York
$160k-$180k/yrHybridFull Time
Socure: Provide AI-driven identity verification and fraud prevention software.
Proven experience building, running, and scaling production systems with deep AWS, Terraform, Kubernetes/EKS, Go or Python, CI/CD, GitHub/ArgoCD, and observability (Datadog, SLIs/SLOs).
SkillzNYSE: SKLZ: Operates a platform for competitive multiplayer mobile gaming.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
Kody: An agentic commerce platform providing integrated in-person payment solutions.
Deep AWS and Git/GitHub experience, strong monitoring/logging and scripting skills, production incident leadership, bilingual Mandarin and English, ownership of CI/CD and deployment practices.
Blackhawk Network: Provider of global branded payment and gift card solutions.
Bachelor's in CS/Engineering or equivalent; experience in platform/DevOps/SRE or similar; strong Linux, AWS, Git; scripting with Python/Bash; production support and incident management; experience with AI-assisted engineering tools.
AWS, Git, Python, Bash, Kubernetes, Docker, Jenkins, Splunk, New Relic, Prometheus, Grafana, OpenTelemetry, ServiceNow, Terraform, CloudFormation, GitHub Copilot, Cursor, Claude
STN: Provides high-performance GPU infrastructure, cloud, and managed IT services.
6+ YOE6+ years in platform/SRE/cloud engineering, deep Kubernetes expertise, Go and/or Python programming, experience operating GPU/AI infrastructure, bachelor's degree or equivalent experience.
Senior Site Reliability Engineer, Production Engineer - ThousandEyes
San Francisco or Seattle or Austin or New York City
$165k-$241k/yrHybridFull Time
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years experience; proficiency in Python or Go; expertise with Kubernetes, cloud (AWS), Unix/Linux; strong SRE principles, incident response, and security-minded engineering.
Python, Go, Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD, CNCF, AWS, Unix, Linux
San Francisco or New York City or Colorado or California or Washington
$220k-$331k/yrOnsiteFull Time
AmplitudeNasdaq: AMPL: Develops digital analytics software for tracking customer product behavior.
8+ YOE8+ years in software/DevOps/SRE, Bachelor\u000degree in Computer Engineering (required), deep Kubernetes and cloud experience, Terraform/IaC and programming (Golang or Python), strong cross-team leadership and communication.
Blackhawk Network: Provider of branded payment and gift card solutions.
Bachelor's in CS/Engineering or equivalent experience; platform/DevOps/SRE experience; strong Linux, AWS, Git; scripting in Python/Bash; experience with major incident management and AI-assisted development tools.
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.