143 sre engineer jobs at 106 companies in Vallejo, CA

2w
Save
Mark Applied
Hide
SRE Engineer (Full Time; Multiple Openings)
Belmont, California, United States
HybridFull Time
RingCentral
RingCentralNYSE: RNG: Sells cloud-based business phone and video conferencing software.
2+ YOEMaintain 24x7 production availability, implement automation/orchestration, partner with development, perform root cause analysis; required experience with cloud, containers, scripting, and monitoring.
Python, Bash, Go, Terraform, Ansible, AWS, GCP, Kubernetes, GitLab, DNS, Docker, CI/CD, TCP/IP, Linux
2mo
Save
Mark Applied
Hide
SRE/Infrastructure Engineer
San Francisco, California, United States
$200k-$350k/yr OnsiteFull Time
E2B
E2B: Open-source cloud infrastructure for running autonomous AI agents.
5+ YOE5+ years producing production cloud infrastructure; strong Terraform and Kubernetes; multi-cloud experience; BYOC deployments; in-person in San Francisco.
Terraform, Kubernetes, Nomad, Google Cloud, Amazon Web Services, Azure, Cloudflare, Go, YAML
2mo
Save
Mark Applied
Hide
Platform Engineer (SRE) - AI Control Plane
San Francisco, California, United States
OnsiteFull Time
Speakeasy
Speakeasy: Automates API SDK and documentation generation for developers.
Platform Engineer (SRE) to own reliability, design deployments, and participate in on-call; strong systems and software engineering.
1w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
1mo
Save
Mark Applied
Hide
Senior SRE Engineer - San Francisco
San Francisco, California, United States
HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
AWS, GCP, Azure, Kubernetes, Go, Python, Java, Cursor, GPT models, Gemini, Claude
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE)
Palo Alto, California, United States
$175k-$229k/yr HybridFull Time
Instrumental
Instrumental: AI-powered software for electronics manufacturing quality and optimization.
5+ YOE5+ years DevOps/SRE experience on public cloud (AWS preferred); expertise in Linux, shell, containers, Kubernetes, terraform, monitoring/logging/APM; strong automation, KPI measurement, and security awareness; U.S. citizenship required for access-controlled work.
AWS, Linux, shell, containerization, Kubernetes, terraform, APM
2mo
Save
Mark Applied
Hide
Staff Cyber Site Reliability Engineer (SRE)
Bethesda or Palo Alto or Dallas or Seattle
$110k-$230k/yr HybridFull Time
GEICO
GEICO: Provides vehicle and property insurance services to consumers.
8+ YOE8+ years in software or site reliability engineering; 5+ years in SRE/DevOps; strong Python; Golang preferred; AWS/Azure/GCP experience; CI/CD and IaC; observability and incident response; security tooling familiarity.
Python, Golang, Grafana, Prometheus, GitHub Actions, Jenkins, Terraform, Ansible
2d
Save
Mark Applied
Hide
Forward Deployed Engineer - SRE
North America or San Francisco
HybridFull Time
Andromeda Cluster
Andromeda Cluster: AI compute orchestration platform for GPU clusters.
Hands-on experience operating GPU clusters, fabric and driver troubleshooting, Kubernetes and Slurm experience, strong systems-level debugging and incident response, proficiency in Python/Go/Bash and IaC tooling.
Slurm, Kubernetes, NCCL, InfiniBand, RoCE, NVLink, CUDA toolkit, NVIDIA drivers, Linux, Python, Go, Bash, Terraform, Helm, Ansible, DCGM, nvidia-smi, VAST, WEKA, Lustre, GPFS
3w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
2mo
Save
Mark Applied
Hide
Senior Software Engineer - SRE
United States or Carson City or San Francisco or Seattle or New York
$160k-$180k/yr HybridFull Time
Socure
Socure: Provide AI-driven identity verification and fraud prevention software.
Proven experience building, running, and scaling production systems with deep AWS, Terraform, Kubernetes/EKS, Go or Python, CI/CD, GitHub/ArgoCD, and observability (Datadog, SLIs/SLOs).
AWS, Terraform, Kubernetes, Amazon EKS, Go, Python, GitHub, GitHub Actions, ArgoCD, Datadog
1mo
Save
Mark Applied
Hide
Head - SRE
Bengaluru or Las Vegas or San Francisco
HybridFull Time
Skillz
SkillzNYSE: SKLZ: Operates a platform for competitive multiplayer mobile gaming.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
EKS, EC2, VPC, IAM, Cost Explorer, Savings Plans, Kubernetes, Istio, Datadog, Prometheus, Jaeger, X-Ray, ArgoCD, GitHub Actions, Go, Python, Java
1mo
Save
Mark Applied
Hide
SRE/DevOps Engineer- Palo Alto, the US
Palo Alto, California, United States
HybridFull Time
Kody
Kody: An agentic commerce platform providing integrated in-person payment solutions.
Deep AWS and Git/GitHub experience, strong monitoring/logging and scripting skills, production incident leadership, bilingual Mandarin and English, ownership of CI/CD and deployment practices.
AWS, GitHub, Git
2mo
Save
Mark Applied
Hide
Senior Platform Engineer
Menlo Park or New York or Seattle
$180k-$220k/yr RemoteFull Time
Verantos
Verantos: Generates high-accuracy real-world evidence for clinical and regulatory use
5+ YOE5+ years in DevOps/platform engineering or SRE; strong AWS, Terraform, CI/CD; Python/TypeScript; healthcare or regulated environments.
AWS, Terraform, Docker, Kubernetes, ECS, EKS, Python, TypeScript, Snowflake, GitHub Actions
1mo
Save
Mark Applied
Hide
Platform Engineer I
Pleasanton, California, United States
$32-$41/hr HybridFull Time
Blackhawk Network
Blackhawk Network: Provider of global branded payment and gift card solutions.
Bachelor's in CS/Engineering or equivalent; experience in platform/DevOps/SRE or similar; strong Linux, AWS, Git; scripting with Python/Bash; production support and incident management; experience with AI-assisted engineering tools.
AWS, Git, Python, Bash, Kubernetes, Docker, Jenkins, Splunk, New Relic, Prometheus, Grafana, OpenTelemetry, ServiceNow, Terraform, CloudFormation, GitHub Copilot, Cursor, Claude
1mo
Save
Mark Applied
Hide
Senior Platform Engineer
Pleasanton or California or United States
HybridFull Time
STN
STN: Provides high-performance GPU infrastructure, cloud, and managed IT services.
6+ YOE6+ years in platform/SRE/cloud engineering, deep Kubernetes expertise, Go and/or Python programming, experience operating GPU/AI infrastructure, bachelor's degree or equivalent experience.
Kubernetes, Slurm, Run:ai, Go, Python, NVIDIA GPU Operator, MIG, MPS, NCCL, KubeRay, Istio, Linkerd, CRDs
2mo
Save
Mark Applied
Hide
Senior Production Engineer, Oeprational Excellence
San Francisco or Sunnyvale
$172k-$209k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
5+ YOE5+ years in Production Engineering or SRE; GPU workloads; Linux; IaC; Kubernetes; Go/Python; strong communication.
Prometheus, Grafana, OpenTelemetry, Terraform, Ansible, Kubernetes, AWS, GCP, Linux
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Production Engineer - ThousandEyes
San Francisco or Seattle or Austin or New York City
$165k-$241k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years experience; proficiency in Python or Go; expertise with Kubernetes, cloud (AWS), Unix/Linux; strong SRE principles, incident response, and security-minded engineering.
Python, Go, Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD, CNCF, AWS, Unix, Linux
1mo
Save
Mark Applied
Hide
Staff Platform Engineer
San Francisco or New York City or Colorado or California or Washington
$220k-$331k/yr OnsiteFull Time
Amplitude
AmplitudeNasdaq: AMPL: Develops digital analytics software for tracking customer product behavior.
8+ YOE8+ years in software/DevOps/SRE, Bachelor\u000degree in Computer Engineering (required), deep Kubernetes and cloud experience, Terraform/IaC and programming (Golang or Python), strong cross-team leadership and communication.
Kubernetes, EKS, GKE, AKS, Terraform, Helm, Kustomize, Argo CD, Argo Workflows, Argo Rollouts, GitHub Actions, Datadog, Amplitude, Golang, Python, EC2, IAM, VPC, ALB, S3, Backstage, Envoy, GitOps, LLM
1mo
Save
Mark Applied
Hide
Platform Engineer I
Pleasanton, California, United States
$41/hr HybridFull Time
Blackhawk Network
Blackhawk Network: Provider of branded payment and gift card solutions.
Bachelor's in CS/Engineering or equivalent experience; platform/DevOps/SRE experience; strong Linux, AWS, Git; scripting in Python/Bash; experience with major incident management and AI-assisted development tools.
AWS, Kubernetes, Git, Python, Bash, GitHub Copilot, Cursor, Claude, Docker, Jenkins, Splunk, New Relic, Prometheus, Grafana, OpenTelemetry, ServiceNow, Terraform, CloudFormation, Linux
2w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD