168 software reliability engineer jobs at 104 companies in Santa Rosa, CA

2mo
Save
Mark Applied
Hide
Software Reliability Engineer
Mountain View or San Francisco
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
2+ YOE2+ years in C++, Java, or Python; interest in distributed and production systems; BS degree or equivalent experience; 3+ years preferred.
C++, Java, Python
1mo
Save
Mark Applied
Hide
Reliability Engineer, Supercomputing
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Builds advanced multimodal AI models and model optimization infrastructure.
Bachelor's degree or equivalent experience; proficiency in Python or Rust; experience operating large-scale clusters and orchestration (Kubernetes or Slurm); Linux and kernel debugging; hardware-to-software root-cause analysis; vendor engagement and reliability analysis.
Python, Rust, Kubernetes, Slurm, Linux, BMC, iDRAC, IPMI, Redfish, NVLink, NVSwitch, DCGM, PyTorch, OpenAI Gym, Fairseq, Segment Anything
1mo
Save
Mark Applied
Hide
Software Engineer, Reliability Platforms
San Francisco or Sunnyvale or New York City
$160k-$235k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ years in infrastructure/platform/backend engineering; fluent in Go or similar; AWS, containerization, and IaC experience (Terraform or Pulumi); SRE concepts (SLOs, error budgets); platform engineering mindset and familiarity with AI tools.
Go, AWS, Terraform, Pulumi
1mo
Save
Mark Applied
Hide
Staff Quality & Reliability Engineer
San Francisco or New York City
HybridFull Time
Beast Industries
Beast Industries: Produces digital media and consumer goods for MrBeast brands.
Expert in software quality engineering and site reliability for consumer-scale distributed systems; owns test strategy, SLOs/error budgets, CI/CD test gates, observability, incident response, and reliability tooling.
CI/CD
1w
Save
Mark Applied
Hide
Software Engineer III, Site Reliability Engineering
Sunnyvale or Fremont or Mountain View or San Bruno or San Francisco or San Jose
$147k-$211k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
2+ YOEBachelor's in CS/Engineering or equivalent,2+ years software development experience,ability to design and troubleshoot large-scale distributed systems preferred.
2w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
Braze
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Ruby on Rails, Ruby, Go, Linux, Unix Shell, Docker, Kubernetes, Terraform, Chef, MongoDB, Redis, Kafka, Postgres, PagerDuty
3w
Save
Mark Applied
Hide
Software Engineer, Infrastructure & Reliability
San Francisco, California, United States
HybridFull Time
CrewAI
CrewAI: Platform for orchestrating collaborative multi-agent AI systems.
Experience building and operating production SaaS infrastructure: cloud, containers, CI/CD, observability, secrets, databases, and automation using Python/Ruby/Go/Bash.
AWS, Docker, CI/CD, GitHub Actions, ECS, ECR, Kubernetes, Helm, PostgreSQL, Redis, Celery, FastAPI, Rails, Sentry, OpenTelemetry, Python, Ruby, Go, Bash, Terraform
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$149k-$224k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
Temporal, Airflow, Argo Workflows, Docker, Kubernetes, DNS, HTTP, Grafana, Prometheus, ELK, Splunk, Datadog, Python, Go, Linux, Claude Code, GitHub Copilot, Codex, Cursor, AWS, GCP, MCP
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
1mo
Save
Mark Applied
Hide
Senior Software Engineer - Observability and Reliability
San Francisco, California, United States
$170k-$240k/yr OnsiteFull Time
Sigma Computing
Sigma Computing: Cloud-native analytics platform featuring a spreadsheet-style interface.
5+ YOERequires strong CS fundamentals, 5+ years building and maintaining software, experience with Go, Open Telemetry, Kubernetes, cloud platforms (GCP/AWS/Azure) and on-call/incident management.
Go, Open Telemetry, Kubernetes, GCP, AWS, Azure, SQL, Python
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
1mo
Save
Mark Applied
Hide
GOV Site Reliability Engineer
United States or Kansas or Washington or California or Texas or Illinois or North Carolina or Colorado or Massachusetts or Pennsylvania or Virginia or Oregon or Nevada or Hawaii or New York or Georgia or Ohio or Arizona or Seattle or San Francisco or New York City
$110k-$183k/yr RemoteFull Time
Veeam
Veeam: Data resilience and security for hybrid cloud environments
3+ YOE3+ years in software engineering with 1+ year in SRE/Platform/DevOps, cloud experience (Azure or comparable), observability (Prometheus, Grafana, OpenTelemetry, ELK), IaC (Terraform/Terragrunt/Pulumi), Kubernetes, CI/CD tooling, programming in TypeScript/JS, Go, Java, or C#, and experience in compliance-oriented environments.
VDC, Prometheus, Grafana, OpenTelemetry, ELK stack, Terraform, Terragrunt, Pulumi, Kubernetes, GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, TypeScript, JS, Go, Java, C#, Azure Government, AWS GovCloud
2mo
Save
Mark Applied
Hide
Founding Engineer - Site Reliability
San Francisco or United States
$185k-$285k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Kubernetes, AWS, Prometheus, Grafana, Datadog, Automation, VPC, GPU compute
3mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Scottsdale or San Francisco or Chicago or New York
$194k-$237k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
12+ YOESenior-level SRE with 12+ years in software/technical leadership; strong in cloud, microservices, automation, and observability.
Python, Go, Java, Docker, Microservices, Kafka, SQS, JMS, Oracle, DynamoDB, Aurora, Redis, memcached, Linux, GIT, Chef, Maven, Jenkins, Networking, Kubernetes
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer -AI Infrastructure Operations
Houston or San Francisco or Seattle
$170k-$265k/yr OnsiteFull Time
Nscale
Nscale: Vertically integrated AI infrastructure provider for high-performance computing.
6+ YOE6+ years SRE/systems/software engineering experience with production ownership, strong software skills (Python/Go), Linux, Kubernetes, distributed systems, SLOs and incident management.
Python, Go, Linux, Kubernetes, InfiniBand, RDMA
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Lincoln or San Francisco
$125k-$165k/yr RemoteFull Time
TELCOR
TELCOR: Provides healthcare software for laboratory and point-of-care operations.
2+ YOEExperience with distributed systems, Redis, queuing systems, Kubernetes (2+ yrs), AWS, Terraform (2+ yrs), observability, and production operations.
Redis, Kubernetes, AWS, Terraform
21h
Save
Mark Applied
Hide
Sr. Software Engineer
San Francisco, California, United States
$166k-$200k/yr HybridFull Time
Pilot
Pilot: Software-powered bookkeeping, tax, and CFO services for businesses.
5+ YOE5+ years software engineering experience, production Python, strong engineering fundamentals, ownership of end-to-end systems, production reliability and observability, strong communication and mentoring skills.
Python, JavaScript, TypeScript, Vue.js, Terraform, AWS, Postgres
2mo
Save
Mark Applied
Hide
Senior/Staff Software Engineer, Core Infrastructure
San Francisco, California, United States
$160k-$210k/yr RemoteFull Time
Zip
Zip: AI-powered intake-to-procure platform for enterprise spend management
6+ YOE6+ years software engineering in infrastructure; BS or higher in CS or related; Kubernetes/EKS, multi-region, observability; experience in a small company; quick learner.
Kubernetes, EKS, Networking, Observability, Reliability, Performance Engineering
3mo
Save
Mark Applied
Hide
Site Reliability Engineer, Enterprise Technology Services
Sunnyvale or San Francisco Bay Area
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
5+ YOE5+ years in SRE/DevOps/Software Eng; strong Java or Python/Bash/LUA; experience with Relational/NoSQL databases; Bachelor's or Master's in CS or related field.
Java, Python, Bash, LUA, Oracle, MongoDB, Prometheus, Splunk, Grafana, CloudWatch, Linux, Networking, Git, CI/CD, Kubernetes, AWS, GCP, Nginx, Envoy, NetScaler, SRE observability

Explore Jobs

Expand Your Job Search