150 software reliability engineer jobs at 77 companies in Scotts Valley, CA

3mo
Save
Mark Applied
Hide
Software Reliability Engineer
Mountain View, California, United States
$146k-$219k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
Production software experience; build automation/tools; strong debugging; reliability engineering interest.
Python, Go, Bash, C++, Observability, Telemetry
2mo
Save
Mark Applied
Hide
Software Reliability Engineer
Mountain View or San Francisco
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
2+ YOE2+ years in C++, Java, or Python; interest in distributed and production systems; BS degree or equivalent experience; 3+ years preferred.
C++, Java, Python
5d
Save
Mark Applied
Hide
Software Engineer - NEO Reliability
San Carlos, California, United States
$200k-$300k/yr OnsiteFull Time
1X
1X: Manufacturing safe, general-purpose humanoid robots for home and work.
Experienced software engineer with reliability and testing expertise across cloud, mobile, on-robot services, and embedded systems; strong DevOps and Python/Linux skills; statistical and systems reasoning.
Python, Linux, CI/CD, APIs, HIL
4d
Save
Mark Applied
Hide
Staff Reliability Engineer
Santa Clara, California, United States
$167k-$291k/yr RemoteFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
8+ YOE8+ years SRE/Platform/DevOps experience, strong Kubernetes and cloud-native platform skills, automation and CI/CD expertise, software engineering with Python/Go/Java/Ruby, observability and reliability knowledge.
Kubernetes, Python, Go, Java, Ruby, GitLab CI/CD, Argo CD, Flux, Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit, TestNG, Ansible, Terraform, Helm, Argo Workflows, Kustomize, Istio, Linkerd, Gateway API, Ingress, Prometheus, OpenTelemetry, AWS (EKS), Azure (AKS), Google Cloud (GKE), GitOps
5d
Save
Mark Applied
Hide
Fellow Software Engineer — AI Performance & Reliability
San Jose or Bellevue
$235k-$402k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
PhD or equivalent in AI/ML/CS, strong software engineering, experience profiling and optimizing ML models and AI workloads, proficiency in Python/C++, ML frameworks, customer-facing troubleshooting and performance analysis.
Python, C++, PyTorch, TensorFlow, JAX, ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL
1mo
Save
Mark Applied
Hide
Software Engineer, Reliability Platforms
San Francisco or Sunnyvale or New York City
$160k-$235k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ years in infrastructure/platform/backend engineering; fluent in Go or similar; AWS, containerization, and IaC experience (Terraform or Pulumi); SRE concepts (SLOs, error budgets); platform engineering mindset and familiarity with AI tools.
Go, AWS, Terraform, Pulumi
3w
Save
Mark Applied
Hide
Staff Software Engineer, Reliability Engineering & SDLC Governance
Carlsbad or Germantown or San Jose or San Francisco or New York City
$165k-$261k/yr OnsiteFull Time
Viasat
ViasatNASDAQ: VSAT: Provides global satellite broadband and secure networking communication services.
8+ YOE8+ years software engineering experience with SDLC governance, SRE/DevOps knowledge, observability, incident analysis, and cross-team influence to improve reliability and customer experience.
Prometheus, Grafana, OpenTelemetry, Datadog, AS9115
2w
Save
Mark Applied
Hide
Software Engineer III, Site Reliability Engineering
Sunnyvale or Fremont or Mountain View or San Bruno or San Francisco or San Jose
$147k-$211k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
2+ YOEBachelor's in CS/Engineering or equivalent,2+ years software development experience,ability to design and troubleshoot large-scale distributed systems preferred.
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Repair and Reliability
South San Francisco, California, United States
$180k-$270k/yr OnsiteFull Time
Zipline
Zipline: Operates an autonomous drone delivery system for medical supplies.
Build reliable production systems for complex, real-world operations; own critical software in fleet reliability and maintenance.
Backend, Systems engineering
1mo
Save
Mark Applied
Hide
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems, strong software engineering (Go/Python), SLO program experience, incident response leadership, chaos engineering and failure-injection expertise, ability to influence across teams.
Go, Python, Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
2mo
Save
Mark Applied
Hide
Staff Software Engineer - Reliability
Palo Alto, California, United States
$218k-$328k/yr OnsiteFull Time
Rubrik
RubrikNYSE: RBRK: Secures enterprise data across cloud and on-premises environments.
8+ YOEUS citizen; 8-12+ years software engineering with SRE/DevOps; BS/MS/PhD in CS/CE or related field; proficient in Go/Python/Java; distributed systems; Unix/Linux; on-call; leadership experience.
Go, Python, Java, Kubernetes, MySQL, Terraform, Pulumi, Prometheus, Grafana, OpenTelemetry
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Compute
San Mateo, California, United States
$243k-$295k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years SRE or software engineering experience; Bachelor’s in Computer Science or equivalent; fluency in Go, Java, or C#; experience with Kubernetes, Nomad, Vault, and Consul; strong reliability and observability practices.
Go, Java, C#, Kubernetes, Nomad, Vault, Consul
6d
Save
Mark Applied
Hide
Senior Inference Reliability Engineer
San Mateo, California, United States
OnsiteFull Time
Parasail
Parasail: Provides scalable cloud infrastructure for AI model inference.
5+ YOE5+ years production engineering experience operating customer-facing systems; strong SRE and production diagnostics skills; Kubernetes, Linux, distributed systems, and software engineering proficiency; ability to lead incident response and build observability.
Kubernetes, Linux, Python, Go, Java, C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Mateo, California, United States
$130k-$200k/yr OnsiteFull Time
IXL Learning
IXL Learning: Provides personalized digital learning platforms and educational resources.
6+ YOEBachelor's degree,6+ years SRE/software engineering,experience with OO and scripting languages,cloud (AWS/GCP),Docker/Kubernetes,monitoring,on-call availability,strong troubleshooting and communication skills.
Java, C++, C, Python, Bash, Perl, AWS, GCP, Docker, Kubernetes
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
1mo
Save
Mark Applied
Hide
Sr. Software Reliability & Stability Quality Engineer, Siri Speech
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Work on Siri Speech and Apple Intelligence across Apple platforms with a privacy-first approach; build and ship high-quality, user-centered software for iOS, iPadOS, macOS, watchOS, and visionOS.
iOS, iPadOS, macOS, watchOS, visionOS
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Mountain View, California, United States
$252k-$308k/yr HybridFull Time
EarnIn
EarnIn: Provides immediate access to earned wages through a mobile app.
7+ YOE7+ years in SRE or related field; experience applying AI/LLMs to operations; strong SLO/SLI and incident management; software engineering in Python or Go; observability and IaC proficiency; AI-assisted development tools; fintech/regulated environment experience.
Datadog, CloudWatch, OpenTelemetry, Terraform, Kubernetes, AWS, Python, Go, Cursor, Claude Code, Copilot
1w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer (Starlink)
Hawthorne or Palo Alto or Redmond
$165k-$270k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOEBachelor's in CS/engineering/math with 5 years software experience or 7+ years SRE/DevOps experience; Linux experience required; Kubernetes, Kafka, cloud-native tooling, and programming in Python/Go/Java/C#/Scala preferred.
Linux, Kubernetes, Istio, Apache Kafka, Apache Spark, HBase, HDFS, Apache Flink, Python, C#, Java, Scala, Go
1mo
Save
Mark Applied
Hide
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
Santa Clara, California, United States
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
15+ YOE15+ years systems software or reliability engineering experience at datacenter scale; BS/MS in CS/EE/Statistics or equivalent; expertise in fleet telemetry, statistical failure analysis, burn-in/certification, and strong cross-functional communication.