125 software reliability engineer jobs at 62 companies in Aromas, CA

3mo
Save
Mark Applied
Hide
Software Reliability Engineer
Mountain View, California, United States
$146k-$219k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
Production software experience; build automation/tools; strong debugging; reliability engineering interest.
Python, Go, Bash, C++, Observability, Telemetry
2mo
Save
Mark Applied
Hide
Software Reliability Engineer
Mountain View or San Francisco
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
2+ YOE2+ years in C++, Java, or Python; interest in distributed and production systems; BS degree or equivalent experience; 3+ years preferred.
C++, Java, Python
1w
Save
Mark Applied
Hide
Fellow Software Engineer — AI Performance & Reliability
San Jose or Bellevue
$235k-$402k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
PhD or equivalent in AI/ML/CS, strong software engineering, experience profiling and optimizing ML models and AI workloads, proficiency in Python/C++, ML frameworks, customer-facing troubleshooting and performance analysis.
Python, C++, PyTorch, TensorFlow, JAX, ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL
3d
Save
Mark Applied
Hide
Senior Software Engineer – Application Reliability , Hybrid
San Jose or North Carolina
$203k-$259k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
10+ YOE10+ years in software engineering focused on reliability or observability; bachelor's or master's in a technical discipline; Python, GCP, GKE, BigQuery, BigTable, SLI/SLO, debugging, and application operations expertise.
Python, GKE, Kubernetes, BigQuery, BigTable, Looker, LangGraph, Cloud Logging, Cloud Trace, Cloud Monitoring, SQL, A2A
1mo
Save
Mark Applied
Hide
Software Engineer, Reliability Platforms
San Francisco or Sunnyvale or New York City
$160k-$235k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ years in infrastructure/platform/backend engineering; fluent in Go or similar; AWS, containerization, and IaC experience (Terraform or Pulumi); SRE concepts (SLOs, error budgets); platform engineering mindset and familiarity with AI tools.
Go, AWS, Terraform, Pulumi
4w
Save
Mark Applied
Hide
Staff Software Engineer, Reliability Engineering & SDLC Governance
Carlsbad or Germantown or San Jose or San Francisco or New York City
$165k-$261k/yr OnsiteFull Time
Viasat
ViasatNASDAQ: VSAT: Provides global satellite broadband and secure networking communication services.
8+ YOE8+ years software engineering experience with SDLC governance, SRE/DevOps knowledge, observability, incident analysis, and cross-team influence to improve reliability and customer experience.
Prometheus, Grafana, OpenTelemetry, Datadog, AS9115
1mo
Save
Mark Applied
Hide
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems, strong software engineering (Go/Python), SLO program experience, incident response leadership, chaos engineering and failure-injection expertise, ability to influence across teams.
Go, Python, Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
2mo
Save
Mark Applied
Hide
Staff Software Engineer - Reliability
Palo Alto, California, United States
$218k-$328k/yr OnsiteFull Time
Rubrik
RubrikNYSE: RBRK: Secures enterprise data across cloud and on-premises environments.
8+ YOEUS citizen; 8-12+ years software engineering with SRE/DevOps; BS/MS/PhD in CS/CE or related field; proficient in Go/Python/Java; distributed systems; Unix/Linux; on-call; leadership experience.
Go, Python, Java, Kubernetes, MySQL, Terraform, Pulumi, Prometheus, Grafana, OpenTelemetry
1mo
Save
Mark Applied
Hide
Sr. Software Reliability & Stability Quality Engineer, Siri Speech
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Work on Siri Speech and Apple Intelligence across Apple platforms with a privacy-first approach; build and ship high-quality, user-centered software for iOS, iPadOS, macOS, watchOS, and visionOS.
iOS, iPadOS, macOS, watchOS, visionOS
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Mountain View, California, United States
$252k-$308k/yr HybridFull Time
EarnIn
EarnIn: Provides immediate access to earned wages through a mobile app.
7+ YOE7+ years in SRE or related field; experience applying AI/LLMs to operations; strong SLO/SLI and incident management; software engineering in Python or Go; observability and IaC proficiency; AI-assisted development tools; fintech/regulated environment experience.
Datadog, CloudWatch, OpenTelemetry, Terraform, Kubernetes, AWS, Python, Go, Cursor, Claude Code, Copilot
1w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer (Starlink)
Hawthorne or Palo Alto or Redmond
$165k-$270k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOEBachelor's in CS/engineering/math with 5 years software experience or 7+ years SRE/DevOps experience; Linux experience required; Kubernetes, Kafka, cloud-native tooling, and programming in Python/Go/Java/C#/Scala preferred.
Linux, Kubernetes, Istio, Apache Kafka, Apache Spark, HBase, HDFS, Apache Flink, Python, C#, Java, Scala, Go
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Quota SRE
Sunnyvale, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent,8 years software/systems engineering experience,5 years SRE experience,5 years software design experience,EMR not mentioned; strong troubleshooting and stakeholder management skills.
Google Cloud, Quotaserver, Bouncer, Slicer
1mo
Save
Mark Applied
Hide
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
Santa Clara, California, United States
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
15+ YOE15+ years systems software or reliability engineering experience at datacenter scale; BS/MS in CS/EE/Statistics or equivalent; expertise in fleet telemetry, statistical failure analysis, burn-in/certification, and strong cross-functional communication.
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
New York or Mountain View
$200k-$220k/yr HybridFull Time
ASAPP
ASAPP: Develops AI-powered software to automate enterprise contact center interactions.
10+ YOE10+ years production software experience; distributed systems; automation; AWS; Terraform; Python/Go; Kubernetes; on-call; strong communication.
Terraform, Python, Go, AWS, Kubernetes
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, TikTok Generalized Arch USTO
San Jose, California, United States
$123k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS or related, strong software engineering and Linux knowledge, proficiency in Python/Go/Java/PHP/C/C++, strong problem solving and communication; SRE and AI-ops experience preferred.
Python, Go, Java, PHP, C, C++, Linux
3w
Save
Mark Applied
Hide
Senior Lead Site Reliability Engineer
Palo Alto, California, United States
$171k-$260k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years applied SRE experience, formal SRE training/certification, expertise in observability, distributed systems, cloud-native and AI-assisted reliability workflows.
Grafana, Dynatrace, Prometheus, Datadog, Splunk, Java, Go (Golang), Python, Terraform, LangChain, LangGraph, AutoGen, CrewAI, GitHub Copilot, Claude, Fluentd, Logstash, Vector, Kafka, RabbitMQ, SQS, Neo4j, TigerGraph, Pinecone, Weaviate, Chroma, Docker, Kubernetes, GitOps, MCP (Model Context Protocol), TensorFlow, PyTorch, scikit-learn, Hadoop, Spark, Flink, MongoDB, Cassandra, DynamoDB, InfluxDB, TimescaleDB, Chaos Monkey, Gremlin, LitmusChaos
1mo
Save
Mark Applied
Hide
(USA) Distinguished, Software Engineer-AI/ML Engineer - Agentic Systems & Site Reliability Engineering
Sunnyvale, California, United States
$169k-$338k/yr OnsiteFull Time
Walmart
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
12+ YOESenior AI/ML & SRE engineer with extensive experience designing agentic AI systems, observability, cloud-native platforms, and reliability tooling for mission-critical, large-scale distributed systems.
TensorFlow, PyTorch, Azure, GCP, AWS, Kubernetes, Docker, Jaeger, Zipkin, OpenTelemetry, Prometheus, Grafana, DataDog, ELK stack, Splunk, Fluentd, Terraform, CloudFormation, Pulumi, Istio, Linkerd, MLflow, Kubeflow, Seldon, Kafka, Pulsar
2w
Save
Mark Applied
Hide
Infrastructure Software Engineer
Campbell, California, United States
$180k-$230k/yr RemoteFull Time
Camus Energy
Camus Energy: Software platform for managing renewable energy grid integration.
3+ YOE3+ years software engineering with infrastructure focus; Python 3, Kubernetes, GCP, CI/CD, observability (Prometheus,Grafana); familiarity with SQL; collaborative incident response and reliability practices.
Python 3, Kubernetes, GCP, CI/CD pipelines, Prometheus, Grafana, Bazel, Go, Node.js, SQL
2mo
Save
Mark Applied
Hide
Sr. Software Engineer
Sunnyvale, California, United States
OnsiteFull Time
SingleStore
SingleStore: Real-time distributed SQL database for transactions and analytics.
6+ YOE6+ years in system-level software; Golang or similar; distributed systems; Kubernetes; strong performance and reliability focus; Bachelor's in CS or equivalent.
Golang, Kubernetes, React, Cloud platforms