189 systems reliability engineer jobs at 119 companies in Fairfield, CA
2mo
Save
Mark Applied
Hide
2mo
Systems Reliability Engineer
San Francisco or New York City
$150k-$170k/yrHybridFull Time
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years in SRE/infrastructure/distributed systems; strong Linux/networking; production experience; Kubernetes and cloud platforms; observability tools; multi-layer debugging; on-call readiness.
Form Energy: Developing multi-day batteries for grid-scale energy storage.
4+ YOEBachelor's degree in relevant engineering field with 4+ years industry experience, expertise in designing reliability tests and accelerated lifetime testing, experience with liquid/gas handling systems, strong communication and cross-functional collaboration.
Eight Sleep: Develops smart mattresses and temperature-regulated sleep technology.
5+ YOE5+ years reliability engineering in electromechanical systems; experience with DOE, FMEA, Weibull, SPC; Python/MATLAB; hardware test equipment; strong collaboration.
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
4+ YOEAssociate degree in engineering, 4+ years in FDA/ISO-regulated environment, experience in product failure analysis, hardware/software integration, embedded systems and electrical design, strong problem-solving and project management skills.
Beast Industries: Produces digital media and consumer goods for MrBeast brands.
Expert in software quality engineering and site reliability for consumer-scale distributed systems; owns test strategy, SLOs/error budgets, CI/CD test gates, observability, incident response, and reliability tooling.
Dedalus Labs: Infrastructure for building and deploying AI agent applications.
Fluency in Rust/Go/C/C++; strong software engineering fundamentals and systems knowledge (OS, networking, distributed systems); performance-oriented debugging and reliability focus.
Stuut: Automates business accounts receivable and collections through AI agents.
7+ YOE7+ years in SRE/infrastructure or backend engineering. Experience with AWS, Kubernetes/EKS, Docker, observability, SLOs/SLIs, Python or TypeScript, CI/CD, and production-grade distributed systems.
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
1+ YOEMaster's or Bachelor's in CS/IT (or equivalent experience), 1+ years managing physical infrastructure, on-call experience, experience with large-scale cloud/distributed systems preferred, and ability to pass Microsoft security screening.
Grow Therapy: Platform connecting mental health providers with patients and insurance.
6+ YOE6+ years operating production systems; hands-on AWS, Kubernetes (EKS), Terraform; experience defining SLOs/SLAs and observability (DataDog); strong communication and systems-thinking skills; PostgreSQL experience a plus.
AeroVect: Develops autonomous driving software for airport logistics vehicles.
5+ YOE5–8 years reliability engineering in hardware-focused field; strong FMEA/FTA/RBD; environmental and accelerated life testing; cross-domain systems; data analysis; excellent communication.
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Sight Machine: Developer of an AI-powered manufacturing data analytics platform.
10+ YOE10+ years experience with Kubernetes/Docker and major cloud providers, 10+ years coding (Python/Go/Java), IaC/CI-CD expertise, experience operating LLM/agentic AI systems, strong Linux and networking fundamentals.
Blackstar Computers: Building AI-native personal computing hardware and software systems.
Design and develop embedded systems, firmware, and drivers; collaborate with hardware and software teams; troubleshoot hardware/software integration; ensure system reliability.