189 systems reliability engineer jobs at 119 companies in Fairfield, CA

2mo
Save
Mark Applied
Hide
Systems Reliability Engineer
San Francisco or New York City
$150k-$170k/yr HybridFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years in SRE/infrastructure/distributed systems; strong Linux/networking; production experience; Kubernetes and cloud platforms; observability tools; multi-layer debugging; on-call readiness.
Kubernetes, Linux, Networking, Prometheus, Grafana, OpenTelemetry, AWS, GCP, Azure, VPNs
1mo
Save
Mark Applied
Hide
Senior/Staff Systems Reliability Engineer
Newark, California, United States
HybridFull Time
Queue
Queue: Developing autonomous robotic pharmacies for automated prescription fulfillment.
3+ YOE3+ years Terraform/IaC experience, deep AWS knowledge (VPC, ECS, RDS, IAM, KMS, SQS, CloudWatch), CI/CD (GitHub Actions), Docker, Linux administration; experience with Rust/TypeScript services and security/observability practices.
Terraform, AWS, VPC, ECS, RDS, IAM, KMS, SQS, CloudWatch, GitHub Actions, Docker, Rust, TypeScript, cargo, clippy, mTLS, X.509
1w
Save
Mark Applied
Hide
Senior Reliability Engineer
Berkeley, California, United States
$129k-$161k/yr OnsiteFull Time
Form Energy
Form Energy: Developing multi-day batteries for grid-scale energy storage.
4+ YOEBachelor's degree in relevant engineering field with 4+ years industry experience, expertise in designing reliability tests and accelerated lifetime testing, experience with liquid/gas handling systems, strong communication and cross-functional collaboration.
1mo
Save
Mark Applied
Hide
Senior Reliability Engineer
San Francisco, California, United States
$150k-$180k/yr OnsiteFull Time
Eight Sleep
Eight Sleep: Develops smart mattresses and temperature-regulated sleep technology.
5+ YOE5+ years reliability engineering in electromechanical systems; experience with DOE, FMEA, Weibull, SPC; Python/MATLAB; hardware test equipment; strong collaboration.
Python, MATLAB, LabVIEW, JMP
2w
Save
Mark Applied
Hide
Senior Reliability Engineer
Alameda, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
4+ YOEAssociate's degree (engineering preferred), minimum 4 years in FDA/ISO-regulated environments, experience in failure analysis, hardware/software integration, embedded systems, and strong project management skills.
2w
Save
Mark Applied
Hide
Senior Reliability Engineer
Alameda, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
4+ YOEAssociate degree in engineering, 4+ years in FDA/ISO-regulated environment, experience in product failure analysis, hardware/software integration, embedded systems and electrical design, strong problem-solving and project management skills.
1mo
Save
Mark Applied
Hide
Staff Quality & Reliability Engineer
San Francisco or New York City
HybridFull Time
Beast Industries
Beast Industries: Produces digital media and consumer goods for MrBeast brands.
Expert in software quality engineering and site reliability for consumer-scale distributed systems; owns test strategy, SLOs/error budgets, CI/CD test gates, observability, incident response, and reliability tooling.
CI/CD
3d
Save
Mark Applied
Hide
Customer Reliability Engineer
San Francisco, California, United States
$204k-$284k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience supporting large-scale compute customers, debugging distributed systems across stacks, strong incident communication, and driving engineering fixes.
InfiniBand, RoCE, Slurm, Kubernetes, NCCL
2w
Save
Mark Applied
Hide
Systems Engineer
San Francisco, California, United States
OnsiteFull Time
Dedalus Labs
Dedalus Labs: Infrastructure for building and deploying AI agent applications.
Fluency in Rust/Go/C/C++; strong software engineering fundamentals and systems knowledge (OS, networking, distributed systems); performance-oriented debugging and reliability focus.
Rust, Go, C/C++, Kubernetes, Firecracker, GitHub
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Builds advanced multimodal AI models and model optimization infrastructure.
Bachelor's degree or equivalent experience; distributed systems, cloud infrastructure or SRE; reliability tooling; incident response; strong cross-team communication.
Kubernetes, Docker, Cloud Platforms, Monitoring Tools, Automation
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
San Francisco, California, United States
$200k-$250k/yr OnsiteFull Time
Stuut
Stuut: Automates business accounts receivable and collections through AI agents.
7+ YOE7+ years in SRE/infrastructure or backend engineering. Experience with AWS, Kubernetes/EKS, Docker, observability, SLOs/SLIs, Python or TypeScript, CI/CD, and production-grade distributed systems.
Python, TypeScript, AWS, Kubernetes, EKS, Docker, FastAPI, Vue.js, PostgreSQL (RDS), CI/CD
5d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$149k-$224k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
Temporal, Airflow, Argo Workflows, Docker, Kubernetes, DNS, HTTP, Grafana, Prometheus, ELK, Splunk, Datadog, Python, Go, Linux, Claude Code, GitHub Copilot, Codex, Cursor, AWS, GCP, MCP
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States or San Francisco or New York City
$101k-$199k/yr OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
1+ YOEMaster's or Bachelor's in CS/IT (or equivalent experience), 1+ years managing physical infrastructure, on-call experience, experience with large-scale cloud/distributed systems preferred, and ability to pass Microsoft security screening.
Azure, InfiniBand, GPUs
1mo
Save
Mark Applied
Hide
Senior Platform Reliability Engineer
San Francisco or New York City or Seattle
$182k-$250k/yr HybridFull Time
Grow Therapy
Grow Therapy: Platform connecting mental health providers with patients and insurance.
6+ YOE6+ years operating production systems; hands-on AWS, Kubernetes (EKS), Terraform; experience defining SLOs/SLAs and observability (DataDog); strong communication and systems-thinking skills; PostgreSQL experience a plus.
AWS, Kubernetes, EKS, Terraform, DataDog, PostgreSQL, Gem
2mo
Save
Mark Applied
Hide
Senior Reliability Engineer
South San Francisco or Atlanta
$150k-$180k/yr HybridFull Time
AeroVect
AeroVect: Develops autonomous driving software for airport logistics vehicles.
5+ YOE5–8 years reliability engineering in hardware-focused field; strong FMEA/FTA/RBD; environmental and accelerated life testing; cross-domain systems; data analysis; excellent communication.
FMEA, FTA, RBD, Weibull Analysis, environmental testing, DFR
1w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
3w
Save
Mark Applied
Hide
Customer Reliability Engineer - Infrastructure
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Apache Airflow, AWS, Azure, CI/CD, GCP, Infrastructure as Code (IaC), Kubernetes, Linux, Python
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
San Francisco, California, United States
$200k-$260k/yr HybridFull Time
Sight Machine
Sight Machine: Developer of an AI-powered manufacturing data analytics platform.
10+ YOE10+ years experience with Kubernetes/Docker and major cloud providers, 10+ years coding (Python/Go/Java), IaC/CI-CD expertise, experience operating LLM/agentic AI systems, strong Linux and networking fundamentals.
Kubernetes, Docker, Azure, GCP, AWS, Python, Go, Java, Terraform, OpenTofu, FluxCD, Jenkins, GitHub Actions, Prometheus, Grafana, Loki, Sentry, Signoz, Helm Charts, Elasticsearch, Kafka, Postgres, LLM
2mo
Save
Mark Applied
Hide
Embedded Systems Engineer
San Francisco, California, United States
OnsiteFull Time
Blackstar Computers
Blackstar Computers: Building AI-native personal computing hardware and software systems.
Design and develop embedded systems, firmware, and drivers; collaborate with hardware and software teams; troubleshoot hardware/software integration; ensure system reliability.
C, C++, Embedded systems, RTOS, ARM
2mo
Save
Mark Applied
Hide
Software Reliability Engineer, Waymo Fleet
Mountain View or San Francisco
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
2+ YOE2+ years in C++, Java, or Python; strong distributed systems and observability focus; Bachelor's degree or equivalent experience.
C++, Java, Python, Distributed Systems

Explore Jobs

Expand Your Job Search