60 reliability automation engineer jobs at 48 companies in Fairfax, CA

3mo
Save
Mark Applied
Hide
Software Reliability Engineer
Mountain View, California, United States
$146k-$219k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
Production software experience; build automation/tools; strong debugging; reliability engineering interest.
Python, Go, Bash, C++, Observability, Telemetry
1mo
Save
Mark Applied
Hide
Reliability Engineer, Supercomputing
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Ensure reliability of GPU supercomputing fleet across hardware, firmware, and OS; debug kernel/driver/hardware issues; engage vendors; automate monitoring and runroot-cause analysis.
Python, Rust, Kubernetes, Slurm, Linux, BMC, iDRAC, IPMI, Redfish, DCGM, NVLink, NVSwitch, Linux kernel
1w
Save
Mark Applied
Hide
Senior Reliability Engineer, Labs
San Francisco or Oakland or United States
$139k-$205k/yr OnsiteFull Time
DoorDash
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
5+ YOEFive years of reliability validation or hardware testing experience in robotics or automated vehicles; bachelor's or higher in engineering; experience with test equipment, CAD, shop tools, Python, and hardware-software validation.
Python, CAD, DAQ, FRACAS, FMEA, FTA, HIL
3w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yr RemoteFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
11h
Save
Mark Applied
Hide
Site Reliability Engineer - rednote
Palo Alto, California, United States
OnsiteFull Time
Rednote
Rednote: A lifestyle-focused social media and e-commerce discovery platform.
Experience with large-scale reliability, high-availability architecture, incident response, cross-region disaster recovery, Linux, networking, middleware, cloud-native infrastructure, automation, and Python, Go, or Java.
Linux, MySQL, Redis, Kafka, Kubernetes, Service Mesh, Python, Go, Java
2mo
Save
Mark Applied
Hide
Senior Database Reliability Engineer
San Francisco or New York City or Seattle or Boston or Los Angeles or Chicago or Washington or United States
$145k-$230k/yr HybridFull Time
Scribe
Scribe: Automatically documents digital workflows into step-by-step process guides.
Deep PostgreSQL and ORM expertise, experience with CDC pipelines (AWS DMS), OpenSearch, Redis, message brokers, observability tools, Python/Go automation, Terraform/IaC, and building reliability/scale for data tiers.
Django, PostgreSQL, Aurora Serverless V2, OpenSearch, Redis, ElastiCache, SQS, RabbitMQ, DMS, S3, Parquet, Snowflake, pganalyze, CloudWatch, Honeycomb, OpenTelemetry, Datadog DBM, pg_stat_statements, Kafka, Python, Go, Terraform, Debezium, Fivetran, Airbyte, pgbouncer, RDS Proxy, Snowpipe, BigQuery, Redshift, SQLAlchemy, ActiveRecord
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute
San Francisco or New York or Austin or Seattle
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience owning large GPU/compute fleets, automation of repair/deployment pipelines, firmware/BMC/Redfish familiarity, incident response and paging, observability and metrics tooling, and proficiency with production automation.
Redfish, BMC, IPMI, Temporal, Cadence, Prometheus, Grafana, Go, Python, Kubernetes, Claude Code, Cursor, LLM APIs, MCP servers
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Plenful
Plenful: AI-powered workflow automation platform for healthcare and pharmacy operations.
5+ YOE5+ years SRE or production infrastructure experience; hands-on with observability, incident response, SLOs, AWS, container and serverless platforms; able to write automation scripts.
OpenTelemetry, Datadog, CloudWatch, Grafana, Sentry, AWS Lambda, ECS, Aurora Postgres, ClickHouse, GitHub Actions, Python, Bash, Vanta
1mo
Save
Mark Applied
Hide
Senior AV Automation Engineer
New York City or Seattle or San Francisco
$149k-$246k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years in systems/site reliability/DevOps/AV automation, proficiency in Python and REST API integrations, experience with automation/configuration tools and networking concepts, related technical degree required.
Splunk, Grafana, Slack, NetBox, Python, Ansible, Terraform, Puppet, Chef, Salt, New Relic, Kentik, Google Meet, Logitech, Neat, Cisco, Q-SYS, Google Workspace, Zoom, WebEx, Git, REST, AWS
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Palo Alto or Palo Alto or Washington
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
Kubernetes, Linux, Bash, Python, Bazel, Makefiles, Terraform, Ansible, TCP/IP
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
1mo
Save
Mark Applied
Hide
Lead Database Reliability Engineer - 11606
San Francisco, California, United States
$142k-$199k/yr RemoteFull Time
Coupa
Coupa: Cloud-based platform for managing and optimizing business expenditures.
8+ YOE8+ years hands-on DBA experience, deep MySQL expertise, scripting (Bash/Python/Ruby), cloud (AWS/RDS/Aurora) and automation experience, monitoring and HA/DR skills, ability to lead architecture and mentor engineers.
SQL Server, MySQL, Bash, Python, Ruby, AWS, RDS, Aurora, Azure, GCP, PMM, New Relic, VividCortex, Orchestrator, Chef, Puppet, Terraform, GitHub
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Bellevue or San Francisco
$147k-$226k/yr OnsiteFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
5+ YOE5+ years SRE/DevOps experience; expert AWS multi-account governance; Terraform and Python automation; Kubernetes and observability experience; strong networking, Linux, security and documentation skills.
AWS, AWS Orgs, IAM, Identity Center, StackSets, Terraform, Python, GitLab, GitHub Actions, Kubernetes, Splunk, CloudWatch, Grafana, BGP, IPsec, VPCs, TGWs, VPC endpoints, Linux
2mo
Save
Mark Applied
Hide
Founding Engineer - Site Reliability
San Francisco or United States
$185k-$285k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Kubernetes, AWS, Prometheus, Grafana, Datadog, Automation, VPC, GPU compute
4w
Save
Mark Applied
Hide
Site Reliability Engineer II
Scottsdale or San Francisco or Chicago or New York City
$86k-$126k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
2+ YOEBachelor's or equivalent, minimum 2 years DevOps/Dev/SRE experience, Linux/Unix experience, infrastructure automation (Chef/Ansible/Puppet, Terraform), containerization (Docker,Kubernetes), cloud (AWS/GCP/Azure), on-call rotation.
Linux, Unix, Chef, Ansible, Puppet, Terraform, Docker, Kubernetes, AWS, GCP, Azure, Java, Ruby, Python, JavaScript, Go
15h
Save
Mark Applied
Hide
Database Reliability Engineer
Livermore, California, United States
$146k-$223k/yr HybridFull Time, Contract
Lawrence Livermore National Laboratory
Lawrence Livermore National Laboratory: Develops science and technology for United States national security.
Bachelor's degree or equivalent experience; broad database reliability, automation, administration, monitoring, application server, operating system, storage, backup, and recovery experience; U.S. citizenship and ability to obtain DOE Q clearance.
Oracle, MySQL, Microsoft SQL Server, MongoDB, Cassandra, PostgreSQL, DynamoDB, WebLogic, Tomcat, SQL, Oracle Forms/Reports, PL/SQL, Java, Python, Windows, Linux, Oracle Enterprise Manager, Datadog, Grafana, SQL Server Management Studio, Icinga, Ansible, SCCM, PowerShell
1mo
Save
Mark Applied
Hide
Staff Database Reliability Engineer, DBRE
Palo Alto, California, United States
$165k-$185k/yr RemoteFull Time
Assured
Assured: Software platform for automating insurance claims processing
8+ YOE8+ years in SRE/DevOps/DBA roles; deep PostgreSQL and Amazon Aurora experience; proficiency with JavaScript/TypeScript and Node.js; experience optimizing production databases and building automation; Terraform, Docker/Kubernetes, Prisma, Redshift familiarity a plus.
PostgreSQL, Amazon Aurora, JavaScript, TypeScript, Node.js, Terraform, Terragrunt, Prisma, Docker, Kubernetes, Redshift, CI/CD
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE)
Palo Alto, California, United States
$175k-$229k/yr HybridFull Time
Instrumental
Instrumental: AI-powered software for electronics manufacturing quality and optimization.
5+ YOE5+ years DevOps/SRE experience on public cloud (AWS preferred); expertise in Linux, shell, containers, Kubernetes, terraform, monitoring/logging/APM; strong automation, KPI measurement, and security awareness; U.S. citizenship required for access-controlled work.
AWS, Linux, shell, containerization, Kubernetes, terraform, APM