45 reliability automation engineer jobs at 34 companies in Suisun, CA

1mo
Save
Mark Applied
Hide
Reliability Engineer, Supercomputing
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Ensure reliability of GPU supercomputing fleet across hardware, firmware, and OS; debug kernel/driver/hardware issues; engage vendors; automate monitoring and runroot-cause analysis.
Python, Rust, Kubernetes, Slurm, Linux, BMC, iDRAC, IPMI, Redfish, DCGM, NVLink, NVSwitch, Linux kernel
1w
Save
Mark Applied
Hide
Senior Reliability Engineer, Labs
San Francisco or Oakland or United States
$139k-$205k/yr OnsiteFull Time
DoorDash
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
5+ YOEFive years of reliability validation or hardware testing experience in robotics or automated vehicles; bachelor's or higher in engineering; experience with test equipment, CAD, shop tools, Python, and hardware-software validation.
Python, CAD, DAQ, FRACAS, FMEA, FTA, HIL
2mo
Save
Mark Applied
Hide
Senior Database Reliability Engineer
San Francisco or New York City or Seattle or Boston or Los Angeles or Chicago or Washington or United States
$145k-$230k/yr HybridFull Time
Scribe
Scribe: Automatically documents digital workflows into step-by-step process guides.
Deep PostgreSQL and ORM expertise, experience with CDC pipelines (AWS DMS), OpenSearch, Redis, message brokers, observability tools, Python/Go automation, Terraform/IaC, and building reliability/scale for data tiers.
Django, PostgreSQL, Aurora Serverless V2, OpenSearch, Redis, ElastiCache, SQS, RabbitMQ, DMS, S3, Parquet, Snowflake, pganalyze, CloudWatch, Honeycomb, OpenTelemetry, Datadog DBM, pg_stat_statements, Kafka, Python, Go, Terraform, Debezium, Fivetran, Airbyte, pgbouncer, RDS Proxy, Snowpipe, BigQuery, Redshift, SQLAlchemy, ActiveRecord
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute
San Francisco or New York or Austin or Seattle
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience owning large GPU/compute fleets, automation of repair/deployment pipelines, firmware/BMC/Redfish familiarity, incident response and paging, observability and metrics tooling, and proficiency with production automation.
Redfish, BMC, IPMI, Temporal, Cadence, Prometheus, Grafana, Go, Python, Kubernetes, Claude Code, Cursor, LLM APIs, MCP servers
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Plenful
Plenful: AI-powered workflow automation platform for healthcare and pharmacy operations.
5+ YOE5+ years SRE or production infrastructure experience; hands-on with observability, incident response, SLOs, AWS, container and serverless platforms; able to write automation scripts.
OpenTelemetry, Datadog, CloudWatch, Grafana, Sentry, AWS Lambda, ECS, Aurora Postgres, ClickHouse, GitHub Actions, Python, Bash, Vanta
1mo
Save
Mark Applied
Hide
Senior AV Automation Engineer
New York City or Seattle or San Francisco
$149k-$246k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years in systems/site reliability/DevOps/AV automation, proficiency in Python and REST API integrations, experience with automation/configuration tools and networking concepts, related technical degree required.
Splunk, Grafana, Slack, NetBox, Python, Ansible, Terraform, Puppet, Chef, Salt, New Relic, Kentik, Google Meet, Logitech, Neat, Cisco, Q-SYS, Google Workspace, Zoom, WebEx, Git, REST, AWS
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
1mo
Save
Mark Applied
Hide
Lead Database Reliability Engineer - 11606
San Francisco, California, United States
$142k-$199k/yr RemoteFull Time
Coupa
Coupa: Cloud-based platform for managing and optimizing business expenditures.
8+ YOE8+ years hands-on DBA experience, deep MySQL expertise, scripting (Bash/Python/Ruby), cloud (AWS/RDS/Aurora) and automation experience, monitoring and HA/DR skills, ability to lead architecture and mentor engineers.
SQL Server, MySQL, Bash, Python, Ruby, AWS, RDS, Aurora, Azure, GCP, PMM, New Relic, VividCortex, Orchestrator, Chef, Puppet, Terraform, GitHub
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Bellevue or San Francisco
$147k-$226k/yr OnsiteFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
5+ YOE5+ years SRE/DevOps experience; expert AWS multi-account governance; Terraform and Python automation; Kubernetes and observability experience; strong networking, Linux, security and documentation skills.
AWS, AWS Orgs, IAM, Identity Center, StackSets, Terraform, Python, GitLab, GitHub Actions, Kubernetes, Splunk, CloudWatch, Grafana, BGP, IPsec, VPCs, TGWs, VPC endpoints, Linux
2mo
Save
Mark Applied
Hide
Founding Engineer - Site Reliability
San Francisco or United States
$185k-$285k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Kubernetes, AWS, Prometheus, Grafana, Datadog, Automation, VPC, GPU compute
3w
Save
Mark Applied
Hide
Site Reliability Engineer II
Scottsdale or San Francisco or Chicago or New York City
$86k-$126k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
2+ YOEBachelor's or equivalent, minimum 2 years DevOps/Dev/SRE experience, Linux/Unix experience, infrastructure automation (Chef/Ansible/Puppet, Terraform), containerization (Docker,Kubernetes), cloud (AWS/GCP/Azure), on-call rotation.
Linux, Unix, Chef, Ansible, Puppet, Terraform, Docker, Kubernetes, AWS, GCP, Azure, Java, Ruby, Python, JavaScript, Go
2d
Save
Mark Applied
Hide
Quality & Reliability Engineer
San Francisco or San Mateo
OnsiteFull Time
Beast Industries
Beast Industries: Produces digital media and consumer goods for MrBeast brands.
2+ YOERequires 2+ years building automated test infrastructure for consumer mobile or web products, with test automation frameworks, CI/CD pipelines, monitoring tools, and strong collaboration skills.
Playwright, Appium, Espresso, XCTest, CI/CD, SLO, AI tools
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$117k-$209k/yr OnsiteFull Time
Autodesk
AutodeskNASDAQ: ADSK: Developing software for architecture, engineering, and entertainment industries.
7+ YOEU.S. citizen required. 7+ years SRE/platform/cloud experience; B.S. in CS/Engineering or equivalent; experience with large-scale cloud production systems, SLOs/SLIs, observability, incident management, automation, and IaC. Programming in Python/Go/Java/PowerShell/Bash.
AWS, Azure, Python, Go, Java, PowerShell, Bash, Kubernetes, Splunk, Dynatrace, Datadog, CloudWatch, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yr RemoteFull Time
Bedrock Ocean Exploration
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Python, Go, Bash, Terraform, AWS, Docker, Kubernetes, Prometheus, Grafana, ROS 2, ROS, Jetson, Linux, IAM
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer – Compute Platforms
San Ramon or United States
$82k-$229k/yr RemoteFull Time
Five9
Five9NASDAQ: FIVN: Provides cloud-based software for enterprise contact center operations.
6+ YOE6+ years in infrastructure/DevOps with compute systems and Kubernetes/OpenStack expertise; strong Linux, automation, and scripting skills.
Kubernetes, OpenStack, Linux, PXE, Redfish, Ansible, Terraform, Helm, Git, Python, Bash, Harvester, Ubuntu, KVM, ArgoCD, CI/CD
1mo
Save
Mark Applied
Hide
Software Engineer, Infrastructure & Reliability
San Francisco, California, United States
HybridFull Time
CrewAI
CrewAI: Platform for orchestrating collaborative multi-agent AI systems.
Experience building and operating production SaaS infrastructure: cloud, containers, CI/CD, observability, secrets, databases, and automation using Python/Ruby/Go/Bash.
AWS, Docker, CI/CD, GitHub Actions, ECS, ECR, Kubernetes, Helm, PostgreSQL, Redis, Celery, FastAPI, Rails, Sentry, OpenTelemetry, Python, Ruby, Go, Bash, Terraform
2mo
Save
Mark Applied
Hide
Site Reliability Engineer for Linux administration
Ontario or Greenville or Clearwater or Fremont
OnsiteFull Time
Hyve Solutions
Hyve SolutionsNYSE: SNX: Designs and manufactures custom hardware for hyperscale data centers.
3+ YOE3+ years Linux production administration (RHEL/Ubuntu/Rocky/CentOS); knowledge of core services, storage, backups, monitoring, security hardening, and basic scripting/automation. Bachelor's in CS/IT or equivalent experience; RHCSA/CompTIA Linux+/LPIC-1 are nice-to-have.
RHEL, Ubuntu, Rocky, CentOS, SSH, DNS, DHCP, NTP, LDAP, SSSD, Postfix, NFS, SMB, systemd, cron, yum, dnf, apt, Prometheus, Zabbix, Veeam, Bacula, NetBackup, LVM, mdadm, multipath, iSCSI, FC, ext4, xfs, btrfs, CIS, STIG, iptables, nftables, firewalld, SELinux, AppArmor, Shell, Python, Ansible, Nagios, VMware, KVM, Nutanix, Docker, Git, journalctl, syslog, auditd
1mo
Save
Mark Applied
Hide
Reliability Design Associate (Fall 2026)
San Francisco, California, United States
$2k/wk OnsiteFull Time, Contract
Astranis
Astranis: Builds and operates small geostationary communications satellites.
Bachelor's in electrical/computer engineering or physics (or equivalent), US work authorization (citizen/green card/asylee/refugee), hands-on hardware design and test experience, interest in electronics, ability to write code for test automation.
LTSpice, PSPICE, ADS, Python, oscilloscope, multimeter, power supply