91 reliability automation engineer jobs at 63 companies in Pacifica, CA

3mo
Save
Mark Applied
Hide
Reliability Engineer
Santa Clara, California, United States
$45k-$121k/yr OnsiteFull Time
Wipro
WiproNYSE: WIT: Global technology services and consulting for digital transformation.
2+ YOEBachelor's in electrical engineering, 2+ years electronics or lab reliability experience, knowledge of failure analysis (PFA, CSAM, X-ray), VLSI board design, test automation scripting, and lab instruments; familiar with Windows/Linux/CentOS and Microsoft Office.
Windows, Linux, CentOS, Microsoft Office, Physical Failure Analysis (PFA), C-Scan Acoustic Microscopy (CSAM), X-ray imaging, oscilloscopes, multimeters, curve tracers, VLSI Board Design
3mo
Save
Mark Applied
Hide
Software Reliability Engineer
Mountain View, California, United States
$146k-$219k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
Production software experience; build automation/tools; strong debugging; reliability engineering interest.
Python, Go, Bash, C++, Observability, Telemetry
1mo
Save
Mark Applied
Hide
Reliability Engineer, Supercomputing
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Ensure reliability of GPU supercomputing fleet across hardware, firmware, and OS; debug kernel/driver/hardware issues; engage vendors; automate monitoring and runroot-cause analysis.
Python, Rust, Kubernetes, Slurm, Linux, BMC, iDRAC, IPMI, Redfish, DCGM, NVLink, NVSwitch, Linux kernel
1w
Save
Mark Applied
Hide
Senior Reliability Engineer, Labs
San Francisco or Oakland or United States
$139k-$205k/yr OnsiteFull Time
DoorDash
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
5+ YOEFive years of reliability validation or hardware testing experience in robotics or automated vehicles; bachelor's or higher in engineering; experience with test equipment, CAD, shop tools, Python, and hardware-software validation.
Python, CAD, DAQ, FRACAS, FMEA, FTA, HIL
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEDegree in CS or related field (or equivalent experience), 8+ years SRE/DevOps/Production Engineering, SRE principles, infrastructure automation, production reliability, Python/Go/Perl/Ruby, Prometheus and Grafana, strong communication.
Python, Go, Perl, Ruby, Prometheus, Grafana
2mo
Save
Mark Applied
Hide
Senior Database Reliability Engineer
San Francisco or New York City or Seattle or Boston or Los Angeles or Chicago or Washington or United States
$145k-$230k/yr HybridFull Time
Scribe
Scribe: Automatically documents digital workflows into step-by-step process guides.
Deep PostgreSQL and ORM expertise, experience with CDC pipelines (AWS DMS), OpenSearch, Redis, message brokers, observability tools, Python/Go automation, Terraform/IaC, and building reliability/scale for data tiers.
Django, PostgreSQL, Aurora Serverless V2, OpenSearch, Redis, ElastiCache, SQS, RabbitMQ, DMS, S3, Parquet, Snowflake, pganalyze, CloudWatch, Honeycomb, OpenTelemetry, Datadog DBM, pg_stat_statements, Kafka, Python, Go, Terraform, Debezium, Fivetran, Airbyte, pgbouncer, RDS Proxy, Snowpipe, BigQuery, Redshift, SQLAlchemy, ActiveRecord
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute
San Francisco or New York or Austin or Seattle
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience owning large GPU/compute fleets, automation of repair/deployment pipelines, firmware/BMC/Redfish familiarity, incident response and paging, observability and metrics tooling, and proficiency with production automation.
Redfish, BMC, IPMI, Temporal, Cadence, Prometheus, Grafana, Go, Python, Kubernetes, Claude Code, Cursor, LLM APIs, MCP servers
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
2w
Save
Mark Applied
Hide
Principal Design Automation Engineer
San Jose, California, United States
$176k-$298k/yr OnsiteFull Time
Micron Technology
Micron TechnologyNASDAQ: MU: Designs and manufactures semiconductor memory and data storage solutions.
8+ YOEMS in electrical engineering required, 8+ years NAND design experience, proficiency in analog/mixed-signal design and layout, circuit verification, Python automation, Cadence tools, HSPICE/Fast SPICE/Verilog, and semiconductor reliability understanding.
HSPICE, Fast SPICE, Verilog, Python, Cadence, UNIX, LVS, DRC
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Plenful
Plenful: AI-powered workflow automation platform for healthcare and pharmacy operations.
5+ YOE5+ years SRE or production infrastructure experience; hands-on with observability, incident response, SLOs, AWS, container and serverless platforms; able to write automation scripts.
OpenTelemetry, Datadog, CloudWatch, Grafana, Sentry, AWS Lambda, ECS, Aurora Postgres, ClickHouse, GitHub Actions, Python, Bash, Vanta
1mo
Save
Mark Applied
Hide
Senior AV Automation Engineer
New York City or Seattle or San Francisco
$149k-$246k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years in systems/site reliability/DevOps/AV automation, proficiency in Python and REST API integrations, experience with automation/configuration tools and networking concepts, related technical degree required.
Splunk, Grafana, Slack, NetBox, Python, Ansible, Terraform, Puppet, Chef, Salt, New Relic, Kentik, Google Meet, Logitech, Neat, Cisco, Q-SYS, Google Workspace, Zoom, WebEx, Git, REST, AWS
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Palo Alto or Palo Alto or Washington
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
Kubernetes, Linux, Bash, Python, Bazel, Makefiles, Terraform, Ansible, TCP/IP
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
1mo
Save
Mark Applied
Hide
Lead Database Reliability Engineer - 11606
San Francisco, California, United States
$142k-$199k/yr RemoteFull Time
Coupa
Coupa: Cloud-based platform for managing and optimizing business expenditures.
8+ YOE8+ years hands-on DBA experience, deep MySQL expertise, scripting (Bash/Python/Ruby), cloud (AWS/RDS/Aurora) and automation experience, monitoring and HA/DR skills, ability to lead architecture and mentor engineers.
SQL Server, MySQL, Bash, Python, Ruby, AWS, RDS, Aurora, Azure, GCP, PMM, New Relic, VividCortex, Orchestrator, Chef, Puppet, Terraform, GitHub
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Bellevue or San Francisco
$147k-$226k/yr OnsiteFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
5+ YOE5+ years SRE/DevOps experience; expert AWS multi-account governance; Terraform and Python automation; Kubernetes and observability experience; strong networking, Linux, security and documentation skills.
AWS, AWS Orgs, IAM, Identity Center, StackSets, Terraform, Python, GitLab, GitHub Actions, Kubernetes, Splunk, CloudWatch, Grafana, BGP, IPsec, VPCs, TGWs, VPC endpoints, Linux
2mo
Save
Mark Applied
Hide
Founding Engineer - Site Reliability
San Francisco or United States
$185k-$285k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Kubernetes, AWS, Prometheus, Grafana, Datadog, Automation, VPC, GPU compute