63 service reliability engineer jobs at 46 companies in Mill Valley, CA

1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$149k-$224k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
Temporal, Airflow, Argo Workflows, Docker, Kubernetes, DNS, HTTP, Grafana, Prometheus, ELK, Splunk, Datadog, Python, Go, Linux, Claude Code, GitHub Copilot, Codex, Cursor, AWS, GCP, MCP
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, ASE
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Work on globally scaled, revenue-critical internet services (App Store, Music, Books, Podcasts, Fitness+); ensure reliability and scalability of services used by billions of devices.
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – CloudVision as a Service (CVaaS)
Santa Clara, California, United States
$101k-$161k/yr RemoteFull Time
Arista Networks
Arista NetworksNYSE: ANET: Provides cloud networking solutions and high-speed multilayer Ethernet switches.
5+ YOEBS/MS or equivalent experience,5+ years software engineering, experience with distributed databases/SaaS deployments, proficiency in Python/Golang/Bash, Kubernetes and cloud platform experience preferred.
Golang, Python, Ansible, Pulumi, Bash, Kubernetes, GKE, GCP
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - HPC
Santa Clara or Durham or Austin
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEBS in CS or equivalent with 5+ years supporting critical services; experience with large-scale HPC clusters (Slurm, LSF, Kubernetes), IaC, CI/CD, multi‑cloud (AWS/GCP/OCI), and 2+ languages such as Python or Go.
Slurm, LSF, Kubernetes, AWS, GCP, OCI, Infrastructure as Code (IaC), CI/CD, Python, Go, Perl, Ruby, AIOps
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - HPC
Santa Clara or Austin or Durham
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years building/supporting critical services; experience with large-scale HPC clusters (Slurm, LSF, Kubernetes); IaC and CI/CD proficiency; coding in Python/Go/Perl/Ruby; monitoring, capacity planning, and incident response skills.
Slurm, LSF, Kubernetes, Infrastructure as Code (IaC), CI/CD, AWS, GCP, OCI, Python, Go, Perl, Ruby
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
2mo
Save
Mark Applied
Hide
Site Reliability Engineer for Linux administration
Ontario or Greenville or Clearwater or Fremont
OnsiteFull Time
Hyve Solutions
Hyve SolutionsNYSE: SNX: Designs and manufactures custom hardware for hyperscale data centers.
3+ YOE3+ years Linux production administration (RHEL/Ubuntu/Rocky/CentOS); knowledge of core services, storage, backups, monitoring, security hardening, and basic scripting/automation. Bachelor's in CS/IT or equivalent experience; RHCSA/CompTIA Linux+/LPIC-1 are nice-to-have.
RHEL, Ubuntu, Rocky, CentOS, SSH, DNS, DHCP, NTP, LDAP, SSSD, Postfix, NFS, SMB, systemd, cron, yum, dnf, apt, Prometheus, Zabbix, Veeam, Bacula, NetBackup, LVM, mdadm, multipath, iSCSI, FC, ext4, xfs, btrfs, CIS, STIG, iptables, nftables, firewalld, SELinux, AppArmor, Shell, Python, Ansible, Nagios, VMware, KVM, Nutanix, Docker, Git, journalctl, syslog, auditd
2w
Save
Mark Applied
Hide
Software Engineer - NEO Reliability
San Carlos, California, United States
$200k-$300k/yr OnsiteFull Time
1X
1X: Manufacturing safe, general-purpose humanoid robots for home and work.
Experienced software engineer with reliability and testing expertise across cloud, mobile, on-robot services, and embedded systems; strong DevOps and Python/Linux skills; statistical and systems reasoning.
Python, Linux, CI/CD, APIs, HIL
3mo
Save
Mark Applied
Hide
Staff Infrastructure Reliability Engineer - Database & Storage
Seattle or San Francisco or Detroit or United States
$180k-$279k/yr HybridFull Time
Rocket Companies
Rocket CompaniesNYSE: RKT: Provides digital mortgage, real estate, and personal finance services.
7+ YOE7+ years AWS/cloud infra; 5+ years PostgreSQL/AWS services; Linux admin/scripting; mentoring; infrastructure as code and security; AI code generation tools; on-call readiness.
AWS, PostgreSQL, Aurora/RDS, S3, ElastiCache, OpenSearch, DynamoDB, Linux, Python, Infrastructure as Code, Security practices, AI code generation tools
1mo
Save
Mark Applied
Hide
Field Service Engineer MGC
Sacramento or San Francisco or San Jose or Modesto or Stockton
$80k-$90k/yr FieldFull Time
MGC Diagnostics
MGC Diagnostics: Sells non-invasive cardiorespiratory diagnostic systems and respiratory software.
2+ YOEAssociate or bachelor's in electronics, biomedical engineering, or related; minimum 2 years field service experience servicing medical equipment; valid driver’s license, reliable transportation; strong troubleshooting, communication, and computer skills.
CRM
2mo
Save
Mark Applied
Hide
Member of Technical Staff, Site Reliablity Engineer
San Francisco, California, United States
$200k-$270k/yr HybridFull Time
Vapi
Vapi: Build and deploy AI-powered voice agents via flexible APIs.
Experience running incident command and postmortems, operating SLOs/error budgets, capacity planning and load testing, Kubernetes production ops, KEDA autoscaling, and shipping services in Go or TypeScript.
Go, TypeScript, Bash, Chronosphere, Prometheus, Grafana, Datadog, OpenTelemetry, Kubernetes, EKS, KEDA
2mo
Save
Mark Applied
Hide
Senior Software Engineer - Observability and Reliability
New York City or San Francisco or London or Sydney
$170k-$240k/yr OnsiteFull Time
Sigma Computing
Sigma Computing: Cloud-native analytics platform featuring a spreadsheet-style interface.
5+ YOE5+ years building high-quality software, strong CS fundamentals, experience building observability tools, proficiency with Go, OpenTelemetry, Kubernetes, participation in on-call rotations, cloud service administration (GCP/AWS/Azure) preferred.
Go, OpenTelemetry, Kubernetes, GCP, AWS, Azure, SQL, Python
2w
Save
Mark Applied
Hide
Founding Engineer (Platform)
San Francisco, California, United States
OnsiteFull Time
Polymr
Polymr: AI-native procurement and manufacturing workflow automation platform.
Experience building multi-tenant systems, data layers, integrations, deployments, and reliable production services; customer-facing collaboration; familiarity with AI coding tools preferred.
Claude Code, Cursor
2mo
Save
Mark Applied
Hide
Senior Plant Engineer, Dixon - Site Services
Dixon or South San Francisco
$110k-$203k/yr OnsiteFull Time
Roche
RocheSIX Swiss Exchange: ROG: Provides innovative pharmaceutical and diagnostic healthcare solutions.
8+ YOEBA/BS in engineering; 8+ years infrastructure design/operation experience; 3+ years campus environment experience; proficiency with BAS and CMMS (e.g., SAP, Siemens); knowledge of NFPA, ASHRAE, OSHA; reliability engineering and capital project leadership.
BAS, CMMS, SAP, Siemens, Historians
3d
Save
Mark Applied
Hide
Sr. Piping Engineer
San Francisco, California, United States
OnsiteFull Time
Noble Thermodynamic Systems
Noble Thermodynamic Systems: Developing zero-emissions, ultra-efficient power generation technology.
Owns piping systems across power projects from conceptual design through construction, commissioning, and operations support, ensuring reliability, safety, accessibility, and serviceability.
Argon Power Cycle
1mo
Save
Mark Applied
Hide
Senior Software Engineer - New Revenue
San Francisco or Seattle
$158k-$266k/yr HybridFull Time
DocuSign
DocuSignNASDAQ: DOCU: Provider of e-signature and intelligent agreement management software.
8+ YOE8+ years full-stack development experience with frontend and backend technologies; experience designing backend APIs, CI/CD, service reliability, and strong communication skills.
JavaScript, Typescript, React, C#, Java, REST, GraphQL, Jenkins, GitHub Actions, CI/CD, Git, Microsoft Azure, AWS, Azure SQL Database, SQL Server, Cosmos DB, Scrum
3mo
Save
Mark Applied
Hide
Lead Systems Engineer - Traffic Management
Durham or Miami or Palo Alto or Washington
$15k-$23k/yr HybridFull Time
Nubank
NubankNYSE: NU: Digital financial platform offering banking, credit, and investment services.
Experience operating large-scale cloud distributed systems, Kubernetes, service mesh (Istio/Linkerd/Envoy), AWS networking, IaC with Pulumi or Terraform, strong reliability/observability and AI-assisted workflows.
Kubernetes, Istio, Linkerd, Envoy, Finagle, AWS, ALB, NLB, VPC, Route 53, IAM, Pulumi, Terraform
2mo
Save
Mark Applied
Hide
Global Product Support (GPS) Engineer III
Santa Clara, California, United States
$124k-$171k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Manufacturers of equipment for semiconductor and display production.
Experience developing service procedures and troubleshooting complex equipment, conducting failure analysis, creating technical manuals and training, and participating in quality/reliability and CAPA processes.
1d
Save
Mark Applied
Hide
Product Engineer
San Francisco or Sacramento or West Sacramento or United States
$7k-$11k/mo HybridFull Time
State Controller's Office
State Controller's Office: California's fiscal controller managing state financial operations and assets.
Software engineering experience across applications, databases, APIs, integrations, testing, cloud services, delivery pipelines, and production operations, with strong judgment in security, privacy, accessibility, and reliability.
SQL, Azure, .NET, Terraform, CI/CD