136 platform reliability engineer jobs at 98 companies in Sunnyvale, CA
1mo
Save
Mark Applied
Hide
1mo
Senior Platform Reliability Engineer
San Francisco or New York City or Seattle
$182k-$250k/yrHybridFull Time
Grow Therapy: Platform connecting mental health providers with patients and insurance.
6+ YOE6+ years operating production systems; hands-on AWS, Kubernetes (EKS), Terraform; experience defining SLOs/SLAs and observability (DataDog); strong communication and systems-thinking skills; PostgreSQL experience a plus.
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
8+ YOE8+ years SRE/Platform/DevOps experience, strong Kubernetes and cloud-native platform skills, automation and CI/CD expertise, software engineering with Python/Go/Java/Ruby, observability and reliability knowledge.
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ years reliability validation or hardware testing experience for robotics/unmanned platforms, Bachelor's in engineering, proficiency with environmental test equipment, Python and CAD, strong communication and analytical skills.
Staff Site Reliability Engineer- Developer Platform
Palo Alto, California, United States
$186k-$233k/yrOnsiteFull Time
Rivian and Volkswagen Group Technologies: A joint venture creating cloud, connectivity and software-defined vehicle solutions for electric vehicles.
5+ YOE5+ years in Platform/DevOps/SRE; Terraform, Kubernetes, GitOps (ArgoCD or Flux), cloud (AWS/Azure/GCP), scripting (Python, Bash, Go); strong communication and mentoring skills.
Senior Site Reliability Engineer Platform Cloud Foundations Engineer
San Jose, California, United States
$64k-$130k/yrOnsiteFull Time
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ YOE8+ years SRE/platform engineering experience with AWS multi-account, Terraform, automation (Python/Go/Ruby), cloud governance, and strong documentation and communication skills.
AWS Organizations, IAM, Terraform, Python, Go, Ruby, Control Tower, Account Factory for Terraform, CloudFormation, EventBridge, Lambda, SQS, IAM Identity Center, GCP
Forge GlobalNYSE: FRGE: Marketplace for trading private shares and pre-IPO stock.
8+ YOE5+ Mgmt8+ years software engineering experience with infrastructure/platform focus, 5+ years people leadership, deep cloud/observability/incident response experience, strong distributed systems judgment, and ability to set platform strategy.
Authorium: Cloud-based administrative operations platform for government agencies.
6+ YOE6+ years in platform/infrastructure/DevOps engineering with distributed systems, cloud architecture, security, and reliability; strong collaboration in an in-person SF office (Mon–Thu).
Wispr Flow: Provides AI-powered voice dictation software for computers and mobile devices.
Experienced engineer who has built billing systems, run payments migrations, reconciled billing across platforms, and built testing for high-reliability billing.
GC AI: AI-powered legal platform for in-house legal teams.
Production experience with Google Cloud Platform, IAM, infrastructure as code (Terraform/Pulumi), CI/CD, observability, and backend development (TypeScript preferred). Strong reliability, operational excellence, and developer experience focus.
Google Cloud Platform, Terraform, Pulumi, CI/CD, TypeScript, IAM
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yrHybridFull Time
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
United States or Kansas or Washington or California or Texas or Illinois or North Carolina or Colorado or Massachusetts or Pennsylvania or Virginia or Oregon or Nevada or Hawaii or New York or Georgia or Ohio or Arizona or Seattle or San Francisco or New York City
$110k-$183k/yrRemoteFull Time
Veeam: Data resilience and security for hybrid cloud environments
3+ YOE3+ years in software engineering with 1+ year in SRE/Platform/DevOps, cloud experience (Azure or comparable), observability (Prometheus, Grafana, OpenTelemetry, ELK), IaC (Terraform/Terragrunt/Pulumi), Kubernetes, CI/CD tooling, programming in TypeScript/JS, Go, Java, or C#, and experience in compliance-oriented environments.