65 platform reliability engineer jobs at 33 companies in Prunedale, CA

1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS/Engineering, strong Linux, networking, databases, Kubernetes, SRE/DevOps toolset knowledge, experience with ClickHouse/Spark/Presto/Doris/Hadoop, coding in Python/Shell/Java/Go, strong problem-solving and communication.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Python, Shell, Java, Go
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer Platform Cloud Foundations Engineer
San Jose, California, United States
$64k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ YOE8+ years SRE/platform engineering experience with AWS multi-account, Terraform, automation (Python/Go/Ruby), cloud governance, and strong documentation and communication skills.
AWS Organizations, IAM, Terraform, Python, Go, Ruby, Control Tower, Account Factory for Terraform, CloudFormation, EventBridge, Lambda, SQS, IAM Identity Center, GCP
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Experience with Linux, networking, databases, Kubernetes, ClickHouse/Hadoop/Doris/Spark/Presto, scripting or programming (Python, Shell, Java, Go), incident management, and capacity planning.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Linux, Python, Shell, Java, Go
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
3mo
Save
Mark Applied
Hide
Head of Platform Product Reliability
San Jose, California, United States
OnsiteFull Time
Etched
Etched: Designs specialized AI chips optimized for transformer architectures.
10+ YOE10+ years reliability engineering for complex hardware systems, degree in engineering, hands-on FMEA/Weibull/HALT-HASS, qualification planning, failure analysis, and cross-functional leadership.
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1w
Save
Mark Applied
Hide
Sr. Database Reliability Engineer
San Jose, California, United States
$139k-$258k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
7+ YOE7+ years operating highly available database platforms; strong experience with MongoDB/Cassandra/MySQL/PostgreSQL, cloud (AWS/Azure), managed DB services, IaC (Terraform/Chef/Ansible), Kubernetes/Docker, Python; bachelor's or equivalent experience.
MongoDB, Cassandra, MySQL, PostgreSQL, Percona XtraDB Cluster, MariaDB Galera Cluster, AWS, Azure, Amazon RDS, Keyspaces, DynamoDB, Azure SQL, Cosmos DB, MongoDB Atlas, Terraform, Chef, Ansible, Kubernetes, Docker, Python
3w
Save
Mark Applied
Hide
Principal Tech Lead Manager - Data Platform & Reliability Engineering
Mountain View, California, United States
$215k-$275k/yr OnsiteFull Time
ID.me
ID.me: Provides secure digital identity verification and authentication services.
5+ YOE3+ Mgmt8+ years engineering experience with 3+ years managing teams,5+ years in data/platform/SRE; bachelor\u0002s or equivalent; deep PostgreSQL and data reliability expertise; strong communication and cloud/IaC experience.
PostgreSQL, Neo4j, Amazon Neptune, Kafka, Kinesis, Kubernetes, Terraform, Helm, AWS
3w
Save
Mark Applied
Hide
Senior Platform AI Engineer
Santa Clara or California
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE8+ years building and operating production platform/backend infrastructure; 5+ years ML infrastructure; strong Python and a compiled language; experience with job queues, sandboxed execution, and production reliability.
Python, C, C++, Go, Java, Rust, Kubernetes Jobs, Celery, Sidekiq, Temporal, container runtimes
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Apply SRE principles to mentor teams, ensure reliability for large-scale analytics infrastructure across Hadoop, HBase, Spark, Data Lakes, and Airflow; participate in production on-call.
Hadoop, HBase, Spark, Data Lakes, Airflow
1mo
Save
Mark Applied
Hide
Senior Systems Reliability Engineer II
Mountain View, California, United States
HybridFull Time
ThoughtSpot
ThoughtSpot: AI-powered analytics platform for enterprise business intelligence.
Experience troubleshooting Linux systems and cloud platforms, hands-on with monitoring tools, on-call/incident management experience, scripting in Python/Go/Bash/Java, B.S. in CS or equivalent preferred.
Grafana, Prometheus, Datadog, Splunk, VMware, AWS, Azure, GCP, Python, Go, Bash, Java, Spotter, SpotterViz
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
3d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
New York City or Austin or Sunnyvale or Redmond
$140k-$215k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years building distributed systems, 5+ years developing SaaS microservices, expert programming skills, distributed-systems expertise, architectural leadership, and a Computer Science degree or equivalent experience.
Go, Java, Scala, Kotlin, Python, Node.js, Kubernetes, AWS, Cassandra, Kafka, Elasticsearch, OpenSearch, Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), GitHub, Stack Overflow
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – CloudVision as a Service (CVaaS)
Santa Clara, California, United States
$101k-$161k/yr RemoteFull Time
Arista Networks
Arista NetworksNYSE: ANET: Provides cloud networking solutions and high-speed multilayer Ethernet switches.
5+ YOEBS/MS or equivalent experience,5+ years software engineering, experience with distributed databases/SaaS deployments, proficiency in Python/Golang/Bash, Kubernetes and cloud platform experience preferred.
Golang, Python, Ansible, Pulumi, Bash, Kubernetes, GKE, GCP
1mo
Save
Mark Applied
Hide
Principal Engineer, Model Development Platform
Sunnyvale, California, United States
$296k-$335k/yr HybridFull Time
Wayve
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
10+ YOE10+ years building large-scale distributed systems or ML infrastructure, 3+ years at staff/principal level, experience with Spark, Ray, Kubernetes, Airflow, MLflow, web frameworks, reliability engineering, and mentoring engineers.
Spark, Ray, Kubernetes, Airflow, MLflow, React, Flask, FastAPI
2mo
Save
Mark Applied
Hide
Platform Engineer- GraphQL/API
Dublin or San Jose or New York City or Shanghai
HybridFull Time
eBay
eBayNASDAQ: EBAY: Global online marketplace for buying and selling diverse products.
3+ YOE3+ years software engineering experience; strong Java and Spring Boot skills; GraphQL and API design experience; CI/CD (Maven/Jenkins); Kubernetes and cloud-native deployments; focus on reliability, scalability, and developer experience.
Java, Spring Boot, GraphQL, REST, CI/CD, Maven, Jenkins, Kubernetes
1mo
Save
Mark Applied
Hide
Engineering Manager, Reliability Platform
San Francisco or Sunnyvale or New York City
$194k-$285k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ Mgmt5+ years leading engineering teams and 5+ years in infrastructure/platform/backend roles; strong platform mindset, SRE experience (SLOs/error budgets), AWS and cloud fundamentals, influence and hiring experience, experience with incident/incident response processes.
AWS, Kafka, MCP, Covey Scout, Covey, IDE
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) (Hybrid)
San Francisco or San Jose or New York City or Milpitas or Mountain View or Holmdel or Goleta or Redwood City or Fremont or Sunnyvale or Brooklyn or Palo Alto
$168k-$245k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
4+ YOERequires 7+ years with a bachelor's, 4+ with a master's, or 1 with a PhD; 4+ years in SRE or related engineering, 3+ years operating Kubernetes, Helm, CI/CD, cloud platforms, and Python or Go.
Kubernetes, Helm, CI/CD, AWS, GCP, Python, Go, Terraform, MLOps, VPC, DNS
2mo
Save
Mark Applied
Hide
Principal Systems Software Engineer - Observability and Telemetry Platform
Santa Clara or Remote
$248k-$397k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
15+ YOEBS in CS or related field; 15+ years infra automation; 8+ years observability/platform experience; proficiency with Python/Go/Perl/Ruby; Linux, networking, and containers expertise.
Python, Go, Perl, Ruby, Linux, Networking, Containers, Kubernetes, OpenStack, Docker, Grafana, OpenTelemetry, Prometheus