39 platform reliability engineer jobs at 19 companies in Marina, CA

1w
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS/Engineering, strong Linux, networking, databases, Kubernetes, SRE/DevOps toolset knowledge, experience with ClickHouse/Spark/Presto/Doris/Hadoop, coding in Python/Shell/Java/Go, strong problem-solving and communication.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Python, Shell, Java, Go
1w
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Experience with Linux, networking, databases, Kubernetes, ClickHouse/Hadoop/Doris/Spark/Presto, scripting or programming (Python, Shell, Java, Go), incident management, and capacity planning.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Linux, Python, Shell, Java, Go
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
2mo
Save
Mark Applied
Hide
Head of Platform Product Reliability
San Jose, California, United States
OnsiteFull Time
Etched
Etched: Designs specialized AI chips optimized for transformer architectures.
10+ YOE10+ years reliability engineering for complex hardware systems, degree in engineering, hands-on FMEA/Weibull/HALT-HASS, qualification planning, failure analysis, and cross-functional leadership.
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
6d
Save
Mark Applied
Hide
Senior Platform AI Engineer
Santa Clara or California
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE8+ years building and operating production platform/backend infrastructure; 5+ years ML infrastructure; strong Python and a compiled language; experience with job queues, sandboxed execution, and production reliability.
Python, C, C++, Go, Java, Rust, Kubernetes Jobs, Celery, Sidekiq, Temporal, container runtimes
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Apply SRE principles to mentor teams, ensure reliability for large-scale analytics infrastructure across Hadoop, HBase, Spark, Data Lakes, and Airflow; participate in production on-call.
Hadoop, HBase, Spark, Data Lakes, Airflow
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – CloudVision as a Service (CVaaS)
Santa Clara, California, United States
$101k-$161k/yr RemoteFull Time
Arista Networks
Arista NetworksNYSE: ANET: Provides cloud networking solutions and high-speed multilayer Ethernet switches.
5+ YOEBS/MS or equivalent experience,5+ years software engineering, experience with distributed databases/SaaS deployments, proficiency in Python/Golang/Bash, Kubernetes and cloud platform experience preferred.
Golang, Python, Ansible, Pulumi, Bash, Kubernetes, GKE, GCP
1mo
Save
Mark Applied
Hide
Principal Engineer, Model Development Platform
Sunnyvale, California, United States
$296k-$335k/yr HybridFull Time
Wayve
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
10+ YOE10+ years building large-scale distributed systems or ML infrastructure, 3+ years at staff/principal level, experience with Spark, Ray, Kubernetes, Airflow, MLflow, web frameworks, reliability engineering, and mentoring engineers.
Spark, Ray, Kubernetes, Airflow, MLflow, React, Flask, FastAPI
1mo
Save
Mark Applied
Hide
Platform Engineer- GraphQL/API
Dublin or San Jose or New York City or Shanghai
HybridFull Time
eBay
eBayNASDAQ: EBAY: Global online marketplace for buying and selling diverse products.
3+ YOE3+ years software engineering experience; strong Java and Spring Boot skills; GraphQL and API design experience; CI/CD (Maven/Jenkins); Kubernetes and cloud-native deployments; focus on reliability, scalability, and developer experience.
Java, Spring Boot, GraphQL, REST, CI/CD, Maven, Jenkins, Kubernetes
1mo
Save
Mark Applied
Hide
Engineering Manager, Reliability Platform
San Francisco or Sunnyvale or New York City
$194k-$285k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ Mgmt5+ years leading engineering teams and 5+ years in infrastructure/platform/backend roles; strong platform mindset, SRE experience (SLOs/error budgets), AWS and cloud fundamentals, influence and hiring experience, experience with incident/incident response processes.
AWS, Kafka, MCP, Covey Scout, Covey, IDE
1mo
Save
Mark Applied
Hide
Principal Engineer, Model Development Platform
Sunnyvale, California, United States
$296k-$335k/yr HybridFull Time
Wayve
Wayve: Develops AI software for autonomous vehicle navigation.
10+ YOE10+ years building large-scale distributed systems or ML infrastructure, 3+ years at staff/principal level, experience with Spark, Ray, Kubernetes, Airflow, MLflow, reliability/observability, mentoring, and optimization or scheduling systems.
Spark, Ray, Kubernetes, Airflow, MLflow, React, Flask, FastAPI
3d
Save
Mark Applied
Hide
Senior Backend Platform Engineer
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOEB.S. or equivalent experience, 8+ years software engineering, strong Go, Linux, Kubernetes, cloud and networking experience, distributed-systems expertise, production reliability and observability.
Go, Linux, Kubernetes, Temporal, AWS, GCP, Azure
1mo
Save
Mark Applied
Hide
Staff Production Engineer (Cloud Platform & Reliability – Machine Identity Security) - hybrid
Santa Clara, California, United States
OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
Design, build, and operate highly available cloud infrastructure; drive IaC and CI/CD improvements; implement monitoring and incident response; mentor engineers; strong infrastructure and systems mindset.
Infrastructure as Code (IaC), CI/CD
1mo
Save
Mark Applied
Hide
Lead Systems Engineer, Web Platform
Santa Clara or London or Bangalore or Noida
HybridFull Time
Eightfold
Eightfold: A building an AI-native enterprise talent platform used by large global organizations.
Senior hands-on engineer owning web platform architecture and reliability; strong experience with WordPress/AWS and modern stacks (React/Next.js); integrations with Marketo, Salesforce, GA4/GTM; strong communication and judgment; uses AI tooling.
WordPress, AWS, React, Next.js, JavaScript, HTML, CSS, Salesforce, Marketo, GA4, Google Tag Manager, Search Console, CI/CD
3w
Save
Mark Applied
Hide
(USA) Distinguished, Software Engineer-AI/ML Engineer - Agentic Systems & Site Reliability Engineering
Sunnyvale, California, United States
$169k-$338k/yr OnsiteFull Time
Walmart
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
12+ YOESenior AI/ML & SRE engineer with extensive experience designing agentic AI systems, observability, cloud-native platforms, and reliability tooling for mission-critical, large-scale distributed systems.
TensorFlow, PyTorch, Azure, GCP, AWS, Kubernetes, Docker, Jaeger, Zipkin, OpenTelemetry, Prometheus, Grafana, DataDog, ELK stack, Splunk, Fluentd, Terraform, CloudFormation, Pulumi, Istio, Linkerd, MLflow, Kubeflow, Seldon, Kafka, Pulsar
2mo
Save
Mark Applied
Hide
Engineering Manager, Data Platform Orchestration
Los Gatos, California, United States
$436k-$791k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
Experienced engineering leader with platform engineering background; experience running highly reliable, high-scale systems, building teams, partnering cross-functionally, and making technical trade-offs. Knowledge of orchestration and big data technologies desirable.
Maestro, Iceberg, Apache Spark, Trino/Presto, Snowflake, BigQuery, Amazon RedShift, Druid, LanceDB, Pensive, Nightingale
3w
Save
Mark Applied
Hide
Principal AI Systems Engineer — C++ / Applied AI
San Jose or San Francisco or Seattle or New York City or Chicago or California
$190k-$361k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years professional software engineering with deep C++ expertise, production integration of AI/LLMs, cross‑platform systems design, reliability and observability, architecture leadership, and strong communication skills.
C++, LLMs, GPT, Claude, Gemini, JSON-RPC, gRPC, WebSockets, REST, TLS, CI, Windows, macOS, Linux
4d
Save
Mark Applied
Hide
Engineering Manager, Software Platform
San Jose, California, United States
$210k-$234k/yr HybridFull Time
Muon Space
Muon Space: Designs, builds, and operates low Earth orbit satellite constellations.
5+ YOE3+ Mgmt5+ years software engineering, 3+ years managing engineering teams; hands-on Kubernetes, Terraform, AWS; experience with production reliability, SLOs, on-call and incident response.
CI/CD, Kubernetes, Terraform, Bazel, AWS, IAM, Cloudflare, Artifactory, Okta, InfluxDB, Grafana
4d
Save
Mark Applied
Hide
Platform Reliability, Availability, Serviceability (RAS), and Manageability Software Architect. Principal Engineer
Santa Clara or Austin
$212k-$318k/yr OnsiteFull Time
Qualcomm
QualcommNASDAQ: QCOM: Designs and manufactures semiconductors and wireless telecommunications products.
6+ YOEExpertise in ARM/ARM64, RAS and manageability, Linux kernel, DDR, PCIe, I2C/SPI/MDIO, 6+ years software experience (varies by degree), proficiency in C/C++/Java/Python, strong documentation and communication skills.
C, C++, Java, Python, Linux, ARM, ARM64, DDR, PCIe, I2C, SPI, MDIO, UEFI, SDEI, APEI