38 infrastructure reliability engineer jobs at 16 companies in Marina, CA

3w
Save
Mark Applied
Hide
Site Reliability Engineer - Data Infrastructure
San Jose, California, United States
$156k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOE2+ years SRE/DevOps experience, bachelor’s degree or equivalent, scripting (Python/Go/Bash), Linux and networking knowledge, familiarity with containers and observability tools.
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Docker, Prometheus, Grafana, ELK Stack, Linux
2w
Save
Mark Applied
Hide
Contract Site Reliability Engineer — AI Accelerator Infrastructure
Santa Clara, California, United States
$155k-$235k/yr HybridContract
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
5+ YOE5+ years SRE/infrastructure experience; strong Linux, colocation and bare-metal skills; Terraform/Ansible; Kubernetes; Prometheus/Grafana or DataDog; Python/Bash; incident response and RCA experience.
AWS, Azure, GCP, Terraform, Ansible, Kubernetes, Prometheus, Grafana, DataDog, Python, Bash, Slurm, LSF, InfiniBand, RoCE, NVLink, Go
4d
Save
Mark Applied
Hide
Principal, System Reliability Engineer
San Jose, California, United States
$185k-$290k/yr OnsiteFull Time
Ayar Labs
Ayar Labs: Develops optical interconnect technology for high-speed data movement.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
FPGA
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOEBachelor's or equivalent and 5+ years SRE/production engineering experience; proficiency with Go/Python/Bash, Linux, networking, and large-scale distributed systems.
Kubernetes, Redis, MySQL, Message Queue, Kafka, Flink, Go, Python, Bash
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Jose or Seattle or San Francisco
$159k-$302k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOEBachelor's or equivalent, 5+ years SRE/infrastructure/backend experience; Kubernetes, Docker, Terraform, AWS, Postgres/Redis, observability, incident response, CI/CD, bash, Node.js/TypeScript experience; on-call participation.
Kubernetes, Docker, bash, CircleCI, Node.js, TypeScript, Postgres, Redis, AWS Aurora (Postgres-compatible), Terraform, AWS
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, AI Foundations, F1 Query
San Jose, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's in CS or equivalent,8+ years building infrastructure/distributed systems,5+ years programming in C++ or Go,5+ years reliability engineering,EMR not mentioned,experience with distributed systems and stakeholder collaboration.
C++, Go, Java, GoogleSQL, Google Cloud
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
1mo
Save
Mark Applied
Hide
Senior Core Infrastructure Engineer
Seattle or Santa Clara
$79k-$210k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
3+ YOEDesign and implement scalable, reliable distributed systems; build fault-tolerant components; develop automation/IaC; apply security controls; participate in incident response and runbooks.
DevOps, Java, Infrastructure as Code (IaC)
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Apply SRE principles to mentor teams, ensure reliability for large-scale analytics infrastructure across Hadoop, HBase, Spark, Data Lakes, and Airflow; participate in production on-call.
Hadoop, HBase, Spark, Data Lakes, Airflow
1mo
Save
Mark Applied
Hide
Software Engineer, Reliability Platforms
San Francisco or Sunnyvale or New York City
$160k-$235k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ years in infrastructure/platform/backend engineering; fluent in Go or similar; AWS, containerization, and IaC experience (Terraform or Pulumi); SRE concepts (SLOs, error budgets); platform engineering mindset and familiarity with AI tools.
Go, AWS, Terraform, Pulumi
1mo
Save
Mark Applied
Hide
Staff Production Engineer (Cloud Platform & Reliability – Machine Identity Security) - hybrid
Santa Clara, California, United States
OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
Design, build, and operate highly available cloud infrastructure; drive IaC and CI/CD improvements; implement monitoring and incident response; mentor engineers; strong infrastructure and systems mindset.
Infrastructure as Code (IaC), CI/CD
1mo
Save
Mark Applied
Hide
Distinguished Technologist Mechanical Engineer (Network Infrastructure)
Sunnyvale, California, United States
$163k-$348k/yr OnsiteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Provides edge-to-cloud IT infrastructure and platform services.
15+ YOEBS in Mechanical Engineering, 15+ years electromechanical product development experience; expertise in chassis, packaging, manufacturability, reliability, SolidWorks/PLM, FEA, GD&T, thermal and cooling technologies; strong leadership and troubleshooting skills.
SolidWorks, EPDM, FEA, ECAD, IDF
1mo
Save
Mark Applied
Hide
Distinguished Technologist Mechanical Engineer (Network Infrastructure)
Sunnyvale, California, United States
$163k-$348k/yr OnsiteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Providing global edge-to-cloud infrastructure and IT solutions for businesses.
15+ YOEBS in Mechanical Engineering, 15+ years electromechanical product development experience; expertise in chassis, packaging, manufacturability, thermal and reliability; SolidWorks and PLM experience; tolerance/GD&T, FEA, ECAD-MCAD, and supplier engagement experience.
SolidWorks, EPDM, PLM, FEA, ECAD-MCAD, IDF, GD&T
1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald
1mo
Save
Mark Applied
Hide
Principal Engineer, Model Development Platform
Sunnyvale, California, United States
$296k-$335k/yr HybridFull Time
Wayve
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
10+ YOE10+ years building large-scale distributed systems or ML infrastructure, 3+ years at staff/principal level, experience with Spark, Ray, Kubernetes, Airflow, MLflow, web frameworks, reliability engineering, and mentoring engineers.
Spark, Ray, Kubernetes, Airflow, MLflow, React, Flask, FastAPI
3d
Save
Mark Applied
Hide
Senior/Lead SRE Platform Services Engineer Technical Leader
San Jose, California, United States
$64k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ years SRE/platform engineering experience, strong software skills (Python/Go/Ruby), infrastructure-as-code, Kubernetes, AWS, CI/CD, production reliability, security/compliance translation, and technical leadership.
Python, Go, Ruby, Kubernetes, AWS, CI/CD
1mo
Save
Mark Applied
Hide
Principal Engineer, Model Development Platform
Sunnyvale, California, United States
$296k-$335k/yr HybridFull Time
Wayve
Wayve: Develops AI software for autonomous vehicle navigation.
10+ YOE10+ years building large-scale distributed systems or ML infrastructure, 3+ years at staff/principal level, experience with Spark, Ray, Kubernetes, Airflow, MLflow, reliability/observability, mentoring, and optimization or scheduling systems.
Spark, Ray, Kubernetes, Airflow, MLflow, React, Flask, FastAPI
2mo
Save
Mark Applied
Hide
Senior Staff Engineer, Software (Sunnyvale, California, United States, 94089)
Sunnyvale, California, United States
$118k-$170k/yr HybridMultiple Commitments Available
Vistance Networks
Vistance NetworksNASDAQ: VISN: Vistance Networks provides infrastructure solutions for communications and data networks.
5+ YOE5+ years in Site Reliability Engineering, DevOps, Cloud Infrastructure, or Production Engineering; Python; Linux; GCP; Kubernetes; observability tools.
Python, Linux, Google Cloud Platform (GCP), Kubernetes, Cloud-native infrastructure, Prometheus, Grafana, OpenTelemetry, ELK, ClickHouse