28 infrastructure reliability engineer jobs at 14 companies in Lathrop, CA

1mo
Save
Mark Applied
Hide
Staff Engineer, Reliability Infrastructure Engineer
Milpitas, California, United States
$117k-$193k/yr HybridFull Time
Sandisk
SandiskNasdaq: SNDK: Designs and manufactures flash memory and data storage products.
6+ YOEBachelor's degree in engineering/CS, 6+ years engineering experience in test infrastructure or reliability, hands-on with environmental chambers, failure analysis, vendor management, and global lab alignment.
3w
Save
Mark Applied
Hide
Site Reliability Engineer - Data Infrastructure
San Jose, California, United States
$156k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOE2+ years SRE/DevOps experience, bachelor’s degree or equivalent, scripting (Python/Go/Bash), Linux and networking knowledge, familiarity with containers and observability tools.
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Docker, Prometheus, Grafana, ELK Stack, Linux
2w
Save
Mark Applied
Hide
Contract Site Reliability Engineer — AI Accelerator Infrastructure
Santa Clara, California, United States
$155k-$235k/yr HybridContract
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
5+ YOE5+ years SRE/infrastructure experience; strong Linux, colocation and bare-metal skills; Terraform/Ansible; Kubernetes; Prometheus/Grafana or DataDog; Python/Bash; incident response and RCA experience.
AWS, Azure, GCP, Terraform, Ansible, Kubernetes, Prometheus, Grafana, DataDog, Python, Bash, Slurm, LSF, InfiniBand, RoCE, NVLink, Go
5d
Save
Mark Applied
Hide
Principal, System Reliability Engineer
San Jose, California, United States
$185k-$290k/yr OnsiteFull Time
Ayar Labs
Ayar Labs: Develops optical interconnect technology for high-speed data movement.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
FPGA
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOEBachelor's or equivalent and 5+ years SRE/production engineering experience; proficiency with Go/Python/Bash, Linux, networking, and large-scale distributed systems.
Kubernetes, Redis, MySQL, Message Queue, Kafka, Flink, Go, Python, Bash
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Jose or Seattle or San Francisco
$159k-$302k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOEBachelor's or equivalent, 5+ years SRE/infrastructure/backend experience; Kubernetes, Docker, Terraform, AWS, Postgres/Redis, observability, incident response, CI/CD, bash, Node.js/TypeScript experience; on-call participation.
Kubernetes, Docker, bash, CircleCI, Node.js, TypeScript, Postgres, Redis, AWS Aurora (Postgres-compatible), Terraform, AWS
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, AI Foundations, F1 Query
San Jose, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's in CS or equivalent,8+ years building infrastructure/distributed systems,5+ years programming in C++ or Go,5+ years reliability engineering,EMR not mentioned,experience with distributed systems and stakeholder collaboration.
C++, Go, Java, GoogleSQL, Google Cloud
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer – Compute Platforms
San Ramon or United States
$82k-$229k/yr RemoteFull Time
Five9
Five9NASDAQ: FIVN: Provides cloud-based software for enterprise contact center operations.
6+ YOE6+ years in infrastructure/DevOps with compute systems and Kubernetes/OpenStack expertise; strong Linux, automation, and scripting skills.
Kubernetes, OpenStack, Linux, PXE, Redfish, Ansible, Terraform, Helm, Git, Python, Bash, Harvester, Ubuntu, KVM, ArgoCD, CI/CD
1mo
Save
Mark Applied
Hide
Staff Production Engineer (Cloud Platform & Reliability – Machine Identity Security) - hybrid
Santa Clara, California, United States
OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
Design, build, and operate highly available cloud infrastructure; drive IaC and CI/CD improvements; implement monitoring and incident response; mentor engineers; strong infrastructure and systems mindset.
Infrastructure as Code (IaC), CI/CD
1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald
2w
Save
Mark Applied
Hide
Sr. Director, Technology, Infrastructure & Stabilization
Oakland or United States or Washington or Ohio or District of Columbia or California or Long Beach or Arizona or Colorado or Connecticut or Florida or Georgia or Maryland or Minnesota or Nevada or Oregon or El Dorado Hills or San Diego or Woodland Hills or Alabama or Illinois or Virginia or Wisconsin or Texas or New York
$240k-$359k/yr HybridFull Time
Ascendiun
Ascendiun: Nonprofit parent overseeing health insurance and clinical service organizations.
12+ YOE6+ Mgmt12 years engineering experience, 6 years people management; experience with cloud-native architectures, CI/CD, observability, reliability, and stabilizing legacy systems; strong communication and cross-functional leadership.
4d
Save
Mark Applied
Hide
Senior/Lead SRE Platform Services Engineer Technical Leader
San Jose, California, United States
$64k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ years SRE/platform engineering experience, strong software skills (Python/Go/Ruby), infrastructure-as-code, Kubernetes, AWS, CI/CD, production reliability, security/compliance translation, and technical leadership.
Python, Go, Ruby, Kubernetes, AWS, CI/CD