93 systems reliability engineer jobs at 35 companies in Lathrop, CA

2mo
Save
Mark Applied
Hide
Quality & Reliability Systems Engineer - (E3)
Santa Clara, California, United States
$111k-$152k/yr OnsiteFull Time
Applied Materials
Applied MaterialsNASDAQ: AMAT: Manufacturers of equipment for semiconductor and display production.
Entry-level quality and reliability engineer to develop standards and test plans, perform FMECA and failure analysis, reliability modeling, and document testing. Travel ~20% and relocation eligible.
FMECA, CRAMS, PQP, ERAMS
2w
Save
Mark Applied
Hide
Senior Systems Reliability Engineer
Pune or San Jose or Durham or Mexico City or Bangalore or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
7+ YOE7+ years SRE experience with networking, virtualization (VMware ESXi), Linux, cloud and strong customer-facing troubleshooting and communication skills.
VMware ESXi, VMware, Linux, DevOps, Cloud, Citrix, Microsoft
4d
Save
Mark Applied
Hide
Sr. Reliability Engineer - SEM Systems
Milpitas, California, United States
$167k-$284k/yr OnsiteFull Time
KLA
KLANASDAQ: KLAC: Provides process control and yield management for semiconductor manufacturing.
Identify failure mechanisms, perform root-cause analysis, develop reliability tests and qualification plans for SEM systems; collaborate with suppliers and implement corrective actions.
1mo
Save
Mark Applied
Hide
Senior/Staff Systems Reliability Engineer
Newark, California, United States
HybridFull Time
Queue
Queue: Developing autonomous robotic pharmacies for automated prescription fulfillment.
3+ YOE3+ years Terraform/IaC experience, deep AWS knowledge (VPC, ECS, RDS, IAM, KMS, SQS, CloudWatch), CI/CD (GitHub Actions), Docker, Linux administration; experience with Rust/TypeScript services and security/observability practices.
Terraform, AWS, VPC, ECS, RDS, IAM, KMS, SQS, CloudWatch, GitHub Actions, Docker, Rust, TypeScript, cargo, clippy, mTLS, X.509
1mo
Save
Mark Applied
Hide
Systems Quality and Reliability Engineer - LPU
Santa Clara, California, United States
$136k-$265k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEBS/MS in EE, Physics or related; 5+ yrs systems test/validation; hands-on quality and reliability experience; lab hardware skills; FA techniques/tools knowledge; Python/Perl/C/C++ on UNIX/Linux; PCB/system test and CM coordination.
Oscilloscopes, Logic analyzers, Power analyzers, FIB, SEM, TDR, VNA, CSAM, OBIRCH, DLS/LADA, LVP, LVI, SerDes, PCIe, DDR, Python, Perl, C++, UNIX/Linux
1w
Save
Mark Applied
Hide
Reliability Engineer, Mechanical Systems, NA
Santa Clara or Quincy
$125k-$140k/yr HybridFull Time
Vantage Data Centers
Vantage Data Centers: Provides hyperscale data center campuses for cloud and AI providers.
2+ YOEBachelor's preferred,2+ years critical facility experience,knowledge of mechanical cooling systems,commissioning,maintenance program design,and ability to perform root-cause analysis;travel up to 25%.
1mo
Save
Mark Applied
Hide
Systems Quality and Reliability Engineer - LPU
Santa Clara, California, United States
$136k-$265k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEBS/MS in EE/Physics or equivalent, 5+ years systems test/validation experience, hands-on FA/RMA debugging, lab equipment use, HTOL/Burn in, FA techniques, fault isolation, high-speed interfaces, Python/Perl/C++ on UNIX/Linux.
oscilloscopes, logic analyzers, power analyzers, HTOL, Burn in, FIB, SEM, TDR, VNA, CSAM, OBIRCH, DLS/LADA, LVP, LVI, SerDes, PCIe, DDR, Python, PERL, C++, UNIX, Linux
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
2w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, AI Foundations, F1 Query
San Jose, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's in CS or equivalent,8+ years building infrastructure/distributed systems,5+ years programming in C++ or Go,5+ years reliability engineering,EMR not mentioned,experience with distributed systems and stakeholder collaboration.
C++, Go, Java, GoogleSQL, Google Cloud
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
2w
Save
Mark Applied
Hide
Site Reliability Engineer - Video Infrastructure
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
SRE for global multimedia transport, storage and processing; strong SRE, networking, OS, database, container and distributed systems troubleshooting skills; bachelor's in CS or equivalent.
C, C++, Java, Python, Go, Linux, MySQL, MongoDB, Redis, ELK, AWS, Google Cloud, Azures
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Product - USDS
San Jose, California, United States
$187k-$438k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
5+ YOEBachelor's in CS or related plus 5+ years deploying/administering large-scale distributed systems; strong Unix/Linux and networking knowledge; experience with programming, debugging, and systems like Nginx, Kubernetes, Docker, Hadoop, Spark, Flink, Kafka.
Unix/Linux, TCP/IP, C, C++, Java, Python, Go, Ruby, Rust, JavaScript, Nginx, Kubernetes, Docker, OpenStack, Hadoop, Spark, Flink, Kafka
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
United States or Houston or Santa Clara or Korea or Germany
RemoteFull Time
Qcells
Qcells: Provider of solar modules, energy storage, and EPC services.
8+ YOE8+ years in SRE/DevOps or software engineering; experience with cloud platforms, distributed systems, observability, CI/CD, and incident response; willingness to travel up to 10%.
AWS, Azure, GCP, Docker, Kubernetes, CI/CD, AI/ML
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Global E-Commerce
San Jose, California, United States
$213k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
5+ YOEBachelor's or equivalent,5+ years SRE/infra experience,proficiency in Go/Python/Java,strong Linux,networking and distributed systems knowledge,cloud-native production experience.
Go, Python, Java, Linux
1mo
Save
Mark Applied
Hide
Principal AI Systems Engineer — C++ / Applied AI
San Jose or San Francisco or Seattle or New York City or Chicago or California
$190k-$361k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years professional software engineering with deep C++ expertise, production integration of AI/LLMs, cross‑platform systems design, reliability and observability, architecture leadership, and strong communication skills.
C++, LLMs, GPT, Claude, Gemini, JSON-RPC, gRPC, WebSockets, REST, TLS, CI, Windows, macOS, Linux
3mo
Save
Mark Applied
Hide
Senior Electrical Systems Engineer
Milpitas, California, United States
$130k-$170k/yr OnsiteFull Time
PDF Solutions
PDF SolutionsNASDAQ: PDFS: Data analytics platform for semiconductor yield improvement.
8+ YOEBachelor/Master/PhD in electrical engineering; 8–10+ years experience; expertise in electrical subsystems of semiconductor tools; reliability analysis; supply chain; strong problem-solving.
LT Spice, Altium, ANSYS Sherlock
2mo
Save
Mark Applied
Hide
Head of Platform Product Reliability
San Jose, California, United States
OnsiteFull Time
Etched
Etched: Designs specialized AI chips optimized for transformer architectures.
10+ YOE10+ years reliability engineering for complex hardware systems, degree in engineering, hands-on FMEA/Weibull/HALT-HASS, qualification planning, failure analysis, and cross-functional leadership.
3d
Save
Mark Applied
Hide
Senior HPC Systems Validation Engineer
San Jose, California, United States
$255k-$340k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years in hardware integration validation, fleet reliability, component qualification or performance validation for HPC/data center products; deep system integration testing and benchmarking experience; hands-on NPI and firmware/software compatibility work.
NVIDIA, AMD, Intel, BMC, BIOS, NICs, PLM, BOM, GPU, NPI