76 site reliability engineer jobs at 29 companies in Lathrop, CA

1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
4w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Santa Clara, California, United States
$163k-$214k/yr OnsiteFull Time
IonQ
IonQNYSE: IONQ: Develops and sells trapped-ion quantum computers and cloud services.
7+ YOE7+ years production engineering experience; hands-on AWS/GCP reliability, observability and SLO ownership, incident command, resilience testing, and multi-team technical leadership.
AWS, GCP, Amazon Bedrock Agent Core
4w
Save
Mark Applied
Hide
Site Reliability Engineer
Research Triangle Park or San Jose or Milpitas or Richardson or Santa Clara
$127k-$182k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years SRE/Cloud Ops experience, Docker and Kubernetes proficiency, scripting in Python/Go/Bash, monitoring and incident response experience, Linux and networking knowledge, CI/CD and IaC familiarity, bachelor’s degree or equivalent.
Docker, Kubernetes, Python, Go, Bash, Git
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
22h
Save
Mark Applied
Hide
Site Reliability Engineer (Multiple Positions)
San Jose, California, United States
$213k-$388k/yr OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOERequires a master's degree and 2 years of related experience, or a bachelor's degree and 5 years of progressive experience. Requires cloud systems, Linux, Docker, Kubernetes, software lifecycle, observability, and reliability engineering expertise.
Linux, Docker, Kubernetes
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEDegree in CS or related field (or equivalent experience), 8+ years SRE/DevOps/Production Engineering, SRE principles, infrastructure automation, production reliability, Python/Go/Perl/Ruby, Prometheus and Grafana, strong communication.
Python, Go, Perl, Ruby, Prometheus, Grafana
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOE8+ years supporting live-site production environments; BS/MS or equivalent; strong Kubernetes, AWS, Python; Akamai/CDN and SRE on-call experience required.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS/Engineering, strong Linux, networking, databases, Kubernetes, SRE/DevOps toolset knowledge, experience with ClickHouse/Spark/Presto/Doris/Hadoop, coding in Python/Shell/Java/Go, strong problem-solving and communication.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Python, Shell, Java, Go
2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
2d
Save
Mark Applied
Hide
Senior Site Reliability Engineer Platform Private Cloud Engineer
San Jose, California, United States
$94k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
7+ YOERequires 7+ years designing and operating enterprise or cloud environments, private cloud and Kubernetes expertise, scripting, IaC tools, Unix/Linux knowledge, and a CS or engineering degree.
VMware, AWS, GCP, Kubernetes, Helm, ArgoCD, Python, Bash, Ruby, Scala, Ansible, Terraform, Unix, Linux
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
New York or San Ramon or Reno
$153k-$210k/yr HybridFull Time
Ridgeline
Ridgeline: Cloud-native platform for investment management operations.
3+ YOE3–6 years SRE/DevOps experience, 2+ years on AWS, proficiency with Terraform, observability, CI/CD, Python/Go/Bash, incident response, and strong communication and troubleshooting skills.
Claude Code, Cursor, Terraform, AWS, EC2, ECS, EKS, RDS, S3, IAM, CloudWatch, GitHub Actions, CircleCI, Buildkite, Python, Go, Bash, Kubernetes, Helm, Kotlin, Node.js, TypeScript
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
1w
Save
Mark Applied
Hide
Site Reliability Engineer, AI Infrastructure
San Jose, California, United States
$123k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
1+ YOEBachelor's degree or equivalent experience, 1+ year in SRE, DevOps, or systems engineering, Linux and networking knowledge, distributed systems experience, programming, scripting, CI/CD, and automation skills.
Linux, Go, Python, C, C++, Java, Bash, Kubernetes, AWS, GCP, Azure, Terraform, Prometheus, Grafana, Distributed Tracing, LLMs, Agentic AI
1d
Save
Mark Applied
Hide
Site Reliability Engineer Spring Co-op 2027
Lowell or Durham or San Jose or Austin
$76k-$166k/yr HybridMultiple Commitments Available
IBM
IBMNew York Stock Exchange: IBM: Global technology providing enterprise software, cloud, and consulting.
Actively enrolled in a bachelor's program, available for a 16-week full-time co-op, and knowledgeable in Linux, monitoring, troubleshooting, automation, scripting, cloud platforms, and production support.
Linux, Python, Go, Bash, IBM Cloud, AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, OpenShift, Ansible, Terraform, Jenkins, IBM Continuous Delivery, ArgoCD, Instana, New Relic, Grafana, Prometheus, PostgreSQL, CouchDB, Redis, Kafka, Spark, SQL, NoSQL, CI/CD
1mo
Save
Mark Applied
Hide
Contract Site Reliability Engineer — AI Accelerator Infrastructure
Santa Clara, California, United States
$155k-$235k/yr HybridContract
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
5+ YOE5+ years SRE/infrastructure experience; strong Linux, colocation and bare-metal skills; Terraform/Ansible; Kubernetes; Prometheus/Grafana or DataDog; Python/Bash; incident response and RCA experience.
AWS, Azure, GCP, Terraform, Ansible, Kubernetes, Prometheus, Grafana, DataDog, Python, Bash, Slurm, LSF, InfiniBand, RoCE, NVLink, Go
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
San Jose, California, United States
$119k-$170k/yr HybridFull Time
Zscaler
ZscalerNASDAQ: ZS: Provides cloud-native cybersecurity solutions through zero trust architecture.
5+ YOE5+ years Linux/UNIX admin, Kubernetes/Docker, automation (Ansible), network/security fundamentals, and strong security practices.
Docker, Kubernetes, Ansible, Python, Golang, BASH, Openstack, CEPH, HashiCorp Vault, nftables
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – CloudVision as a Service (CVaaS)
Santa Clara, California, United States
$101k-$161k/yr RemoteFull Time
Arista Networks
Arista NetworksNYSE: ANET: Provides cloud networking solutions and high-speed multilayer Ethernet switches.
5+ YOEBS/MS or equivalent experience,5+ years software engineering, experience with distributed databases/SaaS deployments, proficiency in Python/Golang/Bash, Kubernetes and cloud platform experience preferred.
Golang, Python, Ansible, Pulumi, Bash, Kubernetes, GKE, GCP