181 sre engineer jobs at 95 companies in Tracy, CA

PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
2w
Save
Mark Applied
Hide
Senior SRE Engineer
Santa Clara, California, United States
$148k-$276k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years SRE/platform experience, deep Kubernetes and CI (GitLab/GitHub) expertise, scripting in Python/Go/bash, IaC with Terraform/Helm/Ansible, observability tooling experience, BS/MS in CS or equivalent.
GitLab CI, GitHub Actions, GitLab-runner, Kubernetes, Python, Go, bash, Terraform, Helm, Ansible, Argo CD, Flux, Prometheus, Grafana, Loki, ELK, OpenTelemetry
1mo
Save
Mark Applied
Hide
Senior DevOps/SRE Engineer
San Jose, California, United States
$179k-$306k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experience building and maintaining CI/CD pipelines (GitHub Actions, Jenkins), automation using IaC, Python and Bash, SRE/leadership experience, strong CI/CD and build-system knowledge, BA/BS in CS/CE/EE or equivalent.
GitHub Actions, Jenkins, GitHub, Infrastructure as Code (IaC), Python, Bash, Windows, Linux
2w
Save
Mark Applied
Hide
Senior SRE Engineer
Santa Clara, California, United States
$148k-$276k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years SRE/platform experience, deep Kubernetes administration, GitLab/GitHub CI at scale, scripting in Python/Go/bash, IaC with Terraform/Helm/Ansible, observability tooling, BS/MS in CS or equivalent experience.
Kubernetes, GitLab CI, GitHub Actions, GitLab-runner, Python, Go, bash, Terraform, Helm, Ansible, Argo CD, Flux, Prometheus, Grafana, Loki, ELK, OpenTelemetry, Linux
1mo
Save
Mark Applied
Hide
SRE/Dev Ops Engineer (Hybrid, Sunnyvale)
Sunnyvale, California, United States
$120k-$180k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
8+ YOE8+ years in DevOps/SRE or platform engineering; production Kubernetes; CI/CD with GitHub Actions/Jenkins/Tekton; IaC (Terraform/Pulumi); GitOps (ArgoCD/Flux); Observability (Prometheus/Grafana); multi-cloud or multi-region experience; able to work in Sunnyvale office 2+ days.
Kubernetes, GitHub Actions, Jenkins, Tekton, Terraform, Pulumi, Crossplane, ArgoCD, Flux, Prometheus, Grafana, Jaeger, OpenTelemetry, Temporal, Argo Workflows, Istio, Linkerd, Go
5d
Save
Mark Applied
Hide
SRE Engineer (Full Time; Multiple Openings)
Belmont, California, United States
HybridFull Time
RingCentral
RingCentralNYSE: RNG: Sells cloud-based business phone and video conferencing software.
2+ YOEMaintain 24x7 production availability, implement automation/orchestration, partner with development, perform root cause analysis; required experience with cloud, containers, scripting, and monitoring.
Python, Bash, Go, Terraform, Ansible, AWS, GCP, Kubernetes, GitLab, DNS, Docker, CI/CD, TCP/IP, Linux
2mo
Save
Mark Applied
Hide
Engineering Lead – Platform & SRE
Santa Clara, California, United States
$175k-$215k/yr HybridFull Time
Kerrigan Robotics
Kerrigan Robotics: AI-powered orchestration platform for factory automation systems.
5+ YOE5+ years SRE/DevOps experience; deep Kubernetes knowledge; expertise with Pulumi/Terraform, Helm; proficiency in Go, Linux networking, observability (Prometheus, Grafana), and CI/CD (GitHub Actions).
Pulumi, Kubernetes, AWS, GCP, Azure, GitHub Actions, Prometheus, Grafana, OIDC, CNIs, Terraform, Helm, Go, Linux
3w
Save
Mark Applied
Hide
Sr. SRE Platform Architect
San Jose or Austin
HybridFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
10+ YOE10+ years production SRE/platform engineering or infra-architecture (including ≥3 years architect-level). Hands-on GPU/AI compute, multi-region observability, Kubernetes and cluster platforms, data-center operations, DDD and plugin framework experience; BS/MS CS.
NVIDIA, DCGM, MIG, vGPU, NVLink, NVSwitch, XID, NCCL, InfiniBand, RoCE, Lustre, NetApp, Pure, DDN, VAST, NVMe-oF, Kubernetes, GPU Operator, Slurm, Volcano, Kueue, Ray, KubeRay, ZTP, BMC, IPMI, Redfish, GitOps
1mo
Save
Mark Applied
Hide
Evaluation Reliability SRE
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Work on Siri and Apple Intelligence to build AI-driven assistant capabilities with strong focus on privacy and cross-platform impact.
4w
Save
Mark Applied
Hide
DevOps Engineer / Site Reliability Engineer (SRE)
Bangalore or India or Milpitas or Seattle or Princeton or Cape Town or London or Zurich or Singapore or Mexico City
OnsiteFull Time
Zensar
ZensarNational Stock Exchange of India: ZENSARTECH: Global technology firm providing digital transformation and infrastructure services.
4+ YOEHands-on GCP, Kubernetes (GKE), Docker, Terraform, Jenkins, Git, Linux administration, CI/CD, monitoring (Prometheus/Grafana), SRE practices; 4+ years experience; Bachelor's or equivalent experience.
Google Cloud Platform (GCP), Compute Engine, GKE, Cloud Storage, IAM, VPC, Terraform, Deployment Manager, Docker, Kubernetes (K8s), Jenkins (Pipeline as Code), ArgoCD, Spinnaker, Git (branching, merging strategies, pull requests), GitHub, GitLab, Bitbucket, Ubuntu, RHEL, CentOS, Bash, Python, Go, Prometheus, Grafana, Google Cloud Monitoring / Logging, Stackdriver (Cloud Monitoring), Elasticsearch, Logstash, Kibana, Istio, Linkerd, Vault, Jboss/Wildfly
4d
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Quota SRE
Sunnyvale, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent,8 years software/systems engineering experience,5 years SRE experience,5 years software design experience,EMR not mentioned; strong troubleshooting and stakeholder management skills.
Google Cloud, Quotaserver, Bouncer, Slicer
4d
Save
Mark Applied
Hide
Senior Engineer, Hybrid Services & Reliability (SRE)
Sunnyvale or Austin
$148k-$222k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
Proven SRE/DevOps experience in hybrid cloud, Linux administration, networking (DHCP/PXE/NTP), IaC/configuration management, automation, and mentoring; growth mindset and independent execution.
Python, Go, Linux, DHCP, PXE, NTP, Chef, Ansible, Terraform, Kubernetes (k8s)
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Palo Alto or San Francisco
$170k-$230k/yr HybridFull Time
Mithril
Mithril: Orchestrates global compute capacity specifically for AI workloads.
3+ YOE3+ years in SRE/Production Engineering; Kubernetes; cloud (AWS/GCP/Azure); Python or Go; Linux; RCA; strong communicator.
Kubernetes, Terraform, Pulumi, Python, Go, Linux, Prometheus, Grafana, OpenTelemetry
1mo
Save
Mark Applied
Hide
Staff Cyber Site Reliability Engineer (SRE)
Bethesda or Palo Alto or Dallas or Seattle
$110k-$230k/yr HybridFull Time
GEICO
GEICO: Provides vehicle and property insurance services to consumers.
8+ YOE8+ years in software or site reliability engineering; 5+ years in SRE/DevOps; strong Python; Golang preferred; AWS/Azure/GCP experience; CI/CD and IaC; observability and incident response; security tooling familiarity.
Python, Golang, Grafana, Prometheus, GitHub Actions, Jenkins, Terraform, Ansible
1w
Save
Mark Applied
Hide
Machine Learning Ops Engineer, Global SRE
San Jose, California, United States
$245k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS or equivalent; expertise in Linux, networking, storage; programming in Python, Go, C, C++, or Java; troubleshooting and production operations experience; SRE of ML systems preferred.
Linux, Python, Go, C, C++, Java
4d
Save
Mark Applied
Hide
Job Posting Title AI/ DevOps Engineer
San Jose, California, United States
$174k-$331k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
6+ YOE6–10 years SRE/infrastructure experience; strong datastore, Kubernetes, cloud (AWS/Azure/GCP), observability and incident response skills; interest in AI/ML ops; automation and reliability focus.
Aerospike, FoundationDB, Postgres, CosmosDB, DynamoDB, Kubernetes, AWS, Azure, GCP, Prometheus, Grafana, OpenTelemetry, Copilot, Claude Code, Codex
3mo
Save
Mark Applied
Hide
IT SRE Team Lead
Sunnyvale, California, United States
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
8+ YOE2+ Mgmt8+ years SRE/DevOps/IT engineering experience with 2+ years leadership; hands-on Python or Go; experience with identity (Okta, Entra), endpoint management (Jamf, Intune), Terraform, GitOps, CI/CD; on-call and SLO experience.
Python, Go, Okta, Entra, Jamf, Intune, Terraform, GitOps, CI/CD
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer (SRE) | Dev Ops Engineer
Menlo Park or Durham
$169k-$224k/yr HybridFull Time
GRAIL
GRAILNasdaq: GRAL: Develops blood tests for early-stage multi-cancer detection.
8+ YOE8+ years in SRE/DevOps or platform engineering; strong cloud; IaC; Kubernetes; CI/CD; observability; security basics.
AWS, GCP, Azure, Terraform, CloudFormation, Ansible, Kubernetes, CI/CD (GitLab CI, GitHub Actions, Jenkins), Prometheus, Grafana, Observability/OpenTelemetry, Datadog
3mo
Save
Mark Applied
Hide
Observability Lead - Cloud SRE & Network Reliability
Fremont, California, United States
$114k-$253k/yr HybridFull Time
Lam Research
Lam ResearchNASDAQ: LRCX: Designs and manufactures wafer fabrication equipment for the semiconductor industry.
12+ YOE6+ MgmtSenior SRE leader with 12+ years infrastructure/SRE/DevOps experience and 6+ years leading teams; deep multi-cloud networking, observability, DR/BCP, backup/restore, automation (Ansible/Terraform/Python), Kubernetes, and incident management experience.
Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, Google Cloud Operations, Splunk, Ansible, Terraform, Python, Kubernetes, AKS, EKS, GKE, ServiceNow
1mo
Save
Mark Applied
Hide
SRE/Devops Engineer- Sunnyvale, CA, the US
Sunnyvale, California, United States
OnsiteFull Time
Kody
Kody: An agentic commerce platform providing integrated in-person payment solutions.
Senior SRE with deep AWS and GitHub experience, strong monitoring/logging and scripting skills, incident leadership, and absolute fluency in Mandarin and English.
AWS, GitHub
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – CloudVision as a Service (CVaaS)
Santa Clara, California, United States
$101k-$161k/yr RemoteFull Time
Arista Networks
Arista NetworksNYSE: ANET: Provides cloud networking solutions and high-speed multilayer Ethernet switches.
5+ YOEBS/MS or equivalent experience,5+ years software engineering, experience with distributed databases/SaaS deployments, proficiency in Python/Golang/Bash, Kubernetes and cloud platform experience preferred.
Golang, Python, Ansible, Pulumi, Bash, Kubernetes, GKE, GCP