176 sre engineer jobs at 90 companies in Gilroy, CA

PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
3w
Save
Mark Applied
Hide
Senior SRE Engineer
Santa Clara, California, United States
$148k-$276k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years SRE/platform experience, deep Kubernetes and CI (GitLab/GitHub) expertise, scripting in Python/Go/bash, IaC with Terraform/Helm/Ansible, observability tooling experience, BS/MS in CS or equivalent.
GitLab CI, GitHub Actions, GitLab-runner, Kubernetes, Python, Go, bash, Terraform, Helm, Ansible, Argo CD, Flux, Prometheus, Grafana, Loki, ELK, OpenTelemetry
1mo
Save
Mark Applied
Hide
Senior DevOps/SRE Engineer
San Jose, California, United States
$179k-$306k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experience building and maintaining CI/CD pipelines (GitHub Actions, Jenkins), automation using IaC, Python and Bash, SRE/leadership experience, strong CI/CD and build-system knowledge, BA/BS in CS/CE/EE or equivalent.
GitHub Actions, Jenkins, GitHub, Infrastructure as Code (IaC), Python, Bash, Windows, Linux
3w
Save
Mark Applied
Hide
Senior SRE Engineer
Santa Clara, California, United States
$148k-$276k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years SRE/platform experience, deep Kubernetes administration, GitLab/GitHub CI at scale, scripting in Python/Go/bash, IaC with Terraform/Helm/Ansible, observability tooling, BS/MS in CS or equivalent experience.
Kubernetes, GitLab CI, GitHub Actions, GitLab-runner, Python, Go, bash, Terraform, Helm, Ansible, Argo CD, Flux, Prometheus, Grafana, Loki, ELK, OpenTelemetry, Linux
1d
Save
Mark Applied
Hide
AI-First SRE/DevOps Engineer
San Jose, California, United States
$120k-$160k/yr HybridFull Time
Axiad
Axiad: Identity security platform for passwordless authentication and credential management.
5+ YOE5–8 years SRE/DevOps experience with production Kubernetes, infrastructure-as-code, GitOps, CI/CD, cloud provider experience, observability, Docker, AI-First tooling; Go or Python preferred.
Kubernetes, Docker, GitOps, Go, Python, Claude Code, Cursor, Windsurf
1mo
Save
Mark Applied
Hide
SRE/Dev Ops Engineer (Hybrid, Sunnyvale)
Sunnyvale, California, United States
$120k-$180k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
8+ YOE8+ years in DevOps/SRE or platform engineering; production Kubernetes; CI/CD with GitHub Actions/Jenkins/Tekton; IaC (Terraform/Pulumi); GitOps (ArgoCD/Flux); Observability (Prometheus/Grafana); multi-cloud or multi-region experience; able to work in Sunnyvale office 2+ days.
Kubernetes, GitHub Actions, Jenkins, Tekton, Terraform, Pulumi, Crossplane, ArgoCD, Flux, Prometheus, Grafana, Jaeger, OpenTelemetry, Temporal, Argo Workflows, Istio, Linkerd, Go
2mo
Save
Mark Applied
Hide
Engineering Lead – Platform & SRE
Santa Clara, California, United States
$175k-$215k/yr HybridFull Time
Kerrigan Robotics
Kerrigan Robotics: AI-powered orchestration platform for factory automation systems.
5+ YOE5+ years SRE/DevOps experience; deep Kubernetes knowledge; expertise with Pulumi/Terraform, Helm; proficiency in Go, Linux networking, observability (Prometheus, Grafana), and CI/CD (GitHub Actions).
Pulumi, Kubernetes, AWS, GCP, Azure, GitHub Actions, Prometheus, Grafana, OIDC, CNIs, Terraform, Helm, Go, Linux
1mo
Save
Mark Applied
Hide
Sr. SRE Platform Architect
San Jose or Austin
HybridFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
10+ YOE10+ years production SRE/platform engineering or infra-architecture (including ≥3 years architect-level). Hands-on GPU/AI compute, multi-region observability, Kubernetes and cluster platforms, data-center operations, DDD and plugin framework experience; BS/MS CS.
NVIDIA, DCGM, MIG, vGPU, NVLink, NVSwitch, XID, NCCL, InfiniBand, RoCE, Lustre, NetApp, Pure, DDN, VAST, NVMe-oF, Kubernetes, GPU Operator, Slurm, Volcano, Kueue, Ray, KubeRay, ZTP, BMC, IPMI, Redfish, GitOps
3mo
Save
Mark Applied
Hide
Tech Lead Manager SRE - USDS
San Jose, California, United States
$209k-$438k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
5+ YOE5+ years designing and troubleshooting large-scale distributed systems, prior people leadership with hands-on contributions, Docker and Kubernetes experience, SRE best practices, strong troubleshooting and mentoring skills.
Docker, Kubernetes, Linux
1mo
Save
Mark Applied
Hide
Evaluation Reliability SRE
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Work on Siri and Apple Intelligence to build AI-driven assistant capabilities with strong focus on privacy and cross-platform impact.
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Quota SRE
Sunnyvale, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent,8 years software/systems engineering experience,5 years SRE experience,5 years software design experience,EMR not mentioned; strong troubleshooting and stakeholder management skills.
Google Cloud, Quotaserver, Bouncer, Slicer
1w
Save
Mark Applied
Hide
Senior Engineer, Hybrid Services & Reliability (SRE)
Sunnyvale or Austin
$148k-$222k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
Proven SRE/DevOps experience in hybrid cloud, Linux administration, networking (DHCP/PXE/NTP), IaC/configuration management, automation, and mentoring; growth mindset and independent execution.
Python, Go, Linux, DHCP, PXE, NTP, Chef, Ansible, Terraform, Kubernetes (k8s)
3mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Palo Alto or San Francisco
$170k-$230k/yr HybridFull Time
Mithril
Mithril: Orchestrates global compute capacity specifically for AI workloads.
3+ YOE3+ years in SRE/Production Engineering; Kubernetes; cloud (AWS/GCP/Azure); Python or Go; Linux; RCA; strong communicator.
Kubernetes, Terraform, Pulumi, Python, Go, Linux, Prometheus, Grafana, OpenTelemetry
1mo
Save
Mark Applied
Hide
Staff Cyber Site Reliability Engineer (SRE)
Bethesda or Palo Alto or Dallas or Seattle
$110k-$230k/yr HybridFull Time
GEICO
GEICO: Provides vehicle and property insurance services to consumers.
8+ YOE8+ years in software or site reliability engineering; 5+ years in SRE/DevOps; strong Python; Golang preferred; AWS/Azure/GCP experience; CI/CD and IaC; observability and incident response; security tooling familiarity.
Python, Golang, Grafana, Prometheus, GitHub Actions, Jenkins, Terraform, Ansible
2w
Save
Mark Applied
Hide
Machine Learning Ops Engineer, Global SRE
San Jose, California, United States
$245k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's in CS or equivalent; expertise in Linux, networking, storage; programming in Python, Go, C, C++, or Java; troubleshooting and production operations experience; SRE of ML systems preferred.
Linux, Python, Go, C, C++, Java
1w
Save
Mark Applied
Hide
Job Posting Title AI/ DevOps Engineer
San Jose, California, United States
$174k-$331k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
6+ YOE6–10 years SRE/infrastructure experience; strong datastore, Kubernetes, cloud (AWS/Azure/GCP), observability and incident response skills; interest in AI/ML ops; automation and reliability focus.
Aerospike, FoundationDB, Postgres, CosmosDB, DynamoDB, Kubernetes, AWS, Azure, GCP, Prometheus, Grafana, OpenTelemetry, Copilot, Claude Code, Codex
3mo
Save
Mark Applied
Hide
IT SRE Team Lead
Sunnyvale, California, United States
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
8+ YOE2+ Mgmt8+ years SRE/DevOps/IT engineering experience with 2+ years leadership; hands-on Python or Go; experience with identity (Okta, Entra), endpoint management (Jamf, Intune), Terraform, GitOps, CI/CD; on-call and SLO experience.
Python, Go, Okta, Entra, Jamf, Intune, Terraform, GitOps, CI/CD
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer (SRE) | Dev Ops Engineer
Menlo Park or Durham
$169k-$224k/yr HybridFull Time
GRAIL
GRAILNasdaq: GRAL: Develops blood tests for early-stage multi-cancer detection.
8+ YOE8+ years in SRE/DevOps or platform engineering; strong cloud; IaC; Kubernetes; CI/CD; observability; security basics.
AWS, GCP, Azure, Terraform, CloudFormation, Ansible, Kubernetes, CI/CD (GitLab CI, GitHub Actions, Jenkins), Prometheus, Grafana, Observability/OpenTelemetry, Datadog
3mo
Save
Mark Applied
Hide
Observability Lead - Cloud SRE & Network Reliability
Fremont, California, United States
$114k-$253k/yr HybridFull Time
Lam Research
Lam ResearchNASDAQ: LRCX: Designs and manufactures wafer fabrication equipment for the semiconductor industry.
12+ YOE6+ MgmtSenior SRE leader with 12+ years infrastructure/SRE/DevOps experience and 6+ years leading teams; deep multi-cloud networking, observability, DR/BCP, backup/restore, automation (Ansible/Terraform/Python), Kubernetes, and incident management experience.
Prometheus, Grafana, Datadog, PagerDuty, ThousandEyes, Azure Monitor, CloudWatch, Google Cloud Operations, Splunk, Ansible, Terraform, Python, Kubernetes, AKS, EKS, GKE, ServiceNow
1mo
Save
Mark Applied
Hide
SRE/Devops Engineer- Sunnyvale, CA, the US
Sunnyvale, California, United States
OnsiteFull Time
Kody
Kody: An agentic commerce platform providing integrated in-person payment solutions.
Senior SRE with deep AWS and GitHub experience, strong monitoring/logging and scripting skills, incident leadership, and absolute fluency in Mandarin and English.
AWS, GitHub
1d
Save
Mark Applied
Hide
SRE III (L3 Tech Support+Virtualization+Linux+Network)8-12yrs
Pune or San Jose or Durham or Mexico City or Bangalore or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
7+ YOE7+ years SRE experience with networking, virtualization (VMware ESXi), Linux, cloud/DevOps, strong customer-support skills and degree in engineering/computer science preferred.
VMware ESXi, Linux, DevOps, Cloud, VMware, Citrix, Microsoft