122 site reliability engineer jobs at 57 companies in Tracy, CA

PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Santa Clara, California, United States
$163k-$214k/yr OnsiteFull Time
IonQ
IonQNYSE: IONQ: Develops and sells trapped-ion quantum computers and cloud services.
7+ YOE7+ years production engineering experience; hands-on AWS/GCP reliability, observability and SLO ownership, incident command, resilience testing, and multi-team technical leadership.
AWS, GCP, Amazon Bedrock Agent Core
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automates data movement into cloud data warehouses.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Research Triangle Park or San Jose or Milpitas or Richardson or Santa Clara
$127k-$182k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years SRE/Cloud Ops experience, Docker and Kubernetes proficiency, scripting in Python/Go/Bash, monitoring and incident response experience, Linux and networking knowledge, CI/CD and IaC familiarity, bachelor’s degree or equivalent.
Docker, Kubernetes, Python, Go, Bash, Git
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Mountain View, California, United States
$189k-$232k/yr HybridFull Time
EarnIn
EarnIn: Provides immediate access to earned wages through a mobile app.
3+ YOE3+ years SRE or related experience; hands-on Python/Go coding; experience with observability, incident response, SLOs/SLIs, and distributed systems; strong communication and documentation skills.
Python, Go, Datadog, CloudWatch, logs, metrics, traces, APM, GitHub Copilot, Cursor, ChatGPT, Claude
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Palo Alto or Newport Beach
$165k-$190k/yr OnsiteFull Time
Obsidian Security
Obsidian Security: Provides cybersecurity and threat detection for enterprise SaaS applications.
3+ YOE3+ years DevOps/SRE experience on GCP and/or AWS, Bachelor's in CS or related, proficiency with Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana; programming in Golang or Python; strong communication and critical thinking.
Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana, Golang, Python, Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong, AWS, GCP
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto or Pittsburgh
$179k-$269k/yr OnsiteFull Time
Latitude AI
Latitude AI: Developing automated driving technology for next-generation Ford vehicles.
4+ YOEBachelor's degree in engineering/computer science (or higher) with 4+ years experience (or equivalent), strong Linux, networking, Go/Python development, cloud (AWS/GCP), Kubernetes, IaC, monitoring and SLO experience.
Go, Python, AWS, GCP, Terraform, CloudFormation, Kubernetes, Prometheus, Elasticsearch, Loki, Jaeger, Tempo, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEDegree in CS or related field (or equivalent experience), 8+ years SRE/DevOps/Production Engineering, SRE principles, infrastructure automation, production reliability, Python/Go/Perl/Ruby, Prometheus and Grafana, strong communication.
Python, Go, Perl, Ruby, Prometheus, Grafana
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Jose or Seattle or San Francisco
$159k-$302k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOEBachelor's or equivalent, 5+ years SRE/infrastructure/backend experience; Kubernetes, Docker, Terraform, AWS, Postgres/Redis, observability, incident response, CI/CD, bash, Node.js/TypeScript experience; on-call participation.
Kubernetes, Docker, bash, CircleCI, Node.js, TypeScript, Postgres, Redis, AWS Aurora (Postgres-compatible), Terraform, AWS
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Palo Alto or Palo Alto or Washington
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
Kubernetes, Linux, Bash, Python, Bazel, Makefiles, Terraform, Ansible, TCP/IP
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1w
Save
Mark Applied
Hide
Site Reliability Engineer II
Sunnyvale, California, United States
$141k-$162k/yr OnsiteFull Time
Illumio
Illumio: Provides zero-trust segmentation software to contain cyberattacks.
2+ YOE2+ years SRE/DevOps experience with hands-on Azure experience; experience with AWS/GCP, IaC, CI/CD, scripting (PowerShell, Python, Go), and security best practices.
Azure, AWS, GCP, Terraform, ARM templates, CloudFormation, Azure DevOps, AWS CodePipeline, Jenkins, PowerShell, Python, Go
3w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Berkeley Heights or Alpharetta or Sunnyvale
$128k-$216k/yr OnsiteFull Time
Fiserv
FiservNew York Stock Exchange: FI: Provides financial technology and payment processing services to institutions.
5+ YOE5+ years production experience with AWS, Kubernetes, and Linux; strong Terraform, CI/CD (GitHub Actions), Docker, GitHub, RDBMS/Document storage, and scripting (Python/Bash/Node/Ruby); experience designing scalable cloud systems.
Amazon Web Services, Kubernetes, GitHub Actions, Terraform, New Relic, Dynatrace, Datadog, Docker, GitHub, Python, Bash, Node, Ruby on Rails
9h
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Redwood City, California, United States
HybridFull Time
Luma AI
Luma AI: Develops multimodal AI for video generation and creative production.
5+ YOE5+ years SRE or infrastructure experience, deep Linux and low-level performance debugging, Terraform, Airflow, Ray, AWS or OCI, high-performance networking (InfiniBand/RDMA/RoCE), security/compliance familiarity.
Linux, Python, Go, Bash, Terraform, Airflow, Ray, AWS, OCI, InfiniBand, RDMA, RoCE, DCGM, ROCm, Kubernetes, NVIDIA, AMD
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, ASE
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Work on globally scaled, revenue-critical internet services (App Store, Music, Books, Podcasts, Fitness+); ensure reliability and scalability of services used by billions of devices.
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto, California, United States
$200k-$400k/yr HybridFull Time
Nectar Social
Nectar Social: AI platform for social commerce and community management.
5+ YOE5+ years operating production systems; cloud (AWS); infrastructure as code; programming; startup environment; reliability-focused with cost awareness.
AWS, Pulumi, Postgres, ClickHouse, Turbopuffer, Temporal
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - SDN
San Francisco or San Jose or Bellevue
$240k-$312k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
5+ YOE5+ years SRE/production engineering experience; Kubernetes, Linux networking, observability, on-call/incident response, automation with Python/Ansible; experience with multi-datacenter and hybrid cloud environments.
Kubernetes, SmartNICs, Python, Ansible, Go, C, Helm, Terraform, GitOps, CI/CD, Linux, OpenStack Neutron, OVN, OVS, DPDK, SR-IOV