55 infrastructure reliability engineer jobs at 24 companies in Prunedale, CA

1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEDegree in CS or related field (or equivalent experience), 8+ years SRE/DevOps/Production Engineering, SRE principles, infrastructure automation, production reliability, Python/Go/Perl/Ruby, Prometheus and Grafana, strong communication.
Python, Go, Perl, Ruby, Prometheus, Grafana
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Data Infrastructure
San Jose, California, United States
$156k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOE2+ years SRE/DevOps experience, bachelor’s degree or equivalent, scripting (Python/Go/Bash), Linux and networking knowledge, familiarity with containers and observability tools.
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Docker, Prometheus, Grafana, ELK Stack, Linux
3w
Save
Mark Applied
Hide
Contract Site Reliability Engineer — AI Accelerator Infrastructure
Santa Clara, California, United States
$155k-$235k/yr HybridContract
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
5+ YOE5+ years SRE/infrastructure experience; strong Linux, colocation and bare-metal skills; Terraform/Ansible; Kubernetes; Prometheus/Grafana or DataDog; Python/Bash; incident response and RCA experience.
AWS, Azure, GCP, Terraform, Ansible, Kubernetes, Prometheus, Grafana, DataDog, Python, Bash, Slurm, LSF, InfiniBand, RoCE, NVLink, Go
1w
Save
Mark Applied
Hide
Principal, System Reliability Engineer
San Jose, California, United States
$185k-$290k/yr OnsiteFull Time
Ayar Labs
Ayar Labs: Develops optical interconnect technology for high-speed data movement.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
FPGA
4d
Save
Mark Applied
Hide
Site Reliability Engineer, AI Infrastructure
San Jose, California, United States
$123k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
1+ YOEBachelor's degree or equivalent experience, 1+ year in SRE, DevOps, or systems engineering, Linux and networking knowledge, distributed systems experience, programming, scripting, CI/CD, and automation skills.
Linux, Go, Python, C, C++, Java, Bash, Kubernetes, AWS, GCP, Azure, Terraform, Prometheus, Grafana, Distributed Tracing, LLMs, Agentic AI
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOEBachelor's or equivalent and 5+ years SRE/production engineering experience; proficiency with Go/Python/Bash, Linux, networking, and large-scale distributed systems.
Kubernetes, Redis, MySQL, Message Queue, Kafka, Flink, Go, Python, Bash
2w
Save
Mark Applied
Hide
Staff Network Reliability Engineer (Cloud Operations)
Mountain View, California, United States
HybridFull Time
Skylo
Skylo: Provides direct-to-device satellite connectivity for mobile and IoT devices.
8+ YOE8+ years cloud/infrastructure/SRE experience with Kubernetes, hybrid cloud operations, observability, database and storage reliability, GitOps, and on-call ownership in 24x7 environments.
Kubernetes, GKE, GCP, kubectl, Pub/Sub, Cloud SQL, Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, PostgreSQL, Redis, ArgoCD, Helm, Terraform, Ansible, Ceph, Rook, Harvester, KubeVirt, KVM, Loki, ELK, Flux CD, Go, Python, BGP, VXLAN, EVPN
1w
Save
Mark Applied
Hide
Staff Engineer, Hardware Reliability
Sunnyvale, California, United States
$156k-$255k/yr HybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
6+ YOEBS in CS/CE or equivalent experience,6+ years in Linux-based infrastructure,4+ years hardware troubleshooting,experience with hardware qualification/integration,firmware/BMC knowledge,and automation for infrastructure at scale.
Linux, BMC, BIOS, IPMI, Redfish, SPEC, SPECpower, fio, unixbench, NCCL, CUDA, InfiniBand, RDMA, RoCE, GPFS, HDFS, Kubernetes
2w
Save
Mark Applied
Hide
Infrastructure Software Engineer
Campbell, California, United States
$180k-$230k/yr RemoteFull Time
Camus Energy
Camus Energy: Software platform for managing renewable energy grid integration.
3+ YOE3+ years software engineering with infrastructure focus; Python 3, Kubernetes, GCP, CI/CD, observability (Prometheus,Grafana); familiarity with SQL; collaborative incident response and reliability practices.
Python 3, Kubernetes, GCP, CI/CD pipelines, Prometheus, Grafana, Bazel, Go, Node.js, SQL
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Apply SRE principles to mentor teams, ensure reliability for large-scale analytics infrastructure across Hadoop, HBase, Spark, Data Lakes, and Airflow; participate in production on-call.
Hadoop, HBase, Spark, Data Lakes, Airflow
3d
Save
Mark Applied
Hide
Senior Database Reliability Engineer (DBRE)
San Jose, California, United States
$90k-$130k/yr RemoteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
5+ YOE5+ years PostgreSQL and production data-system experience; 3+ years Linux engineering and infrastructure automation; 2+ years Python, Bash, Go, Ruby, or Perl; cloud, Kubernetes, and distributed data-system experience.
PostgreSQL, Kubernetes, Amazon RDS, AWS, Terraform, Ansible, Chef, Puppet, Python, Bash, Go, Ruby, Perl, Kafka, MSK, ClickHouse, Redis, MySQL, Cassandra, Elasticsearch, Linux, SRE, ETL
4d
Save
Mark Applied
Hide
Staff Site Reliability Engineer (SRE) (Hybrid)
San Francisco or San Jose or New York City or Milpitas or Mountain View or Holmdel or Goleta or Redwood City or Fremont or Sunnyvale or Brooklyn or Palo Alto
$187k-$268k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
6+ YOERequires 8+ years with a bachelor's, 6+ with a master's, or 3+ with a PhD; 6+ years in SRE or infrastructure engineering, 5+ years operating Kubernetes, cloud, CI/CD, and Python or Go.
Kubernetes, AWS, GCP, Python, Go, Terraform, MLOps, CI/CD
1mo
Save
Mark Applied
Hide
Distinguished Technologist Mechanical Engineer (Network Infrastructure)
Sunnyvale, California, United States
$163k-$348k/yr OnsiteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Provides edge-to-cloud IT infrastructure and platform services.
15+ YOEBS in Mechanical Engineering, 15+ years electromechanical product development experience; expertise in chassis, packaging, manufacturability, reliability, SolidWorks/PLM, FEA, GD&T, thermal and cooling technologies; strong leadership and troubleshooting skills.
SolidWorks, EPDM, FEA, ECAD, IDF
1mo
Save
Mark Applied
Hide
Distinguished Technologist Mechanical Engineer (Network Infrastructure)
Sunnyvale, California, United States
$163k-$348k/yr OnsiteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Providing global edge-to-cloud infrastructure and IT solutions for businesses.
15+ YOEBS in Mechanical Engineering, 15+ years electromechanical product development experience; expertise in chassis, packaging, manufacturability, thermal and reliability; SolidWorks and PLM experience; tolerance/GD&T, FEA, ECAD-MCAD, and supplier engagement experience.
SolidWorks, EPDM, PLM, FEA, ECAD-MCAD, IDF, GD&T
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
San Francisco or Sunnyvale or Seattle
$137k-$202k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
6+ YOE6+ years software engineering experience; BS/MS/PhD in CS or equivalent; deep backend fundamentals in Python and distributed systems; experience with LLM inference/fine-tuning, production reliability, observability, and technical leadership.
Python, Claude Code, Codex, Cursor, vLLM, SGLang, TensorRT-LLM, Kubernetes, AWS, GCP, Modal
1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald
5d
Save
Mark Applied
Hide
Software Engineer III – Data Platform
Mountain View or McLean or New York City or Tampa
$173k-$201k/yr OnsiteFull Time
ID.me
ID.me: Provides secure digital identity verification and authentication services.
3+ YOEBachelor's degree in a technical field; 3–5 years in SRE, DevOps, or infrastructure engineering; cloud experience; and 1+ year with a modern programming language.
Prometheus, Grafana, OpenTelemetry, AWS, GCP, Azure, Java, Go, Python, Ruby, JavaScript, Docker, Kubernetes, Terraform, Pulumi, Ansible, GitOps, CI/CD, FedRAMP, NIST 800-53, SOC2
1mo
Save
Mark Applied
Hide
Principal Engineer, Model Development Platform
Sunnyvale, California, United States
$296k-$335k/yr HybridFull Time
Wayve
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
10+ YOE10+ years building large-scale distributed systems or ML infrastructure, 3+ years at staff/principal level, experience with Spark, Ray, Kubernetes, Airflow, MLflow, web frameworks, reliability engineering, and mentoring engineers.
Spark, Ray, Kubernetes, Airflow, MLflow, React, Flask, FastAPI