68 infrastructure reliability engineer jobs at 38 companies in Aptos, CA

1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEDegree in CS or related field (or equivalent experience), 8+ years SRE/DevOps/Production Engineering, SRE principles, infrastructure automation, production reliability, Python/Go/Perl/Ruby, Prometheus and Grafana, strong communication.
Python, Go, Perl, Ruby, Prometheus, Grafana
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Data Infrastructure
San Jose, California, United States
$156k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOE2+ years SRE/DevOps experience, bachelor’s degree or equivalent, scripting (Python/Go/Bash), Linux and networking knowledge, familiarity with containers and observability tools.
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Docker, Prometheus, Grafana, ELK Stack, Linux
3w
Save
Mark Applied
Hide
Contract Site Reliability Engineer — AI Accelerator Infrastructure
Santa Clara, California, United States
$155k-$235k/yr HybridContract
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
5+ YOE5+ years SRE/infrastructure experience; strong Linux, colocation and bare-metal skills; Terraform/Ansible; Kubernetes; Prometheus/Grafana or DataDog; Python/Bash; incident response and RCA experience.
AWS, Azure, GCP, Terraform, Ansible, Kubernetes, Prometheus, Grafana, DataDog, Python, Bash, Slurm, LSF, InfiniBand, RoCE, NVLink, Go
1w
Save
Mark Applied
Hide
Principal, System Reliability Engineer
San Jose, California, United States
$185k-$290k/yr OnsiteFull Time
Ayar Labs
Ayar Labs: Develops optical interconnect technology for high-speed data movement.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
FPGA
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOEBachelor's or equivalent and 5+ years SRE/production engineering experience; proficiency with Go/Python/Bash, Linux, networking, and large-scale distributed systems.
Kubernetes, Redis, MySQL, Message Queue, Kafka, Flink, Go, Python, Bash
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
Ray, Kubernetes, Prometheus, Grafana, Go, Python, AWS, Azure, GCP, Linux kernel
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto, California, United States
$200k-$400k/yr HybridFull Time
Nectar Social
Nectar Social: AI platform for social commerce and community management.
5+ YOE5+ years operating production systems; cloud (AWS); infrastructure as code; programming; startup environment; reliability-focused with cost awareness.
AWS, Pulumi, Postgres, ClickHouse, Turbopuffer, Temporal
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Palo Alto or Palo Alto or Washington
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
Kubernetes, Linux, Bash, Python, Bazel, Makefiles, Terraform, Ansible, TCP/IP
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Redwood City, California, United States
HybridFull Time
Luma AI
Luma AI: Develops multimodal AI for video generation and creative production.
5+ YOE5+ years SRE or infrastructure experience, deep Linux and low-level performance debugging, Terraform, Airflow, Ray, AWS or OCI, high-performance networking (InfiniBand/RDMA/RoCE), security/compliance familiarity.
Linux, Python, Go, Bash, Terraform, Airflow, Ray, AWS, OCI, InfiniBand, RDMA, RoCE, DCGM, ROCm, Kubernetes, NVIDIA, AMD
2w
Save
Mark Applied
Hide
Staff Network Reliability Engineer (Cloud Operations)
Mountain View, California, United States
HybridFull Time
Skylo
Skylo: Provides direct-to-device satellite connectivity for mobile and IoT devices.
8+ YOE8+ years cloud/infrastructure/SRE experience with Kubernetes, hybrid cloud operations, observability, database and storage reliability, GitOps, and on-call ownership in 24x7 environments.
Kubernetes, GKE, GCP, kubectl, Pub/Sub, Cloud SQL, Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, PostgreSQL, Redis, ArgoCD, Helm, Terraform, Ansible, Ceph, Rook, Harvester, KubeVirt, KVM, Loki, ELK, Flux CD, Go, Python, BGP, VXLAN, EVPN
6d
Save
Mark Applied
Hide
Staff Engineer, Hardware Reliability
Sunnyvale, California, United States
$156k-$255k/yr HybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
6+ YOEBS in CS/CE or equivalent experience,6+ years in Linux-based infrastructure,4+ years hardware troubleshooting,experience with hardware qualification/integration,firmware/BMC knowledge,and automation for infrastructure at scale.
Linux, BMC, BIOS, IPMI, Redfish, SPEC, SPECpower, fio, unixbench, NCCL, CUDA, InfiniBand, RDMA, RoCE, GPFS, HDFS, Kubernetes
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer, AI Foundations, F1 Query
San Jose, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's in CS or equivalent,8+ years building infrastructure/distributed systems,5+ years programming in C++ or Go,5+ years reliability engineering,EMR not mentioned,experience with distributed systems and stakeholder collaboration.
C++, Go, Java, GoogleSQL, Google Cloud
2w
Save
Mark Applied
Hide
Infrastructure Software Engineer
Campbell, California, United States
$180k-$230k/yr RemoteFull Time
Camus Energy
Camus Energy: Software platform for managing renewable energy grid integration.
3+ YOE3+ years software engineering with infrastructure focus; Python 3, Kubernetes, GCP, CI/CD, observability (Prometheus,Grafana); familiarity with SQL; collaborative incident response and reliability practices.
Python 3, Kubernetes, GCP, CI/CD pipelines, Prometheus, Grafana, Bazel, Go, Node.js, SQL
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Apply SRE principles to mentor teams, ensure reliability for large-scale analytics infrastructure across Hadoop, HBase, Spark, Data Lakes, and Airflow; participate in production on-call.
Hadoop, HBase, Spark, Data Lakes, Airflow
2d
Save
Mark Applied
Hide
Senior Database Reliability Engineer (DBRE)
San Jose, California, United States
$90k-$130k/yr RemoteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
5+ YOE5+ years PostgreSQL and production data-system experience; 3+ years Linux engineering and infrastructure automation; 2+ years Python, Bash, Go, Ruby, or Perl; cloud, Kubernetes, and distributed data-system experience.
PostgreSQL, Kubernetes, Amazon RDS, AWS, Terraform, Ansible, Chef, Puppet, Python, Bash, Go, Ruby, Perl, Kafka, MSK, ClickHouse, Redis, MySQL, Cassandra, Elasticsearch, Linux, SRE, ETL
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Foster City, California, United States
$250k-$300k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
5+ YOE5+ years operating GitHub Enterprise at scale, monorepo management, CI/CD integration, infrastructure-as-code (Terraform/Pulumi), cloud platform experience, technical leadership and migration planning.
Git, GitHub Enterprise, GitHub Cloud, Buildkite, GitHub Actions, Jenkins, GitLab CI, Terraform, Pulumi, Bazel, Buck, Reviewable, Gerrit
3d
Save
Mark Applied
Hide
Staff Site Reliability Engineer (SRE) (Hybrid)
San Francisco or San Jose or New York City or Milpitas or Mountain View or Holmdel or Goleta or Redwood City or Fremont or Sunnyvale or Brooklyn or Palo Alto
$187k-$268k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
6+ YOERequires 8+ years with a bachelor's, 6+ with a master's, or 3+ with a PhD; 6+ years in SRE or infrastructure engineering, 5+ years operating Kubernetes, cloud, CI/CD, and Python or Go.
Kubernetes, AWS, GCP, Python, Go, Terraform, MLOps, CI/CD