289 infrastructure reliability engineer jobs at 159 companies in California

1w
Save
Mark Applied
Hide
Network Reliability Engineer, Infrastructure Services
California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Extensive software, systems, or infrastructure engineering experience; expertise designing highly available distributed systems, systems programming, networking, SDN, and cross-functional technical leadership.
JSON, Protocol Buffers (ProtoBuf), REST, RPC, XML, Kubernetes (K8s), OpenStack, Infrastructure as Code
2w
Save
Mark Applied
Hide
Reliability Engineer
San Francisco, California, United States
$130k-$170k/yr OnsiteFull Time
OpenMind
OpenMind: Private robotics software building AI infrastructure and robot orchestration for human-facing service and industrial fleets.
Experience in reliability engineering, test infrastructure, SRE, or autonomous-system simulation; strong Python, Go, or C++ fundamentals; robotics simulation experience; and expertise in measurement and failure analysis.
Python, Go, C++, Isaac Sim, Gazebo, MuJoCo, CI/CD
4mo
Save
Mark Applied
Hide
Staff Infrastructure Reliability Engineer - Database & Storage
Seattle or San Francisco or Detroit or United States
$180k-$279k/yr HybridFull Time
Rocket Homes
Rocket Homes: Technology-driven real estate service provider connecting home buyers and sellers with listings and real estate agents.
7+ YOE7+ years AWS/cloud infra; 5+ years PostgreSQL/AWS services; Linux admin/scripting; mentoring; infrastructure as code and security; AI code generation tools; on-call readiness.
AWS, PostgreSQL, Aurora/RDS, S3, ElastiCache, OpenSearch, DynamoDB, Linux, Python, Infrastructure as Code, Security practices, AI code generation tools
1mo
Save
Mark Applied
Hide
Reliability Engineer, R&D
Austin or New York City or San Francisco or Seattle
$203k-$232k/yr OnsiteFull Time
Fluidstack
Fluidstack: Building and operating civilization-scale data center infrastructure for AI.
Experience in reliability engineering for infrastructure or complex hardware, building availability/RAM models, leading cross-discipline FMEAs, and mining field failure data.
5d
Save
Mark Applied
Hide
Customer Reliability Engineer, Infrastructure
United States or Austin or New York City or Boston or San Francisco
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Private software providing managed Apache Airflow data orchestration for enterprise data teams.
5+ YOERequires 5 years of experience with complex cloud infrastructure, 3 years with Kubernetes, production distributed systems, Linux, monitoring, troubleshooting, customer support, DevOps or CI/CD, and Python scripting.
Astro, Apache Airflow, Kubernetes, AWS, GCP, Azure, Linux, Python
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - Data Infrastructure
San Jose, California, United States
$156k-$317k/yr OnsiteFull Time
TikTok
TikTok: Short-form mobile video and social media platform.
2+ YOE2+ years SRE/DevOps experience, bachelor’s degree or equivalent, scripting (Python/Go/Bash), Linux and networking knowledge, familiarity with containers and observability tools.
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Docker, Prometheus, Grafana, ELK Stack, Linux
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
San Francisco, California, United States
$200k-$350k/yr OnsiteFull Time
Console
Console: AI-native IT service-management software platform automating employee support requests for companies.
5+ YOE5+ years infrastructure/platform/backend experience; hands-on AWS and Kubernetes; experience with Pulumi or Terraform; comfortable in TypeScript/Node/React codebases; production reliability, observability, and enterprise deployment experience.
AWS, Kubernetes, Pulumi, Terraform, TypeScript, Node, React, Slack, Microsoft Teams
1w
Save
Mark Applied
Hide
Sr. Staff Engineer Software, Infrastructure Reliability (Chronosphere)
San Francisco or Denver or Austin or Jacksonville or Bridgeport or Seattle or Boston or New York City
$126k-$205k/yr RemoteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Global cybersecurity platform providing network, cloud, and AI-driven security solutions.
8+ YOERequires 8+ years of relevant experience, backend programming proficiency, cloud-native and distributed systems expertise, Linux and networking knowledge, debugging skills, and experience with AWS or GCP and Kubernetes.
Go, Java, Python, Rust, AWS, GCP, Kubernetes, Linux, Terraform, Cursor, Claude
1mo
Save
Mark Applied
Hide
Contract Site Reliability Engineer — AI Accelerator Infrastructure
Santa Clara, California, United States
$155k-$235k/yr HybridContract
d-Matrix
d-Matrix: Private AI infrastructure serving data centers with inference accelerators, networking, and software.
5+ YOE5+ years SRE/infrastructure experience; strong Linux, colocation and bare-metal skills; Terraform/Ansible; Kubernetes; Prometheus/Grafana or DataDog; Python/Bash; incident response and RCA experience.
AWS, Azure, GCP, Terraform, Ansible, Kubernetes, Prometheus, Grafana, DataDog, Python, Bash, Slurm, LSF, InfiniBand, RoCE, NVLink, Go
3d
Save
Mark Applied
Hide
Reliability Infrastructure Technician, Raxium
Fremont, California, United States
$99k-$140k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
2+ YOEAssociate's degree or equivalent experience, 2 years in semiconductor lab or manufacturing, 1 year troubleshooting electromechanical systems, and precision soldering and mechanical assembly experience.
Manufacturing Execution System (MES), Work-in-Progress (WIP), Python, Linux, Unix, Google Workspace, Google Docs, Google Sheets, Google Slides
1mo
Save
Mark Applied
Hide
Principal, System Reliability Engineer
San Jose, California, United States
$185k-$290k/yr OnsiteFull Time
Ayar Labs
Ayar Labs: -packaged optics providing connectivity for hyperscale AI infrastructure.
5+ YOE5+ years in systems/fleet reliability for large-scale infrastructure, BS in EE/CE, experience building test infrastructure, statistical reliability planning, customer-facing qualification, and on-call fleet operations.
FPGA
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yr HybridFull Time
Radix Health
Radix Health: Healthcare technology helping providers achieve fair reimbursement through integrated IDR software, data, and AI.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
AI, HIPAA, PHI, SOC 2, 401(k)
1mo
Save
Mark Applied
Hide
Software Engineer, Infrastructure & Reliability
San Francisco, California, United States
HybridFull Time
CrewAI
CrewAI: Private software providing multi-agent AI orchestration and enterprise workflow automation for businesses.
Experience building and operating production SaaS infrastructure: cloud, containers, CI/CD, observability, secrets, databases, and automation using Python/Ruby/Go/Bash.
AWS, Docker, CI/CD, GitHub Actions, ECS, ECR, Kubernetes, Helm, PostgreSQL, Redis, Celery, FastAPI, Rails, Sentry, OpenTelemetry, Python, Ruby, Go, Bash, Terraform
1w
Save
Mark Applied
Hide
Infrastructure Engineer, Database
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
LangChain
LangChain: Private AI software providing agent-engineering platforms and open-source frameworks for developers and enterprises.
5+ YOERequires 5+ years in infrastructure, platform engineering, or SRE; Kubernetes and cloud infrastructure expertise; scripting or systems programming; infrastructure-as-code, CI/CD, reliability engineering, and production stateful workload experience.
Rust, Kubernetes, Amazon S3, Google Cloud Storage, Azure Blob Storage, Terraform, Helm, AWS, GCP, Azure, Go, Python, ArgoCD, Pulumi, CDK, Postgres, ClickHouse, Redis, Docker
1mo
Save
Mark Applied
Hide
Alibaba Cloud-Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale
Sunnyvale, California, United States
$104k-$171k/yr OnsiteFull Time
Alibaba Cloud
Alibaba CloudNYSE, HKEX: BABA, 9988: Global cloud computing and data intelligence service provider.
2+ YOE2+ years in distributed systems reliability engineering; high-availability architecture, Kafka/RocketMQ, Kubernetes, automation, and proficiency in Python, Go, or Java required. Bachelor's degree listed.
RocketMQ, Kafka, Kubernetes, K8s, Java, Go, Python, Shell, Terraform, Helm, Operator
3w
Save
Mark Applied
Hide
Site Reliability Engineer, AI Infrastructure
San Jose, California, United States
$123k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture LLC
TikTok USDS Joint Venture LLC: Ensuring U.S. data security and content integrity for TikTok.
1+ YOEBachelor's degree or equivalent experience, 1+ year in SRE, DevOps, or systems engineering, Linux and networking knowledge, distributed systems experience, programming, scripting, CI/CD, and automation skills.
Linux, Go, Python, C, C++, Java, Bash, Kubernetes, AWS, GCP, Azure, Terraform, Prometheus, Grafana, Distributed Tracing, LLMs, Agentic AI
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
2w
Save
Mark Applied
Hide
Site Reliability Engineer - rednote
Palo Alto, California, United States
OnsiteFull Time
Xiaohongshu
Xiaohongshu: Lifestyle-focused social media and e-commerce platform.
Experience with large-scale reliability, high-availability architecture, incident response, cross-region disaster recovery, Linux, networking, middleware, cloud-native infrastructure, automation, and Python, Go, or Java.
Linux, MySQL, Redis, Kafka, Kubernetes, Service Mesh, Python, Go, Java
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
5+ YOEBachelor's or equivalent and 5+ years SRE/production engineering experience; proficiency with Go/Python/Bash, Linux, networking, and large-scale distributed systems.
Kubernetes, Redis, MySQL, Message Queue, Kafka, Flink, Go, Python, Bash
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale
Anyscale: AI infrastructure software helping developers and AI teams build, deploy, and scale machine-learning workloads with Ray.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
Ray, Kubernetes, Prometheus, Grafana, Go, Python, AWS, Azure, GCP, Linux kernel

Explore Jobs

Expand Your Job Search