95 infrastructure reliability engineer jobs at 69 companies in Napa, CA

1mo
Save
Mark Applied
Hide
Customer Reliability Engineer - Infrastructure
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Apache Airflow, AWS, Azure, CI/CD, GCP, Infrastructure as Code (IaC), Kubernetes, Linux, Python
3mo
Save
Mark Applied
Hide
Staff Infrastructure Reliability Engineer - Database & Storage
Seattle or San Francisco or Detroit or United States
$180k-$279k/yr HybridFull Time
Rocket Companies
Rocket CompaniesNYSE: RKT: Provides digital mortgage, real estate, and personal finance services.
7+ YOE7+ years AWS/cloud infra; 5+ years PostgreSQL/AWS services; Linux admin/scripting; mentoring; infrastructure as code and security; AI code generation tools; on-call readiness.
AWS, PostgreSQL, Aurora/RDS, S3, ElastiCache, OpenSearch, DynamoDB, Linux, Python, Infrastructure as Code, Security practices, AI code generation tools
2w
Save
Mark Applied
Hide
Reliability Engineer, R&D
Austin or New York City or San Francisco or Seattle
$203k-$232k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience in reliability engineering for infrastructure or complex hardware, building availability/RAM models, leading cross-discipline FMEAs, and mining field failure data.
1mo
Save
Mark Applied
Hide
Infrastructure Engineer
San Francisco, California, United States
$200k-$350k/yr OnsiteFull Time
Console
Console: Automates IT support and internal operations using AI agents.
5+ YOE5+ years infrastructure/platform/backend experience; hands-on AWS and Kubernetes; experience with Pulumi or Terraform; comfortable in TypeScript/Node/React codebases; production reliability, observability, and enterprise deployment experience.
AWS, Kubernetes, Pulumi, Terraform, TypeScript, Node, React, Slack, Microsoft Teams
3w
Save
Mark Applied
Hide
Software Engineer, Infrastructure & Reliability
San Francisco, California, United States
HybridFull Time
CrewAI
CrewAI: Platform for orchestrating collaborative multi-agent AI systems.
Experience building and operating production SaaS infrastructure: cloud, containers, CI/CD, observability, secrets, databases, and automation using Python/Ruby/Go/Bash.
AWS, Docker, CI/CD, GitHub Actions, ECS, ECR, Kubernetes, Helm, PostgreSQL, Redis, Celery, FastAPI, Rails, Sentry, OpenTelemetry, Python, Ruby, Go, Bash, Terraform
4w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
San Francisco, California, United States
$200k-$250k/yr OnsiteFull Time
Stuut
Stuut: Automates business accounts receivable and collections through AI agents.
7+ YOE7+ years in SRE/infrastructure or backend engineering. Experience with AWS, Kubernetes/EKS, Docker, observability, SLOs/SLIs, Python or TypeScript, CI/CD, and production-grade distributed systems.
Python, TypeScript, AWS, Kubernetes, EKS, Docker, FastAPI, Vue.js, PostgreSQL (RDS), CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.
Terraform, Chef, Ansible, Python, Java, Bash, Kubernetes, Helm, Jenkins, Grafana, Prometheus, OCI - DevOps, Oracle Cloud Guard, Oracle Observability and Management
5d
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
$248k-$405k/yr OnsiteFull Time
Ivo
Ivo: AI-powered contract review and intelligence platform for legal teams.
2+ YOEMinimum 2 years infrastructure experience, SRE skills, SLI/SLO/SLA design, disaster recovery, security controls, incident response, strong systems design and failure-mode thinking.
Microsoft Word, LLM, RAG, VPC
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yr RemoteFull Time
Bedrock Ocean Exploration
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Python, Go, Bash, Terraform, AWS, Docker, Kubernetes, Prometheus, Grafana, ROS 2, ROS, Jetson, Linux, IAM
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
Ray, Kubernetes, Prometheus, Grafana, Go, Python, AWS, Azure, GCP, Linux kernel
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
San Francisco, California, United States
$150k-$300k/yr OnsiteFull Time
Reducto
Reducto: AI platform extracting structured data from unstructured documents
5+ YOE5+ years building production infrastructure; proficient in Python; strong cloud, Kubernetes, networking, storage, and automation; focus on reliability.
Python, Kubernetes, Cloud platforms, Networking, Storage, Automation
2mo
Save
Mark Applied
Hide
Founding ML infrastructure Engineer
San Francisco or United States
$200k-$350k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Experience designing and operating large-scale distributed infrastructure; Kubernetes/Slurm; multi-cloud GPU; reliability and scheduling; startup mindset.
Kubernetes, Slurm, Scheduling, TensorRT-LLM, NCCL, InfiniBand, RoCE, CuTe, Triton, TileLang
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Plenful
Plenful: AI-powered workflow automation platform for healthcare and pharmacy operations.
5+ YOE5+ years SRE or production infrastructure experience; hands-on with observability, incident response, SLOs, AWS, container and serverless platforms; able to write automation scripts.
OpenTelemetry, Datadog, CloudWatch, Grafana, Sentry, AWS Lambda, ECS, Aurora Postgres, ClickHouse, GitHub Actions, Python, Bash, Vanta
1w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
4w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
1mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Jose or Seattle or San Francisco
$159k-$302k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOEBachelor's or equivalent, 5+ years SRE/infrastructure/backend experience; Kubernetes, Docker, Terraform, AWS, Postgres/Redis, observability, incident response, CI/CD, bash, Node.js/TypeScript experience; on-call participation.
Kubernetes, Docker, bash, CircleCI, Node.js, TypeScript, Postgres, Redis, AWS Aurora (Postgres-compatible), Terraform, AWS
1mo
Save
Mark Applied
Hide
Staff Security Reliability Engineer
San Francisco, California, United States
$293k-$385k/yr OnsiteFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
10+ YOE10+ years operating mission-critical on-prem/hybrid infrastructure; strong SRE discipline, IaC and config management (Terraform, Chef, Ansible); identity and Azure/Microsoft Entra experience; observability and incident response.
Terraform, Chef, Ansible, Microsoft Entra, Azure, FleetDM, Azure Virtual Desktop
2w
Save
Mark Applied
Hide
Site Reliability Engineer II
Scottsdale or San Francisco or Chicago or New York City
$86k-$126k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
2+ YOEBachelor's or equivalent, minimum 2 years DevOps/Dev/SRE experience, Linux/Unix experience, infrastructure automation (Chef/Ansible/Puppet, Terraform), containerization (Docker,Kubernetes), cloud (AWS/GCP/Azure), on-call rotation.
Linux, Unix, Chef, Ansible, Puppet, Terraform, Docker, Kubernetes, AWS, GCP, Azure, Java, Ruby, Python, JavaScript, Go

Explore Jobs

Expand Your Job Search