48 data reliability engineer jobs at 34 companies in Santa Rosa, CA
6d
Save
Mark Applied
Hide
6d
Reliability Engineer, R&D
Austin or New York City or San Francisco or Seattle
$203k-$232k/yrOnsiteFull Time
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience in reliability engineering for infrastructure or complex hardware, building availability/RAM models, leading cross-discipline FMEAs, and mining field failure data.
Bedrock Robotics: Automates heavy construction machinery with AI retrofit kits.
10+ YOE10+ years hardware reliability engineering for vehicles/heavy machinery; hands-on DFMEA, FTA, HALT/HASS/ALT, Weibull life-data analysis, FRACAS, and reliability test programs; Bachelor's degree or equivalent.
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
San Francisco or New York City or Seattle or Boston or Los Angeles or Chicago or Washington or United States
$145k-$230k/yrHybridFull Time
Scribe: Automatically documents digital workflows into step-by-step process guides.
Deep PostgreSQL and ORM expertise, experience with CDC pipelines (AWS DMS), OpenSearch, Redis, message brokers, observability tools, Python/Go automation, Terraform/IaC, and building reliability/scale for data tiers.
United States or San Francisco or Boston or Atlanta or Austin or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
4+ YOEData engineering background, 4 years Python, 1 year Airflow administration/DAG creation, Kubernetes/Docker experience, cloud provider (AWS/GCP/Azure) experience, troubleshooting, strong communication, and mentoring experience.
The Walt Disney CompanyNYSE: DIS: Produces media content and operates global theme parks.
5+ YOE5+ years data engineering experience, streaming pipeline expertise, strong Python/Java/SQL skills, experience with embedding/vector stores and workflow orchestration, and ability to design scalable, reliable data systems for AI applications.
Ivo: AI-powered contract review and intelligence platform for legal teams.
Senior/Staff Site Reliability Engineer focusing on uptime, SLO/SLI/SLA, disaster recovery, data residency, security controls, observability, and incident response.
Beast Industries: Produces digital media and consumer goods for MrBeast brands.
8+ YOE8+ years building production data platforms, expertise in batch and streaming pipelines, data modeling, ETL/ELT, reliability, observability, and privacy-safe foundations for AI/ML.
Member of Technical Staff – Senior Engineer, Data Infrastructure & Data Operations
San Francisco or Cambridge
$255k-$340k/yrOnsiteFull Time
Walden Robotics: Builds general-purpose robots and develops the teams and infrastructure to scale robot applications and improve quality of life.
Experience building production data infrastructure and high-throughput pipelines, cloud-based data platform development, platform reliability and cost ownership, and collaboration with ML teams.
Effective AI: AI-powered research and analysis platform for insurance P&L teams
Strong computer science foundation and track record shipping large-scale systems; product judgment; experience turning raw external sources into reliable signals; familiarity with agent-based workflows and tooling.
Technical Lead Manager, Data Engineering, Trust & Safety
San Francisco, California, United States
$385k-$490k/yrOnsiteFull Time
OpenAI: Develops artificial intelligence models and generative AI software services.
Proven experience leading data engineering teams; deep technical expertise in data architecture, modeling, pipelines, reliability, and privacy; experience with Spark and Airflow; strong stakeholder partnership and hiring experience.
United States or Illinois or San Francisco or Los Angeles or Indiana
$120k-$194k/yrRemoteFull Time
AllstateNYSE: ALL: Provides insurance products for vehicles, homes, and businesses.
7+ YOE3+ Mgmt7+ years in software/platform/data engineering, 3+ years leading engineering teams in product environments; experience operating production systems with reliability, observability, and security; cloud-native and integration expertise; strong technical and people leadership.
Plaid: Provides financial data connectivity and payment infrastructure via APIs.
8+ YOE8+ years building and operating backend distributed systems; strong fundamentals in data stores, consistency and reliability; experience with Go or Java; ability to lead complex system design and align cross-team stakeholders.
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years professional software engineering with backend focus; proficiency in a backend language (Python, Go, Java), data modeling across SQL and NoSQL, systems thinking, API and schema design, and ownership of production reliability.
Python, Go, Java, SQL, NoSQL, Slack, Microsoft 365
Labelbox: Provides a platform and services for AI training data management.
4+ YOE4+ years building and shipping reliable full-stack systems; strong system and API design judgment; deep proficiency in TypeScript and/or Python; experience with production distributed systems or ML/data infrastructure preferred.
London or New York or Washington or San Francisco or Dublin or Brussels or Singapore or Hong Kong
OnsiteFull Time
Penta Group: Consultancy providing data-driven stakeholder solutions and reputation strategy.
10+ YOEHands-on technical leader with ~10+ years experience in applied AI and data science, strong Python, SQL/PostgreSQL, AWS, Git, production LLM/ML workflows, NLP, MLOps/DevOps, and people management to ship reliable AI systems.
Rox: AI-powered revenue operating system with autonomous sales agents.
Experience building and operating production distributed systems, real-time data infrastructure, low-latency decisioning, and agent orchestration; strong reliability, scalability, and debugging skills.
Lapel: Unifying customer data to power personalized service at scale.
3+ YOE3+ years backend or software engineering; experience designing APIs, data models, permissions, and backend abstractions; ability to build reliable, scalable primitives; customer-focused with ownership.