26 batch data engineer jobs at 23 companies in California

1mo
Save
Mark Applied
Hide
Staff Data Engineer
San Francisco or New York
$185k-$195k/yr HybridFull Time
Butterfly Network
Butterfly NetworkNew York Stock Exchange: BFLY: Public U.S. medical technology making handheld point-of-care ultrasound hardware and AI-powered clinical software for healthcare professionals.
5+ YOE5+ years data engineering experience, strong dbt and Databricks skills, Python/PySpark, streaming and batch ingestion, CDC pipelines, data quality and observability, and experience integrating with operational systems.
dbt, Databricks, Delta Lake, Delta Live Tables, Python, PySpark, Salesforce, NetSuite, Stripe, Mulesoft, Great Expectations, Unity Catalog, Airflow, Cloud Run, Cloud Functions, BigQuery, AWS, GCP, Terraform, Spacelift, Boomi
2mo
Save
Mark Applied
Hide
Digital - Senior Data Engineer / Data Engineer, Digital Technology
Seattle or Vancouver or Toronto or New York City or Los Angeles
$100k-$300k/yr OnsiteFull Time
Aritzia
AritziaTSX: ATZ: Vertically integrated design house creating Everyday Luxury apparel.
Design and optimize scalable batch and real-time data pipelines; build ETL/ELT workflows; deliver governed data models; enforce data quality, governance, privacy, and security; mentor junior engineers.
SQL, Google Cloud Platform (GCP), BigQuery, Snowflake, Redshift, Airflow, Pub/Sub, Customer Data Platform (CDP)
1mo
Save
Mark Applied
Hide
Staff Data Engineer
California, United States
$185k-$220k/yr RemoteFull Time
Prodege
Prodege: Private consumer marketing and insights platform serving brands, marketers, agencies, and online consumers.
5+ YOE5+ years data engineering experience building large-scale batch and streaming pipelines; strong SQL, Python, Snowflake, dbt; expertise in data modeling, governance, observability, and ML feature pipelines.
SQL, Python, Snowflake, dbt, Iceberg, Trino, Kafka, Flink, Kinesis, Spark
8h
Save
Mark Applied
Hide
Data Engineer
San Francisco or New York City
HybridFull Time
Beast Industries
Beast Industries: Privately held creator-led holding producing digital entertainment, snacks, software, and consumer brands for global audiences.
3+ YOE3+ years building and operating high-volume production data pipelines, with streaming and batch experience, event instrumentation, schema design, data quality, event-driven architectures, and major cloud data stacks.
AWS, GCP, Databricks, Snowflake
2w
Save
Mark Applied
Hide
Data Engineer
Santa Clara, California, United States
$110k-$120k/yr RemoteFull Time
QualityAI
QualityAI: AI-first quality engineering and software testing firm.
Build scalable batch and streaming data pipelines using Python, PySpark, Databricks, Snowflake, orchestration, monitoring, CI/CD, DevOps, and cloud optimization practices.
Python, PySpark, Databricks, Snowflake, Terraform, CI/CD, DevOps, QCraft, Microsoft HSA
1mo
Save
Mark Applied
Hide
Senior Data Engineer
San Francisco, California, United States
$180k-$220k/yr HybridFull Time
Metriport
Metriport: Open-source healthcare data infrastructure providing APIs that help digital health organizations access and exchange patient data.
6+ YOE6+ years data engineering experience building and scaling pipelines; experience across ingestion, storage, processing, warehousing; production coding in TypeScript/Python; mentoring experience; located in or willing to relocate to San Francisco Bay Area.
Spark, Parquet, Iceberg, Delta, S3, Snowflake, BigQuery, Redshift, TypeScript, Python, Kafka, Kinesis, dbt, Airflow, Dagster, FHIR, HIE, IHE, EHR/EMR, NPI, TEFCA, ADT, HL7, HEDIS, RAF, SNOMED, LOINC, ICD-10, Node.js, AWS, ECS, Lambda, SQS, SNS, Batch, CDK, PostgreSQL, Aurora, DynamoDB, Athena, SageMaker
1mo
Save
Mark Applied
Hide
Senior Staff Data Engineer
San Jose or Draper
$166k-$245k/yr HybridFull Time
BILL
BILLNYSE: BILL: Public financial operations platform automating payments, receivables, expenses, and cash management for small and midsize businesses.
8+ YOELead architecture and delivery of data platform capabilities (ingest, lake, streaming, feature store, query, graph, search). Requires distributed systems, streaming and batch experience, SQL and Python, and technical leadership.
Starburst, Databricks Feature Store, Neo4j, OpenSearch, Kafka, Flink, Spark Streaming, Airflow, dbt, Spark, Glue, Apache Iceberg, Delta Lake, Trino, Presto, Python, SQL
1mo
Save
Mark Applied
Hide
Software Engineer - X Data
Palo Alto, California, United States
$125k-$400k/yr OnsiteFull Time
xAI
xAI: Artificial intelligence research and development.
3+ YOE3+ years software engineering experience; expertise in Python,Rust,Scala,Go or Java; experience with data pipelines, realtime and batch processing, and distributed systems.
Python, Rust, Scala, Go, Java, BigQuery, Trino, Clickhouse, Flink, Kafka, Spark, Scalding, SQL
3d
Save
Mark Applied
Hide
AI Context & Data Infrastructure Engineer
San Francisco, California, United States
$250k-$300k/yr OnsiteFull Time
Town
Town: Privately held applied-AI software building assistants that help people create and use software at work.
Significant hands-on experience with lexical and semantic search, large-scale realtime and batch data infrastructure, pipelines, indexing, storage, and retrieval tradeoffs; ranking, knowledge graphs, or LLM retrieval are bonuses.
BM25, ANN, LLM
1mo
Save
Mark Applied
Hide
ML Infra Engineer (Data Systems)
San Francisco, California, United States
OnsiteFull Time
Physical Intelligence
Physical Intelligence: AI robotics developing foundation models and learning algorithms for robots and physically actuated devices.
Strong software engineering fundamentals with experience building distributed systems, large-scale data pipelines, object storage and batch/streaming systems; ownership mindset and performance focus.
ClickHouse, Ray, Flink, Spark
1mo
Save
Mark Applied
Hide
Member of Technical Staff - Data Platform
New York City or San Francisco or London
OnsiteFull Time
Reflection AI
Reflection AI: Private AI research lab building open foundation models and agentic software for enterprise, government, and regulated-industry users.
Experienced data engineer with production-grade pipeline experience for large-scale batch and streaming workloads; strong debugging, cost/performance optimization, and data quality/governance skills.
Spark, Flink, Beam, Airflow, Dagster, Kafka, PubSub, Parquet, Iceberg, Delta Lake, BigQuery, Snowflake, Great Expectations
4w
Save
Mark Applied
Hide
Software Engineer, Data Infrastructure
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience with distributed systems, cloud data lakes, batch and streaming pipelines; proficiency in Python or Rust; bachelor's degree or equivalent; collaborative and cross-functional problem solving.
Apache Spark, Spark, Kafka, Beam, Ray, Delta Lake, Python, Rust, dbt, Terraform, Airflow, Parquet
1mo
Save
Mark Applied
Hide
Member of Technical Staff — Data Ingestion & Quality
San Francisco, California, United States
OnsiteFull Time
Causal Labs
Causal Labs: Private AI research building physics foundation models for weather prediction and control for businesses and governments.
Experience building large-scale data pipelines and QA systems, familiarity with streaming and batch ingestion, vendor collaboration, and strong problem-solving and domain learning skills.
Apache Spark, Ray, Beam
3mo
Save
Mark Applied
Hide
Data Engineering Manager, Data & ML Platform
San Francisco or Montreal or Bengaluru
$220k-$330k/yr HybridFull Time
Hinge Health
Hinge HealthNYSE: HNGE: Public digital MSK clinic using AI, wearable technology, and clinicians to treat joint and muscle pain.
5+ YOE2+ Mgmt5+ years data engineering; 2+ years managing engineering teams; 2+ years building ML platform capabilities; experience with batch & streaming systems and tools such as Kafka, Flink, Spark; proficiency with Python, SQL, dbt, Databricks, and AWS.
Kafka, Flink, Spark, Python, SQL, dbt, Databricks, AWS, Delta Lake, MLflow, Unity Catalog
2mo
Save
Mark Applied
Hide
Senior Manager, Content Promotion & Distribution Data Engineering
Los Angeles or Los Gatos
$525k-$950k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ Mgmt7+ years leading data engineering teams; expertise in data modeling, batch/streaming pipelines, data quality, and multi-modal media pipelines for ML/GenAI; strong stakeholder communication and people management.
S3, Spark
2w
Save
Mark Applied
Hide
State Estimation Engineer - Data Collection Systems
San Jose, California, United States
$150k-$300k/yr OnsiteFull Time
Figure
Figure: Develops general-purpose humanoid robots.
4+ YOERequires 4+ years building multi-sensor fusion and state estimation systems, expertise in real-time filtering and batch optimization, advanced kinematics and optimization knowledge, and high-performance C++ and Python skills.
C++, Python, GTSAM, Ceres, Machine Learning (ML)
1mo
Save
Mark Applied
Hide
Software Engineer, Reconciliation & Reporting
California, United States
$142k-$213k/yr RemoteFull Time
Square Financial Services
Square Financial Services: -owned FDIC-insured industrial bank providing loans and savings services to Square sellers and Cash App consumers.
3+ YOE3+ years software or data engineering experience, BS or equivalent, experience with batch/event-driven data pipelines and workflow orchestration, proficiency in a major language and SQL, and experience delivering end-to-end projects.
Airflow, Temporal, Spark, Java, Python, Go, SQL, BigQuery, Snowflake, PySpark, Delta Lake, Kafka, Hadoop, AWS, Kubernetes, Terraform, ISO-8583
3mo
Save
Mark Applied
Hide
Member of Technical Staff (Software Engineer, Data Platform)
San Francisco or New York City or Palo Alto
$220k-$405k/yr OnsiteFull Time
Perplexity AI
Perplexity AI: Private American software offering an AI-powered conversational search engine with real-time web answers and citations.
5+ YOE5+ years software engineering experience (8+ for staff), production data infrastructure and batch/streaming experience, Airflow/Dagster, Python plus another backend language (Go/TypeScript), ML/AI workflow support, data quality and observability knowledge.
Databricks, Snowflake, Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, ClickHouse, Kinesis, PubSub, Python, Go, TypeScript
2mo
Save
Mark Applied
Hide
Senior Technical Solutions Engineer
San Bruno or Boston or Dallas
$162k-$182k/yr HybridFull Time
Verily
Verily: Private precision-health data platform and AI technology serving healthcare, life-science, payer, government, and research customers.
6+ YOEBA/BS in CS or equivalent, 6+ years experience with Python and/or Java, SQL proficiency, experience with life science/biomedical data and batch workflows, strong communication and project management skills.
Python, Java, SQL, Google Cloud Platform, Amazon Web Services, Microsoft Azure, Docker, Linux
3w
Save
Mark Applied
Hide
Engineering Manager, Data Platform
San Jose, California, United States
$209k-$438k/yr OnsiteFull Time
TikTok USDS Joint Venture LLC
TikTok USDS Joint Venture LLC: Ensuring U.S. data security and content integrity for TikTok.
5+ YOE3+ MgmtBachelor's or master's degree or equivalent experience; 5+ years building enterprise data platforms, including 3+ years leading engineering teams. Requires batch and streaming systems, SLA, incident response, architecture, and communication expertise.
Spark, Flink, Hadoop, Kafka, Druid, ClickHouse