18 batch data engineer jobs at 15 companies in American Canyon, CA
2w
Save
Mark Applied
Hide
2w
Principal Data Engineer, People Data
Bellevue or Chicago or San Francisco or Washington
$197k-$271k/yrHybridFull Time
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
10+ YOE10+ years data engineering experience; expert SQL, ETL/ELT, MPP databases, AWS services; HR/People data experience; experience with batch and realtime pipelines, lakehouse architectures, and data quality frameworks.
5+ YOE5+ years data engineering experience, strong dbt and Databricks skills, Python/PySpark, streaming and batch ingestion, CDC pipelines, data quality and observability, and experience integrating with operational systems.
dbt, Databricks, Delta Lake, Delta Live Tables, Python, PySpark, Salesforce, NetSuite, Stripe, Mulesoft, Great Expectations, Unity Catalog, Airflow, Cloud Run, Cloud Functions, BigQuery, AWS, GCP, Terraform, Spacelift, Boomi
Beast Industries: Produces digital media and consumer goods for MrBeast brands.
8+ YOE8+ years building production data platforms, expertise in batch and streaming pipelines, data modeling, ETL/ELT, reliability, observability, and privacy-safe foundations for AI/ML.
Metriport: Provides an open-source API for healthcare data interoperability.
6+ YOE6+ years data engineering experience building and scaling pipelines; experience across ingestion, storage, processing, warehousing; production coding in TypeScript/Python; mentoring experience; located in or willing to relocate to San Francisco Bay Area.
Metriport: Provides an open-source API for healthcare data interoperability.
8+ YOE8+ years building and operating large-scale data platforms; experience with ingestion, storage, processing, warehousing, streaming, and leadership; strong software engineering and mentorship skills.
Physical Intelligence: Creating foundation models for general-purpose robot intelligence.
Strong software engineering fundamentals with experience building distributed systems, large-scale data pipelines, object storage and batch/streaming systems; ownership mindset and performance focus.
Experienced data engineer with production-grade pipeline experience for large-scale batch and streaming workloads; strong debugging, cost/performance optimization, and data quality/governance skills.
Spark, Flink, Beam, Airflow, Dagster, Kafka, PubSub, Parquet, Iceberg, Delta Lake, BigQuery, Snowflake, Great Expectations
Member of Technical Staff — Data Ingestion & Quality
San Francisco, California, United States
OnsiteFull Time
Causal Labs: Building physics-based causal AI models for predictive weather intelligence.
Experience building large-scale data pipelines and QA systems, familiarity with streaming and batch ingestion, vendor collaboration, and strong problem-solving and domain learning skills.
Hinge HealthNYSE: HNGE: Digital provider of musculoskeletal care and physical therapy
5+ YOE2+ Mgmt5+ years data engineering; 2+ years managing engineering teams; 2+ years building ML platform capabilities; experience with batch & streaming systems and tools such as Kafka, Flink, Spark; proficiency with Python, SQL, dbt, Databricks, and AWS.
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
6+ YOE6+ years in AI/ML engineering or applied data science; strong Python in production; ML model development and deployment; data pipelines (ETL/ELT, batch or streaming); APIs and backend systems; experience with LLM-powered systems and agent workflows; familiarity with Spark, Airflow/Dagster, Snowflake/BigQuery.
Duckbill: SaaS platform for enterprise cloud financial planning and analysis.
Proven experience building data systems and ETL for batch and streaming; strong Python and SQL; experience with data warehouses/lakehouses/OLAP, columnar databases, cloud storage, and data quality practices.
Member of Technical Staff (Software Engineer, Data Platform)
San Francisco or New York City or Palo Alto
$220k-$405k/yrOnsiteFull Time
Perplexity: AI-powered search engine providing conversational answers with citations.
5+ YOE5+ years software engineering experience (8+ for staff), production data infrastructure and batch/streaming experience, Airflow/Dagster, Python plus another backend language (Go/TypeScript), ML/AI workflow support, data quality and observability knowledge.
Verily: Developing data-driven technologies for clinical research and precision health.
6+ YOEBA/BS in CS or equivalent, 6+ years experience with Python and/or Java, SQL proficiency, experience with life science/biomedical data and batch workflows, strong communication and project management skills.
Python, Java, SQL, Google Cloud Platform, Amazon Web Services, Microsoft Azure, Docker, Linux
Beacon AI: Developing an AI-powered-pilot for safer flight operations.
Experience designing and operating AWS cloud and LLM/ML infrastructure, building data pipelines, and implementing security and observability for production systems.
Beacon AI: Developing an AI-powered-pilot for safer flight operations.
Experience building and operating AWS cloud and LLM infrastructure, CI/CD for infra and ML pipelines, data pipelines, security and observability for production systems.
The Walt Disney CompanyNYSE: DIS: Produces movies, operates theme parks, and provides streaming services.
5+ YOEExperience building and operating large-scale data/ML systems, strong distributed systems fundamentals, streaming/batch technologies (Kafka, Kinesis, Spark, Flink), Python, plus working knowledge of Java/Scala/Go/C++, cloud-native (AWS, containers, Kubernetes, IaC), observability, and collaboration skills; Bachelor's or Master's required.
Parallel: Build web infrastructure and search APIs for AI agents.
Deep experience in distributed data processing, data modeling, system reliability, batch and streaming pipelines, and building data quality, lineage, and observability systems.
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
10+ YOE10+ years software development experience building large-scale data platforms using Apache Spark, AWS S3, streaming/batch pipelines; strong distributed systems, API, and architecture expertise; MSc/BS or equivalent; Scrum/Kanban experience.