206 spark data engineer jobs at 119 companies in Vacaville, CA
1mo
Save
Mark Applied
Hide
1mo
Data Engineer
New York City or San Francisco
OnsiteFull Time
Bedrock Robotics: Automates heavy construction machinery with AI retrofit kits.
5+ YOE5+ years data engineering experience in large-scale data lake/warehouse environments; strong SQL and distributed query engine experience; pipeline orchestration (Airflow/Prefect) and Spark/Databricks experience; data quality and observability expertise.
Stitch FixNASDAQ: SFIX: Provides personalized apparel styling services through algorithms and stylists.
2+ YOEBachelor's in engineering or computer science, 2+ years data engineering experience, familiarity with Spark, dbt, Fivetran, Airflow, S3; production Python and SQL; experience with AI coding agents; strong communication.
Spark, dbt, Fivetran, Airflow, S3, Claude Code, Codex, Python, SQL
LendingClubNYSE: LC: Digital marketplace bank providing personal loans and banking services.
8+ YOE2+ Mgmt8+ years data engineering experience with 2+ years technical lead experience; strong SQL, data modeling, ETL, orchestration, Spark, AWS, modern data platforms, and experience applying AI tools to engineering workflows.
United States or San Francisco or New York City or Washington or Austin
$118k-$179k/yrRemoteFull Time
SamsaraNYSE: IOT: Connected Operations Cloud platform for physical operations and IoT.
8+ YOEBachelor's in CS or equivalent, 8+ years as a software/data engineer, 5+ years building production data pipelines and Spark/PySpark experience, strong Python and SQL, cloud data warehouse and ETL tooling experience.
Robert HalfNYSE: RHI: Provides specialized staffing and global business consulting services.
5+ YOE5+ years Python/SQL data engineering; 3+ years cloud AWS; ETL/ELT pipelines; Docker/Kubernetes; Terraform/CloudFormation; Spark; CI/CD; collaboration with data science and MLOps; mentoring.
Windfall: Provides consumer financial data and AI-driven go-to-market insights.
4+ YOE4-8 years in data engineering; experience with Apache Beam/Spark/Flink or MapReduce; JVM language proficiency; distributed data processing; experience at a small company; strong communication and ownership.
United States or San Francisco or New York City or Chicago
$170k-$230k/yrHybridFull Time
Komodo Health: Provides AI-driven healthcare data and patient journey analytics software.
Experience building production data pipelines at scale with advanced Python, SQL, Airflow, Spark, and AWS; healthcare data expertise and strong data quality, reliability, troubleshooting, and collaboration skills required.
The Walt Disney CompanyNYSE: DIS: Produces media content and operates global theme parks.
5+ YOEBachelor's or equivalent experience, 5+ years big data engineering, strong Python/Scala/SQL, Spark/Presto/Hive, Databricks and cloud MPP databases, Airflow, AWS S3, CI/CD, on-call and production operations experience.
OpenAI: Develops artificial intelligence models and generative AI software services.
3+ YOE3+ years data engineering; 8+ years software engineering; Python/Scala/Java; Databricks, Snowflake; ETL schedulers; Spark/Hadoop/Flink; S3/HDFS; strong data pipelines and collaboration.
Sapiom: Financial infrastructure for autonomous AI agents.
5+ YOE5+ years building production data pipelines; hands-on SQL, Python, Spark, AWS Glue, EMR, DBT, Airflow; 3+ years with MPP databases (Snowflake/Redshift/Teradata); on-call experience and strong cross-team communication.
Checkr: AI-powered platform for background checks and identity verification.
10+ YOE10+ years building scalable data platforms; expert PySpark, Python, SQL; experience with Kafka, Spark, Iceberg, data lakes, AWS; strong data modeling and security awareness.
Tatari: A platform for buying and measuring TV advertising campaigns.
5+ YOE5+ years building and operating production ETL pipelines; strong Python and SQL skills; Spark/PySpark, Databricks/Delta Lake, Airflow experience; strong data modeling and data quality practices.
Python, SQL, Spark, PySpark, Databricks, Delta Lake, Airflow, ClickHouse
Foundry Robotics: An AI-native robotics manufacturer focused on deploying advanced assembly and production capability for robotics companies and national-security-critical hardware.
5+ YOE5-7 years data engineering experience; TB–PB-scale pipelines; Spark/Flink/Kafka/Airflow/dbt; Python/SQL; Go/Java/TypeScript a plus; cloud data services; architecture and governance; independent, end-to-end ownership.
Scottsdale or Chicago or San Francisco or New York City
$97k-$149k/yrHybridFull Time
Early Warning Services: Operates payment and risk solutions for the financial industry.
2+ YOEBachelor's degree typical; 2+ years experience building data engineering tools with Spark and Python, experience with ML toolkits, Hadoop or AWS/S3, and relational databases; strong communication and background/drug screen required.
AVEVA: Industrial software for engineering and operational performance management.
3+ YOEBachelor's in a STEM field,minimum 3 years data engineering experience,proficiency with data pipelines,ETL/ELT,SQL,and cloud data platforms (Azure preferred).
Microsoft Azure, Azure Synapse Analytics, Azure Data Factory, Azure Data Lake, Python, Spark, SQL
Austin or Boston or Charleston or Charlotte or Chicago or Dallas or Durham or Harrisburg or Houston or Irvine or Kansas City or Los Angeles or Miami or Nashville or New York or Newark or Palo Alto or Pittsburgh or Portland or Raleigh or San Francisco or Seattle or Washington or Wilmington
$128k-$249k/yrHybridFull Time
K&L Gates: Global law firm providing comprehensive legal and regulatory counsel.
5+ YOE5+ years designing enterprise Microsoft Fabric data platforms; strong SQL and Python skills; experience with Spark, CI/CD, Azure DevOps/GitHub Actions; knowledge of data governance and Azure AI Foundry; Bachelor's degree or equivalent.
Microsoft Fabric, OneLake, Fabric Data Factory, Dataflow Gen2, Fabric Notebooks, Azure AI Foundry, Microsoft Copilot, Claude, SQL, Python, Spark, Azure DevOps, GitHub Actions, Microsoft Purview
Houston or Golden or Reno or Oakland or Salt Lake City
HybridFull Time
Fervo Energy: Generates clean electricity using advanced geothermal drilling technology.
2+ YOE2+ years building and operating production data pipelines with Apache Spark, Python/SQL, Databricks, Azure, and Snowflake; experience with streaming, data modeling, data quality, governance, and CI/CD.
Databricks, Delta Lake, Delta Live Tables, Unity Catalog, Databricks Workflows, Databricks SQL, Apache Spark, PySpark, Spark SQL, Structured Streaming, Kafka, Microsoft Event Hubs, Microsoft Azure Data Factory, Microsoft Azure Data Lake Storage (ADLS), Snowflake, Microsoft Power BI, Spotfire, Python, SQL, Git, Azure DevOps, GitHub Actions, dbt, Terraform, Docker, Airflow, Microsoft Entra ID, Microsoft Key Vault, MQTT, OPC UA, SparkplugB, Canary, Ignition, Snowflake Semantic Views, Databricks Metric Views
Arbital Health: Provides infrastructure for adjudicating healthcare value-based care contracts.
5+ YOE5+ years building data-intensive SaaS platforms, deep Spark and distributed processing expertise, strong SQL and data modeling, experience with Airflow, Terraform, Databricks, dbt/Great Expectations, healthcare data and HIPAA, Python proficiency.
Sakata Seed AmericaTokyo Stock Exchange: 1377: Breeding and distributing high-quality vegetable and flower seeds.
5+ YOEDesign and maintain scalable ETL/ELT pipelines, data models, and Microsoft Fabric solutions; strong SQL and Python skills; 5+ years data engineering experience; knowledge of data governance and monitoring.
Microsoft Fabric, Data Factory, Lakehouse, Warehouse, Dataflow Gen2, notebooks, semantic models, SQL Server, PostgreSQL, MySQL, SQL, Python, REST APIs, Spark
PRISM: Provides risk management and insurance for public entities.
Proven data engineering experience designing pipelines and models; strong Python, Spark SQL, T-SQL, and MS SQL Server skills; experience with Microsoft Fabric, Power BI, Azure DevOps, CI/CD, and data governance preferred.
Microsoft Fabric, Fabric Lakehouse, Fabric Data Factory, Fabric Data Warehouse, Fabric Notebooks, Power BI, Python, Spark SQL, T-SQL, MS SQL Server, OneLake, Azure DevOps, Qlik, CI/CD, Sparkhire