584 spark data engineer jobs at 244 companies in California
2mo
Save
Mark Applied
Hide
2mo
Data Engineer
New York City or San Francisco
OnsiteFull Time
Bedrock Robotics: Autonomous construction technology that retrofits heavy equipment into autonomous machines for general contractors.
5+ YOE5+ years data engineering experience in large-scale data lake/warehouse environments; strong SQL and distributed query engine experience; pipeline orchestration (Airflow/Prefect) and Spark/Databricks experience; data quality and observability expertise.
PlusAI: Private Physical AI developing virtual driver software for factory-built autonomous trucks.
1+ YOEMS in CS/Electrical or related, 1+ years software engineering, expertise in Python and SQL, experience with large-scale data processing (Spark, MapReduce, Kafka), AWS (S3, EC2, RDS), and building/maintaining data pipelines.
IlluminaNASDAQ: ILMN: Global leader in DNA sequencing and genomic technologies.
10+ YOE10+ years of data engineering experience with Python, advanced SQL, data modeling, distributed systems, Databricks or Snowflake, Spark, dbt, cloud platforms, governance, and technical leadership.
Python, SQL, Databricks, Snowflake, Spark, Delta Lake, Apache Iceberg, dbt, Parquet, Unity Catalog, Git, REST APIs, JSON, CI/CD, AWS, SAP, Microsoft Power BI, Tableau
RokuNASDAQ: ROKU: TV streaming platform powering the global television ecosystem.
8+ YOE8+ years data engineering experience, strong SQL, Python preferred, expertise with big-data tech (Hadoop, Spark, Kafka, Hive, Airflow), data modeling, and cloud platforms (AWS/GCP).
California or Colorado or Connecticut or Florida or Georgia or Kentucky or Massachusetts or Michigan or New York or New Jersey or North Carolina or Pennsylvania or South Carolina or Tennessee or Texas or Virginia or Washington or Peachtree Corners
RemoteFull Time
Digital Envoy: Private U.S. technology providing IP geolocation, VPN intelligence, and online authentication data to businesses.
7+ YOERequires 7+ years of data engineering experience, 4+ years with Spark, strong Python and advanced SQL skills, Airflow familiarity, AWS data engineering experience, communication skills, and a bachelor's degree in a technical field.
United States or San Francisco or New York City or Washington or Austin
$118k-$179k/yrRemoteFull Time
SamsaraNYSE: IOT: Connected Operations Cloud platform for physical operations.
8+ YOEBachelor's in CS or equivalent, 8+ years as a software/data engineer, 5+ years building production data pipelines and Spark/PySpark experience, strong Python and SQL, cloud data warehouse and ETL tooling experience.
Anduril Industries: Defense technology developing AI-powered autonomous military systems.
3+ YOE3+ years in data engineering, strong Python and SQL skills, experience with Spark/PySpark, dbt, SQLMesh, Palantir Foundry, cloud platforms (AWS/Azure/GCP), data orchestration (Flyte), and data formats (Apache Iceberg).
The Walt Disney CompanyNYSE: DIS: A leading global entertainment and media conglomerate.
5+ YOE5+ years data engineering experience building large data pipelines; strong SQL, Python/PySpark; experience with Snowflake/Redshift, Databricks, Spark, Airflow, AWS; data modeling and performance tuning; Bachelor's degree or equivalent.
Versant Media GroupNasdaq: VSNT: Media and entertainment managing diverse content and digital brands.
5+ YOE5+ years building production data pipelines with Spark-based platforms, expertise in PySpark and SQL, Databricks experience preferred, bachelor's in CS/data engineering or equivalent, strong CI/CD, orchestration, and data governance skills.
Delta Lake, Databricks Auto Loader, Databricks, Structured Streaming, Change Data Capture (CDC), Unity Catalog, Spark, Photon, Databricks Workflows, Delta Live Tables, Airflow, MLflow, Presto, Flink, Git, Dagster, EMR, Dataproc, Lakeflow Spark Declarative Pipelines (SDP), ChatGPT
Hadrian: Manufactures precision components in automated factories.
Production data-model ownership, expert SQL, Spark, dbt, Dagster or equivalent, Python, data modeling, lake and warehouse internals, semantic layers, pipelines, testing, documentation, and CI/CD.
10+ YOERequires 10+ years with a bachelor's, 8+ with a master's, 6+ with a PhD, or equivalent; expertise in Java, distributed systems, Kafka, Data Lakes, Apache Iceberg, Apache Spark, databases, and data ingestion.
8+ YOE3+ Mgmt8+ years data engineering experience, 3+ years people management; proficiency in data modeling, ETL, Python, SQL; experience with Spark, HDFS, S3 and ETL schedulers (Airflow, Dagster, DBT); experimentation and ML familiarity.
Windfall: People intelligence and AI helping go-to-market teams find, engage, and convert customers.
4+ YOE4-8 years in data engineering; experience with Apache Beam/Spark/Flink or MapReduce; JVM language proficiency; distributed data processing; experience at a small company; strong communication and ownership.
TikTok: Short-form mobile video and social media platform.
BS/MS in CS or equivalent; experience with Hadoop, Hive, Spark, Presto, Kafka, ClickHouse, Flink; ETL, data ingestion, schema design and SQL; big data system architecture experience.
McLean or California or Colorado or Hawaii or Illinois or Maine or Maryland or Massachusetts or Minnesota or New Jersey or New York or Vermont or Virginia or Washington or District of Columbia or Cleveland
$130k-$265k/yrOnsiteFull Time
Accenture Federal ServicesNew York Stock Exchange: ACN: Technology and management consulting for U.S. federal agencies.
4+ YOE4+ years data lifecycle and ETL engineering experience; cloud and on-prem storage; Python, SQL, Spark; experience with ElasticSearch, NiFi, Palantir Foundry; Agile/Scrum experience; active TS/SCI with poly required.
United States or San Francisco or New York City or Chicago
$170k-$230k/yrHybridFull Time
Komodo Health: Privately held healthcare AI and real-world data analytics serving Life Sciences and healthcare organizations.
Experience building production data pipelines at scale with advanced Python, SQL, Airflow, Spark, and AWS; healthcare data expertise and strong data quality, reliability, troubleshooting, and collaboration skills required.
Monks: Global digital-first marketing and technology services.
5+ YOERequires 5+ years of data engineering experience, expert SQL, Python, and experience with Airflow, Spark, Trino/Dremio, Apache Iceberg, Kafka, Docker, ETL/ELT, Git, CI/CD, lakehouse architectures, and production pipelines.
Superhuman: Private AI productivity platform for people and teams, combining writing, collaborative docs, email, and proactive AI agents.
3+ YOERequires 3+ years building production data pipelines, proficiency in SQL, Python, Spark, and modern data platforms, plus experience with machine learning workflows, data modeling, orchestration, CI/CD, and data quality.
Spark, Databricks, Google, Meta, LinkedIn, SQL, Python, Delta Lake, dbt, Snowflake, Databricks Workflows, Airflow, Git, Claude Code, Codex, Google Ads
SHEIN: Private global online fashion and lifestyle retailer serving consumers with affordable apparel and everyday products.
4+ YOEBachelor's in Data Science/Computer Science + 4 years data engineering experience building large-scale ETL pipelines with Hive/Presto/Spark/Flink, data warehousing (Redshift), SQL, Amazon S3, AWS EMR, Airflow; on-call participation.
Sapiom: The platform builders use to ship, run, and scale AI agents.
5+ YOE5+ years building production data pipelines; hands-on SQL, Python, Spark, AWS Glue, EMR, DBT, Airflow; 3+ years with MPP databases (Snowflake/Redshift/Teradata); on-call experience and strong cross-team communication.