375 spark data engineer jobs at 179 companies in Plainfield, NJ
1mo
Save
Mark Applied
Hide
1mo
Data Engineer
New York City or San Francisco
OnsiteFull Time
Bedrock Robotics: Automates heavy construction machinery with AI retrofit kits.
5+ YOE5+ years data engineering experience in large-scale data lake/warehouse environments; strong SQL and distributed query engine experience; pipeline orchestration (Airflow/Prefect) and Spark/Databricks experience; data quality and observability expertise.
Sr Data Engineer, Python + Spark (Data Federation skillset - Data Lakehouse - Eg: Starburst) - New York
New York, New York, United States
$40k-$140k/yrOnsiteFull Time
Photon: Global technology services provider.
5+ YOESenior Data Engineer with 5+ years in Python, Spark; expertise in data federation, lakehouse architectures (Delta Lake/Iceberg/Hudi); Starburst/Trino/Dremio; cloud platforms (AWS/Azure/GCP); strong SQL and data modeling.
FanDuelNYSE: FLUT: Offers online sports betting and daily fantasy sports services.
3+ YOE3+ years in data or software engineering with strong SQL, Python/Java/Scala, Databricks, Airflow, dbt, Spark, Kafka, and cloud (AWS/GCP/Azure) experience.
United States or San Francisco or New York City or Washington or Austin
$118k-$179k/yrRemoteFull Time
SamsaraNYSE: IOT: Connected Operations Cloud platform for physical operations and IoT.
8+ YOEBachelor's in CS or equivalent, 8+ years as a software/data engineer, 5+ years building production data pipelines and Spark/PySpark experience, strong Python and SQL, cloud data warehouse and ETL tooling experience.
EXLNASDAQ: EXLS: Provides data analytics and digital operations solutions to businesses.
10+ YOE10+ years data engineering experience with Databricks, Spark, Delta Lake and Azure Data Platform; strong Python/PySpark and SQL skills; experience with CI/CD, REST APIs, CDC/SCD, and data quality frameworks.
Databricks, Spark, Delta Lake, Unity Catalog, Databricks Workflows, SQL Warehouses, PySpark, Spark SQL, SQL, Python, Azure Data Lake Storage (ADLS Gen2), Azure Synapse, Azure Key Vault, Azure DevOps, Git, Ab Initio, REST APIs, JSON, XML
VersantNasdaq: VSNT: Operates cable television networks and digital media entertainment platforms.
5+ YOE5+ years building production data pipelines with Spark-based platforms, expertise in PySpark and SQL, Databricks experience preferred, bachelor's in CS/data engineering or equivalent, strong CI/CD, orchestration, and data governance skills.
Delta Lake, Databricks Auto Loader, Databricks, Structured Streaming, Change Data Capture (CDC), Unity Catalog, Spark, Photon, Databricks Workflows, Delta Live Tables, Airflow, MLflow, Presto, Flink, Git, Dagster, EMR, Dataproc, Lakeflow Spark Declarative Pipelines (SDP), ChatGPT
CognizantNasdaq: CTSH: Provides global information technology and business process outsourcing services.
5+ YOE5+ years data engineering experience; strong Python, SQL, Spark; experience with Snowflake, DBT and AWS/Azure; exposure to real-time streaming (Kafka/Kinesis); collaborate with data science and analytics teams.
Quantexa: Develops decision intelligence software for data analytics and management.
8+ YOE5+ Mgmt8+ years in data engineering with 5+ years in lead roles; strong Scala/Java/Python skills; experience with Spark, Hadoop, Elasticsearch, cloud platforms, CI/CD tooling, production data pipelines, and mentoring teams.
Spark, Hadoop, Scala, Elasticsearch, Google Cloud, Microsoft Azure, Amazon, Java, Python, Git, Gradle, Nexus, Jenkins, Docker, Bash, ScalaTest
5+ YOE5+ years data engineering experience with Databricks, cloud-native data platforms, Python/SQL/Spark, GenAI/LLM experience, data pipeline and governance expertise for clinical/cross-study data.
Reality Defender: Detects deepfakes and AI-generated media to identify fraud.
Hands-on experience with Kubernetes and AWS, distributed computing (Spark, Ray), workflow orchestration (Airflow), Python and SQL; experience building multi-terabyte and streaming data pipelines for audio/video is preferred.
United States or Arlington or Tysons or Washington or New York City or Chicago or Austin or Atlanta or Boston or Boulder
$113k-$188k/yrRemoteFull Time
Guidehouse: Provides management and technology consulting services to diverse organizations.
3+ YOEBachelor's degree,3+ years data engineering experience,proficiency with Python/PySpark/SQL,Databricks/Spark/Delta Lake experience,knowledge of data pipelines,governance,and cloud platforms.
C the Signs: AI-powered clinical decision support platform for early cancer detection.
Bachelor's in CS/Engineering; data engineering experience; Python/Scala/Java; ETL, data warehousing, data modeling; cloud (AWS/GCP/Azure); Spark; strong problem-solving and communication.
Linero: AI-native, venture-backed building data infrastructure software.
5+ YOE5+ years data engineering; Airflow, dbt, Spark; streaming and batch processing; ontology design; multi-tenant, high-security data environments; based in or relocate to New York City.
3+ YOEBachelor's in CS/Math/related or equivalent; 3+ years with data processing (Hadoop, Spark, Pig, Hive), DB administration or data engineering, software experience in Java/C++/Python/Go/JavaScript, and client-facing project experience.
Balyasny Asset Management: Global multi-strategy investment firm managing diverse alternative asset classes.
3+ YOE3+ years data engineering experience; strong Python, SQL, Spark; Snowflake and cloud (AWS/Azure/GCP) experience; pipeline orchestration and data quality expertise.
Python, SQL, Spark, NoSQL, Snowflake, Hive, Hadoop, Airflow, Luigi, Oozie, NiFi, AWS, Azure, Google Cloud, Go
Vistar Media: Operates a programmatic marketplace for digital out-of-home advertising.
4+ YOE4+ years experience building scalable big-data pipelines using Python, Spark, SQL, cloud (GCP), Airflow and Kafka; strong software engineering, data architecture, and geospatial experience preferred.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
5+ YOE5+ years building analytical data models and production data systems; expertise in Python/Scala/Rust, advanced SQL, Spark/Flink, lakehouse technologies, and dimensional modeling for OLAP workloads.
Bristol Myers SquibbNew York Stock Exchange: BMY: Develops and distributes innovative medicines for serious diseases.
5+ YOE5+ years data engineering experience with cloud platforms, Databricks (Delta Lake, Unity Catalog), Python, SQL, Spark/PySpark, GenAI/LLM (RAG, embeddings), and building production-grade ETL/ELT pipelines and data products.
Solidus Labs: Provides market integrity and risk monitoring software for crypto.
8+ YOEBSc in Computer Science; 8+ years in data engineering; 5+ years software engineering with Java, Rust, or Python; deep ClickHouse expertise; experience with Kafka, Spark, Airflow, Kubernetes, Redis, Snowflake, Prometheus/Grafana, and advanced SQL.