46 data lake engineer jobs at 33 companies in Wanaque, NJ
2mo
Save
Mark Applied
Hide
2mo
Staff Data Engineer- Data Lake
New York, New York, United States
$170k-$190k/yrHybridFull Time
H1: Data platform for healthcare professional insights and drug development.
8+ YOE8+ years building and scaling distributed data platforms; technical leadership and mentoring; strong Python (PySpark), Java/Scala, advanced SQL, Apache Spark, cloud-native big data (AWS EMR/Glue/S3/Athena/Redshift), orchestration (Argo/Airflow), containerization, observability, and experience with large-scale ingestion and data quality.
Bedrock Robotics: Automates heavy construction machinery with AI retrofit kits.
5+ YOE5+ years data engineering experience in large-scale data lake/warehouse environments; strong SQL and distributed query engine experience; pipeline orchestration (Airflow/Prefect) and Spark/Databricks experience; data quality and observability expertise.
Municipal Credit Union: Provides banking and financial services to members in New York.
6+ YOEBachelor's degree or equivalent with 6-8 years experience, strong SQL and data pipeline skills (Fabric, Azure Data Factory), cloud data lake/warehouse experience (Azure Synapse/Cosmos), .NET/C#, and data quality/ETL expertise.
Twenty: Develops autonomous systems for offensive cyber operations.
8+ YOE8+ years in data engineering/architecture; expertise in ETL pipelines, data lakes, schema/index design, columnar databases; proven leadership; US citizenship.
SMBC GroupNew York Stock Exchange: SMFG: Provides global banking, investment, securities, and consumer finance services.
Experience designing Databricks solutions in AWS, building production ETL/ELT pipelines, and using Python, SQL, Spark, Delta Lake, CI/CD, Terraform, and data governance controls.
Socotec: Provides testing, inspection, and certification services for building safety.
7+ YOE7-10 years data engineering experience with Databricks, Spark, Delta Lake, strong SQL and Python skills, enterprise ETL/ELT, experience with MDM, cloud data platforms, and data architecture.
Tatari: A platform for buying and measuring TV advertising campaigns.
5+ YOE5+ years building and operating production ETL pipelines; strong Python and SQL skills; Spark/PySpark, Databricks/Delta Lake, Airflow experience; strong data modeling and data quality practices.
Python, SQL, Spark, PySpark, Databricks, Delta Lake, Airflow, ClickHouse
BarclaysLondon Stock Exchange: BARC: Global bank providing retail, corporate, and investment financial services.
Experience building Java-based data pipelines, data warehouses and lakes; strong skills in Java, Spark, Kafka, Kubernetes, OAuth2/JWT, JUnit, Mockito, GitLab/Bitbucket; experience leading technical teams and ensuring data governance and observability.
Long Island City or Queens or New York or New Jersey or Connecticut
$125k-$128k/yrHybridFull Time
The Fund for Public Health in New York City: Incubates and manages public health programs for New York City.
3+ YOERequires 3+ years with Python, SQL databases, models, and data lake platforms; 2+ years with cloud technologies; undergraduate degree or certificate in a relevant quantitative or computing field.
Appgate: Provides zero trust cybersecurity and AI-driven fraud prevention solutions.
Extensive experience building and operating large-scale data platforms and lakes; hands-on expertise with Apache Spark and Apache Flink; strong production engineering skills including Kubernetes and CI/CD; experience with batch and streaming pipelines.
CapgeminiEuronext Paris: CAP: Provides global IT consulting and digital transformation services.
Bachelor's in CS or equivalent; experience building AWS data lakes/warehouses and ingestion pipelines; Python, Shell, SQL; AWS services knowledge; CI/CD and streaming experience; AWS certs preferred.
Movable Ink: Provides AI-powered content personalization for digital marketing campaigns.
5+ YOE5+ years data engineering experience with Apache Spark/PySpark, GCP Dataproc, Python, Parquet/Delta Lake, Kafka, and cloud/CI/CD tooling; strong optimization and collaboration skills.
Apache Spark, PySpark, GCP Dataproc, Parquet, Delta Lake, Kafka, Google Cloud Platform, Python, Docker, Kubernetes, GitHub Actions, Git
EXLNASDAQ: EXLS: Provides data analytics and digital operations solutions to businesses.
6+ YOE6+ years data engineering experience with AWS (Glue, Redshift, S3, Lambda, EMR/Spark, Kinesis, Athena, RDS, API Gateway), Python and SQL proficiency, experience with data lakes/warehouses, ETL/ELT, DevOps/CI/CD, and leadership/mentoring skills.
Sr Data Engineer, Python + Spark (Data Federation skillset - Data Lakehouse - Eg: Starburst) - New York
New York, New York, United States
$40k-$140k/yrOnsiteFull Time
Photon: Global technology services provider.
5+ YOESenior Data Engineer with 5+ years in Python, Spark; expertise in data federation, lakehouse architectures (Delta Lake/Iceberg/Hudi); Starburst/Trino/Dremio; cloud platforms (AWS/Azure/GCP); strong SQL and data modeling.
Chemnitz or New York City or Berlin or London or Sydney or Tokyo or Prague or Minneapolis or Saint Paul
HybridFull Time
Staffbase: Provides an AI-powered internal communications platform for employees.
Proven Staff Engineer-level experience in data engineering; strong architecture skills for data lakes, foundations, governance and orchestration; hands-on with streaming (Kafka) and Python; data modeling, stakeholder management, and people leadership.
Fortitude Re: Provider of legacy reinsurance solutions for insurance companies.
7+ YOEBachelor's in Math/CS/Engineering, 7+ years data engineering (5+ years programming in Python and SQL), strong ETL/ELT, APIs, cloud and data operations experience, knowledge of data lakes/governance, analytical and communication skills.
Senior Software Engineer, Server Fleet Infrastructure
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$153k-$204k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
4+ YOEBachelor's degree, 4+ years data engineering experience, strong SQL, Python/Java/Scala, experience with data lakes, ETL pipelines, orchestration and observability tools (Airflow, Spark), and cloud platforms.
RobinhoodNASDAQ: HOOD: Provides a commission-free platform for investing and financial services.
5+ YOE5+ years building backend or distributed data systems; proficiency with Java/Scala/Python, Kafka/Flink/Spark, Postgres and CDC, AWS S3, and Kubernetes.
ParamountNASDAQ: PSKY: Produces and distributes media content across global entertainment platforms.
Proven experience leading data engineering teams; expertise in data lake/warehouse/lakehouse architectures, ETL/ELT, streaming (Kafka), cloud-native systems, and production reliability.
Proven leadership of data/infra teams, deep experience building web-scale acquisition and ingestion systems, hands-on with distributed compute, orchestration, and data lake formats; familiarity with LLM training and experimentation.