174 pyspark data engineer jobs at 108 companies in New York
1w
Save
Mark Applied
Hide
1w
Sr. Associate, Date Engineer - PySpark
Atlanta or New York City or Philadelphia or Dallas
OnsiteFull Time
KPMG: Global professional services network providing audit, tax, and advisory.
5+ YOE5+ Mgmt5+ years leading data science/engineering teams, Master’s degree in CS/DS/Statistics/Engineering, hands-on Python, PySpark, SQL, Databricks, Snowflake, experience with cloud platforms, ability to travel, U.S. work authorization required.
Python, PySpark, SQL, Databricks, Snowflake, Linux, AWS, GCP, Azure, R
News CorpNasdaq: NWSA: Global media, publishing, and digital information services.
2+ YOEBachelor's or master's degree in computer science or related field, 2+ years of software engineering experience, and proficiency in Python, SQL, PySpark, cloud services, data pipelines, and data warehouses.
Vertex AI, AWS Lambda, Amazon DynamoDB, Google Cloud Functions, Amazon S3, Kubernetes, AWS Glue, Google BigQuery, Amazon Web Services (AWS), Google Cloud Platform (GCP), Python, SQL, PySpark, Apache Airflow, Google Cloud Dataflow, Snowflake, Amazon Redshift, Git, JIRA, Linux, Shell, SSH, crontab, CloudWatch, Datadog, Splunk, REST, GraphQL, Docker
Regard: A that builds AI-powered software to improve clinical documentation and advance patient care.
3+ YOEBS or equivalent, 3+ years data engineering experience, Python and SQL proficiency, experience with PySpark and distributed data processing, AWS experience preferred, LLM-assisted development familiarity, and willingness to participate in on-call support.
EXLNASDAQ: EXLS: Provides data analytics and digital operations solutions to businesses.
4+ YOERequires 4+ years in data engineering, a bachelor's or master's degree, SQL, Python, PySpark, cloud data platforms, ETL/ELT pipelines, data modeling, Airflow, cloud ecosystems, and engineering team leadership.
SQL, Python, PySpark, Snowflake, Databricks, Apache Airflow, AWS, Azure, GCP, Tableau, Microsoft Power BI, Looker, Spark, Hadoop, Hive, HBase, Kafka
5+ YOE5+ years data engineering experience, strong dbt and Databricks skills, Python/PySpark, streaming and CDC ingestion, data quality and observability, experience with operational integrations.
dbt, Databricks, Python, PySpark, Delta Lake, Delta Live Tables, Airflow, Cloud Run, Cloud Functions, BigQuery, GCP, AWS, Salesforce, NetSuite, Stripe, Mulesoft, Great Expectations, Unity Catalog, Terraform, Spacelift, Boomi
New York PostNASDAQ: NWSA: Daily newspaper delivering news, sports, and entertainment.
2+ YOEMS or BS in computer science or related field,2+ years software engineering experience,proficiency in Python,SQL,PySpark,experience with AWS/GCP and data pipeline tools,strong problem-solving and collaboration skills.
NYSTEC: Provides technology consulting services to New York public agencies.
5+ YOEDesign and maintain enterprise data pipelines, warehouses/lakehouses, ETL/ELT processes using Microsoft data services, SQL, Python/PySpark; implement data governance and optimize platform performance.
Microsoft Fabric, Microsoft Azure Data Factory, Azure Data Lake Storage, Azure SQL, SQL Server Integration Services (SSIS), SQL Server, SQL, Python, PySpark
United States or San Francisco or New York City or Washington or Austin
$118k-$179k/yrRemoteFull Time
SamsaraNYSE: IOT: Connected Operations Cloud platform for physical operations and IoT.
8+ YOEBachelor's in CS or equivalent, 8+ years as a software/data engineer, 5+ years building production data pipelines and Spark/PySpark experience, strong Python and SQL, cloud data warehouse and ETL tooling experience.
New York Blood Center: Non-profit organization collecting and distributing blood and stem cells.
6+ YOE6+ years data engineering experience; bachelor’s in CS/Data Science/IT or related; expert SQL and Python (PySpark); deep Azure data platform experience (ADF, Databricks, Synapse, ADLS); data modeling, pipeline ownership, governance, and HIPAA compliance.
SQL, Python, PySpark, Azure Data Factory, Azure Databricks, Azure Synapse Analytics, Azure Data Lake Storage, Microsoft Purview, Microsoft SQL Server, Oracle, CI/CD
Florida or Montana or Michigan or Massachusetts or Rhode Island or West Virginia or Oregon or Oklahoma or Minnesota or New York or South Carolina or Kentucky or Virginia or Colorado or New Jersey or Wisconsin or Utah or Pennsylvania or Ohio or North Carolina or Nebraska or Mississippi or South Dakota or Maryland or Louisiana or Kansas or Iowa or Illinois or Tennessee or Delaware or Connecticut or California or Hawaii or Arkansas or District of Columbia or Georgia or Missouri or Arizona or Alabama or New Hampshire or Texas or Vermont or Indiana or Washington or Nevada or Idaho or New Mexico or Maine
$143k-$160k/yrRemoteFull Time
AmeriLife: Distributes insurance, annuities, and retirement solutions for retirees.
7+ YOE7+ years data engineering experience with 3+ years Databricks; strong Spark/PySpark, Delta Lake, SQL, Azure, Git and CI/CD skills; Bachelor's degree required.
CapgeminiEuronext Paris: CAP: Provides global IT consulting and digital transformation services.
Design Lakehouse architectures on Databricks, develop PySpark jobs, implement governance with Unity Catalog, work with Airflow and CI/CD; strong data governance and streaming knowledge.
United States or Arlington or Tysons or Washington or New York City or Chicago or Austin or Atlanta or Boston or Boulder
$113k-$188k/yrRemoteFull Time
Guidehouse: Provides management and technology consulting services to diverse organizations.
3+ YOEBachelor's degree,3+ years data engineering experience,proficiency with Python/PySpark/SQL,Databricks/Spark/Delta Lake experience,knowledge of data pipelines,governance,and cloud platforms.
D'Addario: Manufacturer of musical instrument strings and accessories.
3+ YOE3+ years data engineering experience; strong Python, PySpark, and SQL; building ELT/ETL pipelines, analytics data modeling, Microsoft Fabric/Power BI familiarity; Bachelor's in CS/Data or equivalent.
Python, PySpark, SQL, Microsoft Fabric, Microsoft Power BI, Azure Synapse, Azure Data Lake, dbt
Tatari: A platform for buying and measuring TV advertising campaigns.
5+ YOE5+ years building and operating production ETL pipelines; strong Python and SQL skills; Spark/PySpark, Databricks/Delta Lake, Airflow experience; strong data modeling and data quality practices.
Python, SQL, Spark, PySpark, Databricks, Delta Lake, Airflow, ClickHouse
VersantNasdaq: VSNT: Operates cable television networks and digital media entertainment platforms.
5+ YOE5+ years building production Spark-based data pipelines in cloud environments; expertise in PySpark and SQL; experience with Databricks, orchestration (Airflow/Dagster), CI/CD, EMR/Dataproc, data governance, and collaboration using Git.
St. Louis or Brentwood or United States or Alaska or Arizona or Arkansas or California or Connecticut or Delaware or Hawaii or Idaho or Louisiana or Maine or Massachusetts or Michigan or Mississippi or Montana or Nebraska or Nevada or New Mexico or New York or North Carolina or North Dakota or Oregon or Rhode Island or South Carolina or South Dakota or Texas or Utah or Vermont or Washington or West Virginia or Wyoming
HybridFull Time
Navitus Health Solutions: Transparent pharmacy benefit management and specialty pharmacy provider.
8+ YOEBachelor's in CS or related (Master's preferred); AWS/Azure/CDMP certifications; 8+ years data engineering experience, 5+ years building lakehouse solutions on Azure Databricks; expertise in Spark, PySpark, SQL, Python, DataOps, CI/CD, and cloud data services.
Munich or New York City or Denver or Chicago or London or Heidelberg or Singapore
OnsiteFull Time
CEPRES: Digital investment analytics platform for private capital markets.
2+ YOE2+ years data engineering experience; strong Python, Pyspark and SQL skills; ETL/ELT and data pipeline architecture; exposure to cloud (AWS/Azure/GCP), Git and CI/CD; good communication and problem-solving.
Movable Ink: Provides AI-powered content personalization for digital marketing campaigns.
5+ YOE5+ years data engineering experience with Apache Spark/PySpark, GCP Dataproc, Python, Parquet/Delta Lake, Kafka, and cloud/CI/CD tooling; strong optimization and collaboration skills.
Apache Spark, PySpark, GCP Dataproc, Parquet, Delta Lake, Kafka, Google Cloud Platform, Python, Docker, Kubernetes, GitHub Actions, Git