247 spark data engineer jobs at 91 companies in Live Oak, CA
🚀PromotedHiringCafe
Founding Machine Learning / AI Search Engineer
Cupertino, CA, US
$160k-$310k/yrOn-SiteFull Time
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Build the ML and AI search behind HiringCafe — ranking, recommenders, retrieval, and LLM agents that surface jobs people would never find on their own.
Qcells: Provider of solar modules, energy storage, and EPC services.
10+ YOE10+ years in data engineering or architecture, deep SQL and Python, Azure data services experience, distributed processing (Spark), data modeling, governance, and leadership experience.
Azure Fabric, Data Lake, Data Factory, Synapse, NetSuite, SAP, Salesforce, APIs, Delta Lake, Delta Tables, Snowflake, Kafka, Event Hub, Spark, Python, SQL
PlusAI: AI-based virtual driver software for factory-built autonomous trucks.
1+ YOEMS in CS/Electrical or related, 1+ years software engineering, expertise in Python and SQL, experience with large-scale data processing (Spark, MapReduce, Kafka), AWS (S3, EC2, RDS), and building/maintaining data pipelines.
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
2+ YOEDesign and operate large-scale data pipelines on GCP using Spark, Airflow, BigQuery; optimize cloud cost; build automation and AI-driven engineering tools; 2+ years data engineering experience.
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
8+ YOE8+ years data engineering experience, strong SQL, Python preferred, expertise with big-data tech (Hadoop, Spark, Kafka, Hive, Airflow), data modeling, and cloud platforms (AWS/GCP).
PayPalNASDAQ: PYPL: Global digital payments platform for consumers and merchants.
8+ YOEMaster's in CS/Engineering plus 8 years (or Bachelor's plus 10 years). Requires data engineering, Python, Shell, SQL, ETL, BigQuery, GCP, Kafka, Airflow, Spark/Hadoop, BI tools, data modeling experience.
Python, Shell, SQL, Tableau, ThoughtSpot, Google BigQuery, Google Cloud Platform, Erwin, Apache Kafka, Apache Airflow, Automic UC4, Spark, Hadoop
Lam ResearchNASDAQ: LRCX: Designs and manufactures wafer fabrication equipment for the semiconductor industry.
8+ YOEBachelor's degree in CS/CE/Information Systems, 8+ years in data engineering/architecture, strong Microsoft Azure experience, proficiency in Python, SQL, Spark, experience building production-grade data pipelines and modern data architectures for AI/ML.
Microsoft Azure, Microsoft Fabric, Python, SQL, Spark, Azure Databricks, Azure Data Factory, Synapse, Neo4j, Azure AI Search, CI/CD
TikTok: Global short-form video hosting and social media platform.
BS/MS in CS or equivalent; experience with Hadoop, Hive, Spark, Presto, Kafka, ClickHouse, Flink; ETL, data ingestion, schema design and SQL; big data system architecture experience.
Boston ScientificNYSE: BSX: Manufacturer of interventional medical devices and technologies.
13+ YOEBachelor's in CS or related, 13+ years data engineering with 6+ years building AWS cloud-native data platforms; expertise with AWS data services, Snowflake, Python, Terraform, Spark/Flink, and governance automation.
3+ YOEBachelor's in CS/Math/related or equivalent; 3+ years with data processing (Hadoop, Spark, Pig, Hive), DB administration or data engineering, software experience in Java/C++/Python/Go/JavaScript, and client-facing project experience.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
5+ YOE5+ years building analytical data models and production data systems; expertise in Python/Scala/Rust, advanced SQL, Spark/Flink, lakehouse technologies, and dimensional modeling for OLAP workloads.
Boston ScientificNYSE: BSX: Developer and manufacturer of innovative medical devices and therapies.
13+ YOEBachelor's degree required; 13+ years data engineering experience including 6+ years designing cloud-native AWS data platforms and 4+ years building scalable data platform solutions. Expertise with AWS data services, Snowflake, Python, Terraform, Bash, Spark/Flink, and governance automation.
VisaNYSE: V: Global payment technology facilitating electronic funds transfers.
2+ YOE2+ years experience with a Bachelor's (or Advanced) degree; strong programming in Java/Python/Go/Scala; experience building scalable data pipelines, Spark/Kafka/Hadoop ecosystem, cloud platforms, GenAI and data modeling.
Chartmetric: Provides a global music data analytics platform for professionals.
6+ YOE6+ years building production data systems; expert in Python, PostgreSQL, Airflow, Clickhouse, Snowflake, Elasticsearch, Spark, and AWS; Bachelor's in CS/Data Engineering or equivalent experience.
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Experience building large-scale production data pipelines, designing storage and schema for timeseries and warehouse data, proficiency with DBT and data processing tools, strong software engineering skills, and startup experience.
Johnson & JohnsonNYSE: JNJ: Global healthcare providing pharmaceuticals and medical technologies.
7+ YOEBachelor's/Master's in CS/Engineering/Data Science, 7+ years data engineering (5+ for exceptional), Databricks/Spark/Delta Lake, Python/PySpark/SQL, AWS S3, OT integrations (OPC UA, MQTT, MES), data governance and GxP/CSV compliance.
Senior Data Engineer- Large Driving Model, Autonomy
Palo Alto, California, United States
$179k-$224k/yrOnsiteFull Time
RivianNASDAQ: RIVN: Designs and manufactures electric vehicles and charging networks.
5+ YOEB.S./M.S. in CS/data/engineering, 5+ years as a data/ML/motion planning engineer in autonomous driving or robotics; advanced Python and SQL; experience with Spark or Ray, cloud infrastructure, dataset curation, large-scale sensor data, and production-grade data pipelines.