Principal Machine Learning Engineer, Accelerated Apache Spark
Santa Clara, California, United States
$272k-$431k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOESenior ML/DS engineer with 12+ years of experience; 5+ years leading ML model development; strong Python and ML tooling; experience with GPU-accelerated Spark; expertise in ML methods and feature engineering.
San Francisco or Seattle or Sunnyvale or New York City
$131k-$285k/yrHybridFull Time
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
24+ YOERequires a CS degree or equivalent, 24+ years operating production distributed systems, Spark at scale, Kubernetes, cloud infrastructure, observability, and professional Python, Go, Scala, or Java with SQL fluency.
San Francisco or Sunnyvale or Seattle or New York City
$131k-$285k/yrHybridFull Time
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
24+ YOEExperience operating production distributed systems and Apache Spark at scale, Kubernetes production experience, familiarity with schedulers and observability stacks, cloud (AWS) experience, programming in Python/Go/Scala/Java, and SQL fluency. BS/MS/PhD in CS or equivalent.
Principal Machine Learning Engineer, Accelerated Apache Spark
Santa Clara, California, United States
$272k-$431k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOEBS/MS/PhD or equivalent; 12+ years ML/DL experience; 5+ years as technical lead; 2+ years with Apache Spark; strong Python and data-science libraries experience; expertise in LLM/GenAI, RL, XGBoost; leadership and deployment experience.
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years software engineering with 2+ years building production LLM/agentic applications; deep Python fluency; experience with agentic frameworks, REST/GraphQL integrations, Databricks/Spark, and technical leadership.
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
10+ YOE10+ years building large-scale distributed systems or ML infrastructure, 3+ years at staff/principal level, experience with Spark, Ray, Kubernetes, Airflow, MLflow, web frameworks, reliability engineering, and mentoring engineers.
5+ YOEBachelor's degree or equivalent experience, 5 years of software development, and 3 years developing large-scale infrastructure, distributed systems, networks, compute, storage, or hardware architecture.
OpenTelemetry, JMX, Spark, Ray, Flink, Presto, Google Cloud
Cerebras SystemsNasdaq: CBRS: Cerebras Systems manufactures specialized computer chips designed for AI.
3+ YOEMaster's degree in EE/CE/CS or related +3 years experience; expertise in large-scale data pipelines, Hive/Spark, Python/SQL, Tableau, Linux, ML for hardware reliability and performance; strong statistical and A/B testing skills.
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
5+ YOE5+ years applied machine learning experience; strong statistical, optimization, and experimental methodology knowledge; proficient in Spark, Python or Java; experience with real-time, low-latency systems and distributed ML frameworks.
Eightfold: An AI-native enterprise talent platform that builds scalable machine learning solutions and agentic AI for talent workflows.
6+ YOEDeep expertise in ML/DL/NLP and LLMs; technical leadership and team mentoring; proficiency with Python, TensorFlow, PyTorch, big data (Hadoop, Spark); experience deploying ML at scale. MS/PhD and 6+ years preferred.
Wayve: Develops AI software for autonomous vehicle navigation.
10+ YOE10+ years building large-scale distributed systems or ML infrastructure, 3+ years at staff/principal level, experience with Spark, Ray, Kubernetes, Airflow, MLflow, reliability/observability, mentoring, and optimization or scheduling systems.
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Apply SRE principles to mentor teams, ensure reliability for large-scale analytics infrastructure across Hadoop, HBase, Spark, Data Lakes, and Airflow; participate in production on-call.
Livingston or New York City or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in platform or infrastructure engineering with Kubernetes, CI/CD, IaC, observability, multi-region systems, and production ownership for high-availability services.
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
2+ YOEDesign and operate large-scale data pipelines on GCP using Spark, Airflow, BigQuery; optimize cloud cost; build automation and AI-driven engineering tools; 2+ years data engineering experience.
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
Experience building end-to-end ML deployment and low-latency inference infra for real-time ad systems; handling large-scale data with Spark; forecasting ad metrics; building scalable simulations; cross-functional collaboration.
Qcells: Provider of solar modules, energy storage, and EPC services.
10+ YOE10+ years in data engineering or architecture, deep SQL and Python, Azure data services experience, distributed processing (Spark), data modeling, governance, and leadership experience.
Azure Fabric, Data Lake, Data Factory, Synapse, NetSuite, SAP, Salesforce, APIs, Delta Lake, Delta Tables, Snowflake, Kafka, Event Hub, Spark, Python, SQL