Materra
Posted 1mo ago

Machine Learning Data Engineer (DataOps), Materra

Materra
Mountain View, California, United States
$166k-$244k/yrOnsiteFull Time
Responsibilities
  • building pipelines
  • implementing validation
  • managing versioning
Requirements
  • 3+ years building scalable data pipelines
  • Strong Python and SQL skills
  • Experience with data validation
  • Dataset versioning, and ML data lifecycle for annotation and training
Technical tools mentioned
PythonPandasNumPySQLBigQueryCloud StorageDataflowDataprocVertex AI Data PipelinesGoogle Cloud ComposerApache AirflowPrefectDagsterGreat ExpectationsDVCTFXData Validation

Job description

About the team:

Materra is on a mission to radically reduce global waste and move to a true circular economy. The team has developed technology that identifies waste material at the molecular level—starting with plastics. Materra works with industry partners to improve the way recycling centers process plastics using AI and robotics, to make recycling more affordable and scalable.

About the Role

We are looking for a Machine Learning Data Engineer (DataOps) to build and unify the data infrastructure that powers our model training pipelines. In this role, you will lead the effort to consolidate fragmented data sources into a cohesive, high-quality data foundation.

Your primary focus will be designing automated ingestion pipelines, establishing data quality validation frameworks, and managing dataset versioning to support our machine learning training loops. You will bridge the gap between operations, remote annotation teams, and machine learning engineers to ensure our models are trained on reliable, well-structured data.

Key Responsibilities

  • Architect and build automated ETL (Extract, Transform, Load) and ELT (Extract, Load, Transform) data pipelines to aggregate, clean, and harmonize data from disparate sources, databases, and operational ingestion flows.
  • Implement DataOps practices, including data quality monitoring, automated schema validation, and anomaly detection to catch corrupt or mislabeled data early.
  • Standardize and integrate third-party annotation workflows and remote labeling feeds into unified datasets ready for model training.
  • Design and maintain dataset versioning and storage systems to allow reproducible machine learning experiments and seamless data retrieval.
  • Collaborate with machine learning engineers and operations teams to translate raw material, form factor, and sensor metadata into structured training features.

Requirements

  • Education: Degree in Computer Science, Data Engineering, Software Engineering, or a related technical field.
  • Data Engineering & Architecture: 3+ years  experience building scalable data pipelines, managing relational and non-relational databases, and unifying fragmented data storage systems.
  • Modern Python Proficiency: Expertise in Python and data manipulation libraries (e.g., Pandas, NumPy, or SQL).
  • Data Quality & DataOps: Practical experience implementing automated data validation, quality control frameworks, and dataset versioning practices.
  • ML Data Lifecycle Understanding: Hands-on experience structuring datasets specifically for machine learning workflows, including handling annotations, metadata tracking, and training set curation.

Preferred Skills

  • Google Cloud Ecosystem: Hands-on experience with Google Cloud platform tools (e.g., BigQuery, Cloud Storage, Dataflow, Dataproc, Vertex AI Data Pipelines).
  • Workflow Orchestration: Experience managing pipelines using Google Cloud Composer or equivalent orchestration frameworks (e.g., Apache Airflow, Prefect, Dagster).
  • Multimodal / Unstructured Data: Experience handling mixed data types, including image datasets, sensor metadata, and unstructured physical property records.
  • Annotation Platform Integration: Familiarity with data labeling platforms, human-in-the-loop workflows, or integrating third-party annotation APIs.
  • Validation & Versioning Tooling: Exposure to data quality and ML versioning tools (e.g., Great Expectations, DVC, or TFX/Data Validation).

The US base salary range for this full-time position is $166,000 - $244,000 + bonus + equity + benefits. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your location during the hiring process.

Please note that the compensation details listed in US role postings reflect the base salary only, and do not include bonus, equity, or benefits.

About Materra

Google's moonshot R&D division invents and launches breakthrough technologies to solve difficult global problems.

Similar jobs

Machine Learning Data Engineer roles near Mountain View, California
1mo
Save
Mark Applied
Hide
Senior Machine Learning Data Engineer
Sunnyvale, California, United States
$150k-$240k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Providing digital infrastructure for physical AI and autonomy.
5+ YOE5+ years experience building ML data pipelines and infrastructure, familiarity with ML infrastructure, data-centric AI, GPUs, microservices and databases; U.S. citizenship and ability to obtain security clearance required.
React, TypeScript, Python, Golang, Docker, Kubernetes, OpenSearch, Postgres, GPUs, LLMs, VLMs
1mo
Save
Mark Applied
Hide
Senior Machine Learning Data Curation Engineer
Santa Clara, California, United States
$175k-$296k/yr OnsiteFull Time
XPeng
XPengNYSE; HKEX: XPEV; 9868: Global AI mobility technology focused on smart EVs.
3+ YOEBachelor's or Master's in a quantitative field,3+ years managing large-scale datasets,proficiency in Python and SQL,experience with cloud platforms and ML frameworks,strong analytical and problem-solving skills.
Python, SQL, AWS, GCP, BigQuery, PyTorch, Hugging Face
1mo
Save
Mark Applied
Hide
Data Scientist/Machine Learning Engineer
Redwood City, California, United States
$125k-$150k/yr OnsiteFull Time
Cathexis
Cathexis: Veteran-owned systems integrator delivering AI, data, modernization, audit, and mission-support services to defense, intelligence, and civilian agencies.
2+ YOEBachelor's in CS/EE/Statistics (MS/PhD preferred), ~2 years relevant experience, strong Python and applied ML skills, scalable ML experience, strong math and communication skills.
Python, MapReduce, Amazon AWS, Microsoft Azure, Google Cloud Services
3mo
Save
Mark Applied
Hide
Senior Data Scientist + Machine Learning Engineer
Palo Alto or San Francisco
$190k-$210k/yr RemoteFull Time
Neo.Tax
Neo.Tax: AI tax software automating R&D tax credits and ASC 350-40 software capitalization for enterprise finance teams.
6+ YOE6+ years in data science / ML shipping production models; strong Python, SQL; experience building data pipelines; production ML, experimentation, and cross-functional collaboration.
Python, NumPy, Pandas, scikit-learn, PyTorch, TensorFlow, SQL, Airflow