Outpost
Posted 2w ago

ML Data Engineer

Outpost
United States
RemoteContract
Responsibilities
  • measuring accuracy
  • curating datasets
  • labeling data
Requirements
  • 3+ years in data quality/ML data engineering
  • Experience with computer vision and annotation tools
  • Python proficiency
  • Strong analytical and communication skills
Technical tools mentioned
RoboflowLabelboxCVATPythonGCP (GCS)PostgreSQLSnowflakeNode.jsTypeScriptOCR

Job description

About Us:

Outpost is building the backbone of freight. We’re reinventing how supply chain infrastructure works in America with carrier agnostic truck terminals. As a vertically integrated real estate, operations, and technology company, we acquire and operate mission-critical real estate across the country to serve the largest logistics providers in the world. Backed by $1B from Greenpoint Partners, we’re scaling and building the most valuable logistics network in the country.

We thrive on accountability, integrity, and a shared drive to raise the bar. If you’re excited to reshape the industry alongside a high-performance team with a championship mindset that executes relentlessly, welcome aboard.

Role Summary:

Our platform combines AI-powered gate automation, computer vision, and operational software to help logistics operators run smarter, faster facilities. We're a small, high-conviction team shipping real software that ends up in real yards, at real gates, moving real freight; and we're growing fast, with revenue set to grow 10X over the next 18 months.

As we onboard more customers, our computer vision system sees more camera layouts, identifier types, and edge cases than ever. We need someone to own accuracy end-to-end: measuring it, understanding why we get it wrong, and turning that into the labeled data that makes our models better. Today that's mostly measurement and curation. Once the pipeline matures and moves into maintenance mode, we expect this role to also contribute fixes to the product itself, not just flag issues for others to resolve.

Key Responsibilities:

  • Own tracking and reporting of CV accuracy metrics, per customer and per identifier type.

  • Investigate misclassifications and false negatives, categorize root causes, and identify patterns across customers and yards.

  • Curate, label, and prioritize datasets for model retraining, partnering closely with our ML and CV engineers.

  • Build and improve the continuous learning pipeline so new models ship weekly with minimal manual engineering effort.

  • Define functional acceptance criteria for CV accuracy per customer and track progress against them.

  • Translate accuracy findings into decisions the engineering team and customer-facing stakeholders can act on.

  • As the pipeline matures, expect to move from flagging issues to fixing them directly; building the labeling/preprocessing tooling, running retraining jobs, and owning fixes for the error patterns you find, not just reporting them.

What You Can Expect:

  • Direct ownership over the metric that decides whether our product works in the real world.

  • A small team that moves fast, argues in good faith, and trusts engineers to make decisions.

  • Real influence on what the ML team builds next; your findings drive the roadmap, not the other way around.

  • Problems grounded in the physical world: gates, cameras, trucks, yards.

Qualifications:

  • 3+ years in a data quality, ML data engineering or applied ML role.

  • Experience working with computer vision or object detection systems in production.

  • Comfortable writing Python for data analysis, pipeline automation, and dataset tooling.

  • Strong analytical rigor, comfortable digging into large volumes of imagery/data to find patterns, not just running a script and reporting a number.

  • Experience with dataset annotation/labeling tools and workflows (Roboflow, Labelbox, CVAT, or similar).

  • Strong communication skills in English — you write clearly and engage well async.

Preferred Qualifications:

  • Experience with continuous learning or active learning pipelines for production ML systems.

  • Familiarity with OCR systems and identifier recognition (plates, container numbers, etc.).

  • Experience partnering with customer success or support teams on quality metrics.

  • Background in QA/test engineering for ML systems.

  • Experience with Roboflow specifically.

Our Stack:

Python · Roboflow · VLM/OCR pipelines · GCP (GCS) · PostgreSQL · Snowflake · Node.js/TypeScript

Outpost is an Equal Opportunity Employer and Prohibits Discrimination of Any Kind.

About Outpost

Operating a national network of tech-enabled truck terminals.

Year founded
2021
Employees
56
Organization type
Private
Latest investment
Raised $1.00B Corporate Round (2025) — led by GreenPoint Partners
Subsidiaries
Headquarters
US

Similar jobs

ML Data Engineer roles
1w
Save
Mark Applied
Hide
Data & ML Engineer
United States
$150k-$200k/yr RemoteFull Time
Red Cell Partners
Red Cell Partners: Incubating technology companies in healthcare and national security sectors.
5+ YOE5+ years in data engineering, architecture, applied ML, ML engineering, or production analytics; strong Python and SQL; US citizenship; active Secret clearance; CAC eligibility; willingness to travel up to 25%.
Python, SQL, PostgreSQL, pgvector, scikit-learn, XGBoost, PyTorch, AWS Glue, Airflow, dbt, Spark, Kafka, NiFi, NIST AI RMF
3w
Save
Mark Applied
Hide
Senior ML/Data Engineer
New York, New York, United States
$158k-$259k/yr OnsiteFull Time
Catapult
CatapultASX: CAT: Develops wearable tracking technology and video analytics for sports.
5+ YOE5+ years production data engineering with time-series, feature stores, streaming ingestion, graph databases, tenant isolation; strong Python, Golang, and SQL; experience with model calibration and evaluation frameworks.
Python, Golang, SQL, AWS, ECS, EC2, Lambda, SNS, SQS, GraphQL, REST, gRPC, Postgres, Mongo
4w
Save
Mark Applied
Hide
Senior Data & ML Engineer
Alpharetta, Georgia, United States
OnsiteFull Time
Fiserv
FiservNew York Stock Exchange: FI: Provides financial technology and payment processing services to institutions.
8+ YOE8+ years data/ML engineering experience; strong Python and SQL; ETL, data pipeline, MLOps, and production ML experience; AWS Glue, Qlik Data Integration, Snowflake experience; bachelor's degree or equivalent.
Python, SQL, AWS Glue, Qlik Data Integration, Snowflake, S3, Lambda, SageMaker, CloudWatch, Step Functions, ECS, EKS
1mo
Save
Mark Applied
Hide
Staff ML Data Engineer (Datagrid)
San Francisco, California, United States
$227k-$313k/yr HybridFull Time
Procore
ProcoreNYSE: PCOR: Cloud-based construction management software for projects and teams.
8+ YOE8+ years building and operating large-scale data systems for ML; strong SQL and Python; experience with data pipelines, dataset curation, quality, and observability; ability to lead technical efforts and mentor engineers.
SQL, Python, Databricks, Spark, Kafka, Pub/Sub, Airflow, Dagster, AWS, GCP, CI/CD
1mo
Save
Mark Applied
Hide
Data & ML Engineer
Mason, Ohio, United States
OnsiteFull Time
EssilorLuxottica
EssilorLuxotticaEuronext Paris: EL: Designs, manufactures and distributes ophthalmic lenses, frames and sunglasses.
Bachelor's in CS/Engineering, proven experience as Data/ML/Platform Engineer, strong Azure and big-data experience, proficiency in Python/Scala/SQL, production-grade data pipelines, and authorization to work in the U.S. (no visa sponsorship).
Azure, AWS, GCP, Python, Scala, SQL, Spark, Kubernetes, Azure Data Factory, Azure Data Lake, Synapse, Delta, Parquet, JSON, CSV, Azure Databricks, MLFlow, Docker, AKS, APIs, Airflow, SAP CDC, CI/CD, MLOps, DevOps
2mo
Save
Mark Applied
Hide
Member of Technical Staff — ML Data Infra
Seattle, Washington, United States
$200k-$300k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: A Series A building photorealistic, real-time AI avatars with emotional intelligence.
Proven experience building and operating large-scale multimodal data pipelines; strong proficiency with distributed processing frameworks (Spark, Ray, Dask); software engineering best practices; familiarity with video/audio processing and data quality/versioning tools.
Spark, Ray, Dask, FFmpeg, decord, DVC, Delta Lake, Apache Iceberg
5mo
Save
Mark Applied
Hide
ML Data Engineer (Contract-to-hire)
Bethesda, Maryland, United States
$120k-$140k/yr OnsiteContract, Full Time
Potomac Fund Management
Potomac Fund Management: A data-driven asset management firm focused on analytics and AI.
4+ YOE4+ years in data engineering; Python and SQL proficiency; building data pipelines; experience with data platforms; API integrations; ETL/ELT
Python, SQL, Airflow, Dagster, Prefect, Data Lakes, Data Warehouses, Lakehouse, APIs
5mo
Save
Mark Applied
Hide
Lead ML Data Engineer, AI Core
Palo Alto or Miami or Durham or Toronto
HybridFull Time
Nubank
NubankNYSE: NU: Digital financial platform offering banking, credit, and investment services.
6+ YOE6+ years building production ML/data systems; experience with data ingestion pipelines, distributed frameworks (Ray, Spark), Python, data quality, model training/tuning, and leadership/mentoring.
Ray, Spark, Python, MLflow, Dagster, Airflow