MaxInsights
Posted 2w ago

Machine Learning Engineer (Video Understanding & Segmentation)

MaxInsights
Santa Clara, California, United States
OnsiteFull Time
Responsibilities
  • building pipelines
  • developing models
  • mentoring engineers
Requirements
  • MS/PhD or equivalent experience,3+ years in computer vision/multi-modal ML,proficiency in Python/PyTorch,experience with CLIP and LLM-based video systems,vector search infrastructure knowledge
Technical tools mentioned
PythonPyTorchCLIPLLMLangChainLlamaIndexFAISSMilvusVideoCLIPRT-1VLA

Job description

We are seeking a highly motivated Machine Learning Engineer to join our core research and development team, focused on video understanding and segmentation. In this role, you will build the systems that let us search, decompose, and describe massive volumes of egocentric and human-robot video at scale — turning raw, unstructured footage into structured, searchable, and richly annotated training data. You will work across video/image embedding models, LLM-based video understanding, and agentic pipelines that orchestrate multiple models into end-to-end workflows. This is a foundational role that directly shapes the data quality and scalability of our entire training data platform.

Responsibilities

  • Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.

  • Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.

  • Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.

  • Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.

  • Architect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.

  • Develop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.

  • Collaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.

  • Evaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.

  • Contribute to internal tooling, documentation, patents, and open-source initiatives where applicable.

  • Mentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.

Minimum Qualifications

  • MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.

  • 3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks.

  • Strong proficiency in Python and PyTorch, with solid software engineering fundamentals.

  • Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.

  • Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization).

  • Familiarity with agentic system design — tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).

  • Experience working with large-scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).

Preferred Qualifications

  • PhD with a research focus in video understanding, multi-modal learning, or vision-language models.

  • Experience with temporal action segmentation, action localization, or instruction-level video chunking algorithms.

  • Experience working with egocentric video datasets or head-mounted-device (HMD) captured data.

  • Track record of deploying production-scale video search or retrieval systems.

  • Experience integrating foundation or vision-language models (e.g., CLIP, VideoCLIP, RT-1/VLA variants) into perception or decision-making pipelines.

  • Publications in top-tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).

  • Experience with humanoid robotics or embodied AI data pipelines is a plus.

About MaxInsights

Provides robot data collection for physical AI development.

Year founded
2024
Employees
30
Organization type
Private
Latest investment
Seed (2025) — led by Top Harvest Capital, Alumni Ventures, South Park Commons, Tola Capital, Plug and Play Tech Center
Headquarters
US

Similar jobs

Machine Learning Engineer roles near Santa Clara, California
2d
Save
Mark Applied
Hide
Founding Machine Learning Engineer
Mountain View, California, United States
$220k-$300k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
3+ YOE3+ years ML engineering experience; proficiency in Python and PyTorch/TensorFlow/JAX; production ML pipelines, LLM fine-tuning, distributed training, cloud (AWS/GCP/Azure), MLflow or Weights & Biases; strong collaboration skills.
Python, PyTorch, TensorFlow, JAX, AWS, GCP, Azure, MLflow, Weights & Biases
2d
Save
Mark Applied
Hide
Staff Machine Learning Engineer, Applied Research
San Francisco or United States
$189k-$390k/yr RemoteFull Time
Pinterest
PinterestNYSE: PINS: Visual discovery engine for finding inspiration and creative ideas.
6+ YOEMS/PhD in CS/ML/NLP/Statistics/Information Sciences,6+ years industry experience,ML/IR research experience,mastery of Java,C++,Python or ML frameworks (Tensorflow,Pytorch,MLFlow),strong communication and problem-solving skills.
Java, C++, Python, Tensorflow, Pytorch, MLFlow
3d
Save
Mark Applied
Hide
Staff Machine Learning Engineer
Chicago or Seattle or Sunnyvale or New York City or San Francisco
$209k-$258k/yr HybridFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
6+ YOE6+ years developing ML models, Bachelor's in CS or related, familiarity with PyTorch, experience with marketplace ML, strong communication and critical thinking.
PyTorch
3d
Save
Mark Applied
Hide
Staff Machine Learning Engineer, Personalization
Mountain View, California, United States
$152k-$277k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides online retail, grocery delivery, and video streaming services.
4+ YOEBachelor's in a technical field,4+ years applied ML experience,proficiency in Python/Java,experience with ML systems,search or recommendation experience preferred.
Python, Java, Apache Spark, Airflow, Kubeflow, MLflow, Vertex AI, BigQuery, SageMaker, TensorFlow, PyTorch, Scikit[1]learn, Keras, XGBoost, LightGMB, H2o.ai, Weights & Biases
3d
Save
Mark Applied
Hide
Machine Learning Engineer, Infra, AI for Drug Discovery
South San Francisco or New York City
$148k-$274k/yr OnsiteFull Time
Roche
RocheSIX Swiss Exchange: ROG: Provides innovative pharmaceutical and diagnostic healthcare solutions.
3+ YOEBS/MS or equivalent,3+ years in software/infrastructure/platform/MLOps,strong Python,cloud (AWS) and Kubernetes experience,Terraform/Pulumi,observability tools,distributed-systems knowledge,ability to deliver production software.
Python, EKS, EC2, S3, IAM, SQS, SNS, CloudWatch, Kubernetes, Helm, Terraform, Pulumi, Git, Datadog, Prometheus, Grafana, OpenTelemetry, KServe, Triton, vLLM, Ray Serve, Prefect, Dagster
3d
Save
Mark Applied
Hide
Principal Machine Learning Engineer
Mountain View, California, United States
$174k-$274k/yr RemoteFull Time
Atlassian
AtlassianNASDAQ: TEAM: Develops software for team collaboration and project management.
Drive technical direction, design and deploy production ML models (ranking, retrieval, LLM systems), conduct experiments, and mentor engineers across teams.
3d
Save
Mark Applied
Hide
Senior/Staff Machine Learning Engineer, Search & Knowledge Platforms
Santa Clara, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experienced machine learning engineer to develop search and Q&A experiences using search technologies and large language models in a high-performance computing environment.
4d
Save
Mark Applied
Hide
Sr Machine Learning Engineer, Adobe Firefly Services
San Jose or Seattle or San Francisco
$152k-$265k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
4+ YOEMS/PhD or equivalent experience,4+ years ML experience,2+ years leading GPU-intensive GenAI systems,experience with PyTorch,CUDA,Triton,TensorRT,Nvidia Dynamo,Python,Kubernetes,distributed systems,strong communication.
PyTorch, CUDA, Triton, TensorRT, Nvidia Dynamo, Python, Kubernetes