简智新创
Posted 1mo ago

具身运动模型算法评测

简智新创
Beijing, Beijing, China
OnsiteFull Time
Responsibilities
  • building metrics
  • designing scenarios
  • analyzing data
Requirements
  • Design and run model evaluation pipelines for embodied motion models
  • Create reproducible scoring scenarios
  • Clean and analyze multimodal evaluation data, and produce diagnostic reports for model iteration
Technical tools mentioned
PythonVLMvalue modelreward model

Job description

1. 建立模型评测指标体系和定期评测流程,支持多模型、多版本、多场景下的评测、对比和报告产出。
2. 将复杂任务拆解为可执行、可复现、可评分的评测场景,定义成功标准、失败类型和评分规则。
3. 负责评测数据清洗、处理和分析,识别无效数据、异常样本、典型失败模式和高价值训练样本。
4. 参与自动化评测系统建设,结合规则、VLM / value model / reward model 和人工抽审,提高评测效率和一致性。
5. 与数据、运营团队协作,将评测结果转化为模型迭代建议、数据采集需求和诊断报告。

1. 有智能驾驶、机器人、多模态模型或其他 AI 系统的评测、数据分析、benchmark 维护经验。
2. 理解模型评测的基本方法,熟悉 A/B test、场景集构建、长尾/OOD 测试、可复现实验等概念。
3. 能把复杂任务拆成清晰的评测场景,定义成功标准、失败模式和评分规则。
4. 熟悉 Python 数据分析与实验工具链,能处理视频、状态、动作、标签、日志等多模态评测数据。
5. 对真实系统评测有基本理解,知道传感器、标定、环境扰动、人工操作、安全事件等因素会影响模型结果。
6. 有较强的问题拆解和写作能力,能把评测结果整理成清晰的诊断报告。
加分项
1. 做过智能驾驶感知/规划/端到端模型评测、仿真场景评测、数据闭环或问题归因。
2. 做过机器人、VLA、robot foundation model、模仿学习、强化学习或多模态模型相关工作。
3. 熟悉自动化评测、人工标注校准、评分一致性分析或模型 leaderboard。
4. 有 value model / reward model / VLM-as-Judge 相关经验。
5. 有将评测结果转化为训练数据、数据筛选或模型迭代建议的经验。

About 简智新创

Specializes in multimodal data infrastructure for embodied AI development.

Similar jobs

Model Evaluation Engineer roles near Beijing, Beijing
5d
Save
Mark Applied
Hide
模型评测工程师 (音频)
Beijing, Beijing, China
OnsiteFull Time
Shengshu Technology
Shengshu Technology: Develops multimodal generative AI models and world model frameworks.
Bachelor's degree or above preferred in audio, music, sound design, or digital media; experience using named AIGC tools; professional listening, audio quality assessment, visual-audio evaluation, reporting, and communication skills.
Keling, Seedance, TapNow, LibTV
1w
Save
Mark Applied
Hide
【2027 届校招】模型评测工程师
Beijing, Beijing, China
OnsiteFull Time
Qianxun Intelligent
Qianxun Intelligent: Developing general-purpose humanoid robots and embodied AI models.
2027 graduate with a bachelor's degree or higher in computer science, artificial intelligence, robotics, automation, or related fields; Python and deep learning framework experience required.
Python, PyTorch, TensorFlow, Transformer
2w
Save
Mark Applied
Hide
大模型评测工程师 - 模型数据工程
Beijing, Beijing, China
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's degree (2027) in CS/AI/Software Engineering, strong logic and learning ability, experience or interest in LLM/Agent evaluation, evaluation set design, data analysis, and automation; self-driven with strong execution.
1mo
Save
Mark Applied
Hide
大模型评测算法
Beijing, Beijing, China
OnsiteFull Time
ModelBest
ModelBest: Develops efficient, on-device large language models and edge AI.
2+ YOEBachelor's in CS/AI, 2+ years in ML/algorithms or evaluation, strong ML/NLP/multimodal knowledge, proficiency in Python/Java/C++ and deep learning frameworks, and experience designing model evaluation metrics.
Python, Java, C++, PyTorch, TensorFlow, OpenCompass, VLMEvalKit, EvalScope
1y
Save
Mark Applied
Hide
AI院-GLM团队-大模型评测算法工程师(实习)
Beijing, Beijing, China
OnsiteInternship
Zhipu AI
Zhipu AISEHK: 2513: Developing generative AI models and intelligent agent solutions.
Bachelor's degree in computer-related field, strong CS fundamentals and programming, interest in LLMs and code, attention to detail, teamwork, familiar with vLLM, sglang, TGI and Docker.
vLLM, sglang, TGI, Docker
1y
Save
Mark Applied
Hide
大模型测评工程师
Beijing, Beijing, China
OnsiteFull Time
Li Auto
Li AutoNASDAQ / HKEX: LI / 2015: Designing and manufacturing premium smart electric vehicles.
Master's degree in CS/AI/math/software engineering, experience in large-model testing and evaluation, strong algorithms background, proficient in Python/C++ and PyTorch/TensorFlow, able to design evaluation systems and build high-quality datasets.
Python, C++, PyTorch, TensorFlow
5d
Save
Mark Applied
Hide
晓天衡宇-大模型评测工程师-应用方向(北京)
Beijing, Beijing, China
OnsiteFull Time
Alibaba Group
Alibaba GroupNYSE: BABA: A global technology specializing in e-commerce and cloud computing.
3+ YOEBachelor's degree or higher in computer science, statistics, data science, or related fields; experience in LLM evaluation, data analysis, or product testing; Python data analysis and statistical evaluation expertise.
Large Language Models (LLM), Elo, BT, Python, bootstrap, QwenWork, WorkBuddy
7mo
Save
Mark Applied
Hide
具身智能模型评测工程师
Beijing, Beijing, China
OnsiteFull Time
Spirit AI
Spirit AI: Developing embodied AI and robotics for autonomous physical systems.
3+ YOEMaster's in computer science and 3+ years experience; experience evaluating/validating complex intelligent systems and real-world/high-complexity systems (autonomous driving, robotics/RL, large multimodal models); on-policy/closed-loop evaluation preferred.