Agibot
Posted 3w ago

机器人真机强化学习算法实习生

Agibot
Shanghai, Shanghai, China
OnsiteInternship
Responsibilities
  • developing algorithms
  • training models
  • validating hardware
Requirements
  • Current master's or PhD student (excellent undergraduates considered)
  • Familiar with PPO/SAC/TD3/GRPO
  • Proficient in Python and PyTorch
  • RL training/framework experience and robot real-world deployment preferred
  • Available ≥4 days/week for ≥6 months
Technical tools mentioned
PythonPyTorchPPOSACTD3GRPORLINFVeRL

Job description

1.参与真机强化学习算法研发,在导师指导下完成 RL 算法的实现、训练与真机验证,协助开展数据配方、训练策略与评测实验。
2.协助多模态建模与训练实验,涉及图像/视频/语言/触觉/动作等信息融合与能力评估。
3.参与灵巧手/双臂等接触丰富(contact-rich)操作任务的策略学习探索。
4.协助维护训练—评测—迭代闭环,参与真机实验与数据分析。
5.跟踪 RL 与具身智能前沿进展,优秀成果可产出学术论文或开源项目。

1.在读硕士/博士(优秀本科生亦可),计算机/机器人/自动化/AI 等相关专业。
2.熟悉 PPO/SAC/TD3/GRPO 等主流RL算法,对offline/online RL等前沿方向有了解。
3.熟练使用 Python 与 PyTorch,具备良好的代码能力。
4.使用过 RLINF、VeRL 等 RL 训练框架者优先。
5.有课程项目、竞赛或科研中的机器人真机训练部署经验,或对机器人操作、灵巧手等方向有浓厚兴趣者优先。
6.有顶会论文、开源项目或竞赛获奖经历者优先。
7.每周可到岗 4 天及以上,实习期 6 个月以上优先。

About Agibot

Develops humanoid robots and embodied AI for industrial applications.

Year founded
2023
Employees
1000
Organization type
Private
Latest investment
Funding Round (2025)
Headquarters
CN

Similar jobs

Reinforcement Learning Algorithm Intern roles near Shanghai, Shanghai
5d
Save
Mark Applied
Hide
日常实习生-强化学习RL算法-Qwen基础模型
Beijing or Hangzhou or Shanghai
RemoteInternship
Alibaba
AlibabaNYSE: BABA: Provides online marketplaces, cloud computing, and digital payment services.
Current master's or doctoral student in computer science, artificial intelligence, automation, or related fields with reinforcement learning and agent experience; familiarity with PPO and GRPO required.
PPO, GRPO, RL, Agent, Reward Modeling
5mo
Save
Mark Applied
Hide
强化学习算法实习生
Shanghai, Shanghai, China
OnsiteInternship
XPeng
XPengNYSE: XPEV: Designs and manufactures smart electric vehicles and AI-driven mobility.
Seeking candidates to research and develop RL algorithms for autonomous driving; strong RL foundations, simulation and sim-to-real experience, Python/C++, and interest in physical AI.
PPO, GRPO, SAC, MuJoCo, Isaac Gym, CARLA, Autoregression, diffusion, flow matching, LoRA, DPO, SFT, VLA, VLM, Python, C++
5mo
Save
Mark Applied
Hide
【实习】大模型强化学习后训练算法实习生-安全数据中心
Shanghai, Shanghai, China
OnsiteInternship
Shanghai Artificial Intelligence Laboratory
Shanghai Artificial Intelligence Laboratory: A leading research laboratory for foundational artificial intelligence advancement.
Strong CS fundamentals, proficient Python, deep RL knowledge (PPO/DPO/GRPO), PyTorch experience, RLHF/RLVR and LLM training familiarity, distributed training and engineering skills.
Python, PyTorch, OpenRLHF, verl, DeepSpeed, GitHub
11mo
Save
Mark Applied
Hide
端到端强化学习算法实习生
Shanghai, Shanghai, China
OnsiteInternship
NIO
NIONYSE: NIO: Manufacturer of smart electric vehicles and battery swapping systems.
Master's or above in CS/automation/EE/AI/vehicle engineering; experience with end-to-end driving, agent planning, or behavior prediction; proficient in DRL algorithms (DQN, DDPG, SAC, PPO, GRPO).
DQN, DDPG, SAC, PPO, GRPO