Tencent
Posted 2w ago

大模型Code/Agent后训练算法研究员-(深圳)or(北京)or

Tencent
Guangdong or Beijing
OnsiteFull Time
Responsibilities
  • constructing data
  • building environment
  • training agents
Requirements
  • Master's in CS/AI required
  • >2 years experience in large-scale RL or model/agent development preferred
  • Strong deep learning foundations
  • Distributed training experience, and good teamwork and communication

Job description

Responsibilities

1.负责Code和Agent相关数据构建与治理,构建高质量、多样化的Code/Agent训练数据集,搭建数据迭代闭环,通过数据飞轮持续优化数据质量;

2.负责Agent运行环境与训练环境的构建与优化,构建高可用、可扩展的Agent仿真环境,保障Agent训练、测试及落地的稳定性与高效性;

3.负责Agentic RL在Code/Agent场景的训练,参与Agentic RL Infra建设及优化、Agentic RL 算法优化,持续提升Agentic RL训练的效率和稳定性。

Requirements

1.计算机、人工智能等相关专业硕士以上学历;

2.有大规模强化学习、大模型Code/Agent研发相关经验者优先;

3.具有扎实的深度学习算法基础,熟悉深度学习框架和分布式训练推理加速,有实操经验者优先;

4.在多模态/CV/NLP等领域顶级会议(期刊)发表过论文、主导/参与业界知名的开源项目者优先;

5.具备极强的学习能力和技术追求,良好的团队合作和沟通能力。

About Tencent

Provides integrated internet services, digital entertainment, and cloud technology.

Year founded
1998
Employees
112771
Organization type
Public
Headquarters
CN

Similar jobs

Algorithm Researcher roles in Guangdong
3d
Save
Mark Applied
Hide
算法研究员(蛋白质结构生成方向)
Beijing, Beijing, China
OnsiteFull Time
北京科学智能研究院
北京科学智能研究院: A research institute focused on AI for Science innovation.
PhD or higher in progress preferred in computer science, software engineering, artificial intelligence, bioinformatics, or related fields; Python and deep learning framework experience required.
Python, PyTorch, JAX, AlphaFold3
6d
Save
Mark Applied
Hide
ATH MaaS大模型&智能体前沿算法研究-阿里星/A Star
Beijing or Hangzhou
RemoteFull Time
Alibaba Group
Alibaba GroupNYSE: BABA: Global technology specializing in e-commerce and cloud computing.
Excellent master's or doctoral background in computer science, AI, software engineering, mathematics, or related fields; strong deep learning, distributed training, model optimization, research, and coding experience.
Alibaba Cloud, Agentic RL, Agentic Framework, AI搜索, AI客服, AI视频生成, NeurIPS, ICML, ICLR, ACL, EMNLP, ICASSP, AutoResearch
6d
Save
Mark Applied
Hide
ATH MaaS大模型&智能体前沿算法研究-阿里星/A Star
Beijing or Hangzhou
OnsiteFull Time
Alibaba
AlibabaNYSE: BABA: Provides online marketplaces, cloud computing, and digital payment services.
Excellent master's or doctoral background in computer science, artificial intelligence, software engineering, mathematics, or related fields; strong deep learning, distributed training, inference, coding, and large-model research experience.
Alibaba Cloud, Agentic RL, Agentic Framework, AI Search, AI Customer Service, AI Video Generation, NeurIPS, ICML, ICLR, ACL, EMNLP, ICASSP, AutoResearch
1w
Save
Mark Applied
Hide
算法-其他方向
Shenzhen or Beijing or Shanghai
OnsiteFull Time
Tencent
TencentHong Kong Stock Exchange: 0700: Multinational technology conglomerate providing internet and entertainment services.
PhD or excellent master's degree in computing, engineering, AI, mathematics, physics, or related fields; proficiency in at least one programming language; research or project experience preferred.
Java, C, C++, C#, Python
2w
Save
Mark Applied
Hide
视频生成基座模型算法专家
Shanghai or Beijing
OnsiteFull Time
Resonate
Resonate: AI-powered consumer intelligence and predictive data analytics platform.
Master's degree or higher in CS/AI, deep ML/DL theory, first-principles understanding of diffusion or autoregressive models, experience leading large-scale generative model training and alignment (SFT/RLHF).
Diffusion, Transformer, SFT, RLHF
2w
Save
Mark Applied
Hide
算法研究员(搜广推大模型) - Data AML
Beijing or Hangzhou
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's degree (2027) in CS/AI preferred; strong ML and NLP foundation; experience or publications in recommender systems or large models preferred; independent problem solving and teamwork.
3w
Save
Mark Applied
Hide
Super Sparks-校招-座舱具身智能体算法研究员
Beijing or Shanghai
OnsiteFull Time
Nio
NioNYSE: NIO: Designs and manufactures smart premium electric vehicles.
PhD required. Expert in statistical learning and deep learning (Transformer, Diffusion, VAE, Flow Matching). Familiar with reinforcement/behavioral policy learning (RLHF, PPO, DPO, GRPO). Experience training large models and multimodal/embodied interaction research.
Transformer, Diffusion, VAE, Flow Matching, RLHF, PPO, DPO, GRPO, TTS
3w
Save
Mark Applied
Hide
顶尖实习-具身大模型算法研究员-机器人事业部-实习1
Beijing, Beijing, China
OnsiteInternship
Xiaomi
XiaomiHKEX: 1810: Design and manufacturing of consumer electronics and smart devices.
PhD in CV/ML/Robotics/AI or research-oriented Master's; experience in VLM/multimodal/3D/Embodied Agent/RL; strong publication or impactful OSS background preferred; proficient in PyTorch and familiar with Transformer/Diffusion/CLIP/SigLIP/Qwen-VL.
PyTorch, Transformer, Diffusion, CLIP, SigLIP, Qwen-VL