Tencent
Posted 2w ago

混元多模态-音视频表征学习

Tencent
Beijing, Beijing, China
OnsiteFull Time
Responsibilities
  • researching representations
  • designing encoder
  • evaluating performance
Requirements
  • Master's or PhD in CS/AI/electronic information/signal processing
  • Strong ML foundation
  • Multimodal research experience
  • PyTorch and Python proficiency
  • Transformer and representation model knowledge
  • Distributed training familiarity
Technical tools mentioned
PyTorchPythonTransformerCLIPSigLIPPerception EncoderRAEAutoEncoderVAEDeepSpeedMegatron

Job description

课题详情

课题背景:

随着多模态大模型的发展,多模态理解任务正逐步演进为对视频、语音、背景声及上下文等多源异构信息的统一建模。面向真实复杂场景,全模态表征已成为多模态智能领域的重要研究方向。

课题挑战:

本课题的主要挑战在于构建面向全模态音视频理解的编码器,从编码效率和表征维度上实现任意时长的音视频联合编码。

具体工作:

参与全模态表征学习的研究,编码器架构的设计,完善编码器的数据生产和性能评测。

课题要求

1、来自计算机科学、人工智能、电子信息、信号处理等相关专业的博士/硕士同学,具备扎实的机器学习理论基础,有多模态方向研究经验者优先。

2、精通 PyTorch 框架与 Python 编程;深度掌握 Transformer 架构及主流图像和音视频表征模型(如 CLIP,SigLIP,Perception Encoder,RAE,AutoEncoder、VAE等) 原理;熟悉图像和音视频自监督算法及 DeepSpeed/Megatron 分布式训练工具。

3、具备较强的问题分解与独立研究能力;能在模糊问题定义上给出清晰的拆解思路和方法。

About Tencent

Provides integrated internet services, digital entertainment, and cloud technology.

Year founded
1998
Employees
112771
Organization type
Public
Headquarters
CN