Tencent
Posted 2w ago

腾讯会议端云协同语音大模型研究

Tencent
Shenzhen, Guangdong, China
OnsiteFull Time
Responsibilities
  • researching speech
  • designing models
  • deploying systems
Requirements
  • Research background in speech/ASR
  • Speaker recognition
  • Multimodal learning
  • Large models or edge/cloud collaboration
  • Strong PyTorch skills
  • Algorithm implementation and experiment ability
  • Interest in meeting speech scenarios
Technical tools mentioned
PyTorch

Job description

课题详情

课题背景:

随着会议字幕、纪要、问答等能力不断升级,会议语音理解正从单点识别走向端云协同的综合智能。会议场景具有多说话人、远场、长时连续输入等特点,单纯依赖云端或端侧都存在局限,因此需要研究端侧感知、预处理、轻量理解与云端深度推理协同的大模型方案。

课题挑战:

会议场景中存在连续长音频、多说话人、交叉发言、多语种、多方言及复杂噪声环境,端侧还受到算力、功耗、时延限制。如何合理划分端云职责,在端侧完成低时延感知、信息压缩与关键特征提取,并与云端深层语义理解、多模态融合高效协同,是本课题核心挑战。

具体工作:

围绕腾讯会议实时字幕、流式纪要、离线总结、录音文件和多语种内容理解等业务,研究端云协同的会议语音理解方案。重点开展端侧低时延前处理、轻量化表征编码、说话人及场景信息建模、关键信息上云策略,以及端云联合优化机制,持续提升整体理解效果与用户体验。

课题要求

1、在语音识别、语音增强、说话人识别、多模态学习、大模型、端侧推理、模型压缩或端云协同等方向有较强研究基础,有相关论文发表者优先;

2、熟悉 PyTorch 等深度学习框架,具备较强的算法实现与实验能力;有音频、语音、大模型训练优化或端侧部署经验者优先;

3、对会议语音理解和智能会议场景有浓厚兴趣,具备较强的问题分析能力、创新意识和团队协作能力,能够推动科研成果向实际业务落地转化。

About Tencent

Provides integrated internet services, digital entertainment, and cloud technology.

Year founded
1998
Employees
112771
Organization type
Public
Headquarters
CN

Similar jobs

Speech Researcher roles near Shenzhen, Guangdong
2w
Save
Mark Applied
Hide
腾讯会议端云协同语音大模型研究
Shenzhen, Guangdong, China
OnsiteInternship
Tencent
TencentHong Kong Stock Exchange: 0700: Multinational technology conglomerate providing internet and entertainment services.
Research background in speech, multimodal learning, or large models; experience with model compression, on-device inference, or end-cloud collaboration; strong PyTorch and algorithm implementation skills; publications preferred.
PyTorch
1mo
Save
Mark Applied
Hide
语音算法研究员
Beijing or New Territories or Shenzhen
OnsiteFull Time
SenseTime
SenseTimeHKEX: 0020: Develops artificial intelligence software and computer vision technology.
Bachelor's or above (class of 2027) in CS or related, familiarity with ASR/TTS and end-to-end speech models, strong Transformer/RNN-T knowledge, PyTorch/TensorFlow experience, model optimization and deployment skills.
PyTorch, TensorFlow, ONNX, TensorRT, Conformer-Transducer, Whisper, VALL-E, MMS, Transformer, RNN-T