Tencent
Posted 2w ago

腾讯会议—语音大模型方向研究

Tencent
Beijing, China
OnsiteInternship
Responsibilities
  • conducting research
  • training models
  • cleaning data
Requirements
  • Research experience in speech/dialogue areas
  • Familiarity with end-to-end multimodal large models
  • Streaming real-time dialogue implementations, and techniques to improve instruction following

Job description

课题详情

课题背景:

近年来, 端到端语音对话大模型在智能客服、虚拟人交互、辅助诊疗等领域得到广泛应用,其具备低延迟、丰富副语言信息等优点,但模型智商的维持仍面临较大挑战。

课题挑战:

本课题研究聚焦模态对齐、跨模态语义迁移、动态融合推理等关键问题,通过自监督学习、专家混合机制等技术提升模型对复杂场景的适配能力。其核心目标是让智能系统精准捕捉多模态信息中的情感倾向与语义关联,实现更贴近人类交互习惯的自然对话。

具体工作:

参与语音大模型的全流程研发,包括跨模态对齐、多模态理解及生成,涵盖文本和语音等训练数据的清洗和制作、基础模型算法选型与优化,聚焦预训练、监督微调及强化学习等关键环节的技术迭代; 提高在远场、低信噪比、多人、音乐等场景下的理解及生成效果,改善模型在方言、副语言信息等方面的理解能力,加强情感对话能力。

课题要求

1、具有语音对话、语音识别、语音合成、语音翻译、语音合成转换、说话人识别等方向的研究经验;

2、掌握端到端多模态大模型的训练和调优技术;

3、熟悉实时对话模型的流式处理实现方案和多种模态之间对齐的技术实现;

4、掌握改善模型遵循复杂指令能力的技术。

About Tencent

Provides integrated internet services, digital entertainment, and cloud technology.

Year founded
1998
Employees
112771
Organization type
Public
Headquarters
CN

Similar jobs

Speech Researcher roles in Beijing
2w
Save
Mark Applied
Hide
微信小微-支持多方言的音频生成模型
Beijing, Beijing, China
OnsiteFull Time
Tencent
TencentHong Kong Stock Exchange: 0700: Multinational technology conglomerate providing internet and entertainment services.
Master's or PhD in CS/AI/electrical engineering/mathematics, strong deep learning and generative/sequence modeling skills, proficiency in Python and PyTorch/TensorFlow, good communication and interest in audio/AIGC.
Python, PyTorch, TensorFlow
6d
Save
Mark Applied
Hide
端到端语音智能体前沿技术研究-阿里星/C-Star
Beijing or Hangzhou or Shanghai
OnsiteFull Time
Alibaba
AlibabaNYSE: BABA: Provides online marketplaces, cloud computing, and digital payment services.
Master's degree or higher in computer science, electronic engineering, artificial intelligence, signal processing, or related fields; expertise in speech AI, multimodal learning, and large-scale model training.
Transformer, PyTorch, Megatron, DeepSpeed, NeurIPS, ICML, ICLR, ACL, INTERSPEECH, ICASSP, VAD
1mo
Save
Mark Applied
Hide
语音算法研究员
Beijing or New Territories or Shenzhen
OnsiteFull Time
SenseTime
SenseTimeHKEX: 0020: Develops artificial intelligence software and computer vision technology.
Bachelor's or above (class of 2027) in CS or related, familiarity with ASR/TTS and end-to-end speech models, strong Transformer/RNN-T knowledge, PyTorch/TensorFlow experience, model optimization and deployment skills.
PyTorch, TensorFlow, ONNX, TensorRT, Conformer-Transducer, Whisper, VALL-E, MMS, Transformer, RNN-T
2mo
Save
Mark Applied
Hide
顶尖应届-语音大模型算法研究员-MiMo
Beijing, Beijing, China
OnsiteFull Time
Xiaomi
XiaomiHong Kong Stock Exchange: 1810: Designs and manufactures smartphones, consumer electronics, and electric vehicles.
Research and develop large-scale speech-modal pretraining, multilingual speech understanding and generation, noise-robustness in complex acoustic scenes, and efficient speech compression methods.
4mo
Save
Mark Applied
Hide
语音大模型算法研究员
Beijing or Shanghai
OnsiteFull Time
Xiaohongshu
Xiaohongshu: Social media and e-commerce platform for lifestyle sharing.
3+ YOEDegree in AI/EE/CS, 3+ years related experience in speech recognition/synthesis, large-scale data processing, strong communication and problem-solving; publications preferred.
6d
Save
Mark Applied
Hide
音视频数字人多模态大模型算法研究和应用-阿里星/A Star
Beijing or Hangzhou
OnsiteFull Time
Alibaba Cloud
Alibaba CloudNYSE: BABA: Global cloud computing infrastructure and services provider
Master's or doctoral degree in computer science, electronic engineering, AI, or related field; deep ASR, TTS, SER, multimodal LLM, speech dialogue, or NLP research; top-tier publications; Python and PyTorch or TensorFlow.
Python, PyTorch, TensorFlow, SpeechGPT, SALMONN, Qwen-Audio, Mini-Omni, GLM-4-Voice, DeepSpeed, Megatron-LM, FSDP, Kaggle
10mo
Save
Mark Applied
Hide
顶尖实习-语音大模型算法研究员-MiMo
Beijing, Beijing, China
OnsiteInternship
Xiaomi
XiaomiHKEX: 1810: Design and manufacturing of consumer electronics and smart devices.
Research speech large language model pretraining, multilingual understanding and generation, noisy acoustic environments, speech compression, and publication at leading NLP or speech conferences.
Scaling Laws, MoE, Agent