ModelBest
Posted 6d ago

大模型数据开发工程师

ModelBest
Beijing, Beijing, China
OnsiteFull Time
Responsibilities
  • building pipelines
  • improving quality
  • labeling data
Requirements
  • Bachelor's degree or higher in computer science or a related field
  • 3+ years of data development experience, and proficiency in Python. Spark and Ray experience preferred
Technical tools mentioned
PythonApache SparkRay

Job description

1.参与大模型预训练数据的构建,不限于数据采集、数据清洗、数据增强,搭建高效的分布式数据pipeline;
2. 参与大预言模型的数据质量提升,不限于数据风控策略、异常数据检测、数据修复、特定领域数据扩充;
3. 参与模型训练数据标签体系的建设,建设数据自动打标服务,完成海量数据打标;
4. 参与模型数据的全生命周期调优工作,实现数据自动处理,提高数据的处理效率,缩减数据交付周期。

1. 本科以上、计算机相关专业,三年以上数据开发经验;
2. 精通Python语言,熟悉Spark、Ray等分布式框架,具备丰富的大规模数据处理实践经验是加分项;
3. 以下方向有实践经验或兴趣浓厚优先:大规模文本数据质量提升、多模态数据处理、文本数据清洗、语音数据处理。

About ModelBest

Develops efficient, on-device large language models and edge AI.

Year founded
2022
Employees
200
Organization type
Private
Latest investment
Raised $5.00B Funding Round (2026) — led by China Telecom, Shenzhen Capital Group, Inovance Capital, Primavera Capital Group
Subsidiaries
Headquarters
CN