Lightwheel
Posted 4mo ago

数据闭环资深工程师

Lightwheel
Beijing or Shanghai
OnsiteFull Time
Responsibilities
  • building pipelines
  • designing models
  • optimizing performance
Requirements
  • 5+ years building data platforms or data warehouses
  • Experience with PB-scale multimodal data
  • Data modeling, ETL
  • MySQL/PostgreSQL tuning
  • Kubernetes and big-data ecosystems
Technical tools mentioned
ClaudeCursorKubernetesSparkHiveMySQLPostgreSQLGitCI/CDGoPython

Job description

1.负责AI人类数据采集平台的数据闭环体系建设,搭建从数据采集、质检、清洗、挖掘、标注到训练交付的全链路自动化Pipeline,支持PB级多模态数据的高效流转与处理;
2.主导数仓建设与数据模型设计,解决海量数据存储、查询和治理瓶颈,构建高效、稳定的数据仓库与数据湖架构;
3.利用大模型与AI工具实现自动标注、难例挖掘、高价值数据筛选及数据质量提升,持续优化数据闭环效能;
4.设计并优化数据处理流水线,包括实时/离线ETL流程、数据治理机制和性能调优,确保数据从采集端到模型训练端的闭环高效运转;
5.参与GPU集群数据调度、模型评估数据支持及Latency分析,配合绩效监控模块提升整体平台数据效能;
6.与产品、研发、标注团队紧密协作,持续迭代数据平台的数据管理、可视化分析和AI数据分析能力。

1.本科及以上学历,计算机、人工智能、数据科学等相关专业,5年以上数据平台或数仓建设经验;
2.AI Native:熟练使用Claude、Cursor等AI编码/辅助工具,能显著提升数据Pipeline设计、SQL优化和文档产出效率;熟悉Agent、Skill等AI应用场景者优先;
3.有具身智能、自动驾驶或类似多模态数据闭环架构经验,深刻理解图像、点云、视频等数据结构及处理流程;
4.熟悉数据标注平台或数据管理系统的建设经验,理解业务数据完整Pipeline;
5.扎实的数仓建设能力:熟练进行数据建模、ETL设计,具备MySQL/PostgreSQL性能优化、索引设计、SQL调优经验;熟悉OLAP、数据湖(Data Lake)、分布式数据库者优先;
6.具备高并发、高可用数据系统设计经验,熟悉消息队列、任务调度、数据治理等技术;
7.熟悉Kubernetes(K8s)、大数据处理架构(Spark/Hive等)者优先;
8.有代码规范意识,熟练使用Git、CI/CD等工程化工具;编程语言以Go或Python为主;
9.较强的业务理解能力和跨团队沟通能力,对高质量AI训练数据闭环有深刻认知。

加分项:
1.有实际PB级数据平台或AI数据闭环项目落地经验;
2.熟悉平台现有模块(任务管理、Workflow编排、人工标注等)者优先。

About Lightwheel

Provides simulation and data infrastructure for physical AI systems.

Similar jobs

Data Engineer roles near Beijing, Beijing
5d
Save
Mark Applied
Hide
数据研发工程师
Beijing or Hangzhou
OnsiteFull Time
Alibaba
AlibabaNYSE: BABA: Provides online marketplaces, cloud computing, and digital payment services.
Bachelor's degree or higher in computer science, statistics, mathematics, or related fields; big data, data warehouse, distributed architecture, ETL, SQL, UDF, data governance, Python, and LLM or agent application experience.
SQL, Python, LLM
1w
Save
Mark Applied
Hide
数据技术及产品部-RL Data工程师-Work领域(法律方向)
Beijing or Hangzhou
OnsiteFull Time
Alibaba Group
Alibaba GroupNYSE: BABA: A global technology specializing in e-commerce and cloud computing.
3+ YOERequires legal domain expertise, RL/RFT training-data knowledge, taxonomy and rubric design, trajectory annotation and root-cause analysis, Python scripting, and experience producing data at scale.
Python, LLM
1w
Save
Mark Applied
Hide
高级数据工程师(座舱AI)
Beijing, Beijing, China
OnsiteFull Time
Xiaomi
XiaomiHong Kong Stock Exchange: 1810: Designs and manufactures smartphones, consumer electronics, and smart hardware.
3+ YOEBachelor's degree or higher in computer science, software engineering, data science, or related field; 3+ years of data engineering experience; Python, SQL, big data frameworks, ETL, and data modeling expertise.
Large Language Model (LLM), Python, SQL, Spark, Flink, Kafka, Hive, Feature Store, Kubernetes, Airflow, DolphinScheduler
1w
Save
Mark Applied
Hide
数据工程
Shenzhen or Beijing or Shanghai or Guangzhou
OnsiteFull Time
Tencent
TencentHong Kong Stock Exchange: 0700: Multinational technology conglomerate providing internet and entertainment services.
Bachelor's degree or higher in mathematics, statistics, operations research, computer science, or related fields; knowledge of big data technologies, programming, and SQL optimization.
Spark, Flink, TensorFlow, PyTorch, Java, C++, Python, Shell, SQL, GitHub, Large Language Model (LLM)
1w
Save
Mark Applied
Hide
数据工程
Shenzhen or Beijing or Shanghai or Guangzhou
OnsiteFull Time
Tencent
TencentHKEX: 0700: Provides integrated internet services, digital entertainment, and cloud technology.
Bachelor's degree or higher in mathematics, statistics, operations research, computer science, or related fields; knowledge of big data technologies, programming, SQL optimization, and AI tools.
Apache Spark, Apache Flink, TensorFlow, PyTorch, Java, C++, Python, Shell, SQL, GitHub, Large Language Model (LLM), AI
1w
Save
Mark Applied
Hide
海外数据研发工程师
Beijing, Beijing, China
OnsiteFull Time
iReader Technology
iReader TechnologyShanghai Stock Exchange: 603533: Operates a global digital platform for e-books and novels.
Requires big data architecture expertise, Hadoop ecosystem experience, data warehouse and database development, real-time processing with Java or Scala, Linux, Python or Shell, and large-scale data optimization skills.
HDFS, Hadoop, YARN, DataX, Sqoop, HBase, Hive, Spark, StarRocks, Impala, Kylin, MySQL, Redis, ES, BI, DMP, Linux, Python, Shell, Java, Scala, Kafka, Flink, DSP, ADX, SSP, RTB, RTA
2w
Save
Mark Applied
Hide
智能体数据工程师(仿真环境方向)
Beijing, Beijing, China
OnsiteFull Time
ModelBest
ModelBest: Develops efficient, on-device large language models and edge AI.
Proficient Python and engineering practices; experience with containers, remote execution, browser/GUI automation; knowledge of Docker, filesystems, networks, isolation; ability to convert business scenarios into executable, verifiable tasks.
Python, Docker, Kubernetes, Ray
2w
Save
Mark Applied
Hide
Data Engineer, AOP - RoW Central Data Engineer Team
Beijing, Beijing, China
OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
1+ YOE1+ years data engineering experience, data modeling and ETL, SQL and scripting (Python, KornShell), Redshift/Aurora administration, AWS services, on-call incident response, stakeholder collaboration.
Amazon Redshift, Aurora RDS, Tableau Server, EDX, SNS, Andes, APIs, S3, Lambda, ECS Fargate, Step Functions, CloudWatch, SQS, CDK, CloudFormation, Secrets Manager, EventBridge, MyUniverse, SQL, PL/SQL, DDL, MDX, HiveQL, SparkSQL, Scala, Python, KornShell, Hadoop, Hive, Spark, EMR, Informatica, ODI, SSIS, BODI, Datastage