ByteDance
Posted 1mo ago

Tech Lead, Machine Learning Engineer - Global E-Commerce (Conversational AI)

ByteDance
Singapore
OnsiteFull Time
Responsibilities
  • setting direction
  • developing staff
  • driving alignment
Requirements
  • Degree in CS/AI/Math or related field
  • Hands-on ML/NLP experience
  • Production LLM/Agent system experience
  • Strong Python and one of C++/Go/Rust
  • Multi-node post-training/fine-tuning experience
Technical tools mentioned
PythonC++GoRustFSDPDeepSpeedMegatronvLLMTensorRT-LLMMoEKV-cache

Job description

About the team
We are building the next generation of conversational AI for Global E-commerce — a unified Agent system that learns from every interaction, runs in 30+ languages, and is deployed across one of the largest e-commerce surfaces on the internet. Our 2026 north star is a self-evolving Agent: post-training, harness, tools, memory, and evaluation form one closed loop, and every served conversation becomes training, evaluation, and retrieval signal for the next iteration.

Business surface — buyer, seller, dispute & appeals, operations — is the substrate. Our work is foundational LLM + Agent engineering: post-training, agent harness, tool design, memory, evaluation, inference, multilinguality. We are hiring people who want to push the SOTA of these systems in production, at scale, with hundreds of millions of users in the loop.

What we work on
- LLM post-training & alignment — large-scale SFT, DPO/IPO/KTO, online RL (RLHF / RLAIF / RLVR), reward modeling, preference data curation, long-context training, distillation, QAT. We train and adapt frontier-class open-weights models (≥7B → ≥70B) and our own continually-pretrained checkpoints on internal infra (FSDP / DeepSpeed / Megatron-style stacks).
- Agent foundations — harness design (context engineering, sub-agents, durable execution, parallel tool use), tool design (ACI principles, namespaced surfaces, poka-yoke, instrumented traces), memory (episodic + semantic + skill-shaped), MCP and Skill-style extensibility. We treat tools and prompts as APIs and iterate against production traces.
- Auto-eval and observability — LLM-as-judge with calibrated human agreement, real-traffic replay, failure-mode taxonomies, regression + safety + cost + latency harnesses. We have moved root-cause analysis on a single case from ~13 engineer-days to ~3 minutes auto.
- Self-evolving systems — every served conversation becomes a candidate for training data, eval set membership, retrieval index, and skill induction, with privacy and quality gates. The flywheel is the product.
- Inference & serving — vLLM / TensorRT-LLM, MoE, speculative decoding, KV-cache reuse and prompt caching, multi-tenant low-latency serving. Cost per resolved conversation is a first-class metric.
- Multilinguality & locale grounding — 30+ languages, low-resource adaptation, faithful translation, locale-aware reasoning, cross-cultural tone.
- Reasoning & long-context modeling — chain-of-thought / planning post-training, reasoning-trace supervision, long-context training and serving, retrieval-augmented reasoning, self-consistency and verifier models.

Responsibilities
- Set technical direction. Own a multi-quarter roadmap across one or more of: post-training, agent harness, evaluation, self-evolving data flywheel, serving. Translate north-star metrics into a sequence of 2-3 high-ROI bets per quarter and ship them.
- Compound the team. Hire and develop 1-3 strong ICs. Design their work surfaces for growth, not just dispatch. Raise the median technical bar through design review, code review, and 1:1 framing.
- Stay in the loop with the model. Tech Lead is not a manager role. You still write the load-bearing PRs, propose the core abstractions, and write the design docs that decide the team's ceiling for the next 2-3 quarters.
- Drive cross-team alignment. Partner with foundation-model, infra, product, and adjacent algorithm teams; own sign-off on cross-cutting technical decisions.
- Observability and rollback. Build the per-turn tracing, tool-call analytics, and failure-mode taxonomies that let the team diagnose any regression within hours, not days.

Minimum Qualifications
- BS / MS / PhD in CS, AI, Mathematics, or related quantitative field.
- Hands-on experience in ML / NLP / applied DL. Top PhDs with strong publication record may qualify at 4+ years.
- Strong Python and at least one of C++ / Go / Rust for production-path code.
- Hands-on post-training or fine-tuning of frontier-class LLMs (≥7B, multi-node). Not API-only.
- Has led at least one production LLM / Agent system from zero to one.

Preferred Qualifications
- LLM post-training — multi-node SFT / DPO / online RL on ≥7B models; reward modeling; preference data construction; RLAIF / RLVR; distillation; QAT; long-context training; continual pre-training.
- Agent engineering (Anthropic-style) — production agent harness, context engineering, sub-agents, durable execution, MCP / Skill-style extensibility, parallel tool use, computer use.
- Reasoning & planning — chain-of-thought / reasoning-trace training, planner/critic decomposition, self-consistency, verifier models, multi-step reasoning evaluation.
- LLM / Agent evaluation — LLM-as-judge with human-agreement calibration, tau-bench / SWE-bench / GAIA / BFCL-style harnesses, regression + safety + cost + latency-aware evaluation.
- Inference & serving systems — vLLM / TensorRT-LLM, MoE, speculative decoding, KV-cache and prompt caching, low-latency multi-tenant serving.
- Multilingual & cross-cultural reasoning — multilingual SFT/DPO, low-resource adaptation, faithful MT, locale-aware reasoning.

About ByteDance

Developing AI-driven content platforms and mobile applications.

Similar jobs

Machine Learning Engineer roles
1d
Save
Mark Applied
Hide
Senior Machine Learning Engineer (AI Agent)
Singapore, Central Singapore, Singapore
OnsiteFull Time
PatSnap
PatSnap: AI-powered platform for intellectual property and R&D intelligence.
4+ YOEBachelor's or master's degree in computer science or a related field; 4+ years of production NLP/ML engineering; experience with document-level extraction, NER, normalization, and large-scale systems.
NLP, NER, LLMs
1d
Save
Mark Applied
Hide
Principal Machine Learning Engineer
Singapore, Singapore, Singapore
RemoteFull Time
BJAK
BJAK: Online platform for insurance comparison and road tax renewal.
Deep learning and transformer expertise; production ML model experience; modern ML framework proficiency; distributed training, GPU optimization, software engineering, and end-to-end ML systems ownership.
PyTorch, JAX, DeepSpeed, FSDP, Megatron, ZeRO, Ray, vLLM, TensorRT-LLM, FasterTransformer, Apache Arrow, Spark, LoRA, QLoRA, SFT, DPO, PPO, ORPO, RLHF
3d
Save
Mark Applied
Hide
Senior Machine Learning Engineer, TikTok BRIC Community Health
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
$162k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOEMaster's degree or above in a relevant technical field and 2+ years of machine learning experience. Requires strong software engineering, machine learning, Python or Java/C++/Go, and Spark, Hadoop, or Hive experience.
Python, Java, C++, Go, Spark, Hadoop, Hive, LLMs
6d
Save
Mark Applied
Hide
Staff Machine Learning Engineer - Moloco Commerce Media
Menlo Park or Seattle or New York City or San Francisco or Seoul or Beijing or Singapore or Gurgaon or Tokyo or Shanghai or London or Tel Aviv or Berlin
$232k-$348k/yr OnsiteFull Time
Moloco
Moloco: AI-powered programmatic advertising and commerce media platform.
8+ YOERequires 8–12 years in applied machine learning or ML engineering, production model experience, Python, modern ML frameworks, experimentation, distributed systems, and technical leadership.
Python, PyTorch, TensorFlow, JAX, feature stores, distributed training, large-scale data processing frameworks, MLOps, A/B testing, transformer-based architectures
6d
Save
Mark Applied
Hide
RD10013 Machine Learning Engineer
Singapore, Singapore, Singapore
OnsiteFull Time
ASUS
ASUSTaiwan Stock Exchange: 2357: Manufacturer of computers, motherboards, and consumer electronic devices.
3+ YOEBachelor's degree in computer science or related field, 3+ years relevant experience, production ML model development, Python or C++, ML frameworks, and NLP or computer vision expertise.
Python, C++
1w
Save
Mark Applied
Hide
Associate Machine Learning Engineer
Singapore, Singapore
HybridFull Time
Johnson & Johnson
Johnson & JohnsonNYSE: JNJ: Provides pharmaceutical products and medical technology healthcare solutions.
Masters in CS/Engineering/Business Analytics, proficiency in Python, SQL, Spark; familiarity with LLMs and cloud deployment; strong software engineering and communication skills.
Python, SQL, Apache Spark, Databricks, Fabric, LLM, Jenkins, Git, BitBucket, Spinnaker, Helm
2w
Save
Mark Applied
Hide
Staff Machine Learning Engineer, 3D - Singapore Efficiency Team
Singapore, Singapore, Singapore
S$150k-S$255k/yr OnsiteFull Time
Riot Games
Riot Games: Developing and publishing competitive multiplayer video games.
5+ YOEMaster's/PhD in CS/Statistics/Math with ML focus,5+ years applying ML to real-world problems,deep 3D ML expertise,SIGGRAPH/CVPR-level work or shipped projects,proficiency in Python/C/C++/C#,PyTorch/TensorFlow/JAX,strong experimental design and communication.
Python, C, C++, C#, PyTorch, TensorFlow, JAX
2w
Save
Mark Applied
Hide
Machine Learning Engineer
Singapore, Singapore, Singapore
HybridFull Time
Endava
EndavaNYSE: DAVA: Provides software engineering and digital transformation consulting services.
Hands-on experience delivering ML from experimentation to production; strong Python, infrastructure-as-code (Terraform), Docker/Kubernetes, CI/CD, data tooling, and cloud platform experience; degree in CS/Math/related or equivalent experience.
Terraform, Python, Git, Scikit-learn, TensorFlow, PyTorch, LangChain, LlamaIndex, SQL, Pandas, NumPy, Docker, Kubernetes, Helm, Argo CD, MLflow, Kubeflow, Vertex AI, Google Cloud, AWS, Azure, GitHub Actions, GitLab CI, Azure DevOps, Cloud Build