Canva
Posted 1mo ago

Research MLE (Training Optimization)大模型训练优化工程师

Canva
Beijing, Beijing, China
HybridFull Time
Responsibilities
  • designing systems
  • optimizing training
  • writing kernels
Requirements
  • Experience with large-model distributed training
  • Proficiency in Python
  • PyTorch/JAX
  • CUDA/Triton
  • Familiarity with Megatron-LM/NeMo/DeepSpeed
  • Strong systems and optimization skills, and full professional English
Technical tools mentioned
Megatron-LMNVIDIA NeMoFSDPTritonPythonC++RustPyTorchJAXDeepSpeedCUDA

Job description

Company Description:

At Canva, we're building a future powered by AI that's as magical as it is impactful. As a Research Scientist at Canva, you'll be responsible for advancing the future of AI by experimenting with cutting-edge techniques, as well as improving models for real-world quality and performance.

Job Description:

About the Group/Team


We're the CORE team within the Generative AI supergroup. Our mission is to invent foundational technologies that will power the future of AI-assisted design. From large-scale models to groundbreaking research, our team builds the technical core of Canva’s creative intelligence engine. We collaborate globally to ship research that makes a real impact—from smart editing to AI video tools—at massive scale.

 

About the Role/Specialty


As a Machine Learning Engineer, you’ll lead efforts to scale and optimize the training system for our large-scale multimodal and foundation models. You’ll design distributed training systems using Megatron-LM, NVIDIA NeMo, FSDP, and Triton—pushing the limits of performance across compute, memory, and communication layers. You'll sit at the intersection of systems and AI research, directly shaping how we train the models that will power Canva’s next generation of products.

 

What you’ll do (responsibilities)

  • You’ll design, implement, and optimize large-scale machine learning systems for training
  • You’ll improve all aspects of performance, including GPU utilization, communication overhead, and memory efficiency.
  • You’ll partner with research and modeling teams to align systems with algorithmic needs.
  • You’ll evaluate and apply best practices for distributed training using industry-leading frameworks.
  • You’ll dive deep into low-level optimization, including custom CUDA or Triton kernels.

• • You’ll debug, profile, and fine-tune training workflows to unlock new levels of scalability.

Qualifications:

What we're looking for

We’re looking for a systems-first engineer who thrives in fast-paced, high-impact environments. You’re deeply familiar with distributed model training at scale and understand the nuances of optimizing compute at every level of the stack. You're excited by challenges that stretch current boundaries, and you’re a strong collaborator who communicates clearly across domains.

  • Strong background in LLMs, multimodal AI, or diffusion models.
  • Proficiency in Python. Familiarity with a system programming language (e.g. C++ or Rust) is a plus.
  • Deep knowledge of PyTorch or JAX as well as libraries such as Megatron-LM, NeMo, or DeepSpeed.
  • Familiarity with common optimization techniques such as FSDP/ZeRO, gradient checkpointing, or low-precision data types.
  • Hands-on experience writing custom GPU kernels in CUDA or Triton.
  • Excellent communication and problem-solving skills, incl. full proficiency in English.
Additional Information:

大模型训练优化工程师(多模态/图像生成),技术要求:算子优化/分布式训练/GPU集群/训练框架。该岗位面向所有经验阶段的候选人开放,包括社会招聘、应届毕业生,同时开放实习生岗位。

About Canva

Online graphic design platform for creating and editing visual content.

Year founded
2013
Employees
13718
Organization type
Private
Latest investment
Raised $560.00M Series D (2025) — led by Franklin Templeton
Headquarters
AU

Similar jobs

Machine Learning Engineer roles near Beijing, Beijing
1d
Save
Mark Applied
Hide
Senior Machine Learning Engineer, MLE
Beijing, Beijing, China
OnsiteFull Time
Grab
GrabNASDAQ: GRAB: Superapp providing transportation, food delivery, and digital financial services.
2+ YOERequires 2+ years of machine learning engineering experience, computer science fundamentals, large-scale microservices and production data pipeline experience, coding ability, and willingness to use Golang.
Golang, Kafka, Flink, Spark Streaming, RAG, AWS, Azure
5d
Save
Mark Applied
Hide
Staff Machine Learning Engineer
Beijing or China
RemoteFull Time
BJAK
BJAK: Online platform for insurance comparison and road tax renewal.
Experience shipping real ML systems, working with large models, writing production-grade code, and owning scalable training, inference, evaluation, and deployment systems.
Python, PyTorch, JAX, GPU, LoRA, QLoRA, SFT, DPO
6d
Save
Mark Applied
Hide
算法工程师(大模型) - TikTok研发
Shanghai or Beijing or Hangzhou or Guangzhou or Shenzhen
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Requires strong coding, data structures, algorithms, and machine learning expertise; proficiency in C/C++ or Python; experience with large language models, agents, or related AI methods.
C, C++, Python, SFT, DPO, PPO, GRPO, NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, ICCV, KDD, WWW
6d
Save
Mark Applied
Hide
算法工程师-大模型后训练 (Post-training)
Beijing or Hangzhou or Shanghai
RemoteFull Time
Alibaba Group
Alibaba GroupNYSE: BABA: Global technology specializing in e-commerce and cloud computing.
Deep understanding of large-model post-training, data construction, reinforcement learning, Python, PyTorch, and LLM frameworks such as vLLM and verl; strong research, engineering, and collaboration skills.
Python, PyTorch, vLLM, verl, Claude Code, Hermes, Claw, CI
6d
Save
Mark Applied
Hide
下一代大模型智能体框架与数据技术-阿里星
Beijing or Hangzhou or Shanghai
OnsiteFull Time
Alibaba
AlibabaNYSE: BABA: Provides online marketplaces, cloud computing, and digital payment services.
Recent PhD or top master's graduate in computer science, AI, mathematics, or related field; expertise in LLMs, reinforcement learning, agents, data processing, and large-scale model or data engineering.
PyTorch, NeurIPS, ICML, ICLR, KDD, ACL, VLDB, SIGMOD
6d
Save
Mark Applied
Hide
算法工程师-机器学习
Beijing or Guangzhou or Hangzhou or Shanghai or China
OnsiteFull Time
Alibaba Cloud
Alibaba CloudNYSE: BABA: Global cloud computing infrastructure and services provider
Strong machine learning and deep learning knowledge; excellent engineering ability; proficiency in C/C++, Java, or Python; statistics, logical analysis, learning ability, communication, and teamwork skills.
C, C++, Java, Python, Linux, KDDCUP, ImageNet, MSCOCO, ICDAR
1w
Save
Mark Applied
Hide
算法-机器学习方向
Shenzhen or Beijing or Shanghai or Guangzhou or Chengdu
OnsiteFull Time
Tencent
TencentHong Kong Stock Exchange: 0700: Multinational technology conglomerate providing internet and entertainment services.
PhD or excellent master's degree in a relevant technical field; strong machine learning, deep learning, reinforcement learning, probability, statistics, and optimization knowledge; proficiency in C/C++, Java, or Python.
C, C++, Java, Python, Spark, XGBoost, Caffe, TensorFlow
1w
Save
Mark Applied
Hide
Staff Machine Learning Engineer - Moloco Commerce Media
Menlo Park or Seattle or New York City or San Francisco or Seoul or Beijing or Singapore or Gurgaon or Tokyo or Shanghai or London or Tel Aviv or Berlin
$232k-$348k/yr OnsiteFull Time
Moloco
Moloco: AI-powered programmatic advertising and commerce media platform.
8+ YOERequires 8–12 years in applied machine learning or ML engineering, production model experience, Python, modern ML frameworks, experimentation, distributed systems, and technical leadership.
Python, PyTorch, TensorFlow, JAX, feature stores, distributed training, large-scale data processing frameworks, MLOps, A/B testing, transformer-based architectures