XPeng
Posted 3d ago

Staff Machine Learning Engineer - LLM Quantization & Deployment

XPeng
Santa Clara or Mountain View
$215k-$364k/yrOnsiteFull Time
Responsibilities
  • developing inference models
  • building deployment pipelines
  • analyzing numerical errors
Requirements
  • Master's in CS, CE, or EE with 3–5 years' industry experience
  • Expertise in Transformer architectures
  • LLM inference
  • Model quantization
  • PyTorch
  • Inference stacks
  • Python, and software engineering
Technical tools mentioned
PythonPyTorchTensorRT-LLMvLLMSGLangllama.cppONNX RuntimeTVMMLIRAWQGPTQSmoothQuantINT8FP4

Job description

XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and landing (eVTOL) aircraft, and robotics. With a strong focus on intelligent mobility, XPENG is dedicated to reshaping the future of transportation through cutting-edge R&D in AI, machine learning, and smart connectivity.
 
Our mission is to build strong foundation for LLM deployment and quality sign-off for next-gen XPENG Turing AI chip. This includes and is not limited to: LLM model fine tuning, PTQ, QAT, on-vehicle inference and related fields.
 

Key Responsibilities

  • Develop VLA inference models, ensure numerical consistency with training models, and productionize LLM quantization methods, including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques.
  • Develop production-quality Python code with strong testing, observability, reproducibility, and failure handling.
  • Build robust model export, calibration, benchmarking, validation, and deployment pipelines.
  • Engage early with the VLA model research team to establish performance estimates and prove model feasibility.
  • Curate evaluation datasets and establish a comprehensive metric suite to systematically benchmark VLA performance.
  • Analyze numerical errors, accuracy regressions, and performance trade-offs.
  • Develop PTQ and QAT orchestration workflows.
  • Serve as the primary interface with field-testing and simulation teams for issue triage and autonomous driving performance sign-off.
  • Collaborate with the in-vehicle software team on latency analysis and issue triage.
  • Collaborate with the training infrastructure team to develop QAT and model distillation.

Basic Qualifications

  • Master in CS/CE/EE, or equivalent, with 3-5 years of industry experience.
  • Strong understanding of Transformer architectures and LLM inference.
  • Hands-on experience quantizing or deploying deep learning models in production.
  • Proficiency with PyTorch and at least one inference or compilation stack.
  • Strong Python programming and software engineering skills.
  • Ability to work effectively across research, systems, infrastructure, and product teams.
  • Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment.

Preferred Qualifications

  • Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization.
  • Experience with AWQ, GPTQ, SmoothQuant, or related methods.
  • Strong numerical analysis and systems engineering skills.
  • Experience with one or more LLM runtimes, such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes.
  • Experience deploying LLMs on resource-constrained or heterogeneous hardware.
  • Contributions to model optimization, inference, compiler, or serving projects.
  • Publications at NeurIPS, ICML, ICLR, ACL, or related conferences.

What We Provide

  • A fun, supportive and engaging environment.
  • Infrastructures and computational resources to support your work.
  • Opportunity to work on cutting edge technologies with the top talents in the field.
  • Opportunity to make a significant impact on the transportation revolution by the means of advancing autonomous driving.
  • Competitive compensation package.
  • Snacks, lunches, dinners, and fun activities.
 
The base salary range for this full-time position is $215,280 - $364,320, in addition to bonus, equity and benefits. Our salary ranges are determined by role, level, and location. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.
 
We are an Equal Opportunity Employer. It is our policy to provide equal employment opportunities to all qualified persons without regard to race, age, color, sex, sexual orientation, religion, national origin, disability, veteran status or marital status or any other prescribed category set forth in federal or state regulations.

About XPeng

Designs and manufactures smart electric vehicles and autonomous technology.

Similar jobs

Machine Learning Engineer roles near Santa Clara, California
8h
Save
Mark Applied
Hide
Machine Learning Engineer Graduate (E-Commerce Governance) - 2027 Start
San Jose, California, United States
$128k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's degree in software development, computer science, computer engineering, or related field; strong Python, Go, or C++ skills; Linux development; data structures, algorithms, machine learning, and deep learning knowledge.
Linux, Python, Go, C++, natural language processing, computer vision, multimodal, graph algorithms, search algorithms, text mining, data mining, LLM, reinforcement learning, operational research
17h
Save
Mark Applied
Hide
2027 - Full Time Assistant Vice President - Corporate Functions, Analytics Lab (Machine Learning Engineer)
Jersey City or Montreal or Paris or Lisbon or London or New York or Chesterbrook or San Francisco or Boston or Chicago or Denver or Miami or Washington
$150k/yr OnsiteFull Time
BNP Paribas
BNP ParibasEuronext Paris: BNP: Provides global banking, financial, and investment services.
Master’s degree in a quantitative field; December 2026–June 2027 graduation; Python or C, SQL, statistical modeling, data analysis, LLM/agentic workflows, ML frameworks, RAG systems, Git, and communication skills.
Python, C, SQL, LangChain, CrewAI, Google ADK, OpenAI Agent SDK, TensorFlow, PyTorch, scikit-learn, Git, DVC, MLflow, Jenkins, Docker
1d
Save
Mark Applied
Hide
Senior Principal Machine Learning Engineer
San Jose, California, United States
$220k-$300k/yr OnsiteFull Time
SambaNova Systems
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
8+ YOEBachelor's degree in computer science, electrical engineering, or related field; 8+ years of ML engineering experience; deep LLM expertise; technical leadership and complex project delivery experience.
Large Language Models (LLMs), RDU, SambaStack, SambaCloud, speculative decoding, reinforcement learning, mixture-of-experts
2d
Save
Mark Applied
Hide
Founding Engineer - Machine Learning
Mountain View, California, United States
$220k-$300k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
3+ YOERequires 3–10 years of ML engineering, applied science, or research engineering experience; Python and PyTorch, TensorFlow, or JAX; distributed systems, cloud ML infrastructure, MLOps, and large-scale data experience.
Python, PyTorch, TensorFlow, JAX, AWS, GCP, Azure, Weights & Biases, MLflow
2d
Save
Mark Applied
Hide
Machine Learning Engineer
San Francisco, California, United States
$170k-$300k/yr HybridFull Time
Clay
Clay: Platform for automated lead enrichment and sales workflows.
5+ YOE5+ years in machine learning engineering or ML-heavy software engineering, production ML features, strong coding, LLMs or classical ML, data-intensive systems, and product-oriented problem solving.
LLMs, Snowflake, dbt, Dagster
2d
Save
Mark Applied
Hide
Machine Learning Engineer - Satellite Capacity Optimization & Planning
Carlsbad or San Jose or San Francisco or New York City
$141k-$222k/yr OnsiteFull Time
Viasat
ViasatNASDAQ: VSAT: Provider of global satellite-based connectivity and secure communication solutions.
7+ YOERequires 7+ years in ML or optimization, optimization techniques, production cloud ML systems, ML frameworks, geospatial visualization, containerized development, SQL, RESTful APIs, and up to 10% travel.
AWS, GCP, TensorFlow, PyTorch, SQL, RESTful APIs, Airflow, Athena, BigQuery, ECS, Batch
3d
Save
Mark Applied
Hide
(USA) Staff, Machine Learning Engineer
Sunnyvale, California, United States
$169k-$338k/yr OnsiteFull Time
Walmart
WalmartNYSE: WMT: Operates a chain of hypermarkets, discount stores, and grocery stores.
4+ YOEBachelor's degree and 4 years, or 6 years of experience, in software engineering, machine learning engineering, AI systems, or a related field. Requires ML, MLOps, cloud, and programming expertise.
TensorFlow, PyTorch, Scikit-learn, Python, SQL, AWS, GCP, Azure, Kubernetes, CI/CD, MLOps, Web Content Accessibility Guidelines (WCAG) 2.2 AA
3d
Save
Mark Applied
Hide
(USA) Staff, Machine Learning Engineer
Sunnyvale, California, United States
$169k-$338k/yr OnsiteFull Time
Walmart
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
4+ YOEBachelor's degree plus 4 years or 6 years of experience in software engineering, machine learning engineering, AI systems, or related areas. Requires ML, MLOps, cloud, Kubernetes, and technical leadership expertise.
TensorFlow, PyTorch, Scikit-learn, Python, SQL, AWS, GCP, Azure, Kubernetes, CI/CD