Intel
Posted 5d ago

Software Engineer — Distributed LLM Inference Systems

Intel
Shanghai, Shanghai, China
OnsiteFull Time
Responsibilities
  • developing inference systems
  • optimizing model execution
  • profiling workloads
Requirements
  • Master’s degree in computer science
  • Artificial intelligence
  • Software engineering, or related field
  • 0–1 years’ experience
  • Python
  • Modern C++
  • PyTorch
  • Performance optimization, and English proficiency
Technical tools mentioned
PythonC++PyTorchvLLMSGLangTensorRT-LLM

Job description

Job Details:

Job Description: 

The Role and Impact:

  • As a Software Engineer on Intel’s Artificial Intelligence Frameworks team, you will contribute to designing, developing, and optimizing distributed inference systems for large language models.
  • Your day-to-day work will involve implementing distributed inference algorithms, optimizing model execution and communication, transforming neural network models, and developing software components that improve inference performance across diverse hardware architectures. You may work on areas such as disaggregated serving, request scheduling, KV cache management, parallel execution, and efficient communication between inference components.
  • By collaborating with researchers and engineers, you will play a key role in advancing Intel's AI capabilities and ensuring industry-leading solutions.

Business Group: Intel's Artificial Intelligence Frameworks team is dedicated to empowering transformative AI solutions by developing and optimizing software frameworks for machine learning and deep learning. This group works on enhancing the performance of AI applications across diverse computing hardware backends while contributing to open-source communities. As part of Intel, this team supports the mission to drive technological innovation and deliver impactful AI advancements globally.

Key Responsibilities

  • Design, develop, and optimize distributed LLM inference systems and related AI framework components.
  • Implement distributed algorithms, including model/data parallel frameworks and asynchronous communication for deep learning.
  • Develop and optimize components such as request schedulers, model workers, communication layers, and KV cache management mechanisms.
  • Profile distributed inference workloads to identify computation, communication, memory, and scheduling bottlenecks.
  • Collaborate with component teams to improve end-to-end latency, throughput, scalability, and resource utilization.
  • Contribute high-quality code, tests, and documentation to internal and open-source projects while following industry engineering standards.

Qualifications:

Minimum Qualifications –

  • Master’s degree in computer science, Artificial Intelligence, Software Engineering, or a related field, with 0-1 years of hands-on experience demonstrated through internships, academic projects, coursework, or training.
  • Proficiency in Python and modern C++ programming.
  • Foundational knowledge of deep learning and AI frameworks, such as PyTorch.
  • Experience debugging and optimizing software for performance.
  • Basic understanding of machine learning algorithms and techniques.
  • Strong problem-solving skills and the ability to learn unfamiliar systems quickly.

Preferred Qualifications

  • Experience or project exposure related to distributed LLM inference and serving.
  • Experience in contributing to open-source projects or collaborating within open-source ecosystems.
  • Understanding of LLM inference concepts such as prefill and decode, KV cache management, continuous batching, parallelism strategies, and disaggregated serving.
  • Familiarity with inference engines or serving frameworks such as vLLM, SGLang, TensorRT-LLM, or similar technologies.
  • Knowledge of large language models and inference optimization techniques.
  • Knowledge of AI Agent architecture and execution workflows, including tool calling, planning, memory, context management, and multi-agent coordination.
  • Effective communication skills, including fluency in written and spoken English. Take the opportunity to be part of Intel's journey in redefining AI frameworks and enabling groundbreaking innovations.


Take the opportunity to be part of Intel's journey in redefining AI frameworks and enabling groundbreaking innovations. Your contributions will shape the future of AI software and its real-world impact. Apply now and be a part of advancing transformative AI capabilities.

          

Job Type:

College Grad

Shift:

Shift 1 (China)

Primary Location: 

PRC, Shanghai

Additional Locations:

Posting Statement:

All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.

Position of Trust

N/A

Work Model for this Role

This role will require an on-site presence. * Job posting details (such as work model, location or time type) are subject to change.

*

ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.

About Intel

Designs and manufactures microprocessors and semiconductor components.

Similar jobs

Software Engineer roles near Shanghai, Shanghai
3h
Save
Mark Applied
Hide
Software Engineer (Data Solutions), IS&T Ai & Data Platforms
Shanghai, Shanghai, China
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experienced software engineer to build scalable, resilient distributed systems for cloud analytics platforms and data pipelines in a fast-paced environment.
Kafka, Spark, Iceberg, Airflow, Presto
12h
Save
Mark Applied
Hide
Senior Software Engineer, Planning (Robotics)
Shanghai or Singapore
OnsiteFull Time
Grab
GrabNASDAQ: GRAB: Superapp providing transportation, food delivery, and digital financial services.
3+ YOEMaster's degree or higher in a relevant field; 3–5 years in autonomous driving or robotics planning algorithms; strong C++, Python, Linux, planning methods, simulation, and vehicle testing experience.
C++, Python, Linux, ROS, ROS2, Apollo, Autoware, CARLA, NVIDIA Isaac, Isaac Sim, Isaac Lab, Lattice, DP/QP, Hybrid A*, POMDP
16h
Save
Mark Applied
Hide
Senior Software Engineer
Shanghai, Shanghai, China
OnsiteFull Time
Optiver
Optiver: Global market maker providing liquidity to financial markets.
3+ YOEAt least 3 years of software engineering experience, large-scale server-side development, high-throughput low-latency optimization, architectural recommendations, and complex technical problem-solving skills.
AI agents, LLMs
1d
Save
Mark Applied
Hide
Software Engineer II
Shanghai, Shanghai, China
OnsiteFull Time
Cadence Design Systems
Cadence Design SystemsNASDAQ: CDNS: Develops computational software and hardware for electronic system design.
2+ YOEBachelor's degree in computer science or a related technical field plus 2+ years' experience, or a master's degree. Requires C/C++, object-oriented programming, algorithms, data structures, and strong problem-solving and communication skills.
C, C++, PCI Express (PCIe), CXL, Verilog, SystemVerilog, OVM, UVM
2d
Save
Mark Applied
Hide
【2027 Campus】Software Engineer
Shanghai, Shanghai, China
OnsiteFull Time
KLA
KLANASDAQ: KLAC: Provides process control and yield management for semiconductor manufacturing.
0+ YOEBachelor's degree with 2 years' experience or master's degree with no experience; object-oriented programming, algorithms, data structures, software design, debugging, teamwork, and fluent English required.
C, C++, C#, Java, PHP, Python, LabVIEW, JavaScript, Windows, Linux
2d
Save
Mark Applied
Hide
Senior Software Engineer - RDMA and Doca
Shanghai or Beijing
OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEBachelor's degree or higher in computer science, computer engineering, or related field; 5+ years' experience; strong C/C++ and Linux skills; networking and virtualization expertise.
C, C++, Linux, RDMA
2d
Save
Mark Applied
Hide
Senior Software Engineer - AI Networking
Shanghai or Beijing
OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEBachelor's degree or equivalent in computer science, computer engineering, or related discipline; 5+ years' experience; strong C/C++, Linux, networking, and virtualization skills.
C, C++, Linux, DPDK, RDMA, UCX, NCCL, DeepEP, SONiC, vLLM, SGlang, VLAN, STP, OSPF, BGP, PIM
3d
Save
Mark Applied
Hide
Eng, Software Engrg
Shanghai, Shanghai, China
OnsiteFull Time
Carrier
CarrierNYSE: CARR: Provides HVAC, building automation, and refrigeration solutions globally.
2+ YOERequires a bachelor's degree or higher in automation, electrical engineering, building automation, or HVAC, with 2+ years of relevant project implementation experience preferred.
PLC, BACnet, Modbus