NIO
Posted 11mo ago

高性能计算实习生

NIO
Beijing, Beijing, China
OnsiteInternship
Responsibilities
  • developing operators
  • optimizing algorithms
  • ensuring efficiency
Requirements
  • Familiarity with ARM/DSP/CUDA parallel programming
  • Strong C/C++ skills
  • Knowledge of algorithms and processor architectures
  • Preferred: performance analysis
  • Debugging
  • Deep learning/CV and quantization
Technical tools mentioned
Nvidia GPUARM CPUCUDACC++

Job description

基于 Nvidia GPU/ARM CPU 等架构特性完成深度学习算子、CV 算法及计算库的开发及优化;
与深度学习引擎前端团队一起,共同完成深度学习引擎在端侧落地,保证算法运行的高效性及实时性。

岗位职责:

基本要求:
熟悉 ARM/DSP/CUDA 并行编程模型;
熟悉 C/C++ 编程,了解常用算法及数据结构;
了解 Nvidia GPU/ARM CPU 处理器体系结构,理解软件与硬件底层映射关系。
加分项:
熟悉程序性能分析、问题定位及调试方法
熟悉深度学习算法或常见 CV 算法;
熟悉深度学习量化方法。

About NIO

Manufacturer of smart electric vehicles and battery swapping systems.

Similar jobs

High Performance Computing Intern roles near Beijing, Beijing
5mo
Save
Mark Applied
Hide
高性能计算实习生
Beijing, Beijing, China
OnsiteInternship
Xiaomi
XiaomiHong Kong Stock Exchange: 1810: Designs and manufactures smartphones, electric vehicles, and consumer electronics.
Familiarity with CUDA, ability to analyze and optimize GPU code, proficiency in Python and C++, knowledge of algorithms and data structures, engineering skills with git/ssh/cmake, and experience with distributed training or NPU deployment is a plus.
CUDA, Python, C++, Git, SSH, CMake
1y
Save
Mark Applied
Hide
高性能计算实习生
Beijing, Beijing, China
OnsiteInternship
北京科学智能研究院
北京科学智能研究院: A research institute focused on AI for Science innovation.
Master's or above in CS/applied math/computational materials (strong bachelors considered). Proficient in CUDA, GPU profiling (Nsight/VTune), PyTorch/TensorFlow optimization, MPI/OpenMP, experience with large-scale distributed training.
CUDA, Nsight, VTune, PyTorch, TensorFlow, MPI, OpenMP, DeePMD-kit
5mo
Save
Mark Applied
Hide
高性能计算研发实习生-Data 语音
Beijing, Beijing, China
OnsiteInternship
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Pursuing 2027 graduation, strong Python and C++ skills, experience in GPU programming or model quantization/sparsity, familiarity with distributed inference and vLLM/Triton ecosystems.
Python, C++, CUDA, Triton, AscendC, TileLang, vLLM, SGLang, PCIe