Thundersoft
Posted 10mo ago

推理部署工程师(北京上海)

Thundersoft
Beijing or Shanghai
OnsiteFull Time
Responsibilities
  • deploying models
  • optimizing inference
  • monitoring services
Requirements
  • Expertise in Python/C++ and deep learning (PyTorch/TensorFlow)
  • Experience deploying and optimizing large-scale model inference
  • Familiarity with Linux, Docker, Kubernetes, CUDA/OpenCL
  • Monitoring and distributed system skills
Technical tools mentioned
PythonC++PyTorchTensorFlowTensorRTONNX RuntimeLinuxDockerKubernetesCUDAOpenCLPrometheusGrafanaELKRPC

Job description

1 负责AI模型在生产环境中的高效推理部署,优化模型推理性能(延迟、吞吐量、资源利用率等)。
2 设计并实现分布式推理架构,支持高并发、低延迟的实时推理服务。
3 与算法团队协作,完成模型压缩、量化、剪枝等优化,提升模型在边缘端/云端部署的可行性。
4 监控推理服务稳定性,定位并解决线上性能瓶颈和故障问题。
5 探索新兴推理框架(如TensorRT、ONNX Runtime等)在生产环境的应用落地。

1 精通Python/C++,熟悉深度学习框架(PyTorch/TensorFlow),了解模型部署全流程(训练→优化→部署)。
2 熟悉Linux系统调优、Docker容器化、Kubernetes集群管理。
3 具备分布式系统设计经验,熟悉RPC、消息队列等技术。
4 了解GPU/TPU加速推理原理,掌握CUDA/OpenCL等并行计算技术。
​工程能力
1 有大规模AI模型推理服务落地经验(如推荐系统、NLP、CV等场景)。
2 熟悉模型监控工具(Prometheus、Grafana)及日志分析(ELK)。
3 具备A/B测试、灰度发布等工程实践经验。

About Thundersoft

Global provider of intelligent operating system technologies and solutions.

Similar jobs

Inference Deployment Engineer roles near Beijing, Beijing
1mo
Save
Mark Applied
Hide
AI推理部署工程师(具身智能方向)
Beijing or Shanghai or Hangzhou
OnsiteFull Time
辉羲智能
辉羲智能: Designs high-performance automotive AI chips and autonomous driving systems.
Proficient in C++ and Python, experience with model deployment (ONNX, TensorRT, SNPE), performance analysis tools (perf, Nsight), quantization (PTQ/QAT) and solving accuracy drop issues.
C++, Python, perf, Nsight, ONNX, TensorRT, PyTorch, SNPE
4mo
Save
Mark Applied
Hide
视觉大模型推理部署工程师-智能创作(北京/上海/杭州/深圳)
Beijing or Shanghai or Hangzhou or Shenzhen
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
3+ YOE3+ years backend/AI/distributed systems experience; computer-related degree; proficient in Python and Go; experience with large-model deployment, GPU/NPU clusters, and distributed high-concurrency systems.
Python, Go, GPU, NPU, ComfyUI, vLLM, SGLang, Ray
1mo
Save
Mark Applied
Hide
【27届校招】模型部署与推理优化工程师
Beijing, Beijing, China
OnsiteFull Time
Dexmal
Dexmal: An AI focusing on model deployment and inference optimization.
Bachelor's degree in CS/AI/software/electronic info, proficiency in Python or C++, familiarity with PyTorch/TensorFlow/ONNX/TensorRT/OpenVINO, knowledge of deep learning, and interest in model engineering and performance optimization.
Python, C++, PyTorch, TensorFlow, ONNX, TensorRT, OpenVINO, CUDA
4mo
Save
Mark Applied
Hide
模型部署与推理优化工程师
Beijing, Beijing, China
OnsiteFull Time
Dexmal
Dexmal: Develops end-to-end embodied AI software and robotics hardware.
3+ YOEDeploy and optimize large embodied-intelligence models for edge/robot platforms; implement quantization, pruning, distillation, toolchains, and hot-update/version rollback mechanisms.
TensorRT, ONNX Runtime, Python, C++
7mo
Save
Mark Applied
Hide
具身大模型推理&部署工程师
Beijing or Hangzhou
OnsiteFull Time
Spirit AI
Spirit AI: Developing embodied AI and robotics for autonomous physical systems.
Expertise in GPU architecture and CUDA/triton, proficiency in C++ and Python, experience with PyTorch/ONNX/TensorRT and inference optimization (quantization, pruning, distillation), and building end-to-end inference systems for edge devices.
CUDA, triton, C++, Python, PyTorch, ONNX, TensorRT, torch compile, Kernel Fusion, Cuda Graph
4mo
Save
Mark Applied
Hide
大模型推理与部署优化工程师
Beijing, Beijing, China
OnsiteFull Time
ModelBest
ModelBest: Develops efficient, on-device large language models and edge AI.
2+ YOE2+ years experience deploying and optimizing large language/multimodal models; familiarity with vLLM, TensorRT-LLM, SGLang, LightLLM, CUDA or CANN; experience with model serving, container orchestration, and performance debugging.
vLLM, TensorRT-LLM, SGLang, LightLLM, Docker, Kubernetes, Nsight Systems, PyTorch Profiler, CUDA, CANN
8mo
Save
Mark Applied
Hide
AI推理部署与应用框架开发资深工程师
Beijing or Shanghai or Nanjing
OnsiteFull Time
Houmo AI
Houmo AI: Design and manufacture of AI-powered computing chips.
Design and develop AI inference service frameworks, adapt to open-source ecosystems, implement large-model Agent frameworks, and optimize AI applications. Proficiency in C++/Python/Go and networking protocols required.
vLLM, SGLang, PyTorch, ONNX Runtime, OpenAI API, C++, Python, Go, HTTP, SSE, WebSocket, Kubernetes