Smart Logic Tech
Posted 1y ago

分布式推理工程师(Distributed Inference Engineer)

Smart Logic Tech
Shanghai or Beijing or Hangzhou or Chengdu
OnsiteFull Time
Responsibilities
  • developing systems
  • optimizing performance
  • deploying models
Requirements
  • Expertise in Python/C++ and deep learning frameworks
  • Distributed systems and CUDA
  • 3+ years AI model deployment or distributed systems experience
  • Bachelor\u0002s in CS/Electrical/Math required
  • Strong performance optimization skills
Technical tools mentioned
PythonC++PyTorchTensorFlowTorchScriptTriton Inference ServerTensorRTONNX RuntimevLLMSglangFastAPIgRPCNCCLGPU Direct RDMACUDAOpenMP

Job description

职位名称
分布式推理工程师(Distributed Inference Engineer)
岗位职责
分布式推理系统开发
设计并实现大规模AI模型(如LLM、多模态模型)的分布式推理框架,优化吞吐量(Throughput)和延迟(Latency)。
开发模型并行(Tensor/Pipeline Parallelism)、动态批处理(Dynamic Batching)和连续批处理(Continuous Batching)策略。
解决多节点、多GPU环境下的负载均衡和通信瓶颈问题。
高性能推理优化
优化计算图执行(如使用TensorRT、ONNX Runtime、vLLM等工具)。
实现量化(FP8/INT8)、KV Cache优化、Flash Attention等加速技术。
针对CPU/GPU/NPU等硬件进行底层性能调优(如CUDA内核优化、内存管理)。
模型部署与工具链建设
构建自动化模型部署流水线(推理的端到端Pipeline)。
开发监控和调试工具,实时跟踪推理性能(显存占用、计算利用率等)。
跨团队协作
与算法团队合作,针对推理场景优化模型架构(如减少计算量、优化算子)。
与运维团队协作,设计高可用、可扩展的推理服务架构。

技术要求
必备技能
精通 Python/C++,熟悉PyTorch/TensorFlow推理生态(如TorchScript、Triton Inference Server)。
深入理解 分布式系统(NCCL/RPC通信、GPU Direct RDMA)和 并行计算(CUDA、OpenMP)。
熟悉 推理加速技术(量化、剪枝、算子融合)和框架(TensorRT、vLLM 、Sglang)。
有 大规模模型服务化 经验(如使用FastAPI、gRPC构建高并发API)。
加分项
熟悉LLM推理优化(PagedAttention、Speculative Decoding)。
有国产硬件适配经验。
参与过开源推理框架(如GGML、TGI)贡献。
论文/专利成果(涉及模型压缩、分布式推理等方向)。

任职要求
学历:计算机/电子工程/数学等相关专业,本科及以上学历。
经验:3年以上AI模型部署或分布式系统开发经验。
特质:对性能优化有极致追求,能快速定位系统瓶颈。

工作地点
地点:北京/上海

About Smart Logic Tech

Designs and develops high-performance reconfigurable processor chip architecture.

Year founded
2016
Employees
500
Organization type
Private
Latest investment
Raised $100.00M Private Equity (2022) — led by China International Capital Corporation (CICC), Jinshi Investment, Zhongxin Juyuan
Headquarters
CN