NVIDIA
Posted 2mo ago

Infrastructure Software Engineer, Deep Learning Libraries

NVIDIA
Shanghai or Beijing
OnsiteFull Time
Responsibilities
  • Designing testing
  • Building automation
  • Deploying AI agents
Requirements
  • Masters in CS/CE or equivalent
  • 3+ years
  • Python and C/C++
  • CI systems (Jenkins
  • GitHub Actions
  • GitLab pipelines
  • Azure DevOps)
  • AI agents experience
  • SCM/build systems (Git, Make, CMake, Bazel)
Technical tools mentioned
PythonC++JenkinsGitHub ActionsGitLab pipelinesAzure DevOpsKubernetesDockerCMakeGitJira

Job description

We are now looking for an Infrastructure Software Engineer for Deep Learning Libraries!

NVIDIA's Deep Learning Libraries Group is seeking excellent software engineers to enable the next wave of NVIDIA’s highest performing deep learning libraries. The role focuses on NVIDIA's open-source products such as CUTLASS. The mission is to design and develop scalable, modular infrastructure that streamlines development, builds, and tests across NVIDIA’s diverse set of platforms, and address the needs from the open-source community, with the cutting-edge AI technology. Join our technically diverse team of software engineers and infrastructure experts to design the systems that enable NVIDIA to stay ahead of the competition as we deliver the world's fastest deep learning platforms.

What you'll be doing:

  • Designing and developing software for testing and analysis of our codebases

  • Building scalable automation for build, test, integration, and release processes for open-source products

  • Developing and deploying AI agents and similar technology to automate the end-to-end software development cycle

  • Configuring, maintaining, and building upon deployments of industry-standard tools (e.g. Kubernetes, Jenkins, Docker, CMake, Gitlab, Jira, etc.)

  • Advancing the state of the art in those industry-standard tools

What we need to see:

  • A Masters Degree in Computer Science or Computer Engineering or equivalent experience.

  • 3+ years of relevant experience

  • Strong programming skills in Python (or similar) and familiarity with C/C++ development

  • Experience setting up, maintaining, and automating continuous integration systems (e.g. Jenkins, GitHub Actions, GitLab pipelines, Azure DevOps)

  • Extensive experience in AI agents technology

  • Fluency in SCM (e.g. Git, Perforce) and build systems (e.g. Make, CMake, Bazel)

Ways to stand out from the crowd:

  • Experience designing and developing automation in Jenkins with Groovy (or similar)

  • Background with distributed systems and cluster/cloud computing, especially with Kubernetes

  • Experience designing and developing unit and integration test frameworks

  • Close follow the latest trend in AI industry

  • Track record of identifying useful new technologies and incorporating them into SW development flows

This is an opportunity to have a wide impact at NVIDIA by improving development velocity across our many AI/DL/Compute Software projects. Are you creative, driven, and autonomous? Do you love a challenge? If so, we want to hear from you!

About NVIDIA

Designs GPU-accelerated computing and artificial intelligence hardware.

Similar jobs

Infrastructure Software Engineer roles near Shanghai, /
1mo
Save
Mark Applied
Hide
AI Infrastructure Software Engineer — CosmosLab
Beijing or Shanghai or Shenzhen
OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years building large-scale distributed or AI training infrastructure, Bachelor's in CS or equivalent, strong debugging across stack, proficiency in Python and software engineering practices, experience with distributed training/inference.
PyTorch, Megatron, Python, C, C++, CUDA, FSDP, DTensor
1d
Save
Mark Applied
Hide
AI Infra软硬件结合开发工程师
Beijing or Hangzhou or Shanghai or Shenzhen or China
OnsiteFull Time
Alibaba Group
Alibaba GroupNYSE: BABA: Global technology specializing in e-commerce and cloud computing.
Requires computer architecture, OS, networking, distributed systems, FPGA/ASIC or embedded development, Verilog/VHDL, C/C++/Python/Go, driver or firmware development, CUDA kernel development, and cloud computing knowledge.
FPGA, ASIC, C-Model, Hardware Abstraction Layer (HAL), CPU, GPU, Linux, SDK, CLI, IDE, Verilog, VHDL, C, C++, Python, Go, CUDA, GitHub, RAG, MCP, vLLM, Ollama, SFT, RL
1d
Save
Mark Applied
Hide
AI Infra软硬件结合开发工程师
Beijing or Hangzhou or Shanghai or Shenzhen
OnsiteFull Time
Alibaba
AlibabaNYSE: BABA: Provides online marketplaces, cloud computing, and digital payment services.
Requires computer architecture, operating systems, networking, distributed systems, FPGA/ASIC or embedded development, Verilog/VHDL, C/C++, Python or Go, driver and firmware development, distributed computing, and CUDA kernel experience.
FPGA, ASIC, C-Model, Hardware Abstraction Layer (HAL), CPU, GPU, Linux, SDK, CLI, IDE, Verilog, VHDL, C, C++, Python, Go, CUDA, GitHub, vLLM, Ollama, RAG, MCP, SFT, RL
1d
Save
Mark Applied
Hide
AI Infra软硬件结合开发工程师
Beijing or Hangzhou or Shanghai or Shenzhen
RemoteFull Time
Alibaba Cloud
Alibaba CloudNYSE: BABA: Global cloud computing infrastructure and services provider
Graduate role requiring hardware/software fundamentals, FPGA or ASIC or embedded development, Verilog/VHDL, C/C++/Python/Go, Linux, drivers, firmware, distributed computing, CUDA, and AI infrastructure knowledge.
FPGA, ASIC, C-Model, Hardware Abstraction Layer (HAL), CPU, GPU, Linux, SDK, CLI, IDE, Verilog, VHDL, C, C++, Python, Go, CUDA, GitHub, RAG, MCP, vLLM, Ollama, SFT, RL