2 infrastructure reliability engineer jobs at 1 company in Orange, MA

1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Austin or Westford or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE6+ years systems programming experience; BS/MS/PhD in CS or EE (or equivalent); experience building RCA pipelines for HPC/cloud; deep CPU/GPU architecture knowledge; strong C++ and Python; familiarity with Slurm/LSF/Kubernetes.
C++, Python, Slurm, LSF, Kubernetes, CUDA, DCGM, NVML, CRIU, /dev/mcelog, dmesg, journald, Linux kernel
1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald