3 infrastructure reliability engineer jobs at 2 companies in Hooksett, NH

1mo
Save
Mark Applied
Hide
Senior DevOps Site Reliability Engineer
North Andover, Massachusetts, United States
HybridFull Time
Reynolds and Reynolds
Reynolds and Reynolds: Provides software and services for automotive retailers.
5+ YOE5+ years in DevOps/SRE with hands-on AWS, CI/CD (Jenkins), automated deployments for on-prem and cloud, Windows IIS and Linux administration, scripting (Bash, Python, PowerShell), and infrastructure-as-code (Terraform/Ansible).
Jenkins, AWS, EC2, S3, RDS, Lambda, VPC, IAM, Windows IIS, Linux, Azure DevOps, Bash, Python, PowerShell, Terraform, Ansible, Web Deploy (MSDeploy), PowerShell DSC, Docker, Kubernetes, Datadog, Grafana, CloudWatch
1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Austin or Westford or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE6+ years systems programming experience; BS/MS/PhD in CS or EE (or equivalent); experience building RCA pipelines for HPC/cloud; deep CPU/GPU architecture knowledge; strong C++ and Python; familiarity with Slurm/LSF/Kubernetes.
C++, Python, Slurm, LSF, Kubernetes, CUDA, DCGM, NVML, CRIU, /dev/mcelog, dmesg, journald, Linux kernel
1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald