4 infrastructure reliability engineer jobs at 4 companies in Cary, NC

1mo
Save
Mark Applied
Hide
Customer Reliability Engineer - Infrastructure
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Apache Airflow, AWS, Azure, CI/CD, GCP, Infrastructure as Code (IaC), Kubernetes, Linux, Python
2w
Save
Mark Applied
Hide
Site Reliability Engineering (SRE) & DevOps
Raleigh, North Carolina, United States
$91k-$135k/yr OnsiteFull Time
NTT DATA
NTT DATA: Global provider of IT and business consulting services.
5+ YOE5+ years experience; expertise in SRE/DevOps, cloud engineering on GCP, infrastructure automation, observability (SLI/SLO/Error Budgets), Terraform, Python, PowerShell, and Git-based workflows.
Google Cloud Platform (GCP), Terraform, GitHub, Git, Python, PowerShell, Kubernetes, Docker, Jenkins, Maven, Ant, Puppet, Gradle, Ansible, TeamCity, uDeploy, Dataproc, Java, C#
2mo
Save
Mark Applied
Hide
Associate Software Engineer (Site Reliability Engineering)
Raleigh, North Carolina, United States
HybridFull Time
Relay
Relay: Provides cloud-connected smart communication devices for frontline workers.
Associate Software Engineer role in Site Reliability Engineering; requires BS in CS/Math/Physics; hands-on with Linux, programming, and infrastructure tools.
AWS, Terraform, Ansible, Grafana, PagerDuty, Jenkins, Packer, Python, Elixir
1mo
Save
Mark Applied
Hide
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).
C++, Python, Slurm, LSF, Kubernetes, NVIDIA DCGM, NVIDIA Management Library (NVML), CRIU, CUDA, /dev/mcelog, dmesg, journald