6 site reliability engineer jobs at 3 companies in Danville, VA

2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - HPC
Santa Clara or Austin or Durham
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years building/supporting critical services; experience with large-scale HPC clusters (Slurm, LSF, Kubernetes); IaC and CI/CD proficiency; coding in Python/Go/Perl/Ruby; monitoring, capacity planning, and incident response skills.
Slurm, LSF, Kubernetes, Infrastructure as Code (IaC), CI/CD, AWS, GCP, OCI, Python, Go, Perl, Ruby
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - HPC
Santa Clara or Durham or Austin
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEBS in CS or equivalent with 5+ years supporting critical services; experience with large-scale HPC clusters (Slurm, LSF, Kubernetes), IaC, CI/CD, multi‑cloud (AWS/GCP/OCI), and 2+ languages such as Python or Go.
Slurm, LSF, Kubernetes, AWS, GCP, OCI, Infrastructure as Code (IaC), CI/CD, Python, Go, Perl, Ruby, AIOps
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Durham, North Carolina, United States
OnsiteFull Time
Fidelity Investments
Fidelity Investments: Provides investment management, retirement planning, and brokerage services.
5+ YOEBachelor's in a tech field, 5+ years deploying/supporting distributed systems, 3+ years AWS, 3-5 years software development (Python/NodeJS/Java), observability, performance and chaos testing, Kubernetes, IaC (Terraform/Chef), CI/CD.
AWS, Python, NodeJS, Java, Datadog, Splunk, Kibana, Prometheus, Grafana, ELK/OpenSearch, Open Telemetry, K6, JMeter, Kubernetes, Chaos Monkey, Gremlin, Go, Bash, IAM, ARM, Terraform, Chef, S3, EC2, RDS, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Durham, North Carolina, United States
OnsiteFull Time
Fidelity Investments
Fidelity Investments: Provides investment management, retirement planning, and brokerage services.
5+ YOEBachelor’s degree in a technology-related field and 5+ years supporting distributed systems; 3+ years AWS; software development, observability, performance testing, Kubernetes, CI/CD, DevOps, and infrastructure as code experience.
AWS, Python, NodeJS, Java, Datadog, Splunk, Kibana, Prometheus, Grafana, ELK, OpenSearch, Open Telemetry, K6, JMeter, Kubernetes, Chaos Monkey, Gremlin, Go, Bash, IAM, ARM, Terraform, Chef, S3, EC2, RDS
2d
Save
Mark Applied
Hide
Site Reliability Engineer Spring Co-op 2027
Lowell or Durham or San Jose or Austin
$76k-$166k/yr HybridMultiple Commitments Available
IBM
IBMNew York Stock Exchange: IBM: Global technology providing enterprise software, cloud, and consulting.
Actively enrolled in a bachelor's program, available for a 16-week full-time co-op, and knowledgeable in Linux, monitoring, troubleshooting, automation, scripting, cloud platforms, and production support.
Linux, Python, Go, Bash, IBM Cloud, AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, OpenShift, Ansible, Terraform, Jenkins, IBM Continuous Delivery, ArgoCD, Instana, New Relic, Grafana, Prometheus, PostgreSQL, CouchDB, Redis, Kafka, Spark, SQL, NoSQL, CI/CD
3d
Save
Mark Applied
Hide
Site Reliability Engineer Intern 2027
Lowell or Durham or San Jose or Austin
$76k-$166k/yr HybridInternship, Temporary
IBM
IBMNew York Stock Exchange: IBM: Global technology providing enterprise software, cloud, and consulting.
Requires high school diploma or GED and knowledge of monitoring, troubleshooting, automation, Linux, production support, scripting, and cloud providers; bachelor's degree, Kubernetes, and CI/CD experience preferred.
Linux, Python, Go, Bash, IBM Cloud, AWS, Azure, GCP, Kubernetes, OpenShift, Ansible, Terraform, Jenkins, IBM Continuous Delivery, ArgoCD, Instana, New Relic, Grafana, Prometheus, SQL, NoSQL, PostgreSQL, CouchDB, Redis, Kafka, Spark