74 systems reliability engineer jobs at 45 companies in Washington, DC

1mo
Save
Mark Applied
Hide
Senior Reliability Engineer
Laurel, Maryland, United States
$100k-$245k/yr OnsiteFull Time
Johns Hopkins University Applied Physics Laboratory
Johns Hopkins University Applied Physics Laboratory: Conducts research and engineering for national security and space.
5+ YOEBachelor's in engineering/math/physics, 5+ years reliability (RAM) engineering for complex weapon systems, experience with FMEA/FMECA/FTA/PRA, strong communication, ability to travel periodically, active Secret clearance and ability to obtain Top Secret.
Reliasoft, Windchill, Saphire, MADe, JMP, R, SAS, Matlab, Python, CAMEO, DOORS
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Jersey City or McLean or Richmond
HybridFull Time
Exiger
Exiger: AI-powered supply chain risk and compliance management software.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
AWS, Codex, Claude, Chaos Monkey, Gremlin, LitmusChaos, Snowflake, Redshift, Apache Iceberg, Go, C, Java, Linux/Unix
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Lorton, Virginia, United States
$87k-$198k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Consulting and technology services for government and commercial clients
5+ YOE5+ years building and maintaining reliable, scalable systems; experience with physical servers, storage, networking, system upgrades, VMware, and data center design; active Secret clearance and Bachelor's degree required.
VMware
3d
Save
Mark Applied
Hide
Reliability Engineer, Mechanical, NA (Design)
Shackelford County or Texas or Atlanta or Abilene or Dallas or Phoenix or Ashburn or Wisconsin
HybridFull Time
Vantage Data Centers
Vantage Data Centers: Provides hyperscale data center campuses for cloud and AI providers.
2+ YOEMechanical reliability engineer for data center cooling systems; 2–3 years critical facility experience preferred; bachelor’s degree preferred; experience with commissioning, maintenance program design, RCA, and technical support.
2mo
Save
Mark Applied
Hide
Infrastructure Reliability Engineer
Manassas or Sterling or Portland or Chicago or Dallas Fort Worth
OnsiteFull Time
STACK Infrastructure
STACK Infrastructure: Developer and operator of sustainable wholesale data center infrastructure.
5+ YOE5–8 years in critical infrastructure; strong fluency in electrical systems; RCA/forensic troubleshooting; bachelor’s in engineering or equivalent.
Power distribution equipment, Waveform analysis, Fault analysis tools
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Lorton or California
$87k-$198k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Provides technology and management consulting services to diverse organizations.
5+ YOE5+ years building and maintaining reliable, scalable on-prem systems including servers, storage, and network infrastructure; 3+ years with VMware and storage/SAN; Secret clearance and Bachelor’s degree required.
VMware, SAN
1mo
Save
Mark Applied
Hide
Senior Reliability Engineer
Washington, District of Columbia, United States
OnsiteFull Time
Barbaricum
Barbaricum: Provides technology and mission support to federal national security agencies.
10+ YOE10+ years SRE/systems administration experience, Bachelor\u0002s in CS/IT/related (Master's preferred), DoD Secret clearance, expertise in monitoring, automation, cloud (AWS, Microsoft Azure, Google Cloud), scripting (Python, Shell, PowerShell), and configuration management tools.
Ansible, Puppet, Chef, Python, Shell, Microsoft PowerShell, AWS, Microsoft Azure, Google Cloud, Windows, Linux
1d
Save
Mark Applied
Hide
Sr. Project Engineer- Electrical Systems Reliability
Towson, Maryland, United States
$90k-$145k/yr HybridFull Time
Stanley Black & Decker
Stanley Black & DeckerNYSE: SWK: Manufacturer of power tools, hand tools, and outdoor equipment.
5+ YOE5+ years reliability/quality experience in NPD for electrical systems; bachelor in electrical or mechanical engineering required; leadership, reliability analysis, and compliance experience required.
Weibull, FRACAS
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Reston, Virginia, United States
$136k-$184k/yr HybridFull Time
Verisign
VerisignNASDAQ: VRSN: Provides internet domain name registry and DNS infrastructure services.
8+ YOE8+ years deploying and operating mission-critical systems; strong Linux, automation, and Kubernetes expertise; excellent communication.
Linux, Kubernetes, Ansible, Docker, OpenStack, Jenkins, ServiceNow, Python, Networking
1mo
Save
Mark Applied
Hide
Systems Administrator / Site Reliability Engineer
Reston or Washington
OnsiteFull Time
Assured Consulting Solutions
Assured Consulting Solutions: An equal opportunity employer delivering technology solutions for government and industry.
8+ YOE8+ years cloud infrastructure management (AWS, Kubernetes); strong security framework knowledge; Red Hat/Linux admin; 8140 (Security+) and AWS certs; BS or higher or equivalent experience.
AWS, OpenShift, Kubernetes, Linux, Red Hat, PKI, Secrets Manager, DynamoDB, S3, RDS, ABAC
3w
Save
Mark Applied
Hide
Site Reliability Engineer, Intelligence Systems
Reston, Virginia, United States
$146k-$194k/yr OnsiteFull Time
Anduril Industries
Anduril Industries: Defense technology building autonomous military hardware and software.
Active U.S. Top Secret/SCI clearance; mid-level systems administration/DevOps/SRE experience with Linux, IaC (Terraform/Nix), CI/CD, scripting, GCP, containers/Kubernetes, and test automation; operational mindset.
Lattice OS, Terraform, Nix, GCP, Kubernetes, GCR, GitOps, Linux, CI/CD
3w
Save
Mark Applied
Hide
Performance & Reliability Engineer
Washington, District of Columbia, United States
$71k-$137k/yr HybridFull Time
Accenture Federal Services
Accenture Federal ServicesNYSE: ACN: Provides technology and consulting services to U.S. federal agencies.
3+ YOE3+ years SRE experience, 3+ years automation/IaC and scripting (Python), 2+ years SRE observability, 1+ year Generative AI; experience defining SLOs/SLIs, troubleshooting distributed systems; AWS experience preferred; Public Trust clearance.
Python, Amazon Web Services (AWS)
2mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Washington, District of Columbia, United States
OnsiteFull Time
Tiger Analytics
Tiger Analytics: Provides AI and advanced analytics consulting for global enterprises.
Kubernetes, Docker; MLOps tools; Python, Bash; Go a plus; data systems; networking; IaC; CI/CD; monitoring; on-call readiness.
Kubernetes, Docker, Kubeflow, Vertex AI, MLflow, DVC, Python, Bash, Go, BigQuery, Pub/Sub, Pinecone, Milvus, Istio, Anthos, Terraform, Pulumi, GitHub Actions, ArgoCD, Prometheus, Grafana, Google Cloud Operations Suite, SLA/SLO, K8s (GKE), Vertex AI Endpoints, Kubeflow Pipelines
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Seattle or Austin or Reston
$81k-$187k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
5+ YOETS/SCI with Polygraph and U.S. citizenship required. Bachelor's/Master's in CS or related, 5+ years SRE/Systems experience, expertise with Oracle Linux, Ansible, Terraform, Python, Bash, observability and distributed storage.
Oracle Linux, Ansible, Terraform, Python, Bash, Prometheus, Grafana, GlusterFS, VMware, Kubernetes, Docker, Jenkins, PostgreSQL, Active Directory, LDAP, Kerberos, NFS, SMB, iSCSI, NVMe-oF
4w
Save
Mark Applied
Hide
Customer Reliability Engineer - Infrastructure
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Apache Airflow, AWS, Azure, CI/CD, GCP, Infrastructure as Code (IaC), Kubernetes, Linux, Python
1mo
Save
Mark Applied
Hide
Site Reliability Engineer II
Falls Church or South Carolina or Raleigh or Nashville or Louisiana or Pennsylvania or Plain City or South Bend or Orlando or Detroit
OnsiteFull Time
Kastle Systems
Kastle Systems: Managed security services provider for commercial and residential properties.
4+ YOE4+ years SRE/Platform experience owning production systems. Hands-on with Azure/AKS, Kubernetes, Terraform/OpenTofu/Pulumi, GitOps/ArgoCD, observability (Prometheus/Grafana/OpenTelemetry/ELK), Python/Go/Bash, and feature-flag/CI/CD practices.
ArgoCD, Flux, Crossplane, LaunchDarkly, Flagsmith, Terraform, OpenTofu, Pulumi, Prometheus, Grafana, OpenTelemetry, ELK, OpenSearch, Python, Go, Bash, C#, SQL, AKS, Azure Container Registry, Azure Monitor, Cosmos DB, Key Vault, Azure Front Door, GitOps
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Columbia, Maryland, United States
HybridFull Time
Cogent People
Cogent People: A government consulting and technology services firm delivering secure, scalable digital solutions for mission-critical federal and commercial programs.
Bachelor's degree or equivalent, experience in system reliability/DevOps/production support, observability and monitoring tools, incident management, cloud and automation, strong troubleshooting and communication skills.
AWS, Terraform, Splunk, Datadog, Prometheus, CI/CD
3w
Save
Mark Applied
Hide
Site Reliability / Operations Engineer (TS/SCI)
Herndon or Colorado Springs or Melbourne
$82k-$132k/yr OnsiteFull Time
Maxar Intelligence
Maxar Intelligence: Providing satellite imagery and geospatial intelligence for global security.
2+ YOEActive TS/SCI clearance, Bachelor's in CS/IS/Engineering or equivalent, Security+ certification, 2+ years systems/automation/DevOps experience, Linux and Bash proficiency, ELK Stack experience, strong troubleshooting and communication skills.
RHEL, CentOS, Ubuntu, Bash, Elasticsearch, Logstash, Kibana, Python
3w
Save
Mark Applied
Hide
Site Reliability / Operations Engineer (TS/SCI)
Herndon or Colorado Springs or Melbourne
$82k-$132k/yr OnsiteFull Time
Vantor
Vantor: Providing AI-powered spatial intelligence and high-resolution Earth observation.
2+ YOEActive TS/SCI clearance, Bachelor's in CS/IS/Engineering or equivalent, Security+ certification, 2+ years systems/DevOps/automation experience, Linux and Bash proficiency, ELK Stack experience, strong troubleshooting and communication skills.
ELK Stack, Elasticsearch, Logstash, Kibana, Bash, Python, RHEL, CentOS, Ubuntu
2mo
Save
Mark Applied
Hide
Site Reliability Engineer - Chantilly, VA
Chantilly, Virginia, United States
$110k-$130k/yr OnsiteFull Time
ICR
ICR: Provides IT services to government clients.
3+ YOETop Secret/SCI eligible clearance; 3-5 years distributed systems experience; 2-4 years design/architecture; programming experience; GitLab, containerization, Kubernetes.
GitLab, Containerization, Kubernetes, Linux

Explore Jobs

Expand Your Job Search