51 site reliability manager jobs at 38 companies in Frederick, MD

2w
Save
Mark Applied
Hide
Site Reliability Manager
Herndon, Virginia, United States
OnsiteFull Time
Karsun Solutions
Karsun Solutions: Delivers enterprise IT modernization solutions to federal government agencies.
5+ YOE10+ MgmtLead SRE team, ensure application reliability and observability with Datadog, implement IaC and DevSecOps practices, manage platform lifecycle; AWS and containerization experience required.
Datadog, AWS, AWS Cloudwatch, Docker, Kubernetes, Terraform, Ansible, ArgoCD, CI/CD
2w
Save
Mark Applied
Hide
Site Reliability Manager
Herndon, Virginia, United States
OnsiteFull Time
Karsun Solutions
Karsun Solutions: Provider of IT modernization and software development for government agencies.
10+ YOE3+ MgmtBachelor's in CS/Engineering preferred,10+ years SRE experience with AWS,3+ years managing SRE teams,expertise with Datadog,AWS Cloudwatch,Docker,Kubernetes,Terraform/Ansible/ArgoCD,CI/CD;Public Trust eligible.
Datadog, AWS Cloudwatch, AWS, Docker, Kubernetes, Terraform, Ansible, ArgoCD, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Columbia, Maryland, United States
HybridFull Time
Cogent People
Cogent People: A government consulting and technology services firm delivering secure, scalable digital solutions for mission-critical federal and commercial programs.
Bachelor's degree or equivalent, experience in system reliability/DevOps/production support, observability and monitoring tools, incident management, cloud and automation, strong troubleshooting and communication skills.
AWS, Terraform, Splunk, Datadog, Prometheus, CI/CD
4w
Save
Mark Applied
Hide
Site Reliability Engineer
Jersey City or McLean or Richmond
HybridFull Time
Exiger
Exiger: AI-powered supply chain risk and compliance management software.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
AWS, Codex, Claude, Chaos Monkey, Gremlin, LitmusChaos, Snowflake, Redshift, Apache Iceberg, Go, C, Java, Linux/Unix
1w
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
Reston or Austin
OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOE1+ Mgmt8+ years in software engineering or infrastructure (or fewer with relevant degrees), 3–5 years automation/programming experience, data analysis skills, 1 year leadership experience preferred.
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States or Washington
$130k-$160k/yr RemoteFull Time
Berkeley Research Group
Berkeley Research Group: Provides expert testimony and specialized business consulting services.
5+ YOEBachelor's in computer science or similar, 5+ years SRE or similar, programming in Golang/Ruby/Python, Kubernetes, cloud experience (Azure/AWS/GCP), observability tools, and incident management expertise.
Microsoft Azure Cloud Services, GitHub Actions, GitLab CI, Golang, Ruby, Python, Kubernetes, AWS, GCP, Datadog, OpsGenie, PagerDuty
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Palo Alto or Palo Alto or Washington
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
Kubernetes, Linux, Bash, Python, Bazel, Makefiles, Terraform, Ansible, TCP/IP
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Tysons Corner, Virginia, United States
HybridFull Time
Cvent
Cvent: Cloud-based software for event and venue management.
Deep SRE/DevOps experience with CI/CD, scripting (Ruby, Groovy, Bash, PowerShell, TypeScript, Python), AWS, config management (Chef/Puppet/Ansible), containerization (Docker, ECS, EKS, Kubernetes), observability tools, and incident response skills.
Ruby, Groovy, Bash, PowerShell, TypeScript, Python, AWS, Chef, Puppet, Ansible, Windows, Linux/Unix, Datadog, New Relic, Splunk, Docker, ECS, EKS, Kubernetes, Jenkins, MongoDB, Couchbase, Postgres, F5, Nexus, Artifactory, Claude Code, Claude
2w
Save
Mark Applied
Hide
Site Reliability Engineer, Lead
Chantilly, Virginia, United States
$99k-$225k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Consulting and technology services for government and commercial clients
8+ YOE8+ years SRE/DevOps experience with Prometheus, Grafana, ELK, Linux, AWS, Python, Terraform/Terragrunt, Kubernetes; TS/SCI with polygraph; ability to obtain listed security certifications.
Prometheus, Grafana, ELK, Python, Terraform, Terragrunt, Kubernetes, OpenTelemetry, AWS CloudWatch, AWS EKS, Rancher, Jenkins, Git, Docker, Nessus, JIRA, Confluence
1d
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Vienna or Pensacola or Winchester
$128k-$188k/yr HybridFull Time
Navy Federal Credit Union
Navy Federal Credit Union: Offers banking and financial services to the military community.
7+ YOEMaster's degree or equivalent,7+ years SRE experience,expertise in monitoring,incident response,automation,software development,advanced programming in Python/Java/Go,and strong communication and problem-solving skills.
Python, Java, Go
2w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE) / Service Availability Manager
Bethesda, Maryland, United States
$96k-$145k/yr HybridFull Time
Marriott International
Marriott InternationalNASDAQ: MAR: Operates and franchises a global network of hotels and resorts.
5+ YOE5+ years IT experience, 3+ years IT operations and incident/change/release management, undergraduate degree or equivalent, on-call/24x7 availability, proficiency with Python and Shell, familiarity with Ansible, Jenkins, cloud platforms, IaC and containers.
Python, Shell, Ansible, Jenkins, AWS, Azure, GCP, ServiceNow
2mo
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering (Federal)
Washington, District of Columbia, United States
$207k-$285k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
3+ YOE3+ Mgmt3+ years technical leadership; Agile/DevOps; AWS/public cloud; Kubernetes; IaC (Terraform); CI/CD; observability; CS degree.
Kubernetes, Terraform, CI/CD, Grafana, Splunk, APM, AWS
1w
Save
Mark Applied
Hide
Lead DevOps (Site Reliability Engineer)
McLean or Wilmington
HybridFull Time
Anza Mortgage Insurance
Anza Mortgage Insurance: Provides mortgage insurance and credit risk protection for lenders.
Lead SRE with team leadership experience; strong AWS (EKS,Fargate,Aurora), Terraform/Terragrunt, Kubernetes, Argo CD/Workflows, GitHub Actions, Docker, Git; experience with monitoring, incident response, and disaster recovery.
Argo CD, Argo Workflows, Terraform, Terragrunt, Kubernetes, GitHub Actions, AWS, EKS, Fargate, Aurora, Docker, Datadog, Cloudflare, Camunda, GuardDuty, Security Hub, Git
1mo
Save
Mark Applied
Hide
Reliability Manager – Maintenance Planning
Rockville, Maryland, United States
$113k-$151k/yr OnsiteFull Time
Samsung Biologics
Samsung BiologicsKorea Exchange: 207940: Global contract development and manufacturing organization for biopharmaceuticals.
Lead maintenance planning and reliability initiatives, serve as IBM Maximo CMMS site owner, manage maintenance planners, ensure cGMP compliance, optimize schedules, and drive continuous improvement.
IBM Maximo CMMS
1d
Save
Mark Applied
Hide
Site Reliability Engineering (SRE) Director
Washington, District of Columbia, United States
$142k-$158k/yr HybridFull Time
AARP
AARP: Non-profit organization advocating for Americans aged 50 and older.
8+ YOE8+ years SRE/DevOps experience with leadership of enterprise-scale reliability, cloud-native architectures, CI/CD, and operational frameworks; bachelor\u0002s degree or equivalent; strong communication and decision-making skills.
CI/CD, Secure Software Development Lifecycle (SSDLC)
1d
Save
Mark Applied
Hide
Site Reliability Engineering (SRE) Director
Washington, District of Columbia, United States
$142k-$158k/yr HybridFull Time
AARP
AARP: Nonprofit advocacy organization serving Americans aged 50 and older.
8+ YOEBachelor's or equivalent experience in CS/IT,8+ years SRE/DevOps/cloud operations with enterprise leadership,5+ years cloud-native and CI/CD experience,SSDLC and reliability program experience; U.S. work authorization required.
CI/CD pipelines, SSDLC, QA
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Steaming HUB - FreeWheel
Reston, Virginia, United States
OnsiteFull Time
Comcast
ComcastNASDAQ: CMCSA: Provides global telecommunications, media content, and entertainment services.
5+ YOE5+ years relevant experience; strong AWS and Kubernetes experience; Terraform, Ansible, Python/Go scripting; infrastructure, networking, and production incident management skills.
Amazon Web Services (AWS), Oracle Cloud Infrastructure (OCI), Amazon EKS, Kubernetes, Docker, Terraform, Ansible, Jenkins, Python, Go (Golang), Route 53, IAM, VPCs, Load Balancers
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer (US Federal)
Reston, Virginia, United States
$147k-$221k/yr HybridFull Time
Workday
WorkdayNASDAQ: WDAY: Provides cloud-based software for financial and human capital management.
5+ YOE5+ years managing large-scale cloud infrastructure with automation, CI/CD, Kubernetes, Terraform, security and strong collaboration; bachelor’s or equivalent and ability to obtain U.S. security clearance.
Terraform, Argo CD, Kubernetes, Amazon Web Services, C#, Python, Ruby, Rust, Go
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
3d
Save
Mark Applied
Hide
Site Reliability Engineer , Engineering Enablement (Remote)
Boston or Atlanta or Chicago or Washington or New York City or Herndon
$138k-$198k/yr RemoteFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
3+ YOEBachelor's+5 or Master's+3 experience; 3+ years writing production Python or Ruby; 3+ years managing Terraform/Ansible IaC; Unix/Linux experience; containerization and CI/CD experience; ability to operate at scale and participate in on-call rotations.
Python, Ruby, Terraform, Ansible, Unix/Linux, Docker, Kubernetes, DORA, SPACE