33 platform reliability engineer jobs at 17 companies in Rosedale, MD
2w
Save
Mark Applied
Hide
2w
Reliability Engineer
Arlington, Virginia, United States
$80k-$160k/yrOnsiteFull Time
Decision Technologies: Provides engineering and technical support for defense programs.
4+ YOEBachelor's degree in a technical discipline, 4+ years RM&A experience on US Navy platforms or weapon systems, eligible for DoD Secret clearance, strong communication and presentation skills.
Exiger: AI-powered supply chain risk and compliance management software.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
Falls Church or South Carolina or Raleigh or Nashville or Louisiana or Pennsylvania or Plain City or South Bend or Orlando or Detroit
OnsiteFull Time
Kastle Systems: Managed security services provider for commercial and residential properties.
4+ YOE4+ years SRE/Platform experience owning production systems. Hands-on with Azure/AKS, Kubernetes, Terraform/OpenTofu/Pulumi, GitOps/ArgoCD, observability (Prometheus/Grafana/OpenTelemetry/ELK), Python/Go/Bash, and feature-flag/CI/CD practices.
United States or North America or San Francisco or Los Angeles or New York City or Washington or London or Singapore
RemoteFull Time
TRM Labs: An AI-powered intelligence technology helping agencies investigate crime and disrupt illicit activity.
U.S. citizenship, distributed OLAP or serving-layer operations experience, query tuning, data pipeline reliability, incident response, AI tool fluency, independent infrastructure ownership, and on-call readiness.
Guidehouse: Provides management and technology consulting services to diverse organizations.
4+ YOEBA/BS or equivalent experience, 4+ years IT/admin/software/platform experience with AWS, 1+ years cloud deployment experience, proficiency with CI/CD and IaC tools (Terraform, Ansible, GitLab, Artifactory, Packer), scripting (Python, PowerShell, Bash), Windows/Linux, Agile, and ability to obtain Public Trust.
Boston or Miami or New Jersey or New York City or Princeton or Raleigh or Washington or Toronto or North America
$127k-$249k/yrHybridFull Time
MongoDBNASDAQ: MDB: Cloud-based document database platform for software application development.
6+ YOE6+ years software development experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud (AWS, Google Cloud Platform, Microsoft Azure) expertise; Linux and networking knowledge.
Argo Workflows, ArgoCD, Kubernetes, Python, Go, AWS, Google Cloud Platform (GCP), Microsoft Azure, Linux
Peregrine: Data integration and analytics platform for public safety agencies.
8+ YOE8+ years building and operating cloud infrastructure, hands-on ownership of platform/networking/traffic systems, strong security and reliability experience, on-call and incident response experience, degree or equivalent.
Site Reliability Engineer (SRE) / Service Availability Manager
Bethesda, Maryland, United States
$96k-$145k/yrHybridFull Time
Marriott InternationalNASDAQ: MAR: Operates and franchises a global network of hotels and resorts.
5+ YOE5+ years IT experience, 3+ years IT operations and incident/change/release management, undergraduate degree or equivalent, on-call/24x7 availability, proficiency with Python and Shell, familiarity with Ansible, Jenkins, cloud platforms, IaC and containers.
Anduril Industries: Defense technology building autonomous military hardware and software.
Strong backend engineering experience building production platforms; deep expertise in LLM agent framework design and evaluation; familiarity with model post-training workflows (SFT, RL); emphasis on reliability and partnering with ML/product teams.
Lattice OS, Langchain, Deepagents, Claude SDK, Kubernetes, Docker
WEXNYSE: WEX: Provides global payment processing and business information management services.
Staff/lead level engineer with deep experience architecting autonomous/agentic AI systems, cross-platform architecture, security-by-design, CI/CD, cloud reliability, and mentoring senior engineers.
Inovalon: Cloud platform for data-driven healthcare analytics and insights.
5+ YOE5+ years cloud/systems engineering experience; expertise in AWS/Azure/GCP/OCI/Snowflake, infrastructure as code, scripting, and platform governance; bachelor's degree required; strong communication and reliability focus.
AWS, Azure, GCP, OCI, Snowflake, Microsoft Azure DevOps, Microsoft Azure Pipelines, Terraform, Linux Shell, Microsoft PowerShell, Python, Microsoft Active Directory, GPO, RDS, Microsoft Windows, Linux
Senior Lead Software Engineer-AI Foundation Services
Plano or Jersey City or Wilmington or McLean
$171k-$260k/yrOnsiteFull Time
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years software engineering experience building cloud-native AI/ML platform services with Kubernetes, CI/CD, and infrastructure-as-code; proficiency in Python/Java/Go; strong production reliability and secure-by-design practices.
HiltonNYSE: HLT: Global hospitality providing hotel accommodation and lodging services.
7+ YOESeven years in technology; five years with CI tools (GitLab preferred); four years automating delivery; four years administering Kubernetes; four years on AWS; scripting in Bash; on-call reliability; API platforms in Kubernetes; travel up to 10%; hybrid role near US offices.
Leonardo DRSNASDAQ: DRS: Designs and integrates advanced defense technologies for military customers.
5+ YOERequires 5+ years in infrastructure, platform, or site reliability engineering, including 3+ years in MLOps or model serving, hands-on self-hosted AI, cloud development, security engineering, and a bachelor's degree or equivalent.
Beavercreek or Frederick or Fort Walton Beach or Dayton
OnsiteFull Time
Leonardo DRSNASDAQ: DRS: Manufactures advanced electronic systems for defense and military applications.
5+ YOERequires 5+ years in infrastructure, platform, or site reliability engineering, including 3+ years in MLOps or model serving; hands-on self-hosted AI, security engineering, cloud development, bachelor's degree or equivalent, U.S. citizenship, and clearance eligibility.