General Dynamics Mission Systems
Posted 1w ago

Senior Software / Site Reliability Lead Engineer

General Dynamics Mission Systems
United States
$143k-$158k/yrRemoteFull Time
Responsibilities
  • defining standards
  • building observability
  • leading incidents
Requirements
  • Bachelor's degree plus 8 years or master's degree plus 6 years
  • Production SRE/DevOps
  • Observability
  • Python or Bash
  • Infrastructure-as-code
  • Docker
  • Kubernetes, SLOs
  • Incident response
  • U.S. citizenship, and Secret clearance eligibility
Technical tools mentioned
PrometheusGrafanaDatadogELKCloudWatchPythonBashTerraformCloudFormationDockerKubernetes

Job description

Bachelor's degree in Software Engineering, or related Science, Technology, Engineering or Mathematics field, plus a minimum of 8 years of relevant experience; or Master's degree, plus 6 years relevant experience.

CLEARANCE REQUIREMENTS: Ability to obtain a Department of Defense Secret security clearance is required at time of hire. Applicants selected will be subject to a U.S. Government security investigation and must meet eligibility requirements for access to classified information. Due to the nature of work performed within our facilities, U.S. citizenship is required.

What You Will Own

  • Cross-pod reliability standards. Set the reliability bar and ensure it is met consistently across applications. Collaborate with Functional SREs to connect technical reliability metrics to business-side outcomes. You own the engineering signal; together you tell the full reliability story.
  • SLOs and reliability metrics. Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.
  • Monitoring and observability. Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards. You will know when something is degrading before users do.
  • Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong. Alerts you build catch real problems; they don't cry wolf.
  • Incident response. Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog. When something breaks, you coordinate the response and ensure it doesn't break the same way again.
  • Production Readiness. Define and enforce the criteria that determine whether an AI service is ready for production. You are the gate between "it works in dev" and "it's ready to ship."
  • Toil elimination. Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.

What You Won't Own

  • Infrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works
  • Business process decisions or backlog prioritization
  • Business-side reliability metrics - you partner with the Functional SRE on those, but they own that domain

What Makes This Role Different

  • AI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.
  • You are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform.
  • Your production readiness criteria directly determine whether AI services go live. You have real authority to say "not ready."
  • You operate across projects simultaneously — embedded deeply enough to understand large-scale systems, while maintaining consistent standards across all projects.
  • Your software engineering background means you can engage directly with development teams at the design level — catching reliability problems before they become operational ones.

Required Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience
  • Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines
  • Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.
  • Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)
  • Experience with containerized environments — Docker, Kubernetes, container orchestration at scale
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production
  • U.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.

Preferred Qualifications

  • Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines
  • Software engineering fundamentals — you can read, write, and meaningfully review production-quality code. You understand how architectural and design decisions made early translate into operational problems later.
  • Software design experience — you have participated in or led design reviews, defined service interfaces or APIs, and pushed back on design decisions using reliability and operability as criteria
  • Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that have caught real problems.
  • Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)
  • Experience with containerized environments — Docker, Kubernetes, container orchestration at scale
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production

What Sets You Apart

  • You build things that work. Your default response to a problem is code, not a document.
  • You have shipped AI systems that real users depended on in production.
  • You are comfortable working without detailed specs — you can take a problem statement and figure out the right approach.
  • You care about reliability as much as capability — you monitor what you deploy.
  • You move fast without being reckless. You know when to iterate and when to get it right the first time.

Details

  • Remote — 100% telework
  • 9/80 schedule
  • Defense industry experience is not required

This estimate represents the typical salary range for this position based on experience and other factors (geographic location, etc.). Actual pay may vary. This job posting will remain open until the position is filled.


USD $142,696.00 - USD $158,303.00 /Yr.

General Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!


Equal Opportunity Employer / Individuals with Disabilities / Protected Veterans

About General Dynamics Mission Systems

Designs and manufactures defense and mission-critical technology solutions.

Year founded
2015
Employees
12000
Organization type
Public
Headquarters
US

Similar jobs

Site Reliability Engineer roles
8h
Save
Mark Applied
Hide
Site Reliability Engineer Intern (Data Infra) - 2027 Fall
Seattle, Washington, United States
OnsiteInternship
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Currently pursuing a bachelor's degree in computer science or related field; programming experience in C, C++, Java, Python, Go, or Rust; knowledge of Unix/Linux internals, networking, and distributed systems.
C, C++, Java, Python, Go, Rust, Unix, Linux, Kubernetes, Redis, MySQL, Flink, Nginx, Docker, OpenStack, Hadoop, Spark
12h
Save
Mark Applied
Hide
Site Reliability Engineer
New York City, New York, United States
HybridFull Time
Chariot
Chariot: Payment infrastructure for charitable giving and Donor Advised Funds.
4+ YOERequires 4+ years software development, 2+ years backend application development, bachelor's degree preferred, and proficiency with Go, Node, Docker, Terraform, Kubernetes, Postgres, REST APIs, gRPC, and AWS.
Go, Node, Docker, Terraform, Kubernetes, Postgres, REST APIs, gRPC, AWS
13h
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – eCommerce & Google Cloud Platform (GCP)
United States
$50k-$70k/yr OnsiteFull Time
Cognizant
CognizantNasdaq: CTSH: Provides global information technology and business process outsourcing services.
8+ YOERequires 8+ years in SRE, cloud engineering, or DevOps; strong GCP, Kubernetes, Docker, scripting, Terraform, CI/CD, Linux, networking, distributed systems, and observability experience.
Google Cloud Platform (GCP), Google Kubernetes Engine (GKE), Compute Engine, Cloud Storage, Cloud Monitoring, Cloud Operations Suite, Kubernetes, Docker, Python, Bash, Terraform, Git, Jenkins, GitHub Actions, Linux, Prometheus, Grafana, Datadog, Splunk
13h
Save
Mark Applied
Hide
Staff Site Reliability Engineer
New York City, New York, United States
$241k-$270k/yr RemoteFull Time
Garner Health
Garner Health: Identifies high-quality doctors to lower employer healthcare costs.
7+ YOE7+ years operating production cloud infrastructure at scale; deep Kubernetes and Terraform expertise; Python or Go skills; reliability practice design, mentoring, cost optimization, and regulated-environment experience preferred.
Amazon Web Services (AWS), Kubernetes, Terraform, Istio, Python, Go, TypeScript, Postgres, NATS, Datadog, GitLab, Claude
14h
Save
Mark Applied
Hide
Site Reliability Engineer Intern (Global SRE) - 2027 Summer
San Jose or Los Angeles or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteInternship, Full Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Currently pursuing a bachelor's degree in computer science or related field; Unix/Linux, IP networking, and Python, Go, C, C++, or Java programming experience required.
Unix/Linux, IP networking, Python, Go, C, C++, Java
14h
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – eCommerce & Google Cloud Platform (GCP)
United States
$50k-$70k/yr OnsiteFull Time
Cognizant
CognizantNASDAQ: CTSH: Provides IT consulting and technology services to global enterprises.
8+ YOERequires 8+ years in SRE, cloud engineering, or DevOps; strong GCP, Kubernetes, Docker, Python/Bash, Terraform, CI/CD, Linux, networking, distributed systems, and observability experience.
Google Cloud Platform (GCP), Google Kubernetes Engine (GKE), Compute Engine, Cloud Storage, Cloud Monitoring, Cloud Operations Suite, Terraform, Kubernetes, Docker, Python, Bash, Git, Jenkins, GitHub Actions, Linux, Prometheus, Grafana, Datadog, Splunk
15h
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Ann Arbor, Michigan, United States
HybridFull Time
Sight Machine
Sight Machine: Developer of an AI-powered manufacturing data analytics platform.
10+ YOE10+ years with Kubernetes/Docker, cloud infrastructure, coding, IaC, and CI/CD; production agentic AI/LLM experience; strong Linux, networking, observability, documentation, mentorship, and cross-team leadership skills.
Kubernetes, Docker, Azure, GCP, AWS, Python, Go, Java, Terraform, OpenTofu, FluxCD, Jenkins, GitHub Actions, Linux, TCP/IP, Prometheus, Grafana, Loki, Sentry, Signoz, Helm Charts, Elasticsearch, Kafka, Postgres
17h
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE), Akrites
United States
$170k-$190k/yr RemoteFull Time
The Linux Foundation
The Linux Foundation: Non-profit organization supporting open source software development.
8+ YOERequires 5+ years with Linux and cloud infrastructure, 8+ years programming in a compiled language, complex project leadership, incident response, security practices, communication, mentoring, and remote-work experience.
Linux, Go, Python, Google Cloud Platform (GCP), Terraform