3,601 site reliability jobs at 1,569 companies in United States

PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States
RemoteFull Time
Chess.com
Chess.com: Operates a global platform for playing and learning chess.
5+ YOE5+ years in site reliability, DevOps, or infrastructure; UNIX/Linux; cloud with IaC; monitoring; distributed systems.
UNIX/Linux, Cloud platforms (GCP, AWS, or Azure), Terraform, CloudFormation, Ansible, Docker, Kubernetes, Datadog, Prometheus, Grafana, ELK stack, DNS, TCP/IP, Puppet, Chef
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore
VantageScore: Provides credit scoring and data analytics solutions.
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
3d
Save
Mark Applied
Hide
Site Reliability Engineer
Charlotte, North Carolina, United States
HybridFull Time
Electrolux Group
Electrolux GroupNasdaq Stockholm: ELUX-B: Global manufacturer of household appliances and consumer kitchen equipment.
6+ YOE6+ years in infrastructure/site reliability/cloud engineering; experience with cloud platforms, IaC, CI/CD, observability, troubleshooting, and strong collaboration skills.
Microsoft Azure, AWS, Google Cloud Platform, Akamai CDN, Terraform, CloudFormation, Ansible, Puppet, Chef, Microsoft Azure DevOps, GitHub, Argo CD
3mo
Save
Mark Applied
Hide
Site Reliability Engineer, Manager
United States
$135k-$216k/yr RemoteFull Time
Peraton
Peraton: Provides advanced technology and mission support for government agencies.
10+ YOEExtensive experience in site reliability engineering in multi-vendor environments; cloud-native infra, Kubernetes, automation tools, observability, SLI/SLOs, leadership.
Kubernetes, Terraform, Ansible, Chef, Prometheus, Grafana, OpenTelemetry, Python, Go
2mo
Save
Mark Applied
Hide
Founding Engineer - Site Reliability
San Francisco or United States
$185k-$285k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Kubernetes, AWS, Prometheus, Grafana, Datadog, Automation, VPC, GPU compute
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States
$160k-$180k/yr HybridFull Time
Fortress Information Security
Fortress Information Security: AI-powered platform for supply chain cybersecurity and risk management.
4+ YOE4-5 years in SRE/DevOps; Linux, AWS; Kubernetes, Ansible, Jenkins, Terraform; scripting in Bash/Python/JS; on-site Patuxent River, MD; active clearance preferred/required.
Kubernetes, Ansible, Jenkins, Terraform, AWS, Bash, Python, JavaScript, CloudWatch, CloudFormation, Lambda, API Gateway, HashiCorp Vault
3mo
Save
Mark Applied
Hide
Site Reliability Engineer, Manager
United States
$135k-$216k/yr RemoteFull Time
Peraton
Peraton: National security and mission-critical government technology services provider.
10+ YOE10+ years in site reliability or related roles; cloud-native infra, Kubernetes, automation tools; observability tech; programming/scripting; SLIs/SLOs; strong collaboration and leadership.
Kubernetes, Terraform, Ansible, Chef, OpenTelemetry, Prometheus, Grafana, Python, Go
3w
Save
Mark Applied
Hide
Reliability Site Leader
Houston, Texas, United States
OnsiteFull Time
G-3 Chickadee
G-3 Chickadee: Manufacturer of synthetic rubber and engineered polymer solutions.
5+ YOE5+ MgmtBachelor's in engineering required; 5+ years reliability engineering management; experience with mechanical integrity, RCM, TAR planning; ability to implement reliability data systems; CMRP or CRL preferred.
SAP, MES
1mo
Save
Mark Applied
Hide
Databricks Site Reliability Engineer
Golden, Colorado, United States
$103k-$136k/yr OnsiteFull Time
CoorsTek
CoorsTek: Manufacturer of engineered technical ceramics for global industrial markets.
5+ YOE5+ years in site reliability engineering or similar; experience with Databricks, Delta Lake, Unity Catalog, SQL, Python/PySpark; Azure and CI/CD; incident management and production readiness; governance and support pattern development.
Databricks, Delta Lake, Unity Catalog, SQL, Python, PySpark, Notebooks, Jobs, Clusters, Azure, CI/CD, Git
3mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE) – II
Easton, Ohio, United States
HybridFull Time
Huntington
HuntingtonNASDAQ: HBAN: Provides regional commercial, consumer, and mortgage banking services.
3+ YOE3+ years in site reliability/DevOps; incident response; automation; monitoring; collaboration.
Terraform, Ansible, CloudFormation, Prometheus, Dynatrace, Splunk, OpenShift, Windows Server, AWS, GCP, PowerShell, Bash, Python, SQL
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
Braze
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Ruby on Rails, Ruby, Go, Linux, Unix Shell, Docker, Kubernetes, Terraform, Chef, MongoDB, Redis, Kafka, Postgres, PagerDuty
2d
Save
Mark Applied
Hide
Site Reliability Manager
Herndon, Virginia, United States
OnsiteFull Time
Karsun Solutions
Karsun Solutions: Delivers enterprise IT modernization solutions to federal government agencies.
5+ YOE10+ MgmtLead SRE team, ensure application reliability and observability with Datadog, implement IaC and DevSecOps practices, manage platform lifecycle; AWS and containerization experience required.
Datadog, AWS, AWS Cloudwatch, Docker, Kubernetes, Terraform, Ansible, ArgoCD, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Columbia, Maryland, United States
HybridFull Time
Cogent People
Cogent People: A government consulting and technology services firm delivering secure, scalable digital solutions for mission-critical federal and commercial programs.
Bachelor's degree or equivalent, experience in system reliability/DevOps/production support, observability and monitoring tools, incident management, cloud and automation, strong troubleshooting and communication skills.
AWS, Terraform, Splunk, Datadog, Prometheus, CI/CD
1w
Save
Mark Applied
Hide
Site Reliability Engineer
San Antonio, Texas, United States
OnsiteFull Time
Infosys
InfosysNYSE: INFY: Provides IT consulting, software development, and business outsourcing services.
Experience with reliability engineering, Terraform IaC, observability (Datadog), disaster recovery, vulnerability management, CI/CD (Harness, Helm), plus bachelor's degree or equivalent experience.
Terraform, Datadog, Harness, Helm Charts, AWS
3mo
Save
Mark Applied
Hide
Site Reliability Engineer
New York, New York, United States
$177k-$209k/yr HybridFull Time
Peloton
PelotonNASDAQ: PTON: Sells connected fitness equipment and streaming exercise classes.
Kubernetes expertise; observability/monitoring mindset; CI/CD experience; IaC (Terraform/Pulumi); cloud (AWS); security and reliability focus; programming (Python/Go/Java/C).
Kubernetes, Observability, Jenkins, ArgoCD, Harness, Tekton, Terraform, Pulumi, AWS, Python, Golang, Java, C, Nginx, Ubuntu, Chef
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States
$175k-$200k/yr RemoteFull Time
Tern
Tern: Software platform for travel advisors to manage business operations.
Proven production reliability ownership, end-to-end cloud migration experience (GCP preferred), strong observability and monitoring skills, incident leadership, infrastructure-as-code familiarity, and experience coaching engineers.
Ruby on Rails, Hotwire, Postgres, Heroku, Google Cloud Platform, Fivetran, BigQuery, Hex, AppSignal, Bugsnag, Canny, Claude Code
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Lorton or California
$87k-$198k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Provides technology and management consulting services to diverse organizations.
5+ YOE5+ years building and maintaining reliable, scalable on-prem systems including servers, storage, and network infrastructure; 3+ years with VMware and storage/SAN; Secret clearance and Bachelor’s degree required.
VMware, SAN
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Lorton, Virginia, United States
$87k-$198k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Consulting and technology services for government and commercial clients
5+ YOE5+ years building and maintaining reliable, scalable systems; experience with physical servers, storage, networking, system upgrades, VMware, and data center design; active Secret clearance and Bachelor's degree required.
VMware
5d
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yr RemoteFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
London or New York City or Singapore or Sydney or Lisbon
OnsiteFull Time
Thought Machine
Thought Machine: Provides cloud-native core banking and payments software for banks.
Experience delivering reliability and scalability work; production-level Python, Golang or Java; Kubernetes; Terraform/Puppet/Chef/Ansible; observability (Prometheus, Jaeger); GCP or AWS; design patterns for hosting and networking; on-call and documentation skills.
Python, Golang, Java, Kubernetes, Terraform, Puppet, Chef, Ansible, Prometheus, Jaeger, GCP, AWS, Vault