1,050 site reliability engineering jobs at 552 companies in United States

2w
Save
Mark Applied
Hide
Director, Site Reliability Engineering
Denver or United States
$175k-$220k/yr HybridFull Time
Vertafore
VertaforeNYSE: ROP: Provides cloud-based software solutions for the insurance industry.
15+ YOE8+ MgmtBachelor's degree and 15+ years in software engineering/SRE with 8+ years leadership; expertise in CI/CD, observability, incident response, AWS, container orchestration, and operating SaaS products.
GitLab, Jenkins, Ansible, LaunchDarkly, ALB, F5, AWS
1w
Save
Mark Applied
Hide
Senior Engineering Manager, Site Reliability
United States
$195k-$270k/yr RemoteFull Time
Upstart
UpstartNasdaq: UPST: AI-powered lending marketplace for consumer and automotive loans.
7+ YOE5+ Mgmt5+ years reliability engineering management and 7+ years software engineering experience; strong SRE, incident management, observability, distributed systems, cloud, and leadership skills.
Datadog, Grafana, Prometheus, OpenTelemetry, Kubernetes, AWS
2w
Save
Mark Applied
Hide
Director, Site Reliability Engineering
Frisco or Eagan
$159k-$295k/yr HybridFull Time
Thomson Reuters
Thomson ReutersNASDAQ: TRI: Provides professional software, data, and news services globally.
10+ YOE10+ years in SRE or related tech leadership with experience leading global teams, observability, incident management, automation, and resilience engineering.
2mo
Save
Mark Applied
Hide
Site Reliability Engineering Lead
Florida or Chicago or Boca Raton or Alpharetta
$118k-$220k/yr RemoteFull Time
LexisNexis Risk Solutions
LexisNexis Risk SolutionsNYSE: RELX: Provides data and analytics for risk management and compliance.
Lead SRE teams; implement infrastructure as code and DevOps practices; manage production reliability; cloud (AWS/Azure); Kubernetes and Docker; security tooling; incident management; FinOps cost optimization; collaboration with cross-functional teams.
Amazon Web Services, Microsoft Azure, Kubernetes, Docker, GitHub Advanced Security, Qualys, Wiz, Trufflehog
3w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
United States
$157k-$301k/yr RemoteFull Time
DocuSign
DocuSignNASDAQ: DOCU: Provider of e-signature and intelligent agreement management software.
10+ YOE4+ Mgmt10+ years in Infrastructure/SRE/Software Engineering,4+ years managing engineering teams; experience with automation, incident management, SLOs/Error Budgets, and cloud multi-region architectures.
Go, Python, Kubernetes, Azure, Azure Kubernetes Service/AKS, Internal Developer Platform (IDP)
3w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
United States
$157k-$301k/yr RemoteFull Time
DocuSign
DocuSignNASDAQ: DOCU: Provides electronic signature and agreement management software solutions.
10+ YOE4+ Mgmt10+ years in Infrastructure/SRE/Software Engineering, 4+ years managing engineering teams, experience with automation, incident management, SLOs/Error Budgets, and modern languages (Go or Python).
Go, Python, Kubernetes, Azure Kubernetes Service (AKS), Azure, Internal Developer Platform (IDP)
1mo
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
United States
$230k-$255k/yr RemoteFull Time
Aya Healthcare
Aya Healthcare: Provides healthcare staffing and workforce management software solutions.
10+ YOE4+ Mgmt10+ years in SRE/DevOps/Platform roles, 4+ years people management, deep Azure/AKS and observability experience (Datadog), incident command, AIOps/automation experience, and executive communication skills.
Azure, AKS, Kubernetes, Datadog, New Relic, Dynatrace, AppDynamics, Cloudflare CDN, Cloudflare WAF, Cloudflare Workers, Cloudflare Access (ZTNA), Cloudflare Tunnel, Cloudflare Turnstile, Terragrunt, Terraform, GitHub Actions, OPA, Conftest, Helm, ServiceNow, Okta, Entra ID, M365, Jira, OIDC, AIOps
3w
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
San Francisco, California, United States
$204k-$306k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
United States
$122k-$264k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
10+ YOE10+ years experience leading SRE teams with capacity planning, incident management, automation, and cross-functional collaboration for reliable scalable infrastructure.
2mo
Save
Mark Applied
Hide
Site Reliability Engineering
Foster City, California, United States
$140k-$230k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
5+ YOE5+ years SRE/Distributed systems; cloud platforms (AWS, GCP, or Azure); IaC (Terraform, Ansible, Salt, CloudFormation); Kubernetes; Python/Go/C/C++/Java.
AWS, GCP, Azure, Terraform, Ansible, Salt, CloudFormation, Kubernetes, Python, Go, C/C++, Java
2w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
United States
$155k-$170k/yr RemoteFull Time
Claritas Rx
Claritas Rx: Provides data analytics for specialty biopharmaceutical product performance.
7+ YOE3+ Mgmt7+ years SRE/DevOps or infrastructure engineering experience with 3+ years managing teams; deep AWS, IaC, SLOs, incident management, CI/CD, and compliance (HIPAA/SOC2/HITRUST) experience required.
AWS, ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, CloudFront, AWS CDK, Terraform, GitHub Actions, CloudWatch, Sentry, OpenFeature, Tableau, NestJS, TypeScript, React, PostgreSQL, Turborepo, pnpm, AWS Glue, Lake Formation, PySpark, Kinesis, Python, Jira, Claude, Golang
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore
VantageScore: Provides credit scoring and data analytics solutions.
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
1mo
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
United States or United Kingdom or Hong Kong or New Zealand or North America
$187k-$243k/yr RemoteFull Time
Counterpart Health
Counterpart HealthNASDAQ: CLOV: AI-powered physician enablement platform for value-based care.
10+ YOE6+ Mgmt6+ years managing SRE teams and 10+ years SRE/infrastructure experience; deep experience with Kubernetes, GCP, Terraform, Helm, ArgoCD, PostgreSQL, Prometheus/Grafana; strong Python/Go skills; CI/CD (GitHub Actions); FinOps and platform engineering experience.
Kubernetes, GCP, GKE, Cloud SQL, Pub/Sub, GCS, Terraform, Helm, ArgoCD, PostgreSQL, Prometheus, Grafana, Python, Go, GitHub Actions, Claude Code
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Charlotte, North Carolina, United States
HybridFull Time
Electrolux Group
Electrolux GroupNasdaq Stockholm: ELUX-B: Global manufacturer of household appliances and consumer kitchen equipment.
6+ YOE6+ years in infrastructure/site reliability/cloud engineering; experience with cloud platforms, IaC, CI/CD, observability, troubleshooting, and strong collaboration skills.
Microsoft Azure, AWS, Google Cloud Platform, Akamai CDN, Terraform, CloudFormation, Ansible, Puppet, Chef, Microsoft Azure DevOps, GitHub, Argo CD
1w
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
Salt Lake City, Utah, United States
OnsiteFull Time
O.C. Tanner
O.C. Tanner: Provides employee recognition software and corporate award manufacturing services.
5+ YOE2+ Mgmt5+ years in SRE/DevOps or platform engineering with 2+ years in technical leadership; hands-on AWS and Kubernetes experience; observability (OpenTelemetry, Datadog, Coralogix); incident management and SLO/SLI expertise.
OpenTelemetry, Datadog, Coralogix, AWS, Kubernetes, Terraform, Golang, Python, Playwright, PostgreSQL, OpenSearch, Redis, ElastiCache, Aurora, Kafka, ActiveMQ, SNS, SQS
6d
Save
Mark Applied
Hide
Site Reliability Engineering Manager II
Chicago, Illinois, United States
$160k-$200k/yr HybridFull Time
Flywire
FlywireNASDAQ: FLYW: Platform for processing complex global payments across specialized industries.
5+ YOE2+ Mgmt5+ years SRE experience, 2+ years managing SRE teams; programming experience; familiarity with containers, cloud, CI/CD, and testing methodologies; strong communication and incident response skills.
AWS, Ruby, Java, Kotlin, Go, Node, Python, EC2, ECS, Lambda, Cloudwatch, SQS, RDS, Kinesis, S3, ElasticSearch, DocumentDB, Linux, Docker, Terraform, Make, Chef, Gitlab, Jenkins, Sentry, Sumologic, Honeycomb
1w
Save
Mark Applied
Hide
Senior Engineering Manager, Site Reliability
United States
$260k-$280k/yr RemoteFull Time
Horizon3.ai
Horizon3.ai: Autonomous penetration testing platform for continuous security assessment.
Proven experience building and leading SRE teams, defining incident management and on-call programs, SLO/SLA and runbooks, strong observability and cloud (AWS/GCP/Azure) knowledge, hiring and people management skills.
NodeZero, PagerDuty, FireHydrant, AWS, GCP, Azure
1mo
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
United States or United Kingdom or Hong Kong or New Zealand
$187k-$243k/yr RemoteFull Time
Clover Health
Clover HealthNasdaq: CLOV: Provide Medicare Advantage plans and AI-powered clinical decision tools.
10+ YOE6+ Mgmt6+ years managing SRE teams and 10+ years hands-on SRE/infrastructure experience; strong with Kubernetes, GCP (GKE, Cloud SQL, Pub/Sub, GCS), Terraform, Helm, ArgoCD, PostgreSQL, Prometheus/Grafana; Python or Go; GitHub Actions; leadership across time zones.
Kubernetes, GCP, GKE, Cloud SQL, Pub/Sub, GCS, Terraform, Helm, ArgoCD, PostgreSQL, Prometheus, Grafana, Python, Go, GitHub Actions, Claude Code
2mo
Save
Mark Applied
Hide
Technical Lead - Site Reliability Engineering
Saint Louis, Missouri, United States
OnsiteFull Time
London Stock Exchange Group
London Stock Exchange GroupLondon Stock Exchange: LSEG: Provides financial market infrastructure and global data analytics services.
10+ YOESenior SRE/Platform Engineer with 10+ years of hands-on experience in Azure, Kubernetes, observability, and security; strong leadership and collaboration.
Azure, AKS, Azure Container Apps, Virtual Machines, Virtual Networking, Azure managed services, Kubernetes, Linux, Datadog, Prometheus, Grafana, ELK, OpenTelemetry, SRE, AWS, Terraform, CloudFormation, Generative AI, LLM
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineering - Network
Palo Alto or Columbus
$152k-$215k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE10+ MgmtFormal network engineering training, 5+ years applied experience, 10+ years leading technologists, advanced network reliability skills, SD-WAN and cloud (AWS, Azure) proficiency, major network vendor experience, observability tooling and incident leadership.
SD-WAN, AWS, Azure, Palo Alto, Juniper, F5, Broadcom, Arista, Cisco, Grafana, SevOne, Prometheus, Kibana, ThousandEyes, Splunk, Jenkins, GitLab, Terraform, eBPF, TCP/IP, HTTPS, BGP