Site Reliability Engineer (SRE) - AI Platform & Cloud
Alpharetta, Georgia, United States
OnsiteFull Time
Morgan StanleyNYSE: MS: Global financial services firm providing investment and wealth management.
5+ YOESenior SRE with 5+ years production experience; programming in Python/Go/Java; Kubernetes, Docker, cloud (AWS/Azure/Google), IaC (Terraform/Helm/CloudFormation/Ansible); monitoring (Prometheus/Grafana/ELK/Datadog); networking and GPU/AI compute experience.
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
EquifaxNew York Stock Exchange: EFX: Provides consumer credit reporting and data analytics services globally.
5+ YOEBS in CS or related technical field; 5-7 years in software engineering, systems, database or networking; 2+ years in public cloud; Python/Bash/Java/Go/JavaScript; Terraform and CI/CD; on-call; strong problem solving.
United States or San Francisco or Boston or Atlanta or Austin or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
4+ YOEData engineering background, 4 years Python, 1 year Airflow administration/DAG creation, Kubernetes/Docker experience, cloud provider (AWS/GCP/Azure) experience, troubleshooting, strong communication, and mentoring experience.
United States or Kansas or Washington or California or Texas or Illinois or North Carolina or Colorado or Massachusetts or Pennsylvania or Virginia or Oregon or Nevada or Hawaii or New York or Georgia or Ohio or Arizona or Seattle or San Francisco or New York City
$110k-$183k/yrRemoteFull Time
Veeam: Data resilience and security for hybrid cloud environments
3+ YOE3+ years in software engineering with 1+ year in SRE/Platform/DevOps, cloud experience (Azure or comparable), observability (Prometheus, Grafana, OpenTelemetry, ELK), IaC (Terraform/Terragrunt/Pulumi), Kubernetes, CI/CD tooling, programming in TypeScript/JS, Go, Java, or C#, and experience in compliance-oriented environments.
Unum GroupNYSE: UNM: Provides insurance and employee benefits to businesses and individuals.
5+ YOEBachelor's in Computer Science/Engineering plus 5+ years experience; expertise with observability, cloud-native architectures, incident response, automation scripting, CI/CD, infrastructure-as-code, and collaboration in DevOps environments.
Austin or North Dakota or Montana or Maine or New Mexico or New Hampshire or Kentucky or Alabama or Ohio or Nebraska or Louisiana or South Carolina or Illinois or Texas or Nevada or Hawaii or Georgia or Missouri or Iowa or Mississippi or Tennessee or Colorado or North Carolina or Minnesota or Kansas or Wyoming or Wisconsin or West Virginia or Rhode Island or Delaware or Vermont or Arkansas or Utah or Oregon or South Dakota or Florida or Pennsylvania or Michigan or Indiana or Idaho or Arizona or Oklahoma
$127k-$182k/yrRemoteFull Time
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOESTEM degree or equivalent experience,5+ years Linux production experience,Python or Ruby,Ansible,system debugging,cloud/on-prem experience;Kubernetes and FedRAMP knowledge preferred;must be a U.S. Person.
T-MobileNASDAQ: TMUS: Provides wireless voice, data, and mobile internet services.
2+ YOEBachelor's degree (or equivalent) plus experience, DevOps/SRE experience with CI/CD, cloud-native platforms, containerization, automation, monitoring and incident troubleshooting. Familiarity with languages (C, C#, Java, Perl, Python, Go), CI/CD and DevOps tools is preferred.
LexisNexis Risk SolutionsNYSE: RELX: Provides data and analytics for risk management and compliance.
Lead SRE teams; implement infrastructure as code and DevOps practices; manage production reliability; cloud (AWS/Azure); Kubernetes and Docker; security tooling; incident management; FinOps cost optimization; collaboration with cross-functional teams.
Amazon Web Services, Microsoft Azure, Kubernetes, Docker, GitHub Advanced Security, Qualys, Wiz, Trufflehog