17 platform reliability engineer jobs at 15 companies in Georgia

4w
Save
Mark Applied
Hide
Senior Developer - Platform Engineer
Georgia, United States
OnsiteFull Time
Floor & Decor
Floor & DecorNYSE: FND: Specialty retailer of hard surface flooring and related accessories.
5+ YOE5+ years in data engineering or platform engineering; hands-on Databricks and Azure platform administration; experience with ADLS Gen2, Key Vault, AAD, Azure DevOps, SQL Server; focus on governance, security, reliability, performance, and cost optimization.
Databricks, Unity Catalog, Microsoft Azure, Microsoft ADLS Gen2, Microsoft Key Vault, Microsoft Azure Active Directory, Microsoft Azure DevOps, SQL Server, Alteryx Server
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Atlanta, Georgia, United States
$110k-$209k/yr RemoteFull Time
AbbVie
AbbVieNYSE: ABBV: Develops and sells innovative pharmaceutical and biopharmaceutical medicines.
7+ YOE7+ years in site reliability engineering / information security, cloud platforms (AWS/GCP/Azure), CI/CD, Kubernetes, GitOps, Linux/Windows administration; strong security practices.
AWS, GCP, Azure, Kubernetes, Docker, GitOps, Terraform, Crossplane, ArgoCD, Helm, Prometheus, Grafana, OpenTelemetry, Jenkins, GitHub Actions, Azure DevOps, Python, Go, NoSQL, Relational Databases
1w
Save
Mark Applied
Hide
Platform Delivery & Reliability Engineer (Remote) - 29337
Colorado Springs or Columbia or San Antonio or Boise or Greenville or Augusta
$120k-$195k/yr RemoteFull Time
HII
HIINYSE: HII: Builds naval ships and provides global defense technology solutions.
7+ YOEMust obtain U.S. security clearance; 7+ years' relevant experience (varies by degree); deep Kubernetes, IaC, cloud, CI/CD, scripting, SRE practices, troubleshooting and delivery experience across distributed systems.
Kubernetes, Terraform, AWS, Azure, GCP, GitLab CI, Go, Python, Bash, Spark, Trino, Presto, Kafka, NiFi, Iceberg, Delta, Hudi, YouTrack, Nexus
3mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE) - AI Platform & Cloud
Alpharetta, Georgia, United States
OnsiteFull Time
Morgan Stanley
Morgan StanleyNYSE: MS: Global financial services firm providing investment and wealth management.
5+ YOESenior SRE with 5+ years production experience; programming in Python/Go/Java; Kubernetes, Docker, cloud (AWS/Azure/Google), IaC (Terraform/Helm/CloudFormation/Ansible); monitoring (Prometheus/Grafana/ELK/Datadog); networking and GPU/AI compute experience.
Kubernetes, AWS, Azure, Google, API, REST, Python, Go, Java, Docker, Terraform, Helm, CloudFormation, Ansible, Prometheus, Grafana, ELK, EFK, Datadog, Open Telemetry, Loki, Cortex, Kafka, Spark, Flink, SQL, Redis, Snowflake, Slurm, ModelOps, ML Ops, LLM Op
1mo
Save
Mark Applied
Hide
GOV Site Reliability Engineer
United States or Kansas or Washington or California or Texas or Illinois or North Carolina or Colorado or Massachusetts or Pennsylvania or Virginia or Oregon or Nevada or Hawaii or New York or Georgia or Ohio or Arizona or Seattle or San Francisco or New York City
$110k-$183k/yr RemoteFull Time
Veeam
Veeam: Data resilience and security for hybrid cloud environments
3+ YOE3+ years in software engineering with 1+ year in SRE/Platform/DevOps, cloud experience (Azure or comparable), observability (Prometheus, Grafana, OpenTelemetry, ELK), IaC (Terraform/Terragrunt/Pulumi), Kubernetes, CI/CD tooling, programming in TypeScript/JS, Go, Java, or C#, and experience in compliance-oriented environments.
VDC, Prometheus, Grafana, OpenTelemetry, ELK stack, Terraform, Terragrunt, Pulumi, Kubernetes, GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, TypeScript, JS, Go, Java, C#, Azure Government, AWS GovCloud
1d
Save
Mark Applied
Hide
IT Spec (ENTARCH) "Platform/site Reliability Engineer", GS-2210-14 FPL GS-14 (DH)
San Francisco or Denver or Washington or Atlanta or Chicago or Salt Lake City or Seattle
$128k-$197k/yr HybridFull Time
Federal Student Aid
Federal Student Aid: Administers federal student loans and financial aid programs.
1+ YOEDesign and implement scalable cloud platforms, IaC, CI/CD, containers, observability, SRE practices; lead platform engineering and security; must be U.S. citizen and pass background/fingerprint check.
Infrastructure as Code (IaC), CI/CD, containers
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer II
Louisville or Atlanta or Lehi
OnsiteFull Time
Waystar
WaystarNASDAQ: WAY: Provides cloud-based healthcare payment and revenue cycle management software.
7+ YOE7+ years in SRE/DevOps or infrastructure engineering; cloud platforms (AWS, GCP, Azure); Kubernetes; IaC (Terraform, CloudFormation); observability tools (Prometheus, Grafana, Splunk); CI/CD; data platforms/ETL; Python and PowerShell; AI tooling experience.
AWS, GCP, Azure, Kubernetes, Infrastructure as Code, Terraform, CloudFormation, Prometheus, Grafana, Splunk, CI/CD, Python, PowerShell
2w
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
3mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE) - AI Platform & Cloud
Alpharetta, Georgia, United States
OnsiteFull Time
Morgan Stanley
Morgan StanleyNYSE: MS: Provides global investment banking, wealth management, and advisory services.
5+ YOE5+ years of production SRE/infrastructure experience; strong programming/scripting; containerization (Docker, Kubernetes); IaC (Terraform/CloudFormation/Ansible); cloud platforms; experience with monitoring, security, and compliance in regulated environments.
Kubernetes, Cloud (AWS, Azure, Google), API based development, REST framework, Terraform, Helm, CloudFormation, Ansible, Prometheus, Grafana, ELK, EFK, Datadog, TCP/IP, DNS, load balancing, Kafka, Spark, Flink, Snowflake, Redis, Docker, Open Telemetry, Grafana, Loki, Cortex, Canary deployments, Blue/Green deployments, Slurm
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer – Unified Observability
Atlanta, Georgia, United States
OnsiteFull Time
NCR Voyix
NCR VoyixNYSE: VYX: Provides checkout software and kiosks for retailers and restaurants.
10+ YOE10+ years SRE/Platform/Cloud experience; expertise with Azure, GCP, Kubernetes (AKS,GKE); observability tools (Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic); Terraform; Python/Go/PowerShell; bachelor\u0002s degree or equivalent.
Azure, Google Cloud Platform, Kubernetes, AKS, GKE, Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, ServiceNow, Terraform, Python, Go, PowerShell
1mo
Save
Mark Applied
Hide
Staff Platform Engineer (Fully Remote)
Atlanta or United States
$175k-$200k/yr RemoteFull Time
PadSplit
PadSplit: Operates a marketplace for affordable-living and room rentals.
Extensive experience with distributed systems, cloud infrastructure (AWS), backend architecture, Django, and Postgres; strong systems design, reliability, and mentoring skills required.
Django, AWS, Postgres
4d
Save
Mark Applied
Hide
Software Engineer Manager - Platform Reliability Engineering (Remote)
Georgia, United States
$140k-$240k/yr RemoteFull Time
The Home Depot
The Home DepotNYSE: HD: Retailer of home improvement products, building materials, and tools.
5+ YOEManage and lead a platform reliability engineering team to ensure cloud platform resilience, SLO enforcement, incident response, automation, and platform enablement; 5+ years experience and bachelor's degree typical.
Java, Google Cloud, Kubernetes, GKE, Terraform
4w
Save
Mark Applied
Hide
Software Reliability Engineer
Atlanta, Georgia, United States
$84k-$151k/yr OnsiteFull Time
T-Mobile
T-MobileNASDAQ: TMUS: Provides wireless voice, data, and mobile internet services.
2+ YOEBachelor's degree (or equivalent) plus experience, DevOps/SRE experience with CI/CD, cloud-native platforms, containerization, automation, monitoring and incident troubleshooting. Familiarity with languages (C, C#, Java, Perl, Python, Go), CI/CD and DevOps tools is preferred.
C, C#, Java, Perl, Python, Go, Shell, Jenkins, Cloudbees, Ansible, Chef, Puppet, Docker, Kubernetes, AppDynamics, Splunk, Dynatrace, Grafana, Prometheus, Terraform, CI/CD, APM, DevOps
1mo
Save
Mark Applied
Hide
VP of Site Reliability
Atlanta, Georgia, United States
RemoteFull Time
Titan
Titan: AI platform providing secure models and agents for banking.
10+ YOE10+ years in engineering with 5+ years building SRE/platform operations for enterprise or regulated markets; hands-on leader; incident response and on-call expertise.
1d
Save
Mark Applied
Hide
Senior Software Engineer - Reliability Engineering (Remote)
Georgia, United States
$90k-$170k/yr RemoteFull Time
The Home Depot
The Home DepotNYSE: HD: Retailer of home improvement products, building materials, and tools.
3+ YOE3+ years relevant engineering experience; expertise in reliability/SRE practices, incident/problem/change management, cloud platform (GCP), infrastructure automation, observability, container orchestration, and mentoring.
BASH, Python, Golang, Typescript, Java, YAML, JSON, HCL, Terraform, Ansible, Google Cloud Platform, Prometheus, Grafana, OpenTelemetry, Kubernetes, GKE, Unix, Linux, Wiz
1w
Save
Mark Applied
Hide
Data360 Platform Director
San Francisco or New York City or Atlanta or Seattle
$197k-$314k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
10+ Mgmt10+ years leading engineering teams with platform and DevOps expertise; experience with cloud data platforms, automation, reliability, security, and compliance.
Data 360 (Data Cloud), Snowflake, Tableau, MuleSoft
2w
Save
Mark Applied
Hide
Lead Infrastructure Engineer
Charlotte or Atlanta
OnsiteFull Time
Truist
TruistNYSE: TFC: Offers personal banking, business lending, and investment management services.
10+ YOEBachelor's in CS/Engineering/IS, 10+ years infrastructure engineering experience with advanced knowledge of cloud, network, database, storage, platform, automation, and enterprise-scale reliability.
Python, FastAPI, Java, Kubernetes, OpenShift, Open Policy Agent (OPA), Rego, Apache Kafka, Confluent, GitLab CI/CD, Backstage, HashiCorp Vault, OpenTelemetry, Prometheus, Splunk, Jaeger, Tempo, AWS, Azure, Terraform, Ansible, Cosign, CycloneDX, GitLab Duo, GitHub Copilot

Explore Jobs

Expand Your Job Search