42 cloud reliability engineer jobs at 32 companies in Mountain Park, GA

2w
Save
Mark Applied
Hide
Cloud Site Reliability Engineer
Charlotte or Boone or Irving or North Wilkesboro or Flowery Branch or Aurora or Englewood or Fullerton or Coppell or Windsor Mill
RemoteFull Time
Samaritan's Purse
Samaritan's Purse: Nondenominational evangelical Christian organization providing humanitarian aid.
5+ YOEBachelor's degree or related certificate with 5–7 years of cloud engineering experience, or 7+ years related experience. Requires SRE/DevOps, Kubernetes, Linux, GitOps, observability, and 12 Biblical Studies credits.
Grafana, Prometheus, ELK, Datadog, Dynatrace, ArgoCD, Flux, Git, Kubernetes, Linux, OpenStack, Infrastructure as Code, CI/CD, AI
3w
Save
Mark Applied
Hide
Systems Reliability Engineer
Overland Park or Atlanta or Frisco or Bellevue
$84k-$151k/yr OnsiteFull Time
T-Mobile
T-MobileNASDAQ: TMUS: The Un-carrier providing wireless and home internet services.
2+ YOEBachelor's degree required; 2–4+ years preferred. Requires DevOps, cloud, automation, monitoring, scripting, APIs, cybersecurity, and reliability engineering experience, plus U.S. work authorization.
C, C#, Java, Perl, Python, Go, Jenkins, CloudBees, Ansible, Chef, Puppet, Docker, Kubernetes, AppDynamics, Splunk, Microsoft Graph API, REST API, Microsoft Power Apps, Microsoft Power Automate, Microsoft Entra, SailPoint, ServiceNow, Azure, Microsoft Azure DevOps Pipelines, Linux, Shell
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Private enterprise software providing AI-powered identity security and access governance for global enterprises and government institutions.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Alpharetta, Georgia, United States
OnsiteFull Time
Morgan Stanley Wealth Management
Morgan Stanley Wealth ManagementNYSE: MS: Private wealth-management business serving individuals, families, businesses, institutions and foundations with advice, brokerage and financial planning.
5+ YOE5+ years production experience; strong scripting (Python, Perl, Shell, Ruby, Java, C#); DB2/Sybase/Oracle, Autosys, CI/CD, containers/VMs, Splunk/IP Soft/Sockeye, Jenkins/Train; cloud (Azure/AWS); BS in CS/Engineering required.
Python, Perl, Shell, Ruby, Java, C#, DB2, Sybase, Oracle, Autosys, Splunk, IP Soft, Sockeye, Jenkins, Train, Azure, AWS, MQ, UNIX, Linux, Windows
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE)
Atlanta or United States
$120k-$175k/yr RemoteFull Time
PrizePicks
PrizePicks: Atlanta-based daily fantasy sports platform offering skill-based prediction games and cash prizes to sports fans.
5+ YOE5+ years reliability engineering experience; cloud (AWS/Azure/GCP), IaC (Terraform/Crossplane), Kubernetes, Python/Ruby/Go, monitoring tools, incident response, SLO governance, strong cross-functional and debugging skills.
AWS, Azure, GCP, Terraform, Crossplane, Python, Ruby, Go, Kubernetes, Grafana, New Relic, Datadog, Windows, Mac
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Alpharetta or United States
$129k-$161k/yr RemoteFull Time
Priority Commerce
Priority CommerceNasdaq Capital Market: PRTH: Public U.S. payments and banking fintech serving businesses with merchant acquiring, payables, treasury, and embedded-finance solutions.
5+ YOE5+ years software/systems engineering (including 3+ years SRE), expertise in distributed systems, cloud (AWS preferred), CI/CD, observability, incident management, Java/Node.js/JavaScript, and database experience.
AWS, CI/CD, Java, Node.js, JavaScript
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Atlanta or Alpharetta
OnsiteFull Time
Incident IQ
Incident IQ: Private K-12 workflow management software helping school districts manage IT, assets, facilities, events, resources, and HR.
5+ YOE5+ years SRE/DevOps experience, strong systems fundamentals, SLI/SLO and observability experience, incident management, automation and cloud skills, proficient with modern SRE tooling and AI-accelerated execution.
Grafana, PromQL, Grafana Alloy, Prometheus, Datadog, OpenTelemetry, SigNoz, Uptrace, Tempo, Grafana Faro, k6, PagerDuty, Locust, JMeter, Python, Go, Bash, Terraform, Ansible, Kubernetes, Amazon Web Services (AWS), Google Cloud Platform (GCP), Azure, GitOps, eBPF, Grafana Beyla, OpenTelemetry eBPF Instrumentation, .NET, Real User Monitoring (RUM)
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer I
Boston or Seattle or Atlanta
$134k-$215k/yr HybridFull Time
Axon
AxonNASDAQ: AXON: Develops public safety technologies, devices, and cloud software.
7+ YOEBachelor's in CS/Engineering, 7+ years software engineering experience, expertise in distributed systems, Kubernetes, cloud (Azure/AWS/GCP), observability, Kafka, Terraform/Pulumi, and experience with agentic AI/LLM tooling preferred.
Kubernetes, Terraform, Pulumi, Kafka, Grafana, Datadog, New Relic, MySQL, Cassandra, PostgreSQL, Azure, AWS, GCP
3w
Save
Mark Applied
Hide
Systems Reliability Engineer
Overland Park or Atlanta or Frisco or Bellevue
$84k-$151k/yr OnsiteFull Time
T-Mobile
T-MobileNASDAQ: TMUS: The Un-carrier providing wireless and home internet services.
2+ YOEBachelor's degree required; 2–4+ years preferred. Requires DevOps, cloud, automation, monitoring, Python, APIs, Power Platform, identity governance, and reliability engineering experience.
C, C#, Java, Perl, Python, Go, Shell, Jenkins, CloudBees, Ansible, Chef, Puppet, Docker, Kubernetes, AppDynamics, Splunk, Microsoft Graph API, REST API, Microsoft Power Apps, Microsoft Power Automate, Microsoft Entra, SailPoint, ServiceNow, Azure, Microsoft Azure DevOps Pipelines, Microsoft Power Platform, Linux, VMs
2w
Save
Mark Applied
Hide
Site Reliability Engineer - Intermediate
Alpharetta, Georgia, United States
HybridFull Time
Equifax
EquifaxNYSE: EFX: Global data, analytics, and technology.
2+ YOEBachelor's degree or equivalent experience; 2–5 years in software, systems, database, or networking; 1+ year public cloud experience; Linux/Windows administration, automation, monitoring, and coding skills.
Terraform, Python, Bash, Java, Go, JavaScript, Node.js, Chef, Ansible, Docker, Kubernetes
2w
Save
Mark Applied
Hide
Site Reliability Engineer - Intermediate
Alpharetta or Atlanta
HybridFull Time
Equifax
EquifaxNYSE: EFX: Global data, analytics, and technology.
2+ YOEBachelor's degree or equivalent experience; 2–5 years in software, systems, database, or networking roles; 1+ year with public cloud; Linux/Windows administration, automation, monitoring, and programming experience.
Terraform, Python, Bash, Java, Go, JavaScript, Node.js, Chef, Ansible, Docker, Kubernetes
1mo
Save
Mark Applied
Hide
Sr Site Reliability Engineer
Atlanta, Georgia, United States
$178k-$205k/yr HybridFull Time
Workday
Workday: A provider of cloud-based enterprise software and services focused on managing people and finances.
5+ YOEBachelor's degree plus 5 years experience. 5+ years with Ansible, Terraform, Packer, Python, Shell Scripting, Kafka, CI/CD tools (GIT, Maven/Gradle, Jenkins), Docker, Kubernetes, and cloud provisioning; strong monitoring and automation skills.
Ansible, Terraform, Packer, Python, Shell Scripting, Kafka, GIT, Maven, Gradle, Jenkins, Docker, Kubernetes
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Alpharetta, Georgia, United States
OnsiteFull Time
Morgan Stanley Wealth Management
Morgan Stanley Wealth ManagementNYSE: MS: Private wealth-management business serving individuals, families, businesses, institutions and foundations with advice, brokerage and financial planning.
5+ YOEMinimum 5 years production experience; strong scripting (Python, Perl, Shell, Ruby, Java, C#); DB2/Sybase/Oracle, Autosys, Jenkins/Train, Splunk/IP Soft/Sockeye; cloud (Azure/AWS); BS in CS/Engineering required.
Python, Perl, Shell, Ruby, Java, C#, DB2, Sybase, Oracle, Autosys, Jenkins, Train, Splunk, IP Soft, Sockeye, Azure, AWS, UNIX, Linux, Windows
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer – Unified Observability
Atlanta, Georgia, United States
OnsiteFull Time
NCR Voyix
NCR VoyixNYSE: VYX: Global leader in digital commerce for retail and restaurants.
10+ YOE10+ years SRE/Platform/Cloud experience; expertise with Azure, GCP, Kubernetes (AKS,GKE); observability tools (Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic); Terraform; Python/Go/PowerShell; bachelor\u0002s degree or equivalent.
Azure, Google Cloud Platform, Kubernetes, AKS, GKE, Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, ServiceNow, Terraform, Python, Go, PowerShell
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Atlanta, Georgia, United States
OnsiteFull Time
Inspire Brands
Inspire Brands: Private multi-brand restaurant owning, operating, and franchising six restaurant brands for guests and franchisees worldwide.
5+ YOERequires 5+ years in SRE, software, or platform engineering; 2+ years with Kubernetes and containers; a computer science or related bachelor's degree; programming, cloud, observability, SLO, and incident response expertise.
Kubernetes, Python, Go, Java, Node.js, Azure, AWS, GCP, Terraform, Bicep, CI/CD
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Atlanta, Georgia, United States
OnsiteFull Time
Inspire Brands
Inspire Brands: Private multi-brand restaurant owning, operating, and franchising six restaurant brands for guests and franchisees worldwide.
5+ YOERequires 5+ years in SRE, software, or platform engineering; 2+ years with Kubernetes and containers; a computer science or related bachelor's degree; programming, SLO, incident response, distributed systems, cloud, and observability expertise.
Kubernetes, Python, Go, Java, Node.js, Azure, AWS, GCP, Terraform, Bicep, CI/CD, SLOs, Error Budgets
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States or Atlanta or Chicago
$100k-$120k/yr HybridFull Time
Origami
Origami: Privately held Origami Risk provides SaaS risk, safety, and insurance software to insurers, companies, and public entities.
5+ YOE5+ years SRE experience, strong incident management, observability tooling (New Relic, Data Dog, SumoLogic), cloud (AWS/Azure), coding (JavaScript,.NET,C#,+SQL), CI/CD and IaC familiarity, strong communication and problem-solving.
New Relic, Data Dog, SumoLogic, JavaScript, .NET, C#, SQL, SQL Server, AWS, Azure, Windows, CI/CD, Infrastructure as Code (IaC)
3d
Save
Mark Applied
Hide
Reliability Engineer 3 (Observability Specialist)
Brookfield or Atlanta or Hopkins or Cupertino or Charlotte or New York City or Chicago or Gresham or Englewood or Cincinnati or Irving or Earth City
$98k-$116k/yr HybridFull Time
U.S. Bank
U.S. BankNew York Stock Exchange: USB: Diversified financial services and banking institution.
5+ YOEBachelor's degree or equivalent experience and 5–7 years in IT service management, production support, risk analysis, product/project management, or application development; observability, SRE, cloud, distributed systems, and Kubernetes expertise preferred.
Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, OpenTelemetry, Kubernetes
3w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Austin or Atlanta
$100k-$115k/yr OnsiteFull Time
Atlanticus Services Corporation
Atlanticus Services CorporationNASDAQ: ATLC: Public fintech helping U.S. banks provide credit cards, consumer loans, healthcare financing, and auto financing to underserved consumers.
5+ YOERequires 5+ years supporting production applications, Java, AWS, Kubernetes, Docker, Datadog or Splunk, CI/CD, Python or Bash, Linux, cloud troubleshooting, and incident management experience.
AWS, Amazon EKS, Amazon EC2, ALB/NLB, Amazon RDS, IAM, Amazon Route 53, Amazon CloudWatch, Amazon S3, VPC, Datadog, Splunk, Docker, Kubernetes, Jenkins, GitHub Actions, Argo CD, MySQL, Oracle, Python, Bash, Linux, Helm, Terraform, Prometheus, Grafana, OpenTelemetry, Karpenter, Cluster Autoscaler, Java, JVM
1w
Save
Mark Applied
Hide
Site Reliability Engineer III
Alpharetta, Georgia, United States
$87k-$144k/yr OnsiteFull Time
LexisNexis Risk Solutions
LexisNexis Risk Solutions: Private data analytics helping businesses and governments manage risk, prevent fraud, verify identities, and improve operations.
6+ YOEBachelor's degree required, 6+ years in infrastructure, cloud, DevOps, or SRE, 5+ years with Azure, 3+ years supporting Kubernetes, production systems, security tooling, observability, and infrastructure automation.
Microsoft Azure, AWS, Kubernetes, OTEL, SIEM, Grafana, Checkov, GitHub Actions, NIST, ISO 27001, SOC2, DLP, CIS, Chainguard, MFA, OIDC, Terraform, ArgoCD, Snyk, WIZ, TruffleHog, Qualys, Vault, Python, Bash, PowerShell, AI/ML, Prometheus, Loki