Amazon
Posted 2mo ago

Site Reliability Engineer - Software Ops and Scaling , One Material Handling System - Software, Controls and Science

Amazon
Nashville or Arlington or Bellevue
$123k-$175k/yrOnsiteFull Time
Responsibilities
  • designing automation
  • improving processes
  • monitoring systems
Requirements
  • Experience automating
  • Deploying, and supporting large-scale infrastructure
  • Programming in Python, Ruby, Golang, Java, C++, C#, or Rust
  • Linux/Unix proficiency
  • CI/CD pipeline experience
  • Distributed systems experience preferred
Technical tools mentioned
PythonRubyGolangJavaC++C#RustLinux/UnixCI/CD

Job description

Description

Are you inspired by solving complex automation challenges?

Do you thrive on improving processes and creating efficient, scalable solutions? Are you passionate about bridging the gap between development and operations? Join our team as a Software Operations and Scaling, DevOps Engineer to drive technical excellence in our automation infrastructure.

As a key member of our team, you'll shape the future of automation DevOps at scale, working with exciting systems while driving process improvements that impact operations daily. You'll be instrumental in bridging development and operations, creating efficient, reliable, and scalable solutions for our growing automation infrastructure.



Key job responsibilities
- Design and implement automation frameworks for deployment, monitoring, and maintenance of automated systems
- Architect and drive continuous improvement initiatives for system CI/CD pipelines
- Develop and implement service improvement strategies for system reliability and performance
- Create scalable monitoring and diagnostic solutions for complex systems
- Lead technical process improvements across development and operations teams
- Drive standardization of DevOps practices across automation systems
- Collaborate with other support teams to enhance operational efficiency
- Implement predictive maintenance and monitoring solutions

A day in the life
Amazon offers a full range of benefits that support you and eligible family members, including domestic partners and their children. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include:

1. Medical, Dental, and Vision Coverage
2. Maternity and Parental Leave Options
3. Paid Time Off (PTO)
4. 401(k) Plan

If you are not sure that every qualification on the list above describes you exactly, we'd still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets. If you’re passionate about this role and want to make an impact on a global scale, please apply!

Basic Qualifications

- Experience in automating, deploying, and supporting large-scale infrastructure
- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
- Experience with Linux/Unix
- Experience with CI/CD pipelines build processes

Preferred Qualifications

- Experience with distributed systems at scale

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.



USA, TN, NASHVILLE - 122,800.00 - 166,100.00 USD annually
USA, TN, Nashville - 122,800.00 - 166,100.00 USD annually
USA, VA, ARLINGTON - 129,200.00 - 174,800.00 USD annually
USA, VA, Arlington - 129,200.00 - 174,800.00 USD annually
USA, WA, BELLEVUE - 129,200.00 - 174,800.00 USD annually
USA, WA, Bellevue - 129,200.00 - 174,800.00 USD annually

About Amazon

Global online retail and cloud computing technology provider.

Similar jobs

Site Reliability Engineer roles near Nashville, Tennessee
2d
Save
Mark Applied
Hide
Site Reliability Engineer, US Gov
Denver or Arvada or San Francisco or Nashville or Santa Fe or New Orleans or San Diego or Bozeman or United States
$160k-$200k/yr HybridFull Time
Quindar
Quindar: Cloud-native software for automated satellite mission operations.
3+ YOERequires a bachelor's degree, 3+ years of SRE or infrastructure experience, U.S. citizenship, Secret clearance or higher, and expertise in Kubernetes, AWS, Python, Terraform, networking, and CI/CD.
AWS GovCloud, AWS C2E, Kubernetes, AWS EKS, Rancher, Grafana LGTM, Datadog, Python, Terraform, VPN, NLB, ALB, HTTPS, TLS, VPC peering, CDN, GitLab Workflows, Unix, Linux, Auth0, Keycloak, AWS IAM, Git
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Nashville or Brentwood
$180k-$210k/yr HybridFull Time
360 Privacy
360 Privacy: Protects digital privacy for high-profile individuals and executives.
8+ YOERequires 8+ years in SRE, DevOps, or platform engineering; deep Kubernetes/EKS expertise; ElasticSearch, observability, Terraform, AWS IAM, cloud security, and advanced English proficiency.
Kubernetes, Amazon EKS, ElasticSearch, Kibana, Terraform, GitHub Actions, ArgoCD, AWS IAM, AWS RDS, Python
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Nashville, Tennessee, United States
$85k-$210k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
3+ YOEExpertise in reliability engineering, distributed systems, cloud infrastructure, automation, observability, incident management, and software development; 3+ years experience; strong Python/Java/Go and IaC/CI-CD skills.
Python, Java, Go, OCI, AWS, Azure, GCP, Kubernetes, Docker, Terraform, Ansible, Helm, Pulumi, CI/CD Tools, Prometheus, Grafana, OpenTelemetry, ELK, Splunk, Datadog, Large Language Models (LLMs)
2mo
Save
Mark Applied
Hide
Site Reliability Engineer II
Falls Church or South Carolina or Raleigh or Nashville or Louisiana or Pennsylvania or Plain City or South Bend or Orlando or Detroit
OnsiteFull Time
Kastle Systems
Kastle Systems: Managed security services provider for commercial and residential properties.
4+ YOE4+ years SRE/Platform experience owning production systems. Hands-on with Azure/AKS, Kubernetes, Terraform/OpenTofu/Pulumi, GitOps/ArgoCD, observability (Prometheus/Grafana/OpenTelemetry/ELK), Python/Go/Bash, and feature-flag/CI/CD practices.
ArgoCD, Flux, Crossplane, LaunchDarkly, Flagsmith, Terraform, OpenTofu, Pulumi, Prometheus, Grafana, OpenTelemetry, ELK, OpenSearch, Python, Go, Bash, C#, SQL, AKS, Azure Container Registry, Azure Monitor, Cosmos DB, Key Vault, Azure Front Door, GitOps
2mo
Save
Mark Applied
Hide
Site Reliability Engineer II
Raleigh or Nashville or South Carolina or Louisiana or Pennsylvania or Plain City or South Bend or Orlando or Detroit
OnsiteFull Time
Kastle Systems
Kastle Systems: Provides managed security and property technology for buildings and businesses
4+ YOE4+ years SRE/platform experience; Azure, Kubernetes, GitOps, Terraform/OpenTofu, observability and incident management; strong scripting (Python/Go/Bash) and communication skills.
ArgoCD, GitOps, Terraform, OpenTofu, Pulumi, AKS, Azure Container Registry, Azure Monitor, Cosmos DB, Key Vault, Azure Front Door, Kubernetes, Crossplane, Prometheus, Grafana, OpenTelemetry, ELK, OpenSearch, Python, Go, Bash, C#, SQL, Flux, LaunchDarkly, Flagsmith, Linux