Alibaba Cloud
Posted 1mo ago

Alibaba Cloud-Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale

Alibaba Cloud
Sunnyvale, California, United States
$104k-$171k/yrOnsiteFull Time
Responsibilities
  • maintaining stability
  • responding incidents
  • building automation
Requirements
  • 2+ years in distributed systems reliability engineering
  • High-availability architecture
  • Kafka/RocketMQ
  • Kubernetes
  • Automation, and proficiency in Python, Go, or Java required. Bachelor's degree listed
Technical tools mentioned
RocketMQKafkaKubernetesK8sJavaGoPythonShellTerraformHelmOperator

Job description

Job Description

Alibaba Cloud Native Message Middleware Team is responsible for message products, including RocketMQ and other messaging products. We are committed to creating a more stable, user-friendly, streaming, and large-scale messaging platform for the future.

Cloud Product Operations & Reliability

Oversee stability maintenance, performance tuning, and high-availability architecture design for cloud middleware, including messaging middleware (Kafka/RocketMQ).

Manage the containerized middleware lifecycle on Kubernetes clusters: implement deployments, auto-scaling, version upgrades, and resource optimization in K8s environments.

Incident Response & Root Cause Analysis

Lead the troubleshooting of middleware-related incidents (e.g., message backlog, service registration failures) through log analysis, distributed tracing, and monitoring systems.

Develop diagnostic tools using Java/Go to resolve production issues, performance bottlenecks, and compatibility challenges.

Automation & Operational Excellence

Build Python/Go/Shell automation tools to standardize middleware deployment, monitoring, and disaster recovery workflows.

Implement chaos engineering experiments, capacity planning strategies, and failover mechanisms to enhance system resilience.

Strong scripting skills in Shell/Python and experience with Infrastructure as Code (IaC) tools (Terraform preferred).

Position Requirement

Minimum Qualification:

Experience: Over 2 years of experience in distributed systems reliability engineering, familiar with high-availability architecture design, and proficient in at least one of Python, Go, or Java.

Messaging: Cluster management, message reliability assurance, and performance optimization for Kafka/RocketMQ.

Hands-on experience deploying middleware on Kubernetes (Helm/Operator preferred).

Automation: Ability to convert operations experience into automated solutions and familiarity with various message middleware, e.g., Kafka and RocketMQ.

Preferred Qualification:

SRE Practices: Familiar with core SRE practices (incident review, error budgeting, chaos engineering) and experienced in building automated risk control systems.

The pay range for this position at commencement of employment is expected to be between $104,400 and $171,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

If hired, employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

About Alibaba Cloud

Global cloud computing infrastructure and services provider

Similar jobs

Site Reliability Engineer roles near Sunnyvale, California
1d
Save
Mark Applied
Hide
Site Reliability Graduate (Data Infrastructure) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's degree in computer science, computer engineering, or related field required; master's preferred. Requires scripting, Linux, and networking knowledge. Docker, Kubernetes, data stores, observability, and data center experience preferred.
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Linux, Docker, PostgreSQL, Prometheus, Grafana, ELK Stack
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$160k-$240k/yr OnsiteFull Time
Fiserv
FiservNew York Stock Exchange: FI: Provides financial technology and payment processing services to institutions.
Requires mid-to-senior site reliability, operations, or DevOps experience; shell scripting; GCP, GKE, Kubernetes, IaC, monitoring tools, HAProxy, GitHub Actions, and strong troubleshooting skills.
Google Cloud Platform (GCP), GKE, Kubernetes, Terraform, Ansible, Puppet, Prometheus, Grafana, Datadog, HAProxy, GitHub, GitHub Actions, Python, Go, Java
1d
Save
Mark Applied
Hide
Lead Site Reliability Engineer, Platforms
San Jose, California, United States
$124k-$271k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
8+ YOE8+ years of SRE or DevOps experience, programming proficiency, cloud and Kubernetes expertise, Terraform, CI/CD, observability, incident response, U.S. citizenship or green card, and a computer science degree or equivalent.
Kubernetes, AWS, OCI, Terraform, Git, Jenkins, Argo CD, JFrog, ELK, Prometheus, Grafana, Python, Go, Java, IAM, Teleport, Okta, BrightHire
2d
Save
Mark Applied
Hide
Site Reliability Engineer (US - Pacific time)
San Francisco or United States
$271k-$296k/yr RemoteFull Time
PostHog
PostHog: All-in-one product analytics and developer tools platform
Requires hands-on production Kubernetes and AWS experience, Terraform or Terragrunt automation, Linux systems knowledge, stateful infrastructure experience, production debugging, and end-to-end on-call ownership.
Amazon Web Services (AWS), Kubernetes, Amazon Elastic Kubernetes Service (EKS), Karpenter, Cilium, ArgoCD, Terraform, Terragrunt, GitHub Actions, Linux, Cloudflare, Microsoft SharePoint
2d
Save
Mark Applied
Hide
Senior Site reliability Engineer (Linux) IRC302310
San Jose, California, United States
$130k-$140k/yr RemoteFull Time
GlobalLogic
GlobalLogic: Digital product engineering and software development services provider.
7+ YOERequires 7+ years of Linux production experience, automation skills, debugging expertise, cloud and on-premises experience, and a bachelor's or master's degree in a related field.
AWS, Linux, Python, Ruby, Ansible, Docker, Kubernetes, AMI, C, Go
2d
Save
Mark Applied
Hide
Senior Site Reliability Engineer, BCM - DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOEBachelor's degree or equivalent in computer science or related field, 8+ years in site reliability engineering or software development, Python, Linux, networking, and cluster operations experience.
Python, Linux, C++, Kubernetes, Slurm, InfiniBand, Spectrum-X, BCM, Bright Cluster Manager, Base Command Manager
2d
Save
Mark Applied
Hide
Senior Site Reliability Engineer, BCM - DGX Cloud
Santa Clara or United States
$168k-$334k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEBachelor's degree or equivalent in computer science or related field, 8+ years in site reliability engineering or software development, Python fluency, and in-depth Linux and networking knowledge.
Python, Linux, C++, Kubernetes, Slurm, InfiniBand, Spectrum-X
4d
Save
Mark Applied
Hide
Site Reliability Engineer, US Gov
Denver or Arvada or San Francisco or Nashville or Santa Fe or New Orleans or San Diego or Bozeman or United States
$160k-$200k/yr HybridFull Time
Quindar
Quindar: Cloud-native software for automated satellite mission operations.
3+ YOERequires a bachelor's degree, 3+ years of SRE or infrastructure experience, U.S. citizenship, Secret clearance or higher, and expertise in Kubernetes, AWS, Python, Terraform, networking, and CI/CD.
AWS GovCloud, AWS C2E, Kubernetes, AWS EKS, Rancher, Grafana LGTM, Datadog, Python, Terraform, VPN, NLB, ALB, HTTPS, TLS, VPC peering, CDN, GitLab Workflows, Unix, Linux, Auth0, Keycloak, AWS IAM, Git