Lambda
Posted 9h ago

Senior Site Reliability Engineer - Managed Kubernetes

Lambda
San Francisco or San Jose or Bellevue
$240k-$356k/yrHybridFull Time
Responsibilities
  • operating clusters
  • automating platforms
  • responding incidents
Requirements
  • Requires 6+ years in SRE or operations
  • Deep Linux and production Kubernetes expertise
  • Strong Go and Python skills
  • GitOps, Helm
  • Observability, CI/CD, and Kubernetes provisioning experience
Technical tools mentioned
KubernetesPythonGolangGitOpsArgoCDHelmLinuxEKSGKEPrometheusGrafanaFluentBitCI/CDkubeadmCluster APICRDsCSICNI

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.


Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.


What You’ll Do

  • Operate and maintain bare-metal Kubernetes clusters, scaling up to thousands of nodes

  • Handle cluster degradation, recovery, resizing, and incident response using fleet management tools

  • Participate in a well-managed on-call rotation for critical incidents

  • Assist customers with Kubernetes questions, workload integration, storage, and authentication

  • Work closely with our HPC Ops and Datacenter Ops teams for low-level or cross-functional issues

  • Use Python and Golang to create tooling and automate the validation of platform quality.

  • Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes

  • Develop automation for cluster lifecycle management: provisioning, upgrades, patching, and deletion.

  • Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability.

You

  • 6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge of running Linux clusters and systems

  • Strong programming skills in Go and Python; experience with GitOps (e.g., ArgoCD), Helm, and Kubernetes operators

  • Proven experience operating Kubernetes clusters in production environments (on-prem, EKS, GKE, or similar)

  • Can work either independently with limited direction or as part of a team

  • Can work with customers during incidents either via tickets, live messaging, or as part of a larger call.

  • Familiarity with observability tools like Prometheus, Grafana, FluentBit, and CI/CD pipelines

  • Proven experience provisioning Kubernetes using tools such as kubeadm, Cluster API, or similar

Nice To Have

  • Deep Kubernetes expertise: CRDs, CSI, CNI, Kubernetes Operator Coding experience

  • Exposure to HPC clusters, AI/ML workloads, or large-scale GPU clusters

  • Hybrid or multi-cloud Kubernetes environment experience

  • Contributions to CNCF projects or Kubernetes SIGs

Why Join Us

  • Work on cutting-edge Managed Kubernetes platforms for AI/ML workloads

  • Influence the platform roadmap and help shape operations and reliability best practices

  • Collaborate with a highly skilled engineer

  • Opportunity to mentor and grow within a fast-growing, technology-driven environment

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

About Lambda

Provides high-performance GPU cloud infrastructure for AI development.

Year founded
2012
Employees
734
Organization type
Private
Latest investment
Raised $1.50B Series E (2025) — led by TWG Global, US Innovative Technology Fund
Headquarters
US

Similar jobs

Site Reliability Engineer roles near San Francisco, California
6h
Save
Mark Applied
Hide
Site Reliability Engineer, US Gov
Denver or Arvada or San Francisco or Nashville or Santa Fe or New Orleans or San Diego or Bozeman or United States
$160k-$200k/yr HybridFull Time
Quindar
Quindar: Cloud-native software for automated satellite mission operations.
3+ YOERequires a bachelor's degree, 3+ years of SRE or infrastructure experience, U.S. citizenship, Secret clearance or higher, and expertise in Kubernetes, AWS, Python, Terraform, networking, and CI/CD.
AWS GovCloud, AWS C2E, Kubernetes, AWS EKS, Rancher, Grafana LGTM, Datadog, Python, Terraform, VPN, NLB, ALB, HTTPS, TLS, VPC peering, CDN, GitLab Workflows, Unix, Linux, Auth0, Keycloak, AWS IAM, Git
1d
Save
Mark Applied
Hide
Site Reliability Engineer - rednote
Palo Alto, California, United States
OnsiteFull Time
Rednote
Rednote: A lifestyle-focused social media and e-commerce discovery platform.
Experience with large-scale reliability, high-availability architecture, incident response, cross-region disaster recovery, Linux, networking, middleware, cloud-native infrastructure, automation, and Python, Go, or Java.
Linux, MySQL, Redis, Kafka, Kubernetes, Service Mesh, Python, Go, Java
1d
Save
Mark Applied
Hide
Site Reliability Engineer, Enterprise Technology Services
Sunnyvale, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
The posting describes an enterprise technology platform role but does not state specific education, certification, skills, or experience requirements.
1d
Save
Mark Applied
Hide
K8 Site Reliability SME
San Jose or Austin
RemoteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
5+ YOERequires 5+ years of Kubernetes operations, 2+ years managing GPU workloads, Terraform, Helm, GitOps, SRE practices, monitoring, Go or Python, and multi-tenant platform experience.
Kubernetes, Nvidia GPU operator, Terraform, Helm, ArgoCD, Flux, Prometheus, Grafana, Alertmanager, PagerDuty, Go, Python, Slurm, Ray, Kubeflow, Ironic, MAAS, GitOps
1d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
San Mateo or United States
$240k-$300k/yr RemoteFull Time
Skydio
Skydio: Develops autonomous AI drones for defense and industrial inspection.
8+ YOE8+ years in SRE, platform, DevOps, production engineering, or equivalent; strong Kubernetes and AWS experience; Terraform, CI/CD, networking, observability, and production reliability expertise.
Kubernetes, Amazon Web Services (AWS), Amazon Elastic Kubernetes Service (EKS), Terraform, Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, Jenkins, Linux, Python, Go, Helm, GitOps, Datadog, PostgreSQL
2d
Save
Mark Applied
Hide
Site Reliability Engineer (Multiple Positions)
San Jose, California, United States
$213k-$388k/yr OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOERequires a master's degree and 2 years of related experience, or a bachelor's degree and 5 years of progressive experience. Requires cloud systems, Linux, Docker, Kubernetes, software lifecycle, observability, and reliability engineering expertise.
Linux, Docker, Kubernetes
2d
Save
Mark Applied
Hide
Site Reliability Engineer (Multiple Positions)
San Jose, California, United States
$226k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
1+ YOEMaster's degree and 1 year of related experience, or bachelor's degree and 3 years; requires 1 year supporting critical systems, monitoring, troubleshooting, data operations, error analysis, and runbook creation.
2d
Save
Mark Applied
Hide
Senior Site Reliability Engineer - CLO
San Francisco or United States
$153k-$205k/yr RemoteFull Time
Circle
CircleNYSE: CRCL: Digital currency issuer and blockchain financial infrastructure provider.
3+ YOERequires 3–6 years of software development experience, cloud-native AWS or GCP and Kubernetes experience, distributed systems expertise, and a bachelor's degree or equivalent practical experience.
Golang, Java, JavaScript, TypeScript, Python, Rust, AWS, GCP, Kubernetes, RESTful API, SQL