Lambda
Posted 10h ago

Staff Software Engineer - Compute

Lambda
Bellevue or San Jose or San Francisco
$314k-$465k/yrHybridFull Time
Responsibilities
  • designing control planes
  • leading initiatives
  • mentoring engineers
Requirements
  • Requires 10+ years building compute control-plane distributed systems
  • Expertise in durable execution and cloud provisioning
  • Semiconductor enablement
  • Data-center deployment, and proficiency in C/C++, Rust
  • Python, or Go
Technical tools mentioned
UEFILinuxCC++RustPythonGoKVMQEMUSR-IOVDPDKSPDKKubernetesInfiniBandRoCENVMe-oFNVIDIA AI FactoryConnectXBlueField DPUDOCADOCA SNAPCUDA

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our Bellevue, San Francisco, or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

 

About the Role

As a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.

The position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.

What You'll Do


We are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:

  • Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.

  • Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.

  • Guide design of compute platform multi-tenant security model

  • Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.

  • Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.

  • Work with customers on translating vague customer technical requirements into concrete engineering deliverables.

  • Set engineering standards and lead design reviews for mission-critical cloud software at scale.

Who You are

  • 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.

  • Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.

  • Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.

  • Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.

  • Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).

  • Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.

Nice to Have

  • Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).

  • Knowledge of Nvidia’s AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)

  • Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).

  • Experience with Cloud Service Provider Kubernetes offerings.

  • Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

About Lambda

Provides high-performance GPU cloud infrastructure for AI development.

Year founded
2012
Employees
734
Organization type
Private
Latest investment
Raised $1.50B Series E (2025) — led by TWG Global, US Innovative Technology Fund
Headquarters
US

Similar jobs

Software Engineer roles near Bellevue, Washington
5h
Save
Mark Applied
Hide
Senior Principal Software Engineer IS - Cloud Engineering, Hybrid
Redmond or Alaska or California or Montana or New Mexico or Oregon or Texas or Washington
$58-$152/hr HybridFull Time
Providence
Providence: Non-profit health system providing comprehensive medical and social services.
12+ YOEBachelor's degree or equivalent experience; 12 years of related experience, including 5 years as a lead engineer; expertise in C#, Java, Python, Git, SQL/NoSQL, Agile, Azure DevOps, TFS, Jira, and cloud technologies.
C#, Java, Python, Git, SQL, NoSQL, Azure DevOps, TFS, Jira, Azure, AWS
6h
Save
Mark Applied
Hide
Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
San Diego or San Francisco or Pleasanton or Santa Clara or Kirkland
$150k-$262k/yr HybridFull Time
ServiceNow
ServiceNowNYSE: NOW: Enterprise cloud platform for digital workflow automation.
3+ YOERequires 8+ years software development with a bachelor's, 5+ with a master's, 3+ with a PhD, or equivalent; 5+ years Kubernetes; hyperscaler, Go, containers, CI/CD, GitOps, and infrastructure-as-code experience.
Kubernetes, AWS, Microsoft Azure, Google Cloud Platform (GCP), Go, CI/CD, GitOps, Git, infrastructure-as-code, CNI, service mesh, mTLS
6h
Save
Mark Applied
Hide
Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
San Diego or San Francisco or Pleasanton or Santa Clara or Kirkland or United States
$150k-$262k/yr HybridFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
8+ YOERequires 8+ years software development with a bachelor's, 5+ with a master's, 3+ with a PhD, or equivalent; 5+ years Kubernetes; hyperscaler experience; Go; containers; CI/CD; GitOps; infrastructure-as-code.
Kubernetes, AWS, Azure, GCP, containers, CI/CD, GitOps, infrastructure-as-code, Go, CNI, service mesh, mTLS, observability, metrics, tracing, dashboards, AI
6h
Save
Mark Applied
Hide
Software Engineer - Digital Site Development
Issaquah, Washington, United States
$85k-$225k/yr OnsiteFull Time
Costco Wholesale
Costco WholesaleNasdaq: COST: Operates a global chain of membership-only big-box warehouse clubs.
15+ YOE15+ years in enterprise eCommerce architecture and technical design; Java EE, Spring Boot, React, Node.js, cloud, Kubernetes, databases, CI/CD, Agile/Scrum, leadership, and mentoring experience required.
Java EE, Java, Spring Boot, Spring Framework, React, Node JS, AWS, Azure, GCP, GKE, Git, CI/CD, Spanner, DB2, SQL Server, IBM MQ, Biztalk, GCP PubSub, Kafka, Apigee, APIM, DataPower, Google Workspace, Google Sheets, Google Docs, Google Slides, Gmail, Dynatrace, Splunk, J2EE, DevSecOps, TDD
7h
Save
Mark Applied
Hide
Senior Software Engineer
Redmond, Washington, United States
$120k-$235k/yr OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
4+ YOEBachelor's degree in computer science or related technical field and 4+ years of coding experience; preferred qualifications include advanced degree or 6+ years of experience.
C, C++, C#, Java, JavaScript, Python
9h
Save
Mark Applied
Hide
Senior Principal Software Engineer IS - Cloud Engineering, Hybrid
Redmond or Alaska or California or Montana or Hobbs or Oregon or Levelland or Lubbock or Plainview or Washington
$58-$152/hr HybridFull Time
Providence
Providence: Provides comprehensive healthcare services through hospitals, clinics, and health plans.
12+ YOEBachelor's degree or equivalent, 12 years of related experience including 5 years as a lead engineer, object-oriented programming, SQL/NoSQL, Git, Agile, cloud technologies, and complex enterprise projects.
C#, Java, Python, Git, SQL, NoSQL, Azure DevOps, TFS, Jira, Azure, AWS
9h
Save
Mark Applied
Hide
Software Engineer, Fulfillment Core Services
Seattle, Washington, United States
$128k-$160k/yr HybridFull Time
Lyft
LyftNASDAQ: LYFT: Provides an on-demand ride-hailing and multimodal transportation platform.
2+ YOERequires 2+ years of software engineering experience, BS/MS or equivalent, object-oriented programming, distributed systems, cloud platforms, scripting, CI tools, relational and NoSQL databases, and English communication.
Python, Golang, AWS, GCP, Microsoft Azure, Jenkins, Buildkite, CircleCI, TeamCity
10h
Save
Mark Applied
Hide
Senior Software Engineer, Risk
San Francisco or Denver or New York City or Seattle or San Jose or Scottsdale
$185k-$205k/yr HybridFull Time
Gusto
Gusto: Cloud-based payroll and HR software for small businesses.
8+ YOERequires 8+ years of software development experience, production architecture expertise, end-to-end project ownership, technical leadership, reliable software development, and experience using AI tools.
Ruby, Rails, TypeScript, React