Lambda
Posted 2w ago

Software Engineer - Compute

Lambda
San Francisco or San Jose or Bellevue
$266k-$395k/yrHybridFull Time
Responsibilities
  • designing software
  • maintaining infrastructure
  • troubleshooting systems
Requirements
  • 3+ years Go or Python experience
  • 3+ years bare-metal and virtualization hardware management
  • Linux proficiency
  • Production troubleshooting
  • On-call incident ownership
  • Strong communication
Technical tools mentioned
Go (Golang)PythonLinuxSlurmKubernetesKVMQEMUTemporal

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

 

The Compute Software Engineering role plays a key part in the foundation of both the Public and Private Cloud offerings by enabling Bare-Metal and Virtual Machine provisioning. In this role you will: Develop and integrate workflows for compute instance lifecycle, system health, host and cluster validation, and contribute to and evolve the internal platform tooling, and workflows. As part of the Compute Engineering team, this position focuses on improving the bare-metal and VM developer and customer experiences, streamlining the software delivery lifecycle, building, testing, and shipping software efficiently and reliably.

The position works closely with various engineering teams across the company and contributes directly to Lambda’s ability to scale engineering productivity, maintain high delivery velocity, and support the company’s continued growth by making the right engineering workflows easier, faster, and more reliable.

 

What You’ll Do
We are seeking a Software Engineer with good experience in backbone infrastructure that build, scale and optimize GPU-first cloud infrastructure that enables High Performance Computing and demanding AI workloads. In this role you will be responsible for:

  • Design, develop and maintain software for GPU/CPU compute infrastructure with focus on performance, scalability, and reliability.

  • Implement and develop services for baremetal and VM instancing.

  • Develop distributed systems for managing and orchestrating compute resources across various SKU’s.

  • Troubleshoot and debug complex issues in a production and development environment.

  • On-call and incident ownership

  • Collaboration across multiple teams and drive ambiguity in requirements or solutions on RFC’s.

You

  • 3+ years of experience working with Go (Golang) or Python in production environments.

  • 3+ years of experience with bare metal & virtualization hardware management and configuration.

  • Are comfortable working in Linux environments and debugging issues at the OS, hardware, and networking layers.

  • Can independently troubleshoot complex systems and communicate effectively across software, infrastructure, and vendor teams.

Nice to Have

  • Familiarity with GPU Infrastructure or high-performance computing environments.

  • Experience with Slurm or Kubernetes-based cluster management.

  • Experience with core public cloud internals (Virtualization, KVM, QEMU, Security and Fleet health)

  • Experience with durable execution platforms like Temporal.

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

About Lambda

Provides high-performance GPU cloud infrastructure for AI development.

Year founded
2012
Employees
734
Organization type
Private
Latest investment
Raised $1.50B Series E (2025) — led by TWG Global, US Innovative Technology Fund
Headquarters
US

Similar jobs

Software Engineer roles near San Francisco, California
1h
Save
Mark Applied
Hide
Senior Software Engineer, Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build distributed, large-scale data processing systems, frameworks, and platforms using big data technologies while collaborating with Apple TV and Video teams.
2h
Save
Mark Applied
Hide
Senior Software Engineer, Hyperscale Build Environments and Tools
Santa Clara, California, United States
$180k-$270k/yr OnsiteFull Time
Pure Storage
Pure StorageNYSE: PSTG: Provides all-flash enterprise data storage and management solutions.
5+ YOE5+ years of software engineering experience in infrastructure, developer productivity, build systems, or platform engineering, with C/C++, Linux, Docker, CI, and modern build-system expertise.
C, C++, Make, CMake, Linux, Docker
2h
Save
Mark Applied
Hide
Software Engineer, Linux Kernel / Android
San Jose or Los Angeles or Bellevue
$180k-$240k/yr OnsiteFull Time
Rivet Industries
Rivet Industries: Building integrated task systems for frontline industrial and defense operators.
3+ YOERequires 3+ years developing Android/Linux system software, strong Linux kernel and C/C++ skills, AOSP, HALs, drivers, hardware bring-up, embedded builds, debugging, security mechanisms, and U.S. Person status.
Android, Linux, Linux kernel, AOSP, HALs, C, C++, USB, MIPI, I2C, UART, GPIO, PCIe, U-Boot, Android Bootloader, Yocto, Buildroot, Bazel, Soong, OTA, TPM, AR/XR, NPU, DSP
2h
Save
Mark Applied
Hide
Principal Staff Software Engineer, Systems Infrastructure
Mountain View, California, United States
$226k-$369k/yr HybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
10+ YOEBA/BS or equivalent practical experience, 10+ years in software or reliability engineering, 5+ years in technical leadership, distributed systems expertise, and experience defining reliability standards across teams.
Java, Go, C++, Python, LLM, SLO, SLI
4h
Save
Mark Applied
Hide
Exceptional Software Engineer
Redwood City, California, United States
$180k-$400k/yr OnsiteFull Time
Dyna Robotics
Dyna Robotics: Develops general-purpose robots powered by proprietary embodied AI models.
Exceptional software engineering ability, high tolerance for ambiguity, resilience in changing environments, and clear communication. Robotics experience is welcome but not required.
5h
Save
Mark Applied
Hide
Staff+ Software Engineer, Product Sandboxing
San Francisco or New York City or Seattle or California
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
8+ YOERequires 8+ years building scalable distributed systems, strong service-oriented architecture, networking, and systems design expertise, proficiency in Python, Go, or Rust, and experience with cloud infrastructure and Kubernetes.
Python, Go, Rust, GCP, AWS, Azure, Kubernetes
6h
Save
Mark Applied
Hide
Staff Software Engineer, Middle Office
San Francisco, California, United States
$240k-$300k/yr OnsiteFull Time
Carta
Carta: Software platform for cap table management and fund administration.
10+ YOEExpertise in distributed systems and 10+ years of professional software engineering experience recommended, with high-level technical leadership and experience guiding architecture across Python/Django, React, Postgres, JVM languages, gRPC, and AWS.
Python, Django, React, Postgres, gRPC, AWS
7h
Save
Mark Applied
Hide
Staff Software Engineer, MetalDev
New York City or Sunnyvale
$207k-$275k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOE8+ years in software engineering focused on infrastructure, cloud engineering, and distributed systems; expert Go, REST/gRPC APIs, Kubernetes, observability, CI/CD, GPU fleets, technical leadership, and incident response.
Go, REST, gRPC, Kubernetes, Prometheus, Grafana, PromQL, CI/CD, Kafka, ClickHouse, CRDB, DMTF, RedFish