About the Role
About The Role
As a Staff Software Engineer for the Compute pillar at Lambda, you will define the technical vision for next-generation GPU and CPU host instance lifecycle and compute control plane. This role involves bridging high-level distributed systems with low-level semiconductor architecture to enable reliable cloud provisioning and lifecycle management at scale.
What You'll Do
We are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:
- Designing and implementing a highly available and reliable GPU and CPU “host and instance lifecycle” control plane.
- Guiding technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.
- Guiding design of compute platform multi-tenant security model.
- Providing technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.
- Collaborating with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.
- Working with customers on translating vague customer technical requirements into concrete engineering deliverables.
- Setting engineering standards and leading design reviews for mission-critical cloud software at scale.
Who You Are
You have 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale. You possess deep expertise in durable execution models and distributed systems used in cloud-service provisioning, along with a basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.
Nice to Have
Knowledge of Nvidia’s AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs, and switches) and software offerings (like DOCA, DOCA SNAP, CUDA, et al.) is a plus. Familiarity with Linux kernel internals, device drivers, virtualization technologies, and experience with Cloud Service Provider Kubernetes offerings would be beneficial.
Requirements
Cloud Infrastructure Experience
You should have extensive experience in building and optimizing GPU-first compute systems.
Distributed Systems Knowledge
A deep understanding of compute control plane distributed systems is essential.
Technical Leadership
You will provide mentorship and technical leadership to senior engineers.
Programming Proficiency
Proficiency in programming languages such as C/C++, Rust, Python, or Go is required.
Nice to Have
Familiarity with Nvidia’s AI Factory architectural components and software offerings is a plus.
Knowledge of Linux kernel internals and virtualization technologies would be beneficial.
Experience with Cloud Service Provider Kubernetes offerings is a nice addition.
Benefits
Work from Home Flexibility
Lambda offers a designated work from home day each week.
Office Locations
This position requires presence in Bellevue, San Francisco, or San Jose office locations.