Tech Stack
GoPythonLinux
Job Description, Responsibilities & Requirements
About the Position
Senior Software Engineer - Fleet
Location
- San Francisco, California, United States
- San Jose, California, United States
- Bellevue, Washington, United States
Employment Type
- Full Time
Department
- Data Center Business
Compensation
- San Francisco / San Jose: $225K – $346K
- Bellevue: $203K – $311K
About Lambda
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
What You’ll Do
- Develop and Maintain Production Systems: Design, implement, and improve software that powers GPU fleet lifecycle management and machine configuration at scale.
- Automate Infrastructure: Build and enhance automation frameworks for machine provisioning, configuration management, and deployment.
- Support New Hardware Introduction (NPI): Enable bring-up, validation, and production readiness for new server and accelerator platforms.
- Enhance Machine Lifecycle Processes: Improve and refine workflows for bare metal provisioning, firmware updates, and system health monitoring.
- Support DPU Lifecycle Automation: Help provision, configure, update, and debug DPUs and programmable network accelerators.
- Debug Hardware and Firmware Issues: Investigate failures across BIOS, BMC, firmware, networking, storage, and boot flows.
- Collaborate Across Teams: Work closely with infrastructure, security, and product engineering teams to develop scalable and maintainable solutions.
Requirements
- 6+ years of experience working with Go (Golang) or Python in production environments.
- 6+ years of experience with bare metal hardware management and configuration.
- Comfortable working in Linux environments and debugging issues at the OS, hardware, and networking layers.
- Can independently troubleshoot complex systems and communicate effectively across software, infrastructure, and vendor teams.
Nice to Have
- Experience with Go in infrastructure, systems, or backend development.
- Hands-on experience with bare metal provisioning and lifecycle management, including technologies such as Redfish, BMC, IPMI, DHCP, and PXE.
- Experience with DPUs, NVIDIA BlueField, SuperNICs, or other programmable network accelerators.
- Experience diagnosing issues involving drivers, firmware, and hardware compatibility across GPU servers.
- Experience incorporating AI-assisted development tools into engineering workflows, including code generation, debugging, test development, and documentation.
- Experience building Linux distributions or managing OS customization and imaging.
- Exposure to DPU-adjacent networking, storage, or security technologies such as OVS/OVN, BGP/FRR, SR-IOV/VFs, DOCA, NVMe emulation, telemetry agents, or certificate-based service authentication.
- Exposure to Kubernetes and container orchestration concepts.
We Offer
- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Compensation Range: $203K - $346K