About the Role
About Volta
Volta is the category-defining, fully vertically integrated AI infrastructure platform – from capital to clusters to software, under a founder-led enterprise. Our mission is The Utility of Compute™: AI infrastructure as dependable and available as electricity, for every organization that needs it. Launched with a $10B strategic partnership with one of the leading frontier AI labs, a Series A led by Andreessen Horowitz, and a $5B AI Infrastructure Fund, Volta is building the infrastructure layer of the AI era from the ground up. We are 100+ people across London, Palo Alto, and New York, with rapid growth expectations to hundreds.
About The Role
Volta builds and operates large scale GPU compute infrastructure for AI workloads. Our platform is Kubernetes-native, spans multiple regions, and delivers virtual machines, storage, and networking through a fully automated infrastructure stack built on custom Kubernetes operators.
We are building out several platform engineering teams that together own the full stack, from managed bare metal and IaaS through to higher-order platform services. Each team owns a different part of that stack: compute, networking, storage, the control plane and API layer, confidential computing, and the customer-facing surface. The area you work in depends on the team you join and your prior expertise, so no single engineer is expected to cover all of it.
Platform Engineers work at the intersection of infrastructure and software development. Across every team, you will translate three key inputs into durable platform capabilities: product roadmap requirements from the product team, operational learnings from the bring-up teams, and security guidance from the security engineering team. The output of this role is production platform code, not configuration, not runbooks.
What You Will Be Doing
Common across every platform engineering team:
- Design and implement Kubernetes operators and controllers that manage the lifecycle of platform resources.
- Work closely with the product team to turn roadmap requirements into the platform capabilities that support them.
- Collaborate with the bring-up teams to identify operational pain points and turn them into scalable platform features.
- Integrate security guidance from the security engineering team into platform-level controls, and remediate findings at the platform layer.
- Treat observability as a platform concern: instrument services, define meaningful metrics, and build tooling that gives the team visibility into platform health.
- Own the services you build in production, including participation in an on-call rotation, incident response, and the follow-up work that closes structural gaps rather than only the immediate issue.
- Hold to clean interface and versioning practice on anything other teams or customers depend on, including disciplined handling of breaking changes.
- Participate in code review, technical design discussions, and cross-team collaboration in an Agile (Kanban or Scrum) environment.
Depending on your team and background, you will go deep in some of the following:
- Control plane and APIs: improve and extend the API layer between user-facing services and the underlying platform, with disciplined versioning and backward compatibility.
- Compute: build and operate the lifecycle of virtualized and bare metal compute resources.
- Storage: provisioning workflows, attachment reliability, performance tuning, and failure handling.
- Confidential computing: build and extend confidential computing capabilities across the stack, from secure bare metal and confidential VMs to Confidential Containers (CoCo).
- Customer-facing services: the platform surfaces customers interact with directly, including the APIs and interfaces through which they consume capacity, working alongside product and UX.
What You Bring
3 to 5 years of software engineering experience, with a meaningful portion spent on infrastructure or platform systems.
Requirements
Kubernetes expertise
Experience in designing and implementing Kubernetes operators and controllers.
Software engineering experience
3 to 5 years of experience in software engineering, particularly in infrastructure or platform systems.
Collaboration skills
Ability to work closely with product and operational teams to translate requirements into platform capabilities.
Security integration
Experience in integrating security measures into platform-level controls.
Nice to Have
Familiarity with Agile practices such as Kanban or Scrum.
Experience in implementing observability tools and defining metrics for platform health.
Benefits
Equity
Opportunity to participate in the company's equity program.
Remote work
Flexible remote work options available.
Learning budget
Access to a budget for professional development and learning.