About the Role
About AZX
Our mission is to accelerate positive impact in critical industries through AI transformation. We specialize in physics-informed ML and enterprise AI solutions that directly address climate and sustainability challenges.
We’re growing quickly and already work with category-leaders in real estate (CBRE), energy (LevelTen Energy), logistics (Flexe) and utilities.
We’re a public benefit corporation, founded in 2024, and have been profitable from inception.
We work on challenges in clean energy, decarbonization, climate risk, energy systems, and global economics. We’re building our company for long-term success and aim to create the ultimate place to work for those passionate about AI and making a positive impact.
About This Role
You will be responsible for owning the layer where AI work physically happens: the machines, the isolation boundary, and the models running on them. This role anchors on two systems. The first is our inference control plane — open-weight models and custom task-model zoos, hosted and operated across managed GPU clouds and customer-managed Kubernetes clusters, with scale-to-zero economics, cold-start discipline, and per-token cost accounting that stays correct even when a client disconnects mid-stream — along with the Kubernetes layer those workloads live on: operators, autoscaling, node lifecycle.
The second is our agent-sandboxing platform: hardware-isolated microVMs for running untrusted, agent-generated code securely and compliantly by construction, where agents operate with least privilege, never see a credential, and a human gates anything that writes to a system of record. You'll write Rust in the morning, a Kubernetes controller after lunch, and a FastAPI control-plane endpoint before you go home — building the fork engine, the guest agent, and the multi-substrate model lifecycle.
Responsibilities
- Manage the serving tier for open-weight models: engine deployment and configuration, cold-start strategy, per-model SLOs, and upgrade/canary discipline.
- Administer the Kubernetes layer for inference and sandbox workloads: operators and CRDs, autoscaling (KEDA/Karpenter-class), GPU scheduling and sharing, and node lifecycle.
- Own the stateful data plane: hosted vector stores for semantic memory and graph stores for knowledge graphs — deployed, backed up, scaled, and recovered, with restores that are tested rather than hoped for.
- Oversee the sandbox runtime and its host-side control plane: lifecycle, exec, snapshot/fork, teardown, metering, and the threat model of the isolation boundary.
- Direct the FastAPI control-plane services, Terraform/OpenTofu, Bicep, and the dashboards.
- Run the layer the backend services team builds on, expect to debug into their services, and expect them to read your dashboards.
- Manage the open-source posture: build to OSS standards and release as it matures, with reviewed PRs, real docs, and reproducible builds.
Core Qualifications
- 5+ years of shipping production systems in a systems language. Rust is the house language, but polyglots are welcome — deep Go, C/C++, or Zig with genuine appetite for Rust counts.
- Operated Kubernetes workloads that other people depended on — controllers or operators, scheduling, autoscaling, node lifecycle. You've been paged, and the experience changed how you build.
- Strong ability to threat-model isolation boundaries (namespaces, cgroups, seccomp, hypervisors), including identifying what an untrusted guest could observe, forge, or exhaust, and applying security best practices for agentic execution — least privilege, no credentials in the sandbox, audit trails, and human approval on write actions.
- Hands-on experience deploying or operating open-weight models.
Requirements
Rust programming
5+ years of experience in shipping production systems using Rust.
Kubernetes expertise
Experience managing Kubernetes workloads, including operators and autoscaling.
Systems programming
Strong background in systems programming languages like Go, C/C++, or Zig.
Security best practices
Ability to threat-model isolation boundaries and apply security best practices.
Nice to Have
Experience in building and releasing open-source software.
Hands-on experience with deploying AI models in production.
Benefits
Impactful work
Opportunity to work on AI solutions addressing climate and sustainability challenges.
Growth opportunities
A chance to grow in a rapidly expanding company.