About the Role
About Mirantis
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.
Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen. Learn more at www.mirantis.com.
About the Team
Our engineering team builds next-generation, cloud-native bare-metal orchestration platforms for AI neo-cloud environments. We operate on a T-shaped engineering model: while each engineer brings deep expertise in a primary domain, every team member maintains practical fluency across neighboring systems. This shared foundation enables rigorous design reviews, effective cross-domain code reviews, and dependable on-call coverage. Our development is fundamentally based on open-source projects, and contributing to them is a major part of our work.
Role Overview
We are seeking a Senior Software Engineer who combines strong software architecture principles with deep technical domain expertise in Datacenter Networking and DPU architectures. Fitting into our T-shaped engineering model, you will drive the development of control planes powering high-density, high-throughput infrastructure for AI training and inference workloads. Familiarity with automated bare-metal provisioning and lifecycle management is a strong asset.
Core Responsibilities
- Control Plane Development: Design, build, and maintain production-grade control plane microservices, custom Kubernetes operators, and robust reconciliation engines in Rust and Go.
- Architecture & API Design: Author clear Functional (FR) and Non-Functional Requirements (NFR), architectural specs, system sequence diagrams, and clean gRPC/Protobuf and REST API schemas.
- Network Integration: Develop custom integration modules and integrations for modern network operating systems (SONiC, NVUE, Cumulus) and DPU hardware platforms.
- Observability & Diagnostics: Instrument services end-to-end using OpenTelemetry (OTel) traces and metrics, performing systematic root-cause analysis across polyglot distributed systems.
- Quality & Engineering Excellence: Participate in rigorous, review-gated pull request workflows across Rust, Go, and SQL codebases while ensuring high test coverage via mock-driven testing.
Qualifications
- Language Proficiency: Primary mastery of Rust (Tokio async runtime, Tonic, Axum, sqlx) and strong proficiency in Go.
- Data & APIs: Advanced SQL / PostgreSQL fluency, gRPC/Protobuf contract design, and schema evolution.
- Concurrency & Reliability: Strong background in async concurrency models, lock-free patterns, distributed state handling, and mock-driven testing discipline.
- Datacenter Protocols & OS: Expertise in BGP, MP-BGP, EVPN, VXLAN, L3VNI, route targets, and route server design. Hands-on experience with SONiC, Cumulus Linux, and NVUE (featuring a first-class NVUE client).
- DPU & Fabric Ecosystem: Deep knowledge of NVIDIA DOCA, Host-Based Networking (HBN), BlueField DPU architectures, and the DPF (DOCA Platform Framework) operator model (BFB, DPUSet, D).
Requirements
Rust proficiency
Primary mastery of Rust, including frameworks like Tokio and Axum.
Go programming
Strong proficiency in Go for developing microservices.
Datacenter networking
Expertise in protocols such as BGP, EVPN, and VXLAN.
SQL knowledge
Advanced fluency in SQL and PostgreSQL for data management.
Observability tools
Experience with OpenTelemetry for service instrumentation.
Nice to Have
Knowledge of NVIDIA DOCA and BlueField DPU architectures.
Experience in developing custom Kubernetes operators.
Benefits
Remote work
Flexible remote work options are available.
Open-source contributions
Opportunities to contribute to open-source projects.