About the Role
About the Role
We are looking for a Senior Distributed Systems & AI Infrastructure Engineer to build and optimize the infrastructure powering next-generation AI inference at global scale.
You will work on high-performance distributed systems supporting large-scale model serving across thousands of GPUs and other accelerators. The role sits at the intersection of distributed systems, performance engineering, networking, and AI infrastructure, with a strong focus on building systems that are fast, scalable, and highly reliable.
Key Responsibilities
- Design and develop high-performance distributed systems for large-scale AI inference.
- Build scalable infrastructure for model serving across large GPU and accelerator clusters.
- Optimize system latency, throughput, resource utilization, and reliability.
- Design and improve high-performance networking and I/O components.
- Investigate and resolve complex issues across distributed systems and large-scale infrastructure.
- Develop performance-critical components using Rust, Go, or C++.
- Improve the scalability and reliability of infrastructure supporting next-generation AI workloads.
- Work closely with infrastructure, ML, and systems teams to identify and solve performance bottlenecks.
Requirements
- Bachelor's degree or equivalent experience in Computer Science, Engineering, or a related technical field.
- Strong systems programming experience in Rust, Go, or C++.
- Proven experience designing and building high-performance distributed systems at scale.
- Strong understanding of networking, network protocols, and high-performance I/O.
- Strong debugging and problem-solving skills for complex distributed systems.
- Experience optimizing systems for performance, scalability, and reliability.
Preferred Qualifications
- Experience with AI/ML serving infrastructure or large-scale inference systems.
- Familiarity with disaggregated inference architectures.
- Understanding of GPU programming models and GPU memory hierarchies.
- Experience with GPU networking and interconnect technologies such as NVLink, InfiniBand, or RoCE.
- Experience working with large-scale GPU clusters or accelerator infrastructure.
- Knowledge of performance optimization, profiling, and systems benchmarking.
- Experience supporting large-scale AI model training or inference workloads.
Requirements
Bachelor's degree
A degree in Computer Science, Engineering, or a related technical field is required.
Systems programming
Strong experience in Rust, Go, or C++ is essential.
Distributed systems design
Proven experience in designing and building high-performance distributed systems at scale is necessary.
Networking knowledge
A strong understanding of networking, network protocols, and high-performance I/O is required.
Debugging skills
Strong debugging and problem-solving skills for complex distributed systems are needed.
Performance optimization
Experience optimizing systems for performance, scalability, and reliability is important.
Nice to Have
Experience with AI/ML serving infrastructure or large-scale inference systems is preferred.
Familiarity with disaggregated inference architectures is a plus.
Understanding of GPU programming models and memory hierarchies is beneficial.
Experience with GPU networking technologies such as NVLink, InfiniBand, or RoCE is advantageous.
Experience working with large-scale GPU clusters or accelerator infrastructure is a bonus.