About the Role
About the Role
The Senior Inference Systems Engineer will focus on building advanced infrastructure to enhance LLM inference performance through various optimization techniques. This role involves hands-on work with CUDA, GPU architecture, and Rust systems programming to manage KV cache placement and improve overall system efficiency.
Responsibilities
- Design and implement KV cache offloading, streaming, and memory management infrastructure for large-scale LLM serving.
- Build cache-aware scheduling systems that determine when to keep, evict, prefetch, stream, compress, decompress, or recompute KV cache blocks.
- Optimize inference runtimes such as vLLM and SGLang, including paged attention, prefix caching, schedulers, and cache management systems.
- Develop mechanisms that overlap IO operations with attention execution to maximize GPU utilization and minimize latency.
- Build high-performance components in Rust, C++, and CUDA for scheduling, cache coordination, telemetry, and inference optimization.
- Profile and eliminate bottlenecks across GPU, CPU, memory, networking, storage, and runtime layers.
- Design benchmark frameworks and performance tests for long-context, streaming, multi-turn, and high-concurrency workloads.
- Measure and improve key inference metrics including TTFT, TBT/ITL, GPU utilization, cache hit rates, and cost per token.
- Collaborate closely with Product, Platform, ML, and Engineering teams to deliver production-ready optimization capabilities.
Desired Skills and Experience
- Strong hands-on experience with CUDA programming and GPU performance optimization.
- Deep understanding of transformer inference, attention mechanisms, KV cache architecture, batching, streaming generation, prefill, and decode.
- Experience with vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar LLM serving frameworks.
- Experience designing or optimizing KV cache systems, including cache reuse, eviction, prefix caching, radix caching, or cache offloading.
- Strong systems programming skills in Rust, C++, or both.
- Strong Python skills for experimentation, benchmarking, and performance analysis.
- Experience building performance-sensitive schedulers, async IO systems, or distributed infrastructure.
- Strong debugging and profiling skills using tools such as Nsight, CUDA profiling tools, or custom telemetry systems.
- Experience with GPUDirect, RDMA, NVMe, cache compression, FlashAttention, paged attention, or distributed inference architectures is a strong advantage.
- Bachelor’s or Master’s degree in Computer Science, Software Engineering, Electrical Engineering, or a related field.
Requirements
CUDA programming
Strong hands-on experience with CUDA programming and GPU performance optimization.
Rust systems programming
Strong systems programming skills in Rust, C++, or both.
Transformer inference
Deep understanding of transformer inference, attention mechanisms, and KV cache architecture.
Performance optimization
Experience designing or optimizing KV cache systems and performance-sensitive schedulers.
Nice to Have
Strong Python skills for experimentation, benchmarking, and performance analysis.
Experience building async IO systems or distributed infrastructure.