About the Role
About Flipkart
Flipkart is committed to the cause of transforming commerce in India through our investments in made-in-India technology innovations, customer-centric features and constructs, a diverse category landscape and a world-class supply chain. With a customer base of over 350 million, product coverage of over 150 million across 80+ categories, focus on generating direct and indirect employment and a commitment to empowering generations of entrepreneurs and MSMEs and a sustainable growth strategy – Flipkart is maximizing for our customers, stakeholders, and the planet at large! Flipkart is a part of the Walmart-owned Flipkart Group, which also includes group companies Flipkart Health+, Myntra, and Cleartrip. The Group is also a majority shareholder in PhonePe, one of the leading Payments Apps in India.
About the Team:
You will join the Core Search & AI Platform team responsible for driving intelligent discovery across high-scale surfaces like Search, Minutes, Shopsy, and Cleartrip. We build low-latency, high-throughput model serving infrastructure and deep learning pipelines that execute sub-second inference at massive scale. Our focus is on systems engineering, memory optimization, and hardware acceleration (GPU/NPU)—pushing the boundaries of C++, CUDA, and specialized inference engines to power real-time AI across the entire ecosystem.
About the role:
We are seeking a Senior Systems / Platform Engineer (SDE 3) to build and optimize our high-throughput, low-latency AI Model Inference Infrastructure. In this role, you will focus on core model execution, serving platforms, and hardware acceleration to deliver sub-second response times across high-scale search and platform surfaces.
What You’ll Work On:
- Architecting and tuning high-performance LLM model serving engines for sub-second latency under high QPS.
- Optimizing memory layout, KV-cache execution, continuous batching, and model quantization (FP8, AWQ).
- Embedding AI model inference directly into low-latency vector and hybrid search platforms.
What We Are Looking For:
Must-Have: Strong hands-on C++ or high-performance Java/Rust engineering experience with deep system design fundamentals.
Must-Have: Direct experience with inference servers and frameworks such as vLLM, Triton Inference Server, TensorRT, ONNX Runtime, or CUDA.
Nice-to-Have: Background in embedding models into vector search engines (FAISS, Milvus, ScaNN) or custom NPU/GPU kernel optimization.
5-8 years experience is required.
Requirements
C++ or Java/Rust experience
Strong hands-on experience in C++ or high-performance Java/Rust engineering with deep system design fundamentals is essential.
Inference server experience
Direct experience with inference servers and frameworks such as vLLM, Triton Inference Server, TensorRT, ONNX Runtime, or CUDA is required.
Nice to Have
Background in embedding models into vector search engines like FAISS, Milvus, or ScaNN is a plus.
Experience with custom NPU/GPU kernel optimization would be beneficial.