About the Role
About the Role
The AI Inference Engineer will be responsible for building and running the inference engine that powers Perplexity queries, deploying various model architectures while managing latency and cost. This role involves supporting new models, migrating GPU kernels, developing a Rust-native serving runtime, and optimizing performance and reliability.
What You'll Do
- Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure.
- Port our in-house CUDA kernels to NVIDIA's CuTe DSL.
- Develop our internal Rust-based inference server.
- Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving.
- Build dashboards, alerts, and automated remediation for reliability and observability.
Who We're Looking For
- Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar).
- Understanding of modern LLM architectures and production environments.
- Experience with production distributed systems under real load.
- Comfortable working across languages: Rust, Python, CUDA/CuteDSL.
- Self-directed and able to work in fast-moving environments.
Nice-to-Have
- Experience with ML compilers and framework internals.
- Knowledge of distributed GPU communication.
- Familiarity with low-precision inference techniques.
- Experience with profiling and debugging tools.
- Knowledge of container orchestration.
Qualifications
3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework and understanding of GPU architectures and common LLM architectures.
Compensation
Final offer amounts are determined by multiple factors including experience and expertise.
Benefits
- Equity may be part of the total compensation package.
Requirements
GPU programming
Deep experience with CUDA, Triton, or similar technologies.
LLM architectures
Understanding of modern LLM architectures and their production deployment.
Distributed systems
Experience building and operating production distributed systems.
Multi-language proficiency
Comfortable working with Rust, Python, and CUDA/CuteDSL.
Self-directed
Ability to work independently in fast-paced environments.
Nice to Have
Experience with PyTorch internals and custom operators.
Knowledge of NCCL, NVLink, and model parallelism.
Familiarity with quantization techniques.
Experience with Nsight Compute/Systems and CUDA-GDB.
Knowledge of Kubernetes and GPU scheduling.
Benefits
Equity
Equity may be part of the total compensation package.