About the Role
What You’ll Do
Deploy, instrument, and monitor open-weight models served via the inference engine.
Build new engine features to support novel hardware architectures.
Extend the engine for accuracy, explainability, and accountability in regulated settings.
Contribute upstream to the relevant open-source community.
Optimise inference performance (latency, throughput, hardware utilisation).
Who They’re Looking For
Senior-level Python + low-level programming (C/C++, Rust, CUDA).
Direct contributions to a major inference/serving engine or similar ML infra project (e.g. PyTorch, Ray).
Strong grasp of LLM inference internals (KV caching, batching, quantisation).
Experience upstreaming to active OSS communities.
Hardware accelerator optimisation experience (GPU/TPU/etc.).
Genuine enthusiasm for AI-assisted development (LLMs/agents as core tools).
Bonus Points
Regulated-environment experience, AI safety involvement, MLOps/CI/CD depth.
Requirements
Python programming
Senior-level expertise in Python is essential.
Low-level programming
Experience with C/C++, Rust, or CUDA is required.
ML infrastructure experience
Direct contributions to major inference or serving engines are necessary.
LLM inference knowledge
A strong grasp of LLM inference internals is important.
OSS community contributions
Experience in upstreaming to active open-source communities is needed.
Nice to Have
Experience working in regulated environments is a plus.
Involvement in AI safety initiatives is beneficial.
Depth in MLOps and CI/CD practices is advantageous.