About the Role
About the Role
The Software Engineer for GenAI Frameworks at Meta will focus on optimizing generative AI inference on the MTIA platform. This role involves leading technical programs, designing and implementing features, and ensuring high performance and quality for production workloads.
What You'll Do
- Serve as technical owner and domain expert for a key GenAI inference framework area: serving runtime, distributed inference, graph-mode execution and compilation, or core PyTorch integration.
- Lead ambiguous, multi-quarter technical programs end to end: technical design, execution, test strategy and CI, rollout, and production hardening across teams and org boundaries.
- Design, implement, and ship features in generative AI inference frameworks and the MTIA integration layers beneath them, from prototype through production deployment.
- Own performance for production inference workloads end to end: profile across frameworks, runtime, compiler, and kernel boundaries, identify where time actually goes, and drive the fix to the correct layer rather than the convenient one.
- Build and extend the distributed inference substrate: hierarchical KV caching, cross-host transfer paths, disaggregated prefetch and decode, and parallelism strategies for long-context and mixture-of-experts serving.
- Optimize the runtime and execution path: graph capture and replay, host-side latency, memory allocation and placement, and throughput under real traffic conditions.
- Enable frontier models on MTIA and validate accuracy against GPU baselines, closing correctness gaps and distinguishing real regressions from stale references.
- Own the quality bar for your area, ensuring accuracy, stability, latency, and throughput, and test coverage with regression detection that keeps a win from quietly eroding.
- Partner with kernel, compiler, and silicon teams on hardware/software co-design, producing framework-level evidence that shapes the future silicon while the design can still change.
- Provide technical leadership: mentor other engineers, raise the design quality through reviews, and communicate decisions clearly through design documents and cross-team reviews.
Who You Are
Minimum Qualifications:
- 5+ years of hands-on experience with generative AI inference or training frameworks, or production model-serving systems.
- Proficiency in Python, C++, or Rust, including low-level systems code and performance-critical paths.
- Experience with GenAI inference optimization: prefill and decode optimization, KV-cache management and compression, batching and scheduling, and reasoning about latency and throughput tradeoffs.
- Experience with runtime-level optimization: graph-mode execution, host-side latency reduction, memory allocation and placement.
Requirements
Generative AI Experience
5+ years of hands-on experience with generative AI inference or training frameworks.
Programming Proficiency
Proficiency in Python, C++, or Rust, including low-level systems code.
Inference Optimization
Experience with GenAI inference optimization techniques.
Runtime Optimization
Experience with runtime-level optimization and graph-mode execution.
Nice to Have
Experience in mentoring other engineers and raising design quality.
Ability to communicate decisions clearly through design documents.
Benefits
Health Insurance
Comprehensive health insurance plans.
Remote Work
Flexible remote work options.
Learning Budget
Budget for professional development and learning.