About the Role
About the Role
The Software Engineer will join the MTIA Software Tooling team at Meta, focusing on developing and maintaining tools for AI accelerator ASICs. This role involves designing and building developer tools to enhance debugging, profiling, and monitoring capabilities for AI workloads on MTIA hardware.
What You'll Do
- Own the technical vision and roadmap for key areas of MTIA's developer tooling ecosystem, focusing on debugging, workload error analysis, and product / fleet reliability.
- Design and develop debugging tools — including graph-mode debugging, kernel-level diagnostics, and multi-rank fault isolation.
- Contribute to overall MTIA SW tooling infrastructure — enabling profiling, performance debugging, memory sanitization, and reliability analysis for MTIA training and inference workloads.
- Collaborate closely with MTIA compiler, runtime, kernel, and hardware teams to integrate tooling hooks throughout the MTIA software stack.
- Drive AI-native tooling approaches by leveraging automation and LLM-guided diagnostics to improve developer productivity and reduce time-to-root-cause.
- Partner with internal product teams across advertising, recommendations, and generative AI to understand developer pain points and prioritize tooling investments.
- Advise on tooling best practices, debugging methodologies, and systems-level analysis for accelerator software.
- Communicate architectural decisions clearly through design documents and cross-team reviews.
Who You Are
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
- 6+ years of experience in systems software engineering, performance engineering, developer tooling, or a closely related field.
- Experience building debugging, profiling, or diagnostic tools for complex software/hardware systems.
- Proficiency in C++ and Python, including low-level systems programming and scripting for tool automation.
- Experience working across multiple layers of a system stack (compiler, runtime, OS/driver, hardware).
- Experience leading the technical design and delivery of tooling or infrastructure projects from inception through production deployment.
- Experience using data-driven methods and experimentation to evaluate and validate tooling effectiveness and systems performance improvements.
Bonus Points
- Familiarity with ML framework internals (PyTorch graph execution, torch.compile, operator dispatch) and AI compiler stacks (MLIR, LLVM, TVM, Triton).
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
- Experience with accelerator ecosystems (GPU/CUDA, TPU, custom ASICs) including performance profiling, memory analysis, and runtime debugging using their toolchains (cuda-gdb, nsight-compute, nsight-systems, cuda-memcheck).
- Demonstrated cross-stack debugging ability, including Linux kernel and driver-level debugging, with capacity to trace issues across application, OS, and hardware boundaries.
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies.
- Experience with distributed systems debugging or profiling (multi-device, multi-node/multi-rank).
- Experience with Linux debugging and profiling infrastructure (gdb, perf, eBPF).
Compensation
Salary and benefits information is not provided in the job description.
Requirements
Systems Software Engineering
6+ years of experience in systems software engineering, performance engineering, developer tooling, or a closely related field.
C++ and Python Proficiency
Proficiency in C++ and Python, including low-level systems programming and scripting for tool automation.
Debugging Tools Experience
Experience building debugging, profiling, or diagnostic tools for complex software/hardware systems.
Multi-layer System Stack
Experience working across multiple layers of a system stack (compiler, runtime, OS/driver, hardware).
Technical Design Leadership
Experience leading the technical design and delivery of tooling or infrastructure projects from inception through production deployment.
Nice to Have
Familiarity with ML framework internals and AI compiler stacks.
Demonstrated ability to integrate AI tools to optimize/redesign workflows.
Experience with accelerator ecosystems including performance profiling and memory analysis.
Demonstrated cross-stack debugging ability, including Linux kernel and driver-level debugging.
Experience with distributed systems debugging or profiling.