About the Role
About the Role
The Software Engineer, Systems ML role at Meta involves designing and optimizing machine learning infrastructure to support the company's products at scale. You will work on high-performance ML systems, collaborating with various teams to enhance AI infrastructure efficiency for billions of users.
What You'll Do
- Design, build, and optimize large-scale ML training and inference systems, including distributed computing frameworks and hardware-accelerated pipelines.
- Develop and maintain high-performance ML infrastructure components in C++ and Python, ensuring reliability, scalability, and low-latency execution.
- Identify and resolve performance bottlenecks across the ML stack using profiling, instrumentation, and benchmarking tools.
- Architect and evaluate trade-offs in ML system design, including memory bandwidth, compute utilization, and I/O throughput.
- Partner with research and product teams to translate ML model requirements into efficient infrastructure solutions.
- Define and track system-level metrics and service level objectives to maintain production reliability of ML serving systems.
- Lead technical design reviews and contribute to engineering standards for ML systems across the organization.
- Mentor other engineers on ML infrastructure best practices, debugging methodologies, and performance optimization techniques.
- Drive adoption of AI-augmented development workflows to expand engineering productivity and broaden the scope of deliverables.
- Contribute to staged rollout strategies using feature flagging and experimentation frameworks to safely deploy ML system changes.
Who You Are
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
- 6+ years of experience in software engineering with a focus on machine learning systems, AI infrastructure, or high-performance computing.
- Experience developing and optimizing ML training or inference pipelines using frameworks such as PyTorch, TensorFlow, or equivalent.
- Experience with distributed computing architectures and large-scale systems design for ML workloads.
- Experience programming in C++ and Python for performance-critical systems.
- Experience using profiling and performance analysis tools to identify and resolve bottlenecks in ML or compute-intensive systems.
Bonus Points
- Experience optimizing large-scale ranking and recommendation model inference on AI accelerator hardware.
- Experience with hardware-software co-design, including numerics optimization and SIMD or vectorization techniques.
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
- Experience with GPU programming using CUDA, ROCm, or equivalent hardware accelerator kernel development.
- Experience with ML compiler technologies such as MLIR, LLVM, TVM, XLA, or IREE.
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies.
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews).
About Meta
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology.
Requirements
Machine Learning Systems
6+ years of experience in software engineering with a focus on machine learning systems.
C++ and Python
Experience programming in C++ and Python for performance-critical systems.
Distributed Computing
Experience with distributed computing architectures and large-scale systems design for ML workloads.
Performance Optimization
Experience using profiling and performance analysis tools to identify and resolve bottlenecks.
Nice to Have
Experience with GPU programming using CUDA, ROCm, or equivalent.
Experience with ML compiler technologies such as MLIR, LLVM, TVM, XLA, or IREE.
Benefits
Health Insurance
Comprehensive health insurance plans for employees.
Remote Work
Flexible remote work options to support work-life balance.