About the Role
About the Role
The Principal Software Engineer, Systems at Meta will lead the development of AI systems aimed at enhancing production reliability. This role involves designing and implementing agentic systems that autonomously investigate and mitigate production incidents while ensuring safety and effectiveness.
Responsibilities
- Define the technical vision and architecture for agentic reliability systems across Meta.
- Personally design, code, and ship production agentic systems for complex infrastructure problems.
- Lead the development of an AI agent for production incident investigation and mitigation, advancing it toward accurate, trusted, and safely supervised autonomous action.
- Develop major improvements in agent reasoning, context, tool use, planning, evaluation, learning, and safe execution.
- Explore and apply techniques including automated hill climbing, fine-tuning, reinforcement learning, model routing, distillation, pruning, and inference optimization.
- Identify new high-value applications of AI across incident prevention, detection, mitigation, observability, and infrastructure operations.
- Move rapidly from ambiguous problems to prototypes, validate them against real production workloads, and develop successful approaches into reliable systems at scale.
- Build evaluation and experimentation systems that connect agent quality to outcomes such as investigation accuracy, successful mitigation, incident duration, and reduced operational work.
- Build closed-loop improvement systems that turn production outcomes into evaluations, experiments, and better agent behavior.
- Establish architectures and guardrails for production actions, including authorization, independent validation, auditability, rollback, and human oversight.
- Partner with Infrastructure, AI, Product, and Reliability leaders to integrate agentic capabilities into Meta’s production ecosystem.
- Influence technical strategy across organizations and mentor other engineers working on distributed systems and applied AI.
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
- 12+ years of software engineering experience, including experience building and operating large-scale distributed or infrastructure systems.
- Experience setting technical direction and leading complex, multi-year engineering efforts across organizational boundaries.
- Experience applying AI or machine learning systems to production problems.
- Demonstrated experience moving from technical concept to production deployment and measurable impact.
Requirements
Software Engineering Experience
12+ years of software engineering experience, particularly in large-scale distributed systems.
AI and Machine Learning
Experience applying AI or machine learning systems to production problems.
Technical Leadership
Experience setting technical direction and leading complex engineering efforts.
Production Deployment
Demonstrated experience moving from technical concept to production deployment.
Nice to Have
A degree in Computer Science, Computer Engineering, or a relevant technical field.
Experience building and operating large-scale infrastructure systems.