About the Role
About the Role
Meta is seeking a software engineer to join our AI & Systems Co-Design team to drive the definition of our next-generation compute and storage architectures. As a key member of the team you'll work closely with internal software and platforms engineering teams to drive workload analysis and understand infrastructure requirements. You will drive technology path-finding, roadmap definition and co-design activities to deliver new capabilities and efficient systems for our fleet. Furthermore, you'll work with external industry partners to influence their roadmaps and build the best products for Meta’s Infrastructure. As part of the team, you will have ample opportunities to publish your work in leading conferences, industry journals and make open source contributions.
What You'll Do
- Utilize extensive understanding of hardware architecture - CPUs (x86/ARM), Flash/HDD, memory and storage systems, networking, and GPUs - to identify key platform resource bottlenecks.
- Collaborate closely with software product teams to re-architect services, improve performance through algorithm redesign, reduce resource consumption.
- Develop representative benchmarks (in C++/Rust/Python/Hack) to capture fleet requirements and drive early evaluation of upcoming platforms.
- Drive fleet-wide detailed workload analysis and keep ahead of evolving business needs and impact to compute and storage performance.
- Identify novel hardware/software co-design opportunities based on industry trends and new paradigms.
- Conduct pathfinding activities to quantify the value proposition for Meta and drive the roadmap definition.
- Influence vendor hardware roadmaps and the broader ecosystem to align with Meta's requirements.
- Partner with Product Engineering and Infrastructure Engineering teams to find the optimal way to deliver the hardware roadmap into production and drive adoption.
Who You Are
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
- 5+ years of experience with programming and scripting languages such as C, C++, Rust, Java, PHP, Python.
- Experience with large-scale infrastructure, distributed systems, full-stack analysis of server applications.
- 5+ years of experience with hardware architecture, compute technologies and/or storage systems.
Bonus Points
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews).
- Experience with developing, debugging and analysis of distributed systems such as spark, distributed caching, data-preprocessing, AI training etc.
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies.
- Experience working in a highly cross-functional environment in software organizations, hardware, supply chain, testing, and infrastructure teams.
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements).
- Master’s degree or PhD in Computer Science, or a related technical field.
- 3 years of experience in a technical leadership role, leading projects, driving technical direction and mentoring.
- Deep architectural knowledge of CPU, GPU, Accelerators, Networking, Memory, Flash/HDD Storage systems, DPU. Hands-on experience in building and optimizing server architectures/systems.
Compensation
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is making strides in AI and systems co-design.
Requirements
Hardware Architecture Knowledge
Extensive understanding of hardware architecture including CPUs, memory, and storage systems.
Programming Skills
5+ years of experience with programming languages such as C, C++, Rust, Java, PHP, and Python.
Infrastructure Experience
Experience with large-scale infrastructure and distributed systems.
Technical Degree
Bachelor's degree in Computer Science, Computer Engineering, or a relevant technical field.
Nice to Have
Experience with responsible and ethical AI practices.
Experience in developing and analyzing distributed systems.
Demonstrated ongoing development in AI skills and technologies.
Experience in a technical leadership role with project management.
Benefits
Open Source Contributions
Opportunities to publish work in leading conferences and journals.
Collaborative Environment
Work closely with cross-functional teams and industry partners.