About the Role
Overview
Windows is one of the largest codebases in the world, running on more than a billion devices. The Windows Engineering System team is the backbone that enables Windows engineers to design, build, validate, and ship at scale—securely and with high quality. Windows, and the industry at large, are entering a new era of AI-discovered vulnerabilities. The pace of vulnerability discovery is changing, with advances in AI making it possible to find more issues, faster, across more code, and new mechanisms that can accelerate both discovery and analysis. The fastest way to reduce customer exposure is to find issues before attackers can use them.
As Windows leadership has announced publicly, Windows is expanding its ability across the platform to find issues earlier, accelerate the engineering work to fix them, strengthen validation, and deliver timely, high-quality updates that keep customers protected. Meeting this reality means building autonomous systems that scale vulnerability discovery, proof, and submission to AI-scale volume.
Windows Engineering Systems owns running and operating service infrastructure to tackle this problem called autonomous “autopilots.” The “Hunter” autopilot loop runs Microsoft Security’s multi-model scanning harness (MDASH) over Windows source and proves candidate vulnerabilities exploitable on real VMs, so only the highest-confidence findings ever reach engineers. Windows runs on over a billion devices, so running this at Windows scale has meant standing up dedicated cloud infrastructure for scanning and proving.
As a Principal Software Engineer, you will help scale that backbone to the next order of magnitude to keep pace with AI-scale discovery, driving systems engineering at the scale of one of the world’s largest codebases.
Responsibilities
Why this role exists
As AI-powered discovery expands to more code and more scenarios, the scanning-and-proving infrastructure behind it must grow by an order of magnitude: more throughput, broader coverage, and production-grade reliability. That growth lands squarely on backend compute and orchestration. This role owns the distributed backbone that makes vulnerability discovery and proof run reliably and cost-effectively at Windows scale.
What You'll Own
- AI Evals & Benchmarks: Building AI evals and benchmarks to ensure that the system stays healthy and does not regress across multiple scenarios.
- Throughput & Reliability: Scaling the pipeline by an order of magnitude and improve reliability.
- Compute, Capacity & Orchestration: Managing compute and AI infra at Windows-wide scale; smoothing capacity spikes across model-hosting backends; managing queuing/scheduling and pool prioritization infrastructure.
- Cost & Efficiency: Driving billing/consumption estimation and cost-of-goods modeling.
Who You Are
6+ years building large-scale distributed systems / service backends (queuing, scheduling, autoscaling, multi-tenant compute).
Deep experience with capacity planning, throughput/latency SLAs, and cost-of-goods on cloud infrastructure (Azure preferred).
Track record owning architecture of a complex, ambiguous system end to end, and mentoring senior engineers.
Comfortable negotiating and holding a hard dependency line with partner teams.
Qualifications
Required Qualifications:
Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python, OR equivalent experience.
Preferred Qualifications
Master's Degree in Computer Science.
Requirements
Distributed Systems Experience
6+ years building large-scale distributed systems and service backends.
Capacity Planning
Deep experience with capacity planning and throughput/latency SLAs.
Cloud Infrastructure Knowledge
Experience with cost-of-goods modeling on cloud infrastructure, preferably Azure.
Mentorship Skills
Proven track record of mentoring senior engineers.
Nice to Have
A Master's Degree in Computer Science is preferred.
Comfortable negotiating with partner teams.
Benefits
Growth Opportunities
Opportunities for professional growth and development.
Inclusive Culture
A culture of inclusion where everyone can thrive.