About the Role
Overview
Microsoft is a company where passionate innovators come to collaborate, envision what can be, and take their careers further. This is a world of more possibilities, more innovation, more openness, and sky-is-the-limit thinking in a cloud-enabled world.
Microsoft's Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging and real-time analytics, and business intelligence. The products in our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is to build the data platform for the age of AI, powering a new class of data-first applications, and driving a data culture.
Within Microsoft Fabric, the Azure Monitor team builds services that enable customers to monitor, detect, troubleshoot, and mitigate issues with their services through an increasingly agentic experience. Azure Monitor includes Log Analytics, Application Insights, Container Insights, Hosted Prometheus, Azure Managed Grafana, and more. Additionally, Azure Monitor is the platform upon which Microsoft Sentinel is built. We have a multi-billion dollar business that is growing rapidly, and we run some of the world’s highest scale observability services both for Microsoft internally and for our external customers, processing over 1.5 Exabytes of logs daily and tracking over 100 billion active metrics.
Responsibilities
- Design, implement, test, deploy, and operate highly available distributed services and automation used to configure and migrate large-scale telemetry workloads that will be leveraged across the fleet.
- Build and enhance APIs, tools, and subsystems for telemetry collection, routing, storage, and efficient data access.
- Integrate advanced capabilities (e.g., machine learning–based anomaly detection and data validation) to enhance platform intelligence and insights.
- Implement robust monitoring, alerting, and diagnostics and ensure production services run reliably, including participation in on-call rotations and incident response.
- Collaborate with partner teams to deliver end-to-end observability solutions and contribute to design reviews and best practices that uphold high engineering standards.
- Help evolve how the team builds with AI, championing agentic development practices and sharing repeatable patterns that improve engineering velocity and quality, as well as training and evolving AI to assist with non-coding scenarios such as site reliability engineering, support and administration.
- Drive the next generation of our backend systems and mentor engineers, contribute to technical design reviews, and communicate clearly with cross-organization stakeholders about tradeoffs, risks.
Requirements
Distributed systems design
Experience in designing and implementing highly available distributed services.
API development
Proficiency in building and enhancing APIs for telemetry and data access.
Machine learning integration
Knowledge of integrating machine learning capabilities for enhanced platform intelligence.
Monitoring and diagnostics
Experience in implementing robust monitoring and diagnostics for production services.
Nice to Have
Familiarity with cloud platforms and services, particularly Azure.
Experience mentoring engineers and contributing to technical design reviews.
Benefits
Career development
Opportunities for professional growth and career advancement.
Collaborative culture
A fast-paced, startup-like culture that values collaboration and innovation.