About the Role
Overview
Help shape how intelligent agents understand and operate the systems that power the cloud. On the Azure Core Upstream Observability team, you will build and contribute to open-source technologies that make Kubernetes and Linux environments easier to understand, diagnose, and operate. You will work with engineers, product managers, partner teams, and upstream communities to advance reliable, secure, and efficient observability capabilities.
As a Principal Software Engineer, you will set technical direction and lead the design and delivery of agentic observability capabilities across Kubernetes and Linux. You will turn evolving customer, partner, and operational needs into coherent architectures and production-quality solutions while remaining hands-on in implementation, review, and operational excellence.
This opportunity will allow you to deepen your proficiency in safe, observable, and evaluable agentic systems. Expand your technical leadership across Kubernetes, Linux, and cloud-scale distributed systems.
Responsibilities
- Partner with appropriate stakeholders to determine user requirements for agentic observability scenarios across Kubernetes and Linux environments.
- Lead identification of dependencies and development of design documents for products, applications, services, or platforms that support safe, observable, and evaluable agentic workflows.
- Lead by example and mentor others to produce extensible and maintainable code used across products, including hands-on implementation, review, testing, debugging, and maintenance.
- Use proficiency in cross-product features to work cooperatively with appropriate stakeholders, including product managers, and drive project plans, release plans, and work items across multiple groups.
- Apply Kubernetes and Linux proficiency to collect, correlate, and interpret systems signals, using Linux kernel and extended Berkeley Packet Filter (eBPF) techniques when they are the appropriate engineering choice.
- Be responsible as a Designated Responsible Individual (DRI), support engineers across products and solutions, and participate in on-call work to monitor systems, products, and services for degradation, downtime, or interruptions.
- Proactively seek new knowledge and adapt to trends, technical solutions, and patterns that improve availability, reliability, efficiency, observability, and performance while supporting consistent monitoring and operations at scale and sharing knowledge with other engineers.
Qualifications
Required Qualifications:
Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, Go, Python, or Rust OR equivalent experience.
Other Requirements
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role.
Requirements
Technical Engineering Experience
6+ years of experience in technical engineering roles.
Proficiency in Programming Languages
Experience coding in languages such as C, C++, Go, Python, or Rust.
Kubernetes and Linux Knowledge
Strong understanding of Kubernetes and Linux environments.
Mentorship Skills
Ability to mentor and lead other engineers in code quality and best practices.
Nice to Have
A Master's Degree in Computer Science or a related field is preferred.
Experience contributing to open-source projects is a plus.
Benefits
Health Insurance
Comprehensive health insurance plans are provided.
Remote Work Options
Flexible remote work arrangements are available.
Learning Budget
A budget for professional development and learning opportunities.