About the Role
About the Role
As a Senior Software Engineer on the AIOps platform team at NVIDIA, you will architect and build a mission-critical observability and prediction platform that processes massive telemetry streams from GPU clusters. Your role will involve developing distributed systems and operationalizing predictive AI models at scale, ensuring high performance and reliability.
What You'll Be Doing
- Architect and build an agentic AIOps system that autonomously monitors GPU fleet health, aggregates and correlates massive telemetry streams, surfaces intelligent alerts, and orchestrates multi-step diagnostic workflows and corrective actions.
- Research, evaluate, and prototype data storage strategies and data representations across diverse database technologies and modalities.
- Design distributed systems to handle the extreme telemetry density of large-scale AI clusters.
- Instrument services with deep observability (metrics, logs, traces) to support rapid debugging and continuous performance improvement.
- Build and own the model-serving infrastructure that operationalizes predictive algorithms at scale.
- Contribute to the platform's core libraries and abstractions that accelerate development across the broader AIOps engineering team.
What We Need To See
- B.Sc./M.Sc. in Computer Science, Computer Engineering, or a related technical field.
- 5+ years of software engineering experience building production distributed systems.
- Expert-level proficiency in languages such as Go, C++, or Rust.
- Solid understanding of Kubernetes and container-based deployments for production services.
- Experience deploying, monitoring, and maintaining ML models or data-intensive services in a production environment.
- Comfort working in ambiguous, fast-moving environments.
Ways To Stand Out From The Crowd
- Experience building ML model-serving platforms or MLOps tooling at scale.
- A track record of taking systems from prototype to stable, production-grade platform.
- A "Systems" Thinker: Understanding the full stack from data movement to processing in a distributed cluster.
- Practical Innovation: Ability to simplify complex problems and build internal tools or frameworks.
Compensation and Benefits
With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers.
Requirements
Software Engineering Experience
5+ years of experience building production distributed systems.
Core Systems Programming
Expert-level proficiency in languages such as Go, C++, or Rust.
Kubernetes Knowledge
Solid understanding of Kubernetes and container-based deployments.
ML Model Deployment
Experience deploying and maintaining ML models in a production environment.
Nice to Have
Experience building ML model-serving platforms or MLOps tooling at scale.
Ability to understand the full stack from data movement to processing.
Benefits
Competitive Salaries
NVIDIA offers competitive salaries.
Generous Benefits Package
Includes various perks that make NVIDIA a desirable employer.