About the Role
About the Role
The Senior Software Engineer for AIOps at NVIDIA AI will be responsible for building core distributed systems that manage telemetry streams from GPU clusters and operationalize predictive AI models. This role involves architecting an observability platform, designing high-scale engineering solutions, and contributing to the development of model-serving infrastructure.
What You'll Be Doing
- Architect and build an agentic AIOps system that autonomously monitors GPU fleet health, aggregates and correlates massive telemetry streams, surfaces intelligent alerts, and orchestrates multi-step diagnostic workflows and corrective actions.
- Research, evaluate, and prototype data storage strategies and data representations across diverse database technologies and modalities.
- Design distributed systems to handle the extreme telemetry density of large-scale AI clusters.
- Instrument services with deep observability (metrics, logs, traces) to support rapid debugging and continuous performance improvement.
- Build and own the model-serving infrastructure that operationalizes predictive algorithms at scale.
- Contribute to the platform's core libraries and abstractions that accelerate development across the broader AIOps engineering team.
Who You Are
- B.Sc./M.Sc. in Computer Science, Computer Engineering, or a related technical field.
- 5+ years of software engineering experience building production distributed systems.
- Expert-level proficiency in languages such as Go, C++, or Rust.
- Solid understanding of Kubernetes and container-based deployments for production services.
- Experience deploying, monitoring, and maintaining ML models or data-intensive services in a production environment.
- Comfort working in ambiguous, fast-moving environments.
Bonus Points
- Experience building ML model-serving platforms or MLOps tooling at scale.
- A track record of taking systems from prototype to stable, production-grade platform.
- A "Systems" Thinker who understands the full stack.
- Practical Innovation in simplifying complex problems.
Compensation
With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers.
Requirements
Software Engineering Experience
5+ years of experience building production distributed systems.
Core Systems Programming
Expert-level proficiency in languages such as Go, C++, or Rust.
Kubernetes Knowledge
Solid understanding of Kubernetes and container-based deployments.
ML Model Deployment
Experience deploying and maintaining ML models in production.
Nice to Have
Experience building ML model-serving platforms or MLOps tooling at scale.
Ability to understand the full stack and how data moves across systems.
Benefits
Competitive Salaries
NVIDIA offers competitive salaries.
Generous Benefits Package
Includes various perks that make NVIDIA a desirable employer.