About the Role
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
NVIDIA Cloud Functions (NVCF) is an Open Source Platform that links workloads to GPUs. It lets teams deploy, manage, and serve GPU-accelerated, containerized applications across regions and clusters worldwide. The platform routes inference, streaming, and batch jobs across decentralized GPU clusters. This allows endpoints to scale repeatably, whether hosted on-prem or in the cloud.
What You'll Be Doing
You'll be working in a distributed team that explores innovative ways to make GPU and DPU accelerated applications easier to develop, deploy, and monitor on the latest and greatest NVIDIA hardware.
- Design and ship services in Java, Go and Rust, building in the open on a public repository where your commits, design proposals, and reviews are transparent to the community.
- Work on automating and optimizing build, test, integration, and release processes for cloud native.
- Partner with engineering teams across NVIDIA so the platform integrates with adjacent NVIDIA technologies, including the KAI Scheduler, NVIDIA NIM, Grove and Dynamo.
- Help steward an open-source project. You will triage community issues and pull requests, write docs contributors can build on.
What We Need To See
Bachelor’s or Master’s Degree in Computer Science or equivalent program from an accredited University/College and 8+ years of hands-on software engineering.
- Expert level knowledge in a systems programming language (Go, C, Rust) and proven understanding of Data Structures, Algorithms and Distributed Software Architecture.
- Strong understanding of Container Orchestration Systems (Kubernetes) and Container Technologies with hands-on automation experience in continuous integration frameworks like Gitlab & ArgoCD.
- Expertise in a scripting language (Bash, Python) and knowledge and experience working with System internals of Unix/Unix-like kernels such as Linux.
- Understanding of performance, security and reliability in complex distributed systems.
Ways To Stand Out From The Crowd
- Background with pub-sub models and message queues.
- Experience optimizing for high-throughput network paths, with a working understanding of unary versus streaming and bidirectional protocols across HTTP/2 and gRPC.
- Experience with developing Kubernetes Custom Resources and Operators deployed in Cloud Service Providers.
We have some of the most hard-working and skilled people in the world working for us and our world-class engineering teams are growing fast. If you're a creative and self-motivated engineer with a real passion for technology, we want to hear from you!
Requirements
Systems programming languages
Expert level knowledge in Go, C, or Rust is required.
Container orchestration
Strong understanding of Kubernetes and container technologies is essential.
Scripting languages
Expertise in Bash or Python is necessary.
Distributed systems
Proven understanding of distributed software architecture is needed.
Nice to Have
Background with pub-sub models and message queues is a plus.
Experience optimizing for high-throughput network paths is beneficial.
Experience with developing Kubernetes Custom Resources and Operators is advantageous.
Benefits
Diverse environment
Work in a diverse and supportive environment.
Career growth
Opportunities for learning and professional growth.