About the Role
About the Role
Roles at Hugging Face are very fluid and dynamic -- we're looking for someone who is comfortable taking on different challenges that evolve over time. This role sits on the Xet Storage team, the group responsible for the storage system behind all of Hugging Face. Today we store over 200PB (and growing rapidly!) of the world's most important ML & AI assets. Xet is the underlying storage architecture for the entire Hugging Face platform and community -- from the largest model repositories to the datasets and Spaces that millions of people build on every day.
In this role, you'll work across two closely connected surfaces. You'll contribute to xet-core, our open-source project written in Rust that powers hf-xet -- the Python library underpinning the Hugging Face Hub client and the wider ecosystem of open-source tools the community relies on. And you'll design, build, and operate meaningful and challenging features in the Xet Storage backend, contributing to the broader Infrastructure organization at Hugging Face. We're a small team building and operating incredible things at enormous scale, in a high-trust, low-process, async, and remote environment -- and we lean heavily on the latest AI tools to move fast.
If you love writing low-level, high-performance code and are energized by the challenge of large-scale, scalable services, this is the team for you. This usually means proven experience building and operating production systems software, but we consider every applicant on an individual basis.
Requirements
What we're looking for
- 8+ years building and scaling distributed systems, storage, or networking infrastructure
- Proficiency in a low-level systems language, with Rust strongly preferred (we also work in Python, Typescript, Go, and C++ across the stack)
- A track record of working independently in a high-trust, low-process environment -- you're comfortable with ambiguity and take ownership end to end
- Comfort operating in a fast-moving, async, and fully remote environment
- A passion for building simple, robust, and scalable systems relied on by engineering and science teams around the world
Bonus points if you have
- Experience designing efficient, high-performance, fault-tolerant data storage and retrieval systems
- Experience operating production systems -- monitoring, alerting, and distributed debugging and recovery
- Familiarity with git internals, cloud infrastructure (AWS, Azure, GCP, Kubernetes), databases (relational and non-relational), or networking
About You
If you love open-source, are excited by the intersection of low-level performance work and large-scale services, and want your code to sit at the foundation of the world's largest platform for AI builders, then we can't wait to see your application!
If you're interested in joining us, but don't tick every box above, we still encourage you to apply! We're building a diverse team whose skills, experiences, and backgrounds complement one another. We're happy to consider where you might be able to make the biggest impact.
One more thing
At Hugging Face we believe great AI shouldn't require a massive cluster, we build for everyone, especially the GPU-poor. And because we read every application, here's a small sign that you read this one too: start your answer to the first application question with the words "GPU-poor and proud 🤗". No trick, no catch, it just tells us a real person is on the other side.
Benefits
More about Hugging Face
We are actively working to build a culture that values diversity, equity, and inclusivity. We are intentionally building a workplace where people feel respected and supported—regardless of who you are or what you do.
Requirements
Distributed systems experience
8+ years building and scaling distributed systems, storage, or networking infrastructure.
Proficiency in Rust
Proficiency in a low-level systems language, with Rust strongly preferred.
Independent work
A track record of working independently in a high-trust, low-process environment.
Remote work comfort
Comfort operating in a fast-moving, async, and fully remote environment.
Passion for scalable systems
A passion for building simple, robust, and scalable systems relied on by engineering and science teams.
Nice to Have
Experience designing efficient, high-performance, fault-tolerant data storage and retrieval systems.
Experience operating production systems including monitoring and distributed debugging.
Familiarity with cloud infrastructure (AWS, Azure, GCP, Kubernetes) and databases.