About the Role
About the Role
As an ML Infrastructure Engineer, you will play a pivotal role in building and optimizing the reliable, high-performance ML platform that powers recommendations on X. We're looking for exceptional engineers who are passionate about our mission and have a strong desire to make a meaningful impact.
Responsibilities
- Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
- Developing data pipelines and integrating large-scale data, training, and inference systems
- Collaborating with ML teams to productionize models and ensure seamless integration across the stack
- Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
- Working across the full stack to solve complex problems independently
- Mentoring junior engineers and contributing to the growth of the team
Basic Qualifications
- Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience
- 2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
- 2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
- Strong proficiency with Python and experience with compiled languages such as C++ or Rust
Preferred Skills and Experience
- Deep familiarity with modern ML frameworks such as JAX or PyTorch
- Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
- Comfortable with Linux systems and orchestration tools
- Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling
Compensation and Benefits
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
Requirements
Degree in Computer Science
A Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or a related field is required.
Industry Experience
At least 2 years of experience in high traffic or large-scale production environments.
ML Platforms Knowledge
Experience with ML platforms and training infrastructure is essential.
Proficiency in Python
Strong proficiency in Python and experience with compiled languages like C++ or Rust is necessary.
Nice to Have
Deep familiarity with modern ML frameworks such as JAX or PyTorch is preferred.
A low-level understanding of compute systems, including distributed storage and NVIDIA drivers, is a plus.
Benefits
Equity
Employees receive equity as part of the compensation package.
Comprehensive Health Coverage
Includes medical, vision, and dental coverage.
Retirement Plan
Access to a 401(k) retirement plan is provided.
Disability Insurance
Short & long-term disability insurance is included.