About the Role
About the Role
As a Software Engineer focused on Platform Infrastructure, you will design and implement large-scale distributed systems that support supercomputing clusters. You will work on optimizing performance across various systems and collaborate on hardware and software co-design to enhance AI training capabilities.
What You'll Do
- Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
- Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
- Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
- Maintain and innovate on our codebase to ensure scalability and reliability.
- Develop tools to enhance team productivity and streamline workflows.
Who You Are
- Systems programming experience in C, C++, or Rust.
- Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
- Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations.
Bonus Points
- Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers.
- Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives).
- Proficiency in performance analysis, profiling, and low-level optimization techniques.
- Solid understanding of computer networks and the TCP/IP stack.
- Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark).
- Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
- Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel.
- Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar).
Compensation and Benefits
Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
Requirements
Systems programming
Experience in C, C++, or Rust is essential.
Computer systems fundamentals
A solid understanding of how computers execute code is required.
Kubernetes expertise
Hands-on experience with Kubernetes, including cluster architecture and networking.
Nice to Have
Strong debugging skills across the full stack are a plus.
Deep knowledge of operating systems internals is beneficial.
Proficiency in performance analysis and low-level optimization techniques is advantageous.
Benefits
Equity
Employees receive equity as part of the compensation package.
Health coverage
Comprehensive medical, vision, and dental coverage is provided.
Retirement plan
Access to a 401(k) retirement plan is included.