About the Role
What You'll Do
At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.
We are hiring AI / ML Platform Engineers to build the platform layer that makes AI-for-engineering workflows scalable, reliable, and reproducible. This role focuses on the infrastructure and platform systems that support large-scale agent execution, distributed training and inference, experiment tracking, benchmark automation, artifact management, and GPU cluster utilization.
You will work closely with ML Systems Research Engineers, AI Research Scientists, Applied AI Engineers, and hardware domain experts to operationalize the Blueprint framework across kernel optimization, RTL/PPA optimization, ECO fixing, verification, simulation, and debugging workflows.
This is a platform engineering role, not a pure research role. The focus is to build robust shared systems that allow researchers and engineers to run more experiments, compare results reliably, reduce manual orchestration, and move successful workflows into production engineering use.
Who You Are
You are a strong systems engineer who enjoys building reliable platforms for AI researchers and applied engineers. You understand distributed systems, ML workloads, GPU infrastructure, experiment management, and production reliability. You can turn messy research workflows into reusable services, APIs, dashboards, job systems, and automation.
You care about reproducibility, observability, performance, and developer experience. You are comfortable working across ML, infrastructure, and hardware/software tooling, and you can partner with research teams without requiring every requirement to be fully specified upfront.
Key Responsibilities
- Build and operate the shared AI platform for agentic engineering workflows, including job submission, scheduling, orchestration, retries, logging, artifact storage, and experiment tracking.
- Develop reliable infrastructure for distributed training, distributed inference, batch evaluation, and large-scale agent rollout across GPU clusters.
- Build platform services for benchmark execution, correctness checking, profiling, regression tracking, and reproducible evaluation.
- Maintain artifact systems for generated kernels, RTL edits, traces, logs, profiler outputs, benchmark results, simulator outputs, and formal verification artifacts.
- Support scalable integrations with compilers, ROCm/HIP tooling, profilers, simulators, EDA tools, vLLM, SGLang, and internal engineering systems.
- Improve GPU cluster utilization, scheduling efficiency, reliability, quota management, and workload isolation.
- Build dashboards and observability systems for experiment status, resource usage, failure modes, benchmark trends, regression detection, and team productivity.
- Partner with ML Systems Research Engineers to productionize research workflows for RL systems, inference systems, quantification systems, and evaluation pipelines.
- Partner with Applied AI Engineers to make Blueprint harnesses reusable across kernel optimization, RTL optimization, verification, firmware, and CPU/GPU performance workflows.
- Establish platform standards for reproducibility, data retention, run metadata, artifact lineage, access control, and operational reliability.
Technical Focus Areas
Distributed ML platform infrastructure for training, inference, evaluation.
Requirements
Distributed Systems Knowledge
You should have a strong understanding of distributed systems.
ML Workloads Experience
Experience with machine learning workloads is essential.
GPU Infrastructure Familiarity
Familiarity with GPU infrastructure is required.
Experiment Management Skills
You need skills in managing experiments effectively.
Nice to Have
Experience in developing APIs for platform services is a plus.
Familiarity with automation tools can enhance your application.
Benefits
Career Advancement
Opportunities for career advancement within a leading tech company.
Collaborative Culture
Join a culture that values collaboration and innovation.