About the Role
About LinkedIn
LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. We're also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture that's built on trust, care, inclusion, and fun where everyone can succeed.
Join us to transform the way the world works.
Job Description
At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.
Responsibilities
Join us to push the boundaries of scaling large models together. The team is responsible for scaling LinkedIn’s AI model training, feature engineering and serving with hundreds of billions of parameters models and large scale feature engineering infra for all AI use cases from recommendation models, large language models, to computer vision models. We optimize performance across algorithms, AI frameworks, data infra, compute software, and hardware to harness the power of our GPU fleet with thousands of latest GPU cards.
The team also works closely with the open source community and has many open source committers (TensorFlow, Horovod, Ray, vLLM, Hugginface, DeepSpeed etc.) in the team. Additionally, this team focuses on technologies like LLMs, GNNs, Incremental Learning, Online Learning and Serving performance optimizations across billions of user queries.
Model Training Infrastructure
As an engineer on the AI Training Infra team, you will play a crucial role in building the next-gen training infrastructure to power AI use cases. You will design and implement high performance data I/O, work with open source technologies to identify and resolve issues in popular libraries like PyTorch, Huggingface etc., enable distributed training over 100s of billions of parameter models, debug and optimize deep learning training, and provide advanced support for internal AI teams in areas like model parallelism, tensor parallelism etc.
Finally, you will assist in and guide the development of containerized pipeline orchestration infrastructure, including developing and distributing stable base container images, providing advanced profiling and observability, and updating internally maintained versions of deep learning frameworks and their companion libraries like CUDA, cuTile, cuDNN, NCCL, RDMA, Tensorflow, PyTorch, TorchRec, Flash Attention, PyTorch Lightning and more.
Feature Engineering
This team shapes the future of AI with the state-of-the-art Feature Platform, which empowers AI Users to effortlessly create, compute, store, consume, monitor, and govern features within online, offline, and nearline environments, optimizing the process for model training and serving. As an engineer in the team, you will explore and innovate within the online, offline, and nearline spaces at scale (millions of QPS, multi-terabytes of data, etc), developing and refining the infrastructure necessary to transform raw data into valuable feature insights.
Utilizing leading open-source technologies like Spark, Beam, and Flink and more, you will play a crucial role in processing and structuring feature data, ensuring its most optimal storage in the Feature Store, and serving feature data with high performance.
Model Serving Infrastructure
This team builds low latency high performance applications serving very large & complex models across LLM and Personalization models. As an engineer, you will build compute efficient infra on top of native cloud, enable GPU based inference for a large variety of use cases, cuda level optimizations for high performance.
Requirements
AI model training
Experience in building and optimizing AI model training infrastructure.
Deep learning frameworks
Proficiency in popular deep learning frameworks like PyTorch and TensorFlow.
Feature engineering
Knowledge of feature engineering processes and tools.
Container orchestration
Experience with containerized pipeline orchestration and related technologies.
Nice to Have
Experience contributing to or collaborating with open source projects.
Familiarity with GPU optimization techniques for AI applications.
Benefits
Flexible work
Hybrid work model allowing flexibility in work location.
Growth opportunities
Investment in employee growth and development.