Recommendation Architecture AI/ML Infrastructure Engineer Graduate (Data-Arch-TikTok Live) - 2027 Start

San Jose, California, US Until 10/7/2026 H-1B sponsor history First posted August 8, 2026 Last posted August 8, 2026
Job description

Description

Our team develops the core training and serving infrastructure that powers one of the world's largest recommendation systems, enabling billions of personalized recommendations every day. We are also advancing the next generation of AI infrastructure for foundation models and LLMs, driving innovation in large-scale model training, online inference, and GPU optimization.
As part of the team, you will work on distributed training and inference systems, high-performance GPU computing, and scalable LLM infrastructure. You'll collaborate closely with experienced engineers and researchers to transform cutting-edge AI technologies into production systems that directly impact the experience of hundreds of millions of TikTok users. This role is ideal for candidates who are passionate about LLM systems, distributed computing, GPU programming, and building AI systems at massive scale.We are looking for passionate talented individuals to join our Model Infrastructure team, building the next generation of infrastructure for TikTok's For You recommendation system and Large Language Models (LLMs).

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

Online Assessment
Candidates who pass resume screening will be invited to participate in Our Company's technical online assessment.

Responsibilities
- Build and optimize infrastructure for large-scale model training and online inference.
- Develop distributed systems supporting large recommendation models and LLMs.
- Improve training and inference performance through GPU optimization and efficient communication.
- Collaborate with researchers to develop and deploy LLM training and serving solutions.
- Analyze system bottlenecks and implement performance optimizations.

Requirements

Minimum Qualifications:
- Individuals who are completing or have recently completed a Bachelor's or Master's degree in Artificial Intelligence, Software Development, Computer Science, Computer Engineering or a related discipline.
- Strong programming skills in C++ or Python. •Good understanding of data structures, algorithms, and computer systems.
- Familiarity with PyTorch or TensorFlow.
- Knowledge of Transformer architectures and Large Language Models (LLMs).
- Strong problem-solving skills and a passion for building large-scale AI systems.

Preferred Qualifications:
- Hands-on experience with LLM training or inference through research, internships, or open-source projects. •Familiarity with distributed training concepts (e.g., DP, TP, PP, FSDP, ZeRO).
- Experience with GPU programming using CUDA, Triton, or similar technologies.
- Understanding of LLM serving techniques such as KV Cache, Continuous Batching, or FlashAttention. •Contributions to open-source projects or research in machine learning systems, distributed systems, or LLM infrastructure.

By submitting an application for this role, you accept and agree to our global applicant privacy policy, which may be accessed here: https://careers.tiktok.com/legal/privacy

About this role

Summary

Build and optimize large-scale AI training and inference infrastructure for TikTok.

Job title

Recommendation Architecture AI/ML Infrastructure Engineer Graduate (Data-Arch-TikTok Live) - 2027 Start

Experience level

graduate level

Industry

software

Location requirements

San Jose, California, USA, on-site; remote not specified

Salary

Not specified

Visa sponsorship

H-1B sponsor history

Management role

No

Skills & keywords

Required skills

C++Pythondata structuresalgorithmscomputer systemsPyTorchTensorFlowLLMs

Preferred skills

LLM trainingdistributed trainingCUDATritonKV CacheFlashAttentionopen-source contributions

Specializations

AI infrastructuredistributed systemsGPU programminglarge language modelsrecommendation systems
Locations

Structured locations inferred from the posting.

San Jose, CA, USA

On-site City