视频生成模型 · 训练 Infra 工程师
Embedding VC
Apply to this job我们在训练自研视频生成基础模型(DiT / Flow Matching),需要一位既能搭起训练平台、又能把研究代码变成数百卡集群上稳定结果的工程师。你不只是用平台的人,更是建平台的人。
你会做
• 训练平台搭建:从作业调度、断点续训、监控告警到数据 / 权重流水线,把分散的脚本沉淀为团队可复用的训练基础设施。
• 数百卡规模的分布式训练:FSDP、张量并行、Context Parallel、Ulysses,把 MFU 推到合理水位。
• PB 级视频数据 pipeline:NVDEC 解码、VAE latent 缓存、变分辨率 bucket sampling。
• 显存与性能:FlashAttention、FP8 混合精度、Triton kernel、activation checkpoint 策略。
• 训练稳定性:loss spike 根因分析、断点秒级恢复、慢节点自动剔除。
硬性要求
• 精通 PyTorch distributed 与 CUDA 体系结构。
• 至少一个主流训练框架(Megatron / DeepSpeed / FSDP / TorchTitan)的源码级理解。
• ≥ 256 卡训练实战经验。
• 有从零或半程搭建训练平台 / 集群调度 / 训练工具链的经验。
加分
• DiT / Diffusion / 视频数据处理经验。
• 写过 Triton / CUTLASS kernel。
• 对 HunyuanVideo / Wan / CogVideoX 等开源项目有源码级了解。
我们提供
• 真实数百卡算力、把训练当工程问题的团队、合规边界内的开源 / 发表空间。
Summary
Build and optimize large-scale distributed video model training infrastructure.
Job title
训练 Infra 工程师
Experience level
5+ years
Minimum experience
5+ years exp
Industry
software
Location requirements
San Francisco Bay Area, remote not specified
Salary
Not specified
Management role
No
Required skills
Preferred skills
Specializations
Structured locations inferred from the posting.
Unknown location