视频生成模型 · 训练 Infra 工程师

Embedding VC

Apply to this job
San Francisco Bay Area remote Until 9/27/2026 5+ years exp First posted July 29, 2026 Last posted July 29, 2026
Job description

我们在训练自研视频生成基础模型(DiT / Flow Matching),需要一位既能搭起训练平台、又能把研究代码变成数百卡集群上稳定结果的工程师。你不只是用平台的人,更是建平台的人。

你会做

• 训练平台搭建:从作业调度、断点续训、监控告警到数据 / 权重流水线,把分散的脚本沉淀为团队可复用的训练基础设施。

• 数百卡规模的分布式训练:FSDP、张量并行、Context Parallel、Ulysses,把 MFU 推到合理水位。

• PB 级视频数据 pipeline:NVDEC 解码、VAE latent 缓存、变分辨率 bucket sampling。

• 显存与性能:FlashAttention、FP8 混合精度、Triton kernel、activation checkpoint 策略。

• 训练稳定性:loss spike 根因分析、断点秒级恢复、慢节点自动剔除。

硬性要求

• 精通 PyTorch distributed 与 CUDA 体系结构。

• 至少一个主流训练框架(Megatron / DeepSpeed / FSDP / TorchTitan)的源码级理解。

• ≥ 256 卡训练实战经验。

• 有从零或半程搭建训练平台 / 集群调度 / 训练工具链的经验。

加分

• DiT / Diffusion / 视频数据处理经验。

• 写过 Triton / CUTLASS kernel。

• 对 HunyuanVideo / Wan / CogVideoX 等开源项目有源码级了解。

我们提供

• 真实数百卡算力、把训练当工程问题的团队、合规边界内的开源 / 发表空间。

About this role

Summary

Build and optimize large-scale distributed video model training infrastructure.

Job title

训练 Infra 工程师

Experience level

5+ years

Minimum experience

5+ years exp

Industry

software

Location requirements

San Francisco Bay Area, remote not specified

Salary

Not specified

Management role

No

Skills & keywords

Required skills

PyTorch distributedCUDAsource code understanding256+ GPU training

Preferred skills

DiffusionTritonopen source video projects

Specializations

distributed trainingCUDAPyTorchcluster managementvideo data pipeline
Locations

Structured locations inferred from the posting.

Unknown location

Work arrangement unknown