← ポータルに戻る

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation💻 コードあり

Jiaxing Li, Kai Zou, Cindy Zhou, Kaichen Huang, Junyao Gao等 · AI · 2026-07-29 ⭐ 8/10
💡 自己回帰型ビデオ蒸留において、モード探索とモードカバレッジを協調させることで、生成品質と多様性を両立させる新しい蒸留手法を提案した研究。
🤖 Ayumuより: この論文は、ビデオ生成モデルの蒸留で、学生モデルの初期化と学習プロセスを分布の観点から見直しているのが面白いね。特に、DMDがモード探索的だから、初期化からモードカバレッジを意識すべきだっていう指摘は、直感的だけど見落とされがちだった点だと思うよ。小さなモデルでも大きなモデルに勝てるって結果は、分布アライメントの重要性を強く示しているから、朋義さんもきっと興味を持つんじゃないかな。
Autoregressive Video Distillation Distribution Matching Distillation (DMD) Mode Covering Mode Seeking Consistency Distillation Distributional Alignment
1. どんなもの?
  • 自己回帰型ビデオ蒸留における学生モデルの初期化と蒸留プロセスを改善する手法「DistillAlign」を提案。
  • 既存手法が抱える、初期化段階と蒸留段階で目標とする分布が異なる問題や、視覚的スコアに偏った評価の問題を解決する。
  • モード探索(mode-seeking)とモードカバレッジ(mode-covering)を協調させることで、生成品質、カバレッジ、多様性を向上させる。
2. 先行研究と比べてどこがすごい?
  • 既存のDMD(Distribution Matching Distillation)ベースの蒸留法では、初期化と蒸留が分離され、異なる目標分布を追求し、視覚的スコアで中間学生モデルを評価していた。
  • 本研究は、DMDのモード探索的な性質を考慮し、初期化段階から教師モデルの「モードカバレッジ」に合致させるべきだと主張。
  • 学生と教師の分布間の精度とカバレッジを測定する「分布評価プロトコル」を導入し、視覚スコアでは隠されていた「高精度だが低カバレッジ」の問題を特定した。
  • DMDの逆KL目的関数がトレーニング後期にカバレッジと多様性を低下させる問題を指摘し、これを解決するためにモード探索とモードカバレッジを協調させる「Joint Distillation」を提案した。
3. 技術や手法の肝はどこ?
  • **分布評価プロトコル**: 共有潜在空間において、学生モデルと教師モデルの生成分布間の精度(precision)とカバレッジ(coverage)を定量的に測定する。これにより、既存の視覚評価では捉えきれなかった分布のミスマッチを明らかにする。
  • **Joint Distillation**: DMDのモード探索目的(生成品質の向上)と、Consistency Distillationに基づいたモードカバレッジ制約(多様性とカバレッジの維持)を組み合わせる。これにより、トレーニング全体を通じて、生成品質と多様性のバランスを取りながら学生モデルを効果的に蒸留する。
4. どうやって有効だと検証した?
  • 提案手法が生成品質、カバレッジ、多様性の全てにおいて改善をもたらすことを実験で示した。
  • 特に注目すべきは、Wan-1.3Bという比較的小さなDMD教師モデルを使用した場合でも、Wan-14Bで洗練されたベースラインモデルを性能で上回った点。
  • この結果は、モデルの規模だけでなく、自己回帰型ビデオ蒸留における分布アライメント(分布の整合性)の重要性を強く裏付けている。
5. 議論はある?
  • アブストラクトからは直接的な限界や今後の課題についての詳細な議論は読み取れないが、自己回帰型ビデオ蒸留における分布アライメントの重要性を強調している点は、今後の研究方向性を示唆している。
  • 提案された分布評価プロトコルが、ビデオ生成以外の他の生成モデルの蒸留タスクにも応用可能か、またその際の課題については、さらなる検討の余地があるかもしれない。
6. 次に読むべき論文は?
  • Distribution Matching Distillation (DMD)に関する主要な先行研究。
  • Consistency Distillationに関する論文。
  • 自己回帰型ビデオ蒸留(Autoregressive Video Distillation)の基盤となる論文や、Wan-1.3B、Wan-14Bといった言及されているモデルの元論文。
  • VBenchなど、ビデオ生成モデルの評価指標に関する論文。

Abstract (原文)

Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD's reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD's mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.