返回目录
开源项目其他开源工具类新手

GitHub - KohakuBlueleaf/TTVidT

[NeurIPS 2026] TT-VidT Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining https://arxiv.org/abs/2609.33419 Accepted at NeurIPS 2026 (main track). TT-VidT is a self-supervised video encoder built for motion . A DINOv3 ViT-B/16

0 次阅读2026/10/03 发布
GitHub - KohakuBlueleaf/TTVidT 来源图片

社区作者 · zZz

它解决什么问题

[NeurIPS 2026] TT-VidT

Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

https://arxiv.org/abs/2609.33419

Accepted at NeurIPS 2026 (main track).

TT-VidT is a self-supervised video encoder built for motion .

A DINOv3 ViT-B/16 processes every frame independently (the appearance path), while a compact Temporal Transfer pathway turns each frame into a handful of motion tokens that exchange information across time.

It is pretrained with Diff Compression : a DiT decoder must reconstruct every later frame from the first frame's spatial features plus that frame's motion tokens, so the motion tokens are forced to carry exactly what the first frame cannot explain.

This repository contains everything used for the paper: the encoders (TT-VidT in TT1D / TT3D form and the ViT3D and DisMo-2D3D baselines), the six pretraining objectives, decoder pretraining, the data pipeline, frozen-probe and fine-tuning evaluation, and one config per experiment in the paper.

video (8 frames) ──► DINOv3 ViT-B/16, per frame ──────────────► spatial features z_t │ (interleaved) ▼

Temporal Transfer (12 layers)

Temporal Transfer (12 layers)
K=8 motion tokens per frame,

block-causal attention over [motion tokens ; 4x-downsampled z_t] │ ▼ motion tokens m_t ──┐ first-frame features z_1 ────┴──► DiT decoder ──► frame t (Diff Compression)

Installation

命令
git clone https://github.com/KohakuBlueleaf/TTVidT && cd TTVidT
命令
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126 # match your CUDA
命令
pip install -e .
命令
Python >= 3.10. The DINOv3 weights ( facebook/dinov3-vitb16-pretrain-lvd1689m ) are

gated on the Hugging Face Hub: accept the license there and huggingface-cli login before training. Loading a trained checkpoint does not need them.

Using a trained model

import torch from ttvidt . hub import load_model

model = load_model ( "KBlueLeaf/TTVidT" , device = "cuda" ) # pretrained encoder; or a .ckpt / exported dir video = torch . rand ( 1 , 8 , 3 , 256 , 256 , device = "cuda" ) * 2 - 1 # [B, T, C, H, W] in [-1, 1] with torch . no_grad (), torch .

autocast ( "cuda" , dtype = torch . float16 ): out = model . encoder ( video ) motion = out . motion_output # [B, T, 1, 768] pooled motion embedding per frame

Frames should be resized to 256x256 and normalised with mean = std = 0.5. The pretrained encoder KBlueLeaf/TTVidT also loads with transformers alone ( AutoModel.from_pretrained(..., trust_remote_code=True) ). See docs/model.md for the architecture, parameter counts and loading.

Reproducing the paper

Step Guide

docs/data.md

  1. Prepare pretraining data (OpenVid-1M 384px + Moments-in-Time) and the benchmarks

docs/training.md

  1. (Optional) pretrain the DiT decoders; the pretrained ones on the Hub ( KBlueLeaf/TTVidT-decoders ) are used by default

docs/training.md

  1. Pretrain encoders: one config per table cell

docs/evaluation.md

  1. Frozen-probe evaluation, fine-tuning, diagnostics

Every experiment is a Python config run with KohakuEngine 's kogine run :

TT-VidT (TT3D + Diff Compression, decoder S video-pretrained), 2 GPUs

kogine run scripts/train/pretrain_encoder.py -c configs/pretrain/ttvidt_tt3d_diffcomp.py

frozen evaluation of the result: extract -> check -> probe (3 seeds) -> mean +- std

命令
bash scripts/eval/run_frozen_eval.sh ttvidt/ < run_id > /checkpoints/epoch=7.ckpt

To change a setting, write a config that builds on an existing one (see docs/configs.md , which also lists which config belongs to which table).

Repository layout

src/ttvidt/ model, training and data code (installable package) model/ DINOv3VidTModel (TT-VidT), VideoMAE3DViTModel, DisMo2DPlus3DModel, MotionDecoder (DiT), frame VAE modules/ Temporal Transfer layers (tt.py: TT1D, tt3d.py: TT3D), attention blocks trainer.

py TTVidTrainer: objectives, optimisation, EMA, logging data/ tar-of-JPEG video datasets, augmentation, decoder-pretraining data hub.

py load training checkpoints / exported models src/vidmet/ benchmark dataset loaders (HMDB51, ARID, IARD, Jester, SSv2, Diving48, EK-100) configs/ _base/ shared recipes (encoder pretraining, decoder pretraining) pretrain/ encoder pretraining: sweep (Table 1), decoder ablation (Table 2), final models decoder/ DiT decoder pretraining eval/ feature extraction scripts/ train/ pretrain_encoder.

py, pretrain_decoder.py, export_decoder.py eval/ extract_features.py, probe.py, finetune.

py, diagnostics, V-JEPA 2 baseline data/ video -> tar conversion; benchmarks/: dataset setup scripts analysis/ analytical parameter / FLOP counts tools/ smoke test, checkpoint export docs/ guides

License

Apache-2.0, see LICENSE . DINOv3, the datasets and the V-JEPA 2 baseline are subject to their own licenses.

Acknowledgement

This work is supported by the NVIDIA Taiwan AI Research & Development Center (TRDC).

This work utilize sponsored compute resources from Spellbrush.

— 本文由 AI 根据公开来源辅助整理,命令、版本与许可证请在使用前到原始页面复核。

安装 / 开始使用

Installation

命令
git clone https://github.com/KohakuBlueleaf/TTVidT && cd TTVidT
命令
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126 # match your CUDA
命令
pip install -e .
命令
Python >= 3.10. The DINOv3 weights ( facebook/dinov3-vitb16-pretrain-lvd1689m ) are

gated on the Hugging Face Hub: accept the license there and huggingface-cli login before training. Loading a trained checkpoint does not need them. Using a trained model import torch from ttvidt .

hub import load_model model = load_model ( "KBlueLeaf/TTVidT" , device = "cuda" ) # pretrained encoder; or a .ckpt / exported dir video = torch . rand ( 1 , 8 , 3 , 256 , 256 , device = "cuda" ) * 2 - 1 # [B, T, C, H, W] in [-1, 1] with torch .

no_grad (), torch . autocast ( "cuda" , dtype = torch . float16 ): out = model . encoder ( video ) motion = out . motion_output # [B, T, 1, 768] pooled motion embedding per frame Frames should be resized to 256x256 and normalised with mean = std = 0.5.

The pretrained encoder KBlueLeaf/TTVidT also loads with transformers alone ( AutoModel.from_pretrained(., trust_remote_code=True) ). See docs/model.md for the architecture, parameter counts and loading. Reproducing the paper Step Guide

docs/data.md

docs/training.md

docs/training.md

docs/evaluation.md Every experiment is a Python config run with KohakuEngine 's kogine run :

  1. Prepare pretraining data (OpenVid-1M 384px + Moments-in-Time) and the benchmarks
  2. (Optional) pretrain the DiT decoders; the pretrained ones on the Hub ( KBlueLeaf/TTVidT-decoders ) are used by default
  3. Pretrain encoders: one config per table cell
  4. Frozen-probe evaluation, fine-tuning, diagnostics

TT-VidT (TT3D + Diff Compression, decoder S video-pretrained), 2 GPUs

kogine run scripts/train/pretrain_encoder.py -c configs/pretrain/ttvidt_tt3d_diffcomp.py

frozen evaluation of the result: extract -> check -> probe (3 seeds) -> mean +- std

命令
bash scripts/eval/run_frozen_eval.sh ttvidt/ < run_id > /checkpoints/epoch=7.ckpt

To change a setting, write a config that builds on an existing one (see docs/configs.md , which also lists which config belongs to which table).

Repository layout src/ttvidt/ model, training and data code (installable package) model/ DINOv3VidTModel (TT-VidT), VideoMAE3DViTModel, DisMo2DPlus3DModel, MotionDecoder (DiT), frame VAE modules/ Temporal Transfer layers (tt.py: TT1D, tt3d.

py: TT3D), attention blocks trainer.py TTVidTrainer: objectives, optimisation, EMA, logging data/ tar-of-JPEG video datasets, augmentation, decoder-pretraining data hub.

py load training checkpoints / exported models src/vidmet/ benchmark dataset loaders (HMDB51, ARID, IARD, Jester, SSv2, Diving48, EK-100) configs/ _base/ shared recipes (encoder pretraining, decoder pretraining) pretrain/ encoder pretraining: sweep (Table 1), decoder ablation (Table 2), final models decoder/ DiT decoder pretraining eval/ feature extraction scripts/ train/ pretrain_encoder.

py, pretrain_decoder.py, export_decoder.py eval/ extract_features.py, probe.py, finetune.

py, diagnostics, V-JEPA 2 baseline data/ video -> tar conversion; benchmarks/: dataset setup scripts analysis/ analytical parameter / FLOP counts tools/ smoke test, checkpoint export docs/ guides License Apache-2.0, see LICENSE .

DINOv3, the datasets and the V-JEPA 2 baseline are subject to their own licenses. Acknowledgement This work is supported by the NVIDIA Taiwan AI Research & Development Center (TRDC). This work utilize sponsored compute resources from Spellbrush.

适用场景

学习研究
开源项目实践