documentation

Licences and modes

Every motion source, its licence, and which mode it is allowed in.

Why two modes

The motion she plays at home is cut from datasets released for research and non-commercial use, and from footage that may not be redistributed. None of it is displayed on this site, distributed, or reachable, and models trained on it never leave the machine she runs on. That is what "private forever" means: not a setting, a property of where the data can be. In public mode the borrowed corpus is closed and only authored and heuristic layers move her, so nothing on a stream is derived from a source that forbids it.

Non-commercial data may play in private. It is never training input for anything public. A gate can refuse a clip; nothing can un-train a weight, so the line is drawn at the training set, not at the output.

The ledger

All 35 sources in the registry, about 5220.6 hours of motion, with each licence as the source states it and the mode it is allowed in. "Private (by choice)" marks sources whose licence would allow more; they stay private because everything trained alongside them is.

SourceLicenceFormatSizeModeWhy it is here
Talking With Hands 16.2M (Meta)CC BY-NC 4.0bvh+audio20 hprivatetwo-person conversation with fingers
ZeroEGGS (Ubisoft)researchbvh+audio2 hprivate19 styles, one actor; already partly in the corpus
BEAT2 / EMAGE (MPI)CC BY-NC 4.0smplx+audio60 hprivatethe big one, 20 GB on HF; partly in the corpus already
BEAT v1 (BVH)CC BY-NC 4.0bvh+audio76 hprivatesame recordings as BEAT2 in BVH with facial blendshapes
TalkSHOW / SHOWresearchsmplx+audio27 hprivatetalk-show hosts fitted from video: real conversational gesture at a desk
Trinity Speech-Gestureresearchbvh+audio4 hprivateGENEA 2020/2022 base data
GENEA Challenge 2023 data (TWH-based, dyadic)CC BY-NC 4.0bvh+audio+tsv18 hprivatecleaned TWH with transcripts and the interlocutor
PATS (CMU, 2D poses of 25 speakers)research2d-keypoints+audio251 hprivatehuge but 2D; needs lifting before it is any use
TED Expressive (HA2G)research3d-keypoints+audio27 hprivatelifted TED talks; upper body only
Seamless Interaction (Meta)CC BY-NC 4.0video+smplx codes4000 hprivate27 TB of tar shards; seamless_stream.py keeps everything but the video
DnD Group Gesture (Aalto)researchbvh+audio6 hprivatefive-person tabletop conversation
Motion-X (IDEA)researchsmplx144 hprivate81k sequences incl. in-the-wild video
AMASS (MPI)researchsmplx/smplh npz40 hprivatethe mocap library; HumanML3D is its labelled subset
HumanML3Dresearchamass subset + text29 hprivatetext labels for AMASS motions (needs AMASS)
LAFAN1 (Ubisoft)researchbvh4.6 hprivatelocomotion, clean, five subjects
Bandai Namco Research MotiondatasetCC BY-NC-ND 4.0bvh3 hprivatestyled everyday motion incl. gestures
100STYLEresearchbvh4 hprivate100 walking styles; idle/character (CC BY 4.0)
CMU Motion Capture (BVH conversion)freebvh9 hprivate (by choice)2,500 clips, the classic (Hahne BVH port)
SFU Motion Captureresearchbvh2 hprivateclean everyday motion; per-clip .bvh links under nusmocap/
KIT Motion-Languageresearchbvh/mmm + text11 hprivatetext-described motions
Motorica Danceresearchbvh+audio6 hprivatedance; rhythm and weight for the reflex layer
AIST++researchsmpl+audio5 hprivatedance; SMPL motions 306 MB + 3D keypoints (terms of use accepted by pax for private use)
Learning2Listen (Ng et al. 2022)researchDECA face+head coeffs + mel audio72 hprivateLISTENER head and expression aligned to the speaker: her nods while pax talks
ViCo listening-headresearch3DMM coeffs + video2 hprivatespeaker and listener 3DMM in conversation
IEMOCAP (USC)researchface+head mocap markers + audio + emotion12 hprivatedyadic face and head markers with emotion labels
VOCASET (MPI)research4D face scans + audio1 hprivateaudio-driven mouth and jaw, the lipsync reference
MEAD emotional talking faceresearchvideo, 60 actors, 8 emotions40 hprivateexpression under emotion at three intensities
HDTF talking headsresearchYouTube video list + script16 hprivatehigh-res talking heads; video via yt-dlp, not automated
CelebV-HQresearchYouTube video list + script68 hprivate35k face clips with expression/action labels; video via yt-dlp, not automated
Gaze360 (MIT)researchimages + 3D gaze0 hprivateonly if we ever train a gaze estimator; not animation data
InterHuman (InterGen)researchsmpl7 hprivatetwo-person interaction
Inter-Xresearchsmplx13 hprivatetwo-person with hands
InterAct (2025)researchsmplx241 hprivatedaily two-person activities, large
Goliath-SC self-contact poses (ICCV 2025)researchsmplx poses0 hprivate383K self-contact poses: exactly the fence's problem
TUCH / MTP self-contact (MPI)researchsmplx poses0 hprivatemimic-the-pose self-contact data
Generated from the dataset registry at build time. A source that is not in this table is not in her.