documentation

Body: motion

A VRM humanoid whose motion is matched from real people talking, composed in layers, gated in the units the viewer sees.

Matched, not played

The corpus is cut into 1-second windows of motion capture, with the loudness envelope and 13 spectral coefficients of the speaker's audio stored beside each window. While she speaks, the matcher searches every second for the window whose audio shape, timbre and starting pose best fit her own voice and where she already is, ranking the three separately and adding the ranks so that units never have to be traded against each other. The chosen window crossfades in on a sine. Continuity comes out of the search, not out of segmenting gestures.

Per-bone gains

Every layer writes a delta over her rest pose and the deltas compose by multiplication, so nothing lands exactly where it was authored unless the layers under it yield. The corpus layer scales each bone: head above one, torso well under, arms by a gain that is deliberately proximal-weighted (the upper arm carries more of the performer's motion than the hand), because real conversational gesture is distal-heavy and a uniform scale leaves only the wrists moving.

Gates in on-screen degrees

Two rules refuse a window: a transition jump larger than 55 degrees, and a peak arm rotation faster than 90 degrees a second. Both are measured after the gains, in the degrees the viewer will see, because the same raw motion is a twitch at one gain and invisible at another. A cut that is forced (the window ended and nothing was in reach) gets a crossfade sized to its jump rather than squeezed into the search cadence, and no new window is picked while a crossfade is running.

Slow springs

The torso and head sit on a spring 3 times slower than the arms. A change of window is a frequency problem, not an amplitude one: lowering a weight shrinks the motion and keeps the reversals; a slower spring keeps the size and removes them. Measured: the head's direction reversals went from over one a second to none at the same size.

The reflex layer, fitted not trained

Under the corpus, a small network of coupled oscillators with state: breath, a weight shift between feet with stance and dwell, torso sway that is the sum of unrelated periods rather than one sine. Its shape was designed; its numbers were fitted from how real streamers actually shift their weight (how long on one foot, how far they lean, whether speech moves them), as histograms a person can read. When the corpus is busy the reflex stands down, on a slow ease of a smoothed signal, because a stand-down that follows a per-window number step for step was measured as a 9 mm hip step once a second.

Held postures

Clasped hands and crossed arms are solved onto her own measured body: a two-bone analytic arm solver with an elbow pole, wrist limits from human ranges, and targets found by searching against her skin map rather than authored in the air. A transition into a held posture curves around her torso instead of through it.