Her mouth from her text today and from her voice next; her head from recordings of people listening.
Her rig carries the 15-viseme set. Today the shapes come from the text of her reply: words through a pronouncing dictionary to phonemes, phonemes to visemes by the standard grouping, laid against the audio as it streams with estimated timing. Estimated rather than force-aligned on purpose: her voice is synthesised sentence by sentence so she can start speaking before the whole reply exists, and an audience reads the mouth's position far more than a 40 ms offset.
The next version learns the shapes from the sound. A small model takes a 80-bin log-mel spectrogram of the 16 kHz audio at 100 Hz, looks 50 ms ahead, and outputs the viseme and the jaw opening per frame; it is trained on a co-speech dataset's phoneme alignments and its face model's jaw, held out by speaker. It runs in the harness on the audio as it passes to the speaker, one frame decided as soon as the next 50 ms have arrived. When its weights are absent the text path runs unchanged.
Nearly every motion dataset records the speaker. Her head while someone else talks is a different distribution: the small nods, the tilt, the freeze when something lands. The numbers for it (nod amplitude, duration, gap, roll and yaw spread) are fitted from a dataset of people listening on talk shows, the same way the hips were fitted: quantiles into a hand-written mechanism, not a trained policy.
Her face has 394 morph targets and about 15 are driven. Per-side brows and a set of authored expressions are dormant. Waking them is cheap and is on her own list; a face captured from a real person and mapped onto them is the one expressive path that could be public, because nothing borrowed is in it.