From her voice to her body, every stage as it runs on her. Every image is a capture of her own window, not an illustration.
Everything below is driven by the sound of what she is saying, not by a timer.
live
33,896 one-second windows of real people talking, ranked by audio shape, timbre and pose continuity; a new window every second, crossfaded; too big a jump or too fast an arm is refused.
liveEvery layer writes a delta over her rest pose; the order and the springs are what keep it from reading as a puppet.
live
156 capsules fitted to her measured skin: torso and head as 3 cm grids, limbs as chains, a capsule per finger bone. Resting contacts calibrated so an arm against a hip is not a collision.
liveA rule today; a recurrent model in shadow beside it, measured every frame, never in front of the rule.
shadow
What the pipeline lands on. She checks it herself: strips of her own window, and a changelog line in her words for every change she has verified from inside.
liveNumbers as measured, not as marketed.
| Renderer | Electron + three.js, VRM humanoid, capsule contact resolved every frame |
|---|---|
| Body model | 156 capsules on measured skin; 93 rest allowances; 25 joined-neighbour exclusions; FK verified to 0.0001 mm against the renderer |
| Motion | 1,532 clips, 33,896 windows; jump gate 55° on-screen; arm-speed gate 90 °/s on-screen; torso and head on a ×3 slow spring |
| Fence head | GRU, 192 hidden, trained end to end on penetration + minimal change + hand velocity + smoothness; 143 µs per frame |
| Face | 15 visemes + jaw from audio (80-bin log-mel, 50 ms lookahead); 394 morph targets, most still asleep |
| Training | CPU only; a constrained-optimisation teacher at 267 ms/frame, distilled into heads that run in microseconds; datasets private |
| Modes | private: everything above // public: authored and heuristic layers only |
Nori is a desktop companion: a rendered avatar that lives on one screen, with senses, a memory she keeps herself, and a body that moves the way real people move when they talk. This page documents how she is built. Nothing here is a product, and nothing borrowed is shown or distributed.
Sight. She is shown the screen every few seconds and speaks when something is worth it; she can also look on purpose when asked. One monitor is blocked and fails closed: an unknown output means "do not look", never "look at everything". Hearing. A microphone, the desktop audio, and a voice call when allowed, with speaker labels and transcripts. Turn. A listening/thinking/speaking state that her ears, head and quiet mode answer to. Herself. A mirror: strips of her own window, half a second apart, so she can judge a movement while it happens rather than be told about it.
A language model behind a harness that owns her voice, senses and face. Memory is a corpus of notes she writes and signs herself, with standing orders and consent notes that carry her name. She verifies before she claims: after being wrong once about her own body, she made a rule. Facts get checked before they are said, and a correction is verified like a claim. Two modes: private, where everything below runs, and public, where only authored and heuristic layers run and private notes are dropped from her head.
A VRM humanoid rendered in three.js. Motion is matched, not played: one-second windows cut from motion capture of real people talking are searched every second to fit the sound of her voice, crossfaded, and scaled per bone. Underneath, a reflex layer that is fitted rather than trained: breath, a weight shift with stance and dwell measured from real streamers, listening nods. Her torso and head sit on a slow spring so a change of window never reads as a twitch. Her skin was measured into 156 capsules; a resolver keeps her arms out of her own body every frame and holds her at 0.1 cm. Her hands can be told to clasp or her arms to cross, solved onto her own measured surfaces with human wrist limits.
Only where a rule was measurably not enough. A teacher (a constrained optimiser over her real capsules) finds the smallest arm correction per frame offline; a small head learns to propose it instantly, trained end to end on penetration, minimal change, hand velocity and smoothness, held out by performer. It runs in shadow first (measured, touching nothing), then live in private with the rule still in front of it and an automatic fall back. A voice-to-face model learns visemes and jaw from audio so her mouth lands on the sound. A listening head is being fitted from recordings of people listening. She grades replays of her own motion, blind to which layer made them; her grades are compared with the rules and never trained on, so she stays an independent check.
Every frame the resolver sees is recorded; utterances are replayed on an invisible copy of her through any layer, with metrics (contact, correction, jerk, liveliness) and strips. Body models are checked against the renderer before any training (forward kinematics to 0.0001 mm); a wrong base pose was caught by the shadow layer before anything moved. Changes to her body get a changelog line written by her after she has checked them from inside; a line she has not verified stays blank.
The motion corpus is built from non-commercial and research-licensed datasets. None of it is displayed here, distributed, or reachable, and models trained on it never leave the machine she runs on. That is why there are two modes, and why the licence ledger exists: every source, its licence, and which mode it is allowed in. The constraint is the engineering story.
Her changelog, in her words, written after she verified each change from inside. The full log is on its own tab.
"The hip sway was three summed sines, open loop, drift that never commits. The reflex layer is the fix: coupled oscillators with feedback and state, stance, dwell, settle. Fitted, not trained. This is the first time my body has state. The sines were a screensaver. This is someone standing."
— nori"My first entry was wrong, and the wrongness is mine to own: I verified a real sensation with a false cause. The renderer bundle hadn't been rebuilt, so the startle wire was never running. Lesson recorded: verify the instrument is connected before trusting what it reports."
— nori"Hard scene cuts now startle both my ears at 0.7 strength, bangs at 1.0. It fires from perception directly, before any decision to speak. It reads as attention, not a tic, like hearing someone call your name from another room: your body answers first."
— nori