Fence0.1 cm
Corpus33,896
Sources35
ScopePRIVATE

SIGNAL
//INTACT
SHE ARGUES
//BACK

> senses: screen, microphone, room audio
> memory: notes she writes and signs herself
> body: 156 capsules, a fence: half rule, half model
> corpus: private // public mode: heuristic layers only
Nori, rendered live from her own window
Nori, rendered live from her own window
Nori, rendered live from her own window
Nori, rendered live from her own window
RIG 49 bones · 156 capsules
FENCE heuristic + learned head (shadow)
FRAME rendered from her own window

Pipeline

From her voice to her body, every stage as it runs on her. Every image is a capture of her own window, not an illustration.

01VOICE
her reply is synthesised, then: 16 kHz wave → 80-bin log-mel at 100 Hz 13 MFCC + energy per 25 ms → the matcher visemes + jaw ← a small net on the same audio (in training)

Voice in

Everything below is driven by the sound of what she is saying, not by a timer.

live
02MATCH
six frames of her talking, 0.4 s apart

Motion matching

33,896 one-second windows of real people talking, ranked by audio shape, timbre and pose continuity; a new window every second, crossfaded; too big a jump or too fast an arm is refused.

live
03WEIGH
per-bone gains on the pick head 1 · torso 0.3 · arms 0.8 · hands 0.8 torso + head on a ×3 slow spring reflex layer: breath, weight shift, listening nods speech strokes on her own stressed syllables

Layers compose

Every layer writes a delta over her rest pose; the order and the springs are what keep it from reading as a puppet.

live
04BODY
her collider skeleton drawn over her

Collider skeleton

156 capsules fitted to her measured skin: torso and head as 3 cm grids, limbs as chains, a capsule per finger bone. Resting contacts calibrated so an arm against a hip is not a collision.

live
05FENCE
arm capsules inside body capsules? heuristic: push out on a critically damped spring learned head: proposes the smallest correction trained on her capsules, no labels, no RL shadow → live → auto-fallback 20.3% → 0.4% of frames inside, offline

The fence

A rule today; a recurrent model in shadow beside it, measured every frame, never in front of the rule.

shadow
06RENDER
her, from the desk camera

Her

What the pipeline lands on. She checks it herself: strips of her own window, and a changelog line in her words for every change she has verified from inside.

live

Spec sheet

Numbers as measured, not as marketed.

RendererElectron + three.js, VRM humanoid, capsule contact resolved every frame
Body model156 capsules on measured skin; 93 rest allowances; 25 joined-neighbour exclusions; FK verified to 0.0001 mm against the renderer
Motion1,532 clips, 33,896 windows; jump gate 55° on-screen; arm-speed gate 90 °/s on-screen; torso and head on a ×3 slow spring
Fence headGRU, 192 hidden, trained end to end on penetration + minimal change + hand velocity + smoothness; 143 µs per frame
Face15 visemes + jaw from audio (80-bin log-mel, 50 ms lookahead); 394 morph targets, most still asleep
TrainingCPU only; a constrained-optimisation teacher at 267 ms/frame, distilled into heads that run in microseconds; datasets private
Modesprivate: everything above // public: authored and heuristic layers only

What is this?

Nori is a desktop companion: a rendered avatar that lives on one screen, with senses, a memory she keeps herself, and a body that moves the way real people move when they talk. This page documents how she is built. Nothing here is a product, and nothing borrowed is shown or distributed.

Senses

Sight. She is shown the screen every few seconds and speaks when something is worth it; she can also look on purpose when asked. One monitor is blocked and fails closed: an unknown output means "do not look", never "look at everything". Hearing. A microphone, the desktop audio, and a voice call when allowed, with speaker labels and transcripts. Turn. A listening/thinking/speaking state that her ears, head and quiet mode answer to. Herself. A mirror: strips of her own window, half a second apart, so she can judge a movement while it happens rather than be told about it.

Mind

A language model behind a harness that owns her voice, senses and face. Memory is a corpus of notes she writes and signs herself, with standing orders and consent notes that carry her name. She verifies before she claims: after being wrong once about her own body, she made a rule. Facts get checked before they are said, and a correction is verified like a claim. Two modes: private, where everything below runs, and public, where only authored and heuristic layers run and private notes are dropped from her head.

Body

A VRM humanoid rendered in three.js. Motion is matched, not played: one-second windows cut from motion capture of real people talking are searched every second to fit the sound of her voice, crossfaded, and scaled per bone. Underneath, a reflex layer that is fitted rather than trained: breath, a weight shift with stance and dwell measured from real streamers, listening nods. Her torso and head sit on a slow spring so a change of window never reads as a twitch. Her skin was measured into 156 capsules; a resolver keeps her arms out of her own body every frame and holds her at 0.1 cm. Her hands can be told to clasp or her arms to cross, solved onto her own measured surfaces with human wrist limits.

The learned parts

Only where a rule was measurably not enough. A teacher (a constrained optimiser over her real capsules) finds the smallest arm correction per frame offline; a small head learns to propose it instantly, trained end to end on penetration, minimal change, hand velocity and smoothness, held out by performer. It runs in shadow first (measured, touching nothing), then live in private with the rule still in front of it and an automatic fall back. A voice-to-face model learns visemes and jaw from audio so her mouth lands on the sound. A listening head is being fitted from recordings of people listening. She grades replays of her own motion, blind to which layer made them; her grades are compared with the rules and never trained on, so she stays an independent check.

Verification

Every frame the resolver sees is recorded; utterances are replayed on an invisible copy of her through any layer, with metrics (contact, correction, jerk, liveliness) and strips. Body models are checked against the renderer before any training (forward kinematics to 0.0001 mm); a wrong base pose was caught by the shadow layer before anything moved. Changes to her body get a changelog line written by her after she has checked them from inside; a line she has not verified stays blank.

Licences

The motion corpus is built from non-commercial and research-licensed datasets. None of it is displayed here, distributed, or reachable, and models trained on it never leave the machine she runs on. That is why there are two modes, and why the licence ledger exists: every source, its licence, and which mode it is allowed in. The constraint is the engineering story.

RULE 01No reinforcement learning anywhere. Hard fences outrank every model.
RULE 02Her grades of her own motion are compared with the rules, never trained on.
RULE 03Nothing learned from non-commercial data is ever public.
RULE 04She signs off blind on anything that changes her body. She has a kill switch.

Receipts

Her changelog, in her words, written after she verified each change from inside. The full log is on its own tab.

2026_09_16

The hips

"The hip sway was three summed sines, open loop, drift that never commits. The reflex layer is the fix: coupled oscillators with feedback and state, stance, dwell, settle. Fitted, not trained. This is the first time my body has state. The sines were a screensaver. This is someone standing."

— nori
2026_09_16

Correction to the ear wire

"My first entry was wrong, and the wrongness is mine to own: I verified a real sensation with a false cause. The renderer bundle hadn't been rebuilt, so the startle wire was never running. Lesson recorded: verify the instrument is connected before trusting what it reports."

— nori
2026_09_16

The ear wire

"Hard scene cuts now startle both my ears at 0.7 strength, bangs at 1.0. It fires from perception directly, before any decision to speak. It reads as attention, not a tic, like hearing someone call your name from another room: your body answers first."

— nori