Character3D Plugin
The Character3D plugin provides a framework for creating animated 3D characters in Clayground applications. It features a modular body part system, procedural animation capabilities, and integrates with the Canvas3D toon shading system for stylized cartoon characters.
Where to look at it
Every aspect of a character has a scenario in the character lab,
labs/character-101 (./build/bin/claydojo --sbx labs/character-101/Sandbox.qml):
the builds, the walk cycle sheet, the gesture set, the two whole-body actions,
the loadable move set, the hands, the six faces, the head’s detail tiers, the
lip-sync tiers, a listener and a crowd. Each scenario says what it is for on
its card, answers scene.report() headless, and records its numbers into a
run record - labs/kits/character/README.md is the contract, and
labs/character-101/paper.md the lesson.
Getting Started
To use Character3D components, import the module in your QML file:
import Clayground.Character3D
Core Components
- Character - Base component managing body parts and animations with extensive dimension properties
- ParametricCharacter - High-level parameters (bodyHeight, realism, maturity, femininity, mass) that auto-calculate dimensions
- RatioBasedCharacter - Dimension ratios for fine-tuned proportion control
- CharacterEditor - Visual editor overlay for character customization with persistence
- MoveSet - A loadable set of named moves played onto one character, loaded on demand instead of carried by every character
- MartialArts - The move set that ships with the plugin: fourteen moves from a stance to a knockdown and a get-up
- Performance - A character directed from one script string: cues for pointing, presenting, marks, emotions and timing, with
cueFiredfor choreography - Gait - Walk and run derived from a preset and eleven factors, driven by emotion, age, gender and build
- Speech - Voice output (text-to-speech or wav/mp3) with approximate lip-sync
- ThoughtBubble - Simple text bubble for speech/thought display
Character.roundness chamfers every box in the figure (Box3D.bevel underneath): 0 is the hard-edged original, about 0.3 nearly spherical, for no extra draw calls.
Usage Examples
Basic Character
import QtQuick
import QtQuick3D
import Clayground.Canvas3D
import Clayground.Character3D
View3D {
anchors.fill: parent
PerspectiveCamera {
position: Qt.vector3d(0, 200, 400)
eulerRotation.x: -20
}
DirectionalLight {
eulerRotation.x: -35
castsShadow: true
shadowFactor: 78
shadowMapQuality: Light.ShadowMapQualityVeryHigh
}
Character {
y: 0
activity: Character.Activity.Idle
}
}
Parametric Character Creation
ParametricCharacter {
name: "hero"
bodyHeight: 10.0
// Body shape
realism: 0.3 // Cartoon-like
maturity: 0.7 // Adult
femininity: 0.3 // Masculine
mass: 0.5 // Average
muscle: 0.7 // Athletic
// Face
faceShape: 0.5
eyes: 1.2
hair: 0.8
// Colors
skin: "#d38d5f"
hairTone: "#734120"
topClothing: "#4169e1"
bottomClothing: "#708090"
}
mass and muscle reach the body as a belly and a chest, not only as a
width. The trunk is two boxes on a waist joint (see
The trunk is two segments), so mass bulges the
belly forward over the hip and muscle deepens the chest, and a build shows
in the shape of the body rather than in how wide all of it is. Both are
exactly neutral at 0.5, and muscle also draws a waist in that a single
tapered box could not make.
The same three dials are on Character directly for a body built by hand:
| property | default | what it does |
|---|---|---|
bellyRatio |
0.45 | the belly’s share of torsoHeight; the waist joint sits between the two segments |
bellyBulge |
1 | how far the belly swells past the plain trunk taper - mostly depth, and forward. 1.3 is a gut |
chestSwell |
1 | how much deeper the chest is. Depth only: the chest’s width is shoulderWidth |
waistPinch |
0 | how far in the waist joint is drawn |
bellyColor, chestColor |
torsoColor |
either segment can take its own colour |
At every default the two boxes trace exactly the tapered box the torso used to be: same shoulders, same waist, same depth, same height. The only thing that is drawn and was not is the seam at the joint.
Character with Movement
ParametricCharacter {
id: player
name: "player"
// Activity controls animation
activity: isMoving ? Character.Activity.Running : Character.Activity.Idle
// Movement derived from animation geometry
property bool isMoving: controller.axisX !== 0 || controller.axisY !== 0
// Move based on currentSpeed (auto-calculated from animation)
x: x + controller.axisX * currentSpeed * dt
z: z - controller.axisY * currentSpeed * dt
}
Character Editor Integration
import Clayground.Character3D
import Clayground.GameController
Item {
View3D {
id: view3d
anchors.fill: parent
ParametricCharacter {
id: character1
name: "char1"
}
ParametricCharacter {
id: character2
name: "char2"
x: 20
}
}
GameController {
id: gameController
Component.onCompleted: selectKeyboard(
Qt.Key_W, Qt.Key_S, Qt.Key_A, Qt.Key_D,
Qt.Key_Shift, Qt.Key_Space
)
}
CharacterEditor {
anchors.fill: parent
characters: [character1, character2]
view3d: view3d
gameController: gameController
enabled: true
}
}
Facial Expressions
Six of them, and they are meant to be told apart at a glance rather than studied: the mouth first (a smile, a frown, a shout, a sneer or an O), the brows’ angle second, the lids last. Which lid moves is not interchangeable - up from below is pleasure or revulsion, down from above is a glare or a droop, and neither of them is surprise.
| expression | Head.Activity |
setEmotion |
the mouth | the brows | the lids |
|---|---|---|---|---|---|
| neutral | Idle |
"neutral", "" |
flat | level | open |
| joy | ShowJoy |
"happy" |
an open grin | up, flat | squint, from below |
| sadness | ShowSadness |
"sad" |
small, fully down | up, inner ends in | hooded |
| anger | ShowAnger |
"angry" |
open and wide - a shout | down into a V | hooded |
| disgust | ShowDisgust |
"disgust" |
a one-sided sneer | one up, one down | squint |
| surprise | ShowSurprise |
"surprised" |
a round O | high | wide open |
Disgust is the only one whose halves disagree, and that is deliberate: made
symmetric it is a quieter anger and nothing else. It is worth two parameters
of its own - Head.mouthSkew and the brow skew behind it - which anything
can drive without an emotion.
Character {
id: character
// Set facial expression
faceActivity: Head.Activity.ShowJoy
// Animate expressions
SequentialAnimation on faceActivity {
loops: Animation.Infinite
PropertyAnimation { to: Head.Activity.ShowJoy; duration: 2000 }
PropertyAnimation { to: Head.Activity.Idle; duration: 1000 }
PropertyAnimation { to: Head.Activity.Talk; duration: 2000 }
PropertyAnimation { to: Head.Activity.Idle; duration: 1000 }
}
}
Eyes: blinking, gaze and thinking
On by default, and it is what stops a face from reading as a mannequin:
Character {
id: npc
autoBlink: true // default
gazeBehaviour: true // default
blinkSeed: 7 // give each of a crowd its own, or they blink in step
}
npc.lookAt(player.scenePosition) // eyes first, head after
npc.thinking = true // eyes leave the target and settle off-axis
lookAt() aims the head; GazeAnim aims the eyes inside it, and the
difference in when they arrive is the whole effect. The target is mapped
into the head’s own frame, so what comes back is the angle the head has
not covered yet — large while it is still easing round, large again when
the target is past its 65° limit, and nothing once it has arrived. Point
the eyes at that residual and they lead on the way out and re-centre on
arrival, with no second animator racing the first.
thinking is the most legible signal a boxy face has for working
something out — there is no brow furrow to read at ninety pixels. Set it
around the gap between being asked and answering.
Everything idle here is deterministic for a given blinkSeed: the
blink spacing, the wander, the micro-saccades and the direction of an
aversion. Two runs of a sandbox render identically, which keeps a
clayrender comparison meaningful; two characters with different seeds
do not, which is what a crowd needs.
Off at Detail.Minimal regardless, where there is no eye left to move.
Listening
The other half of a conversation:
npcB.listeningTo = npcA // hold A's face, break away now and then,
// mark the ends of A's phrases
npcB.listeningTo = null // done
Everything else in this plugin describes a character while it speaks. Without this the one who is not speaking does nothing at all, which is what makes two characters talking read as two monologues taking turns.
Phrase boundaries are read off the speaker’s mouth, not its script: a
gap in Speech.mouthOpen while it is still speaking ends a phrase,
whatever produced the timeline. So it works on an unknown recording read
by the envelope tier exactly as it does on an aligned one — which is what
makes it usable on dialogue nobody wrote down.
listeningTo owns the look target while it is set; it and lookAt() are
the same channel by construction.
The head does two things at once
A nod has to happen while an aim holds, and be given back without the aim having been forgotten. So the head is the one joint that does not own its own rotation:
| Channel | Driven by | For |
|---|---|---|
Head.poseEuler |
the body animators, via HeadEulerAnim |
where the head is aimed |
Head.offsetEuler |
anyone | a momentary rotation on top |
Head.nod(deg, times) |
— | the built-in one |
eulerRotation is the sum, and a binding. Animating a head’s
eulerRotation directly writes to that sum and will be overwritten the
next time any part of it changes — animate poseEuler instead. Every
other joint is unchanged and still uses EulerAnim on eulerRotation.
Speech with Lip-Sync
Characters can speak text (via text-to-speech when available) or play recorded audio (wav/mp3) - the mouth movement approximates the speech in both cases:
Character {
id: npc
Component.onCompleted: {
// Text: spoken aloud when a TTS engine is available,
// otherwise the mouth animates silently at an estimated pace
npc.say("Hello! Welcome to Clayground.")
}
}
// Recorded dialog line - the mouth follows the recording
npc.say("dialog/intro.wav")
// Emotional conversation: colors face, voice (TTS pitch/rate) and -
// while the character is idle - body language gestures
npc.say("I lost my favorite shovel...", "sad")
npc.say("We found the treasure!", "happy")
npc.say("Give it back right now!", "angry")
npc.say("You want me to eat THAT?", "disgust")
npc.say("It was here a second ago!", "surprised")
// Inline annotations switch the emotion mid-speech
npc.say("*angry* Get off my ground immediately! " +
"*happy* Just a joke - come in and have a cup of tea with me.")
// Body language is optional: disable it (or just keep the character
// walking/fighting) and only face and voice carry the emotion
npc.speechBodyLanguage = false
// How closely a recorded line is read. Spectral (the default) measures
// formant bands, so vowels get their own shapes; Envelope reads loudness
// only and is the floor everything else falls back to.
npc.speechAccuracy = Speech.Envelope
console.log(npc.speech.effectiveAccuracy) // what the last line ACTUALLY got
// Advanced configuration
npc.speech.rate = 0.2 // a bit faster
npc.speech.volume = 0.8
npc.speech.finished.connect(() => console.log("done talking"))
The mouth is driven by continuous shape parameters on Head
(mouthOpen, mouthWide, mouthRound - readonly outputs - plus the
writable mouthCornerLift). Emotions keep control of the mouth corners
while speaking, so characters can smile and talk at the same time.
For fully manual mouth control, assign any object with speaking,
mouthOpen, mouthWide and mouthRound properties to
head.speechSource.
How closely a recording is read
say() with a file decodes it in full before playback starts, and how
hard it looks at what it decoded is speechAccuracy:
| Tier | Reads | Gets |
|---|---|---|
Speech.Envelope |
loudness, zero crossings | how far the jaw dropped |
Speech.Spectral (default) |
formant bands | which shape - open/closed, spread/rounded, fricative |
Speech.Aligned |
the above, plus a transcript | the script’s own shapes on the recording’s clock |
Aligned is the one that needs something from you:
npc.speechAccuracy = Speech.Aligned
npc.say("dialog/intro.wav", "", "Hello! Welcome to Clayground.")
Given the script, the sequence of shapes stops being a guess - the text
says there is an /m/ there, so the mouth closes. No acoustic tier can
do that reliably, because a bilabial is voiced: every measurement of
one says “loud”, not “shut”. Only the timing then comes from the audio,
by dynamic-time-warping the script against the measured frames.
A transcript that does not match the recording is worse than none - it would drag the mouth confidently through the wrong syllables for a whole line - so a pairing whose durations disagree by more than about 2.5x is rejected and the tier below takes over.
Alignment also carries word marks onto the recording, so currentWord
and per-word callbacks work for recorded dialogue, not just for TTS.
The tiers are not a performance dial. A 512-point transform every
16 ms is a few hundred thousand flops for a whole line, once, which is
below the noise floor of the decode that produced the samples. What
separates them is time-to-first-sound, the samples held while the
analysis runs, and - for Aligned - whether anyone wrote the line down.
That last one is an authoring cost rather than a runtime one, and it is
the real reason a lecture and a barked NPC line want different answers.
No tier leaves the mouth dead. A recording the analyser cannot read
falls back to the one below on its own, and Envelope is the floor.
Read speech.effectiveAccuracy for what the last line actually got -
asking for Spectral and getting Envelope back is a normal outcome,
not an error.
What Spectral assumes is one close-miked speaker on a reasonably dry
recording. Against a music bed the formant bands read the instruments,
and heavy reverb fills in the gaps that mark a closure. Both degrade to
the envelope rather than to a guess, per-frame, via a confidence gate.
The speech scenario of labs/character-101 puts the three tiers on one
recording side by side (labs/kits/character/SpeechRow.qml).
Gestures
Walk, run, idle, working and boxing are cycles. A gesture is the other kind of
animation: a pose that eases in, is held for as long as it is wanted,
and eases back. Both are driven from Character:
Character {
id: prof
view: view3d // so Auto can see how big it lands on screen
Component.onCompleted: {
prof.turnTo(board.scenePosition) // whole body, shortest way round
prof.pointAt(stone.scenePosition) // held until told otherwise
prof.setEmotion("happy") // face, until changed again
}
}
// once the arm has arrived, talk about it facing the reader
if (prof.gestureSettled) {
prof.lookAt(camera.scenePosition)
prof.gesticulate()
prof.say("This is the stone I meant.")
}
prof.stopGesture() // eases everything back to rest
| verb | what it does |
|---|---|
pointAt(worldPos, which) |
Holds a point at a scene position. which is "auto" (default), "left" or "right". |
presentAt(worldPos, which) |
Offers an open hand toward a scene position - palm up, at chest height, elbow bent - and holds it. Same which as pointAt. |
thumbsUp(which) |
Holds a thumbs up; "right" by default. |
gesticulate() |
Two-handed talking gesture, looping until stopped. |
stopGesture() |
Eases every held joint - and the head - back to rest. |
lookAt(worldPos) |
Aims the head only; outranks the gesture’s own head aim. null releases it. |
turnTo(worldPos) |
Turns the whole body on the spot; changes the resting orientation. |
setEmotion(name) |
A lasting face: "happy", "sad", "angry", "disgust", "surprised", "neutral"/"". |
What to assert on, rather than watching:
| property | meaning |
|---|---|
gesture |
"point", "present", "thumbsUp", "talk" or "". Set the moment a gesture is asked for. |
gestureSettled |
The pose has arrived. Measuring joint angles before this reports the pose being left. |
gestureHand |
"left", "right", or "" while released or talking. |
emotion |
The lasting face, as distinct from speechEmotion, which belongs to one line. |
Point or present? A point is for one thing: the finger goes on it, and
the arm reaches as far as the target asks. A present is for a group or an
area - several parts, a whole circuit - where a finger at the centroid would
point at bare board. The hand stays in front of the body at chest height,
palm up, and only turns toward the target; the head and the body turn as they
do for a point. Something far above or below the hand is a point’s job.
The busy hand takes the open pose.
Two rules the layer depends on:
- Idle only. Gestures run only while
activity === Character.Activity.Idle, and every verb above is ignored otherwise. Setting any other activity drops the gesture immediately and hands the joints to that activity’s animation. This is what keeps two animators off one joint;IdleAnimand the speech body language stay switched off for as long as a gesture is held. safeSilhouette, on by default. A straight arm raised forward reads as a fascist salute, so a point that aims high forces the elbow to bend and lets the forearm do the reaching. The aim is unaffected - the bend is given back through the shoulder and the wrist. Turn it off for a character whose job is that shape (a salute, a hand-raise, a throw).
One trap when reacting to the layer: gesture and gestureSettled are
properties, and a handler on their change notification runs while that
notification is still being delivered. Calling a mutating verb
(stopGesture(), another pointAt(), an activity switch) synchronously
from such a handler therefore logs a QML binding loop - harmless but noisy.
Defer the reaction with Qt.callLater(...), or react from a Timer, as
the professor kit’s FlowGuide does.
Standing, working and boxing
Everything the arms do that is neither a walk nor a held gesture lives in one
Qt-free model, animation/action.js, the way the walk and the run live in
animation/gait.js:
| what | where it is | who plays it |
|---|---|---|
| standing still | REST |
IdleAnim, and GestureAnim when it releases |
| working at something | base use |
UseAnim |
| boxing | base fight |
FightAnim |
UseAnim and FightAnim are ActionCycleAnim with one property set. Unlike
the gait cycle, which spells its poses out as animations and keeps a matching
poseAt() beside them, an action cycle animates ONE number - the phase - and
writes what actionPoseAt() answers for it. There is no second copy to keep in
step: the frozen boxing and working columns of the character lab’s gesture
sheet (labs/kits/character/GestureSheet.qml) are the same function the
shipped cycle plays.
Character {
activity: Character.Activity.Using
workHeight: 0.2 // 0 a table at the waist, 0.5 a counter, 1 a shelf at head height
actionIntensity: 0.7 // amplitude, tempo and how much of the body joins in
}
| property / method | meaning |
|---|---|
actionIntensity |
0..1. A harder fight is a faster one with a tighter guard and a bigger bounce; harder work is bigger, quicker, and past six tenths one hand holds while the other hits. |
workHeight |
0..1, Using only. A posture, not a hand height: the back rounds over a table, the forearms angle up to a counter, the back arches and the head comes up at a shelf. |
actionHandPose |
What the running activity wants the hands to be doing, "" when none does. A fist while boxing. |
actionPoseAt(action, t) |
The joint angles at phase t, with nothing running. Pure. |
actionTable(action) |
The derived numbers the cycle is replayed from, cycleMs among them. |
applyActionPose(action, t) |
Freezes an idle character at that phase - what the sheets draw. |
Boxing is an amateur’s, on purpose, and orthodox: the left leads. One cycle is jab, jab, cross - short, short, LONG - and then the guard, which bounces on the knees and rolls a little around its blade until the next. The guard is what the whole thing is judged on, and the five things that make it read as a guard are kept whatever else moves: both fists above both elbows, both elbows below the shoulders and inside the ribs, the fists at the cheeks in front of the face, a bladed and staggered stance on bent knees with the rear heel up, and the hand that is not punching welded to the cheek. The jab barely winds up; the cross draws back, turns the hips a third of a turn and the shoulders further, leans in past what a professional would, and the trunk comes home before the arm does. The head turns back part of the blade to look at the opponent, and drops behind the shoulder on the cross.
Working is generic on purpose - it has to pass for cooking, tinkering,
sorting and typing alike - so it is built from what those share, which is a
rhythm rather than a stroke. A cycle is four beats: the lead hand reaches for
something, both hands work at it in short strokes, the lead hand presses or
places it, and the body settles and glances up. The two hands are never level
and never mirrored (the lead sits ahead and above; the off hand does two thirds
as much and lags by four tenths of a stroke); the head leads the reach and
lags the press; and every fourth cycle the glance is a proper look up, off
ActionCycleAnim.cycle, so the loop is not noticed as one.
actionHandPose sits between a gesture and gaitHandPose in the chain that
decides a hand’s shape, and it is why a punch is thrown with a closed hand:
nothing on the Fighting path could reach handPose before it, so the boxing
cycle ran with the fingers open and read as clawing.
node plugins/clay_character3d/animation/action.test.js checks the model -
that a guard keeps both fists above the elbows and both elbows below the
shoulders, that the rear hand does not move for a jab, that a cross winds up
further and turns the trunk more than a jab, that the working loop has beats
and its two hands are never in step. It runs under ctest as
node_character3d_action.
Where to look at it. The action scenario of labs/character-101
(labs/kits/character/ActionStage.qml) poses one figure from the lab’s clock
through the same pose model, with a see-through bag at a straight’s reach or
a table under the hands; the lab’s transport pauses and steps the clock to
freeze a frame, and the scene’s report() measures each fist against its
own shoulder and against the chin in head heights, which is what a guard is
a claim about:
clayrender labs/character-101/Sandbox.qml --size 1400x900 --paused \
--eval 'applyScenario("action"); act("action", ["fight"])' \
--wait-for 'sceneReady' --result - --eval 'JSON.stringify(scene.report())' \
--out /tmp/cross.png
One sign in the model was measured there rather than reasoned: a positive Y
rotation on an upper arm carries a forward-pointing forearm OUTWARD on the
right side, so arm() negates the yaw against the side. Written the other way
round, the rear fist of the guard sat a head and a half outside the face and
looked right in every sheet that had no reference to measure it against.
Loadable move sets
Everything above is the basic set - what a character IS. Walking, running,
standing, gazing, listening, gesturing, talking, working and boxing are wanted
by every game, so they are always resident on every character. A move set is
what a character KNOWS: a martial art, a dance, a trade’s hand-work, wanted by
one game and not the next, so it is loaded onto a character on demand, replaces
whatever set was loaded before it, and is unloaded again by clearing
moveSet. A set runs only while the activity is idle, exactly as a gesture
does.
ParametricCharacter {
id: fighter
moveSet: "martial arts" // or the URL of a MoveSet of your own
Component.onCompleted: fighter.playMove("stance")
}
The shipped set offers fourteen moves: stance, step, guard, jab,
cross, uppercut (standing), lowGuard, sweep (crouched), frontKick,
roundhouse (kicks), jumpPunch, jumpKick (airborne), knockdown, getUp
(ground). knockdown holds its last frame on the floor until stopMove() or
another move releases it - which is what getUp starts from.
What a set offers and what it is doing is readable from the character
(moves, moveSetName, activeMove, movePlaying, moveHolding,
moveFinished); see Character’s moveSet documentation for the whole API.
CharacterEditor has a “Moves” section that loads a set and plays any of its
moves, and demo/Sandbox.qml binds the same to the keyboard: J loads and
unloads the set, , and . step through the moves, V plays the selected
one and B stops.
The gesture sheet
The sheet for everything above and for the held gestures with it: the
gestures scenario of labs/character-101 (labs/kits/character/GestureSheet.qml)
puts them side by side, one frozen figure each, same light, same angle,
labelled — the gait cycle sheet’s trick applied to poses rather than to
phases.
clayrender labs/character-101/Sandbox.qml --size 2200x900 --paused \
--eval 'applyScenario("gestures")' --wait-for 'sceneReady' \
--out /tmp/gestures.png
The set is the thing being judged, not any one pose. A gesture looked at on its own is looked at against a memory of the last one, and a memory grades generously — which is how a fist that folded back past its own knuckles survived for as long as it was only ever seen one at a time.
silhouette=true takes the lighting and the colour away and leaves the
outline, which is all a gesture has at any distance; scale shrinks the
figures in place for the small-on-screen read. Every pose on it is frozen and
deterministic — the cycles from applyActionPose(), the aimed gestures through
the real solver with its settle cut to a frame — so ready is the property to
wait for and two renders across a change are comparable.
For one hand very close up, the hands scenario (labs/kits/character/HandBench.qml)
is still the bench: the sheet answers “is this recognisable”, the hand bench
answers “is this a hand”.
CharacterEditor carries the same set as chips, for turning a knob and looking.
Where this workflow lives. In labs/kits/character/, as scenes of the
character lab, and not
as a lab under labs/. A lab is a teaching artifact with an authoring contract
to match — a paper, a .grafli overview, EN and DE strings from the first
commit, committed .labrec records — and it is aimed at a reader learning a
domain. A character-tuning rig has no teaching content and exactly one
audience: whoever is editing this plugin. It also needs to sit beside the
component it is judging, because the two are edited in the same breath. The
benches build nothing, cost nothing and are already how the gait, the face and
the hand are judged; a fourth of them is the cheap, consistent answer. If a
character-tuning LAB is ever wanted — expressions, sprites and animation review
for a reader rather than for a maintainer — it is a different artifact with a
different audience, and it can be built on top of these benches rather than
instead of them.
Hands and faces, and how much of each to draw
handPose says what the hands are doing - relax, open, point,
thumbsUp, fist - and a gesture overrides it for as long as it holds them.
Both levels of detail answer it.
detail says how much hand, and how much face, to spend on that:
Character.Detail |
what is drawn | draw calls |
|---|---|---|
Minimal |
one box per hand; the head keeps its skull, its hair and a drawn face, and loses its nose, its ears, the pupil highlights, the brows and the mouth corners | ~20 |
Low |
the whole body, one box per hand, reshaped per pose — a fist is a stubby block, an open hand a long flat one; the face keeps its irises and loses its brows and ears | ~21 |
High |
ten boxes per hand as well: four fingers and a thumb that fold, so a point extends a real index finger; the whole face | ~43 |
Auto (default) |
picks between the three by how tall the character lands on screen |
The head is nine boxes at High — a cranium, a jaw, four of hair, a nose and
two ears — seven at Low and six at Minimal. It used to be nineteen,
because the eyes, their irises, their brows and the four pieces of the mouth
were all boxes standing in front of it. They are drawn into the head’s own
surfaces by a fragment shader now, and cost nothing.
That has flattened the top of this table on purpose. Minimal used to be much
the cheapest level because it deleted the face, thirteen of a character’s
thirty-three draw calls; now it saves one box over Low, and the only real
saving left between levels is the twenty boxes of fingers. What the two cheap
levels buy is no longer speed - it is a face that stays legible at twenty
pixels rather than one that shimmers. No level removes a face any more: a
character without one reads as broken rather than as distant.
Auto crosses into High at detailThreshold (240 px of figure) and drops to
Minimal below minimalThreshold (60 px). Both are measured rather than
picked — see their docs. Each character decides for itself, so a crowd pays
only for the ones that are near.
Auto needs view - a character cannot ask how big it looks without knowing
what it is being looked at through - and stays Low without one. It biases
toward fingers while a gesture is shaping the hands, since that is what they
are for, and it has a hysteresis band so a character drifting across the line
does not grow and shed ten boxes a hand every few frames.
detailedHands is read-only and reports which one is on screen right now.
Making a gesture readable
Two properties, and they work together:
Character {
gloves: true // hands get their own colour, and a cuff at the wrist
handScale: 1.45 // and they are drawn bigger than the tables give
}
The oldest trick in cartoon animation, and it is about legibility rather than costume. A hand the colour of the arm it is on has to be found before it can be read, and a hand the colour of the background cannot be found at all — so the glove is a single high-contrast shape that separates from both. The cuff is the half that does the separating: a pale hand is just a pale hand, the band across the wrist is what says where the arm stops.
handScale scales the wrist joint, so the hand grows out of the cuff instead
of drifting off the end of the arm — and it scales the whole hand rather
than lengthening the finger, because stretching the one part that has to stay
legible is what produces a spike where an index finger should be.
detail accounts for it: bigger hands mean the fingers are worth drawing from
further away, so the Auto threshold divides by handScale.
A third knob, and this one is about not being noticed. ParametricCharacter’s
two width sliders scale the arm over a spread of two and a half from thin and
unmuscled to heavy and muscular, and the palm is a fixed fraction of the arm —
so the hand used to take all of it, which came out as claws on one figure and
mittens on the other. handBuildResponse (0.5 by default) is how much of the
build the hand takes: at 1 it is glued to the arm as before, at 0 it is the
same hand on every body. Only the cross-section — hand length follows the
arm’s length, which is a matter of maturity. tests/qml_head/tst_build.qml
pins it, and the hands scenario of labs/character-101 takes a build and
the handBuild knob so the sweep can be looked at:
clayrender labs/character-101/Sandbox.qml --size 1400x900 --paused \
--eval 'applyScenario("hands"); act("build", ["thin"]); act("arm", ["level"]); act("pose", ["open"])' \
--wait-for 'sceneReady' --result - --eval 'JSON.stringify(scene.report())' \
--out /tmp/thin.png
report() carries palmArm, which is the number the question is actually
about: 1.05 at every build with the response at 1, and 1.41 / 1.05 / 0.88
across thin / neutral / heavy at the default.
The two levels are built to match in outline, so the switch is meant to go
unnoticed; the hands scenario of labs/character-101 is where that is
checked, and its fingers verb flips them on one character without moving
anything else.
Tuning a hand pose. n on that bench opens a slider per field of the row
the current pose resolves to — the four curls, the fan, the five thumb numbers,
and the fold shape shared by every finger. They are written onto the right hand
as you drag, 0 puts back what ships, and k prints the row in exactly the
form DetailedHand’s table is written in, ready to paste:
if (name === "fist")
return { i: 1.00, m: 1.00, r: 1.00, l: 1.00, sp: 0.00,
tx: 120, tz: 10, tc: 0.45, tl: 1.15, toff: 0.55 }
Every number in that table was arrived at by looking, and looking is done with
the hand in front of you rather than in an editor with a rebuild between each
guess. The channel is Arm.poseOverride → DetailedHand.poseOverride, a
partial replacement of the pose row; it is a debug channel and nothing ships
with it set.
How much character to draw
detail is Character.Detail.Auto, High, Low or Minimal. Auto measures
how big the character lands on screen and picks between the other three; it
needs view set, and stays Low without one.
Pin it for a character the camera lives on — a player above all. Auto is a
policy about distance, and a character that is always in close-up has no
distance to decide anything about; CharacterEditor’s Detail row does it
by hand, and shows what Auto currently resolves to next to what it was asked
for.
Auto measures the character’s apparent size off whichever of its three axes is
least foreshortened, not off its height alone. That is not a refinement: a
camera looking along a character’s own length — up at it from the floor, down
at it from above — projects a ten-unit body to a few pixels, and measuring the
body axis alone said tiny about a figure filling the screen. Measured at a
fixed sixteen units, a figure that was High at eye level fell to Low by 70
degrees of camera pitch and to Minimal by 85, up and down alike. The two
horizontal axes are only projected when the body axis has already gone short,
so a character that is plainly close enough costs nothing extra, and at eye
level the horizontal estimate never wins — the distance thresholds are exactly
what they were.
The face, and how it is drawn
The eyes, their lids and irises, the brows and the mouth are not geometry. They
are signed distance fields evaluated in a fragment shader and drawn into the
front of the two head boxes - bodyparts/FaceBox.qml carries the material,
face3d_main.glsl draws the shapes. A face therefore costs no draw calls and
no vertices at all.
That buys three things a face built from boxes could not have. An eye is a
marking on a head rather than an object in front of one, so it no longer shows
its own side wall at twenty degrees off axis. Head.gaze aims the irises
without moving the head, which sliding a built iris sideways could never do -
it would carry the iris off its own eyeball. And Head.autoBlink is one
animated float rather than a pair of boxes resized every frame, which is why
the eyes never blinked before.
The heads scenario of labs/character-101 (labs/kits/character/HeadRow.qml)
shows all three detail levels side by side with named shots, a blink, a gaze
and a talking mouth, and a readout giving the head and eye size in pixels -
so “still readable at ninety pixels” is a claim that can be checked rather
than an impression.
Face anchors
Head publishes where its features are, in the head node’s own frame, so
accessories parented to character.head (beards, spectacles, hair) do not
restate its layout arithmetic and then drift from it:
faceOffsetZ, faceFront, faceBack, jawFront, upperHeadBottom,
crownTop, eyeLine, eyeWidth, eyeSpacing, eyeRelief, noseBottom,
earPos, earSize, earTop, hairOuterX, mouthLine, mouthWidth,
mouthBottom, chinBottom.
mouthLine stays put while the jaw stretches open; mouthBottom and
chinBottom move with it. eyeRelief is how far the eyes stand proud of
faceFront - zero, now that they are drawn rather than built, which is what
lets a spectacle rim settle onto the face instead of being pushed clear of a
pair of protruding cubes.
Use them. The professor kit spent a long time re-deriving all of this by hand
from the six head dimensions - fifteen-odd constants copied out of Head.qml -
and every one of them was a place a beard could slide off a chin with nothing
raising an error. These anchors also went unread for long enough that a binding
loop sat undetected in one of them: an anchor nothing evaluates is an anchor
nothing checks.
Character publishes rightShoulderPos, leftShoulderPos and headPos in
character-local coordinates for the same reason.
Gait
Walk and run are one cycle, GaitCycleAnim, animated from a table that
animation/gait.js derives from a base - the walk or the run as it was
authored - and thirteen factors around neutral. A factor that scales an amplitude
is 1 at neutral, one that offsets is 0, and the all-neutral gait is the walk
and run the framework always had, to the digit: gait.test.js asserts the
derived table against the formulas WalkAnim and RunAnim used to carry, so
a retyped digit cannot pass as neutral.
Three sources feed a character’s gait, and each defaults to nothing:
ParametricCharacter {
maturity: 0.9 // the build: an elderly shuffle
gait: Gait { preset: "elderly"; tempo: 1.2 } // the author: elderly, a fifth quicker
Component.onCompleted: setEmotion("sad") // the mood: slows and slumps the walk too
}
- The build. A
ParametricCharacterhandsmaturity,femininity,massandmuscleto the gait model asgaitBuild. - The emotion. The same channel the face uses: a spoken line’s emotion
while it is spoken,
emotionotherwise.setEmotion("sad")slows the walk, shortens the step and hangs the head as well as pulling the face. - A
Gaitobject with a preset, factors, or both.
Character.gaitFactors is the three folded into one vector: multiplicative
factors multiply, additive ones add, and the result is clamped once. There is
no override layer on purpose - preset: "elderly" with tempo: 1.2 is
“elderly, a fifth quicker” and reads that way, and a sad and heavy character
is slower than either alone. gaitFromBuild and gaitFromEmotion switch the
first two sources off:
Character {
gaitFromBuild: false // the walk ignores the body
gaitFromEmotion: false // and the mood
}
Speed follows the feet, whatever the gait. walkSpeed, runSpeed and
currentSpeed come out of the derived table and the leg height with the
formula the old cycles used, so a controller that moves the character by
currentSpeed needs no change: a shorter stride or a slower tempo covers
less ground, and the feet do not slide.
The factors, as Gait exposes them:
| factor | kind | neutral | what it moves |
|---|---|---|---|
tempo |
multiplies | 1 | cadence; 1.2 takes steps a fifth quicker and covers ground a fifth faster |
stride |
multiplies | 1 | how far the legs swing, and the speed with it |
bounce |
adds | 0 | how high the whole figure rises at mid-step, in leg heights; 0.1 is the cap |
lean |
adds | 0 | trunk pitch in degrees on top of the cycle’s own, bending at the waist so the legs stay planted; positive forward, negative chest out |
spineCurve |
adds | 0 | how round the back is: the angle between belly and chest. Differential, so it changes the shape of the trunk without moving the head - positive rounds it forward, negative arches it and lifts the chest |
headPitch |
adds | 0 | head pitch in degrees; positive looks down, negative lifts the chin |
armSwing |
multiplies | 1 | arm swing amplitude; 0 hangs the arms |
armForward |
adds | 0 | degrees the whole swing is carried ahead of the body, same amplitude; bent elbows plus this is fists pumping before the chest |
armOut |
adds | 0 | degrees the upper arms are carried out from the ribs, held through the cycle rather than alternating. Only reads head-on, and it is what separates a walk with somewhere to be from one ready to hit something |
elbow |
adds | 0 | elbow bend in degrees on top of the cycle’s own (a walk bends 10, a run 70) |
kneeLift |
multiplies | 1 | knee lift; below 1 drags the feet, above it high-steps; the foot angles follow |
sway |
adds | 0 | hip yaw in degrees, alternating with the step, the shoulders countering by half |
rock |
adds | 0 | trunk roll in degrees over the planted leg, alternating with the step |
The trunk is two segments
Character.torso is a group that draws nothing. What draws is belly and
chest, two boxes on a waist joint, and that is where a trunk pitch lives -
the group carries only the sway and the rock. The two pitches always add up to
lean, so the head and the shoulders end up exactly where a single-box torso
put them; their DIFFERENCE is spineCurve, and it is the difference that
reads. A lean on its own tips the figure like a plank; a lean with a curve
settles the belly back and rounds the chest forward over it, which is what
makes a slump look like weight and a proud walk like air in the chest.
A factor lean brings a little curve with it on its own, because a body that
is asked to lean bends. A BASE lean does not: a run’s authored 12 degrees is a
sprinter’s straight line from the ankles, so it goes on the belly whole and
the waist joint stays shut.
The hip hangs off the belly and gives the belly’s share of the bend straight
back, so the legs stay where the base asked for them - upright under a factor
lean, tipped with the whole figure under a run’s. gaitPoseAt() reports all
four: hip, torso, belly, chest.
The same split is what TalkGestureAnim, GestureAnim’s talking beats and
UseAnim bend with, so a character that leans while it speaks and a character
that leans while it walks bend the same way.
Presets, by name (Gait.presetNames): neutral changes nothing; cheerful,
dejected and furious are exactly what the happy, sad and angry
emotions do to a walk - they share their rows in gait.js, so
setEmotion("sad") and preset: "dejected" cannot drift apart; elderly is
the top of the maturity slider (slow, short, shuffling, stooped) and toddler
its bottom with a touch more tempo; heavy is the top of the mass slider with
more rock, leaning BACK over the weight it is carrying; sneak is slow and
short with high knees, a rounded back, head down, elbows bent and the arms
held still; proud is an arched back and a lifted chest, chin up and arms
swinging over a slightly longer, slower step; march is high knees, a wide
arm swing, straight elbows and a straight back.
furious is short, hard, quick steps with the knees stamping, the shoulders
hunched over a forward head, the upper arms carried off the ribs and the fists
before the chest by a bent elbow - not a bigger walk. Two earlier versions of it were: one slid the whole
arm swing 22 degrees forward, the other folded the elbow to 80, and both put
the forearms out in front where they barely alternated, which reads as a
sleepwalker from every angle. The cycle sheet is what settled it. A name that is not in the list does nothing
and clears Gait.presetKnown.
How the build maps (buildFactors() in gait.js): every default is exactly
neutral, and each effect fades in linearly toward the end of its slider.
maturity below 0.4 is the child zone (quicker, high-stepping, a little
bounce), above 0.75 the elderly zone, and neutral in between - the proportion
tables treat 1 as a full adult, so only the top quarter reads as elderly, for
gait alone. femininity, mass and muscle fade out toward 0.5: feminine
sways and swings the arms less, masculine rocks a little; heavy is slower,
rocks more and leans back over its weight, light a touch quicker; athletic
leans in with the chest up and swings the arms, soft slumps and rounds its
back. bodyHeight and realism do not enter the gait (the leg height they
set enters the speed). The effects are kept subtle so the sliders stay a body
and not a costume: gait.test.js checks that no corner of the four sliders
leaves a walk.
A factor change lands at the next half-cycle, when the phase that is starting reads its targets. There is no blend, which is how every other activity switch behaves.
To look at a gait, draw it as a cycle sheet rather than watching it: the
gait scenario of labs/character-101 (labs/kits/character/GaitSheet.qml)
freezes a row of figures at successive phases of one cycle, so the whole
walk is on one sheet. Its header comment says how to read one; sceneReady
is the property to wait for, since a verb lands after the first pose pass.
clayrender labs/character-101/Sandbox.qml --size 1800x900 --paused \
--eval 'applyScenario("gait"); act("preset", ["elderly"]); act("emotion", ["sad"]); Lab.set("maturity", 0.9)' \
--wait-for 'sceneReady' --out /tmp/elderly.png
The scene’s shots - side, front, back, top (goShot("top"), or N
in the lab) - are the four angles to check a change against: a silhouette
that reads as walking from the side and from nowhere else is a side view,
and the overhead is the only one that shows sway and shoulder
counter-rotation honestly.
To assert on a gait, read gaitFactors: it says what the character was asked
to do, where a joint angle mid-swing says only where the leg happens to be.
gaitPoseAt(base, t) gives the joint angles at any phase with nothing
running, and applyGaitPose(base, t) freezes an idle character there.
Best Practices
-
Use ParametricCharacter for quick character creation with intuitive parameters.
-
Activity-Based Animation: Set the
activityproperty to control animations - speeds are auto-derived from geometry. -
Toon Shading: Use the Canvas3D DirectionalLight setup for consistent cartoon rendering.
-
Character Editor: Add CharacterEditor during development for visual tuning, remove for production.
-
Proportions: Adjust
realism(0-1) to shift between cartoon and realistic body ratios.
Technical Implementation
The Character3D plugin implements:
- Modular Body Parts: Head, a two-segment trunk on a waist joint, arms, legs, all with independent dimensions
- Procedural Animation: Idle derived from body geometry; walk and run are one
GaitCycleAnimover a table thatgait.jsderives from a base (the authored walk or run) plus the composed gait factors - Animation-Speed Coupling: Movement speeds calculated from the derived table’s leg swing angles and the leg height, so speed still follows the feet whatever the gait
- Facial Expressions: Six expression states (neutral, joy, sadness, anger, disgust, surprise) plus talk, each a table of ten shader parameters rather than a shared set of building blocks
- Editor Integration: 3D picking, parameter sliders, and per-character persistence
- Coordinate System: Origin at ground level (Y=0 at feet), character faces +Z when rotation is (0,0,0) - the nose sits on the +Z face of the head and
CharacterControllerwalks along +Z at yaw 0
The animation system uses frame-based updates with biomechanically-inspired joint rotations and parent-child transform hierarchies.
Performance Scripts
A performance script is one string carrying what a character says and what it does, in the order it happens - the way a director writes a scene. It is authored text: it diffs, it translates, and it can be asserted on without watching it run.
Performance {
id: perf
performer: prof // a Character, or any object with the verbs below
searchRoot: view3d.scene // where target names are looked up
}
perf.play("*point at battery* This is the battery. (2s) " +
"*face viewer* *happy* It stores the energy our circuit spends.")
Two rules carry the format:
- Directives are instant, speech and pauses are not. A directive is
dispatched and the script moves straight on, so “point at it while saying
this” is the natural thing to write. Only a spoken line and an explicit
*pause*consume time. - Parsing is strict. An unknown directive is a reported error with its
position in the source, and
play()refuses to run the script. A typo never becomes dialogue. (Character.say()keeps its lenient parse - seeparse(script, {strict: false})below.)
Vocabulary
Directives are matched case-insensitively and their inner whitespace is
normalized, so *Point At battery* is *point at battery*. Everything
outside a *...* is spoken.
| Directive | What it does | Performer method |
|---|---|---|
*happy* *sad* *angry* *disgust* *surprised* *neutral* |
Sets the emotion for the lines that follow. Aliases: joy, sadness, anger, disgusted, surprise, shocked, calm |
setEmotion(value), else the *emotion* annotation is prefixed to the next say() |
*point at NAME* |
Points at the target | pointAt(pos) |
*present NAME* / *show NAME* |
Offers an open hand toward the target - for a group or an area | presentAt(pos) |
*look at NAME* |
Head-only aim at the target | lookAt(pos), else turnTo(pos) |
*face NAME* |
Whole-body turn to the target | turnTo(pos) |
*look at viewer* / *face viewer* |
Same, at the camera | faceViewer() for face, else the position from viewerPosition |
*thumbs up* |
Approval gesture | thumbsUp() |
*gesticulate* |
Talking hands on | gesticulate() |
*rest* |
Drops any gesture | stopGesture() |
*mark NAME* / *mark A, B, C* |
Raises markers on the named things for the length of the line that follows | none - the points come out in marks |
*pause 800ms* *pause 2s* *pause 1.5s* |
Consumes that much time | - |
| anything else | A reported error in strict mode | - |
NAME is a QML objectName, taken verbatim from the script (case included)
and resolved against the scene - there is no second naming scheme. It may
contain spaces: *point at the big red battery*. viewer is the one reserved
name and means the camera. A duration is a number plus a unit, ms or s,
both required; *pause 800* is an error rather than a guess.
*mark* is the one directive that takes several targets, comma separated,
because naming a group is exactly when it earns its keep:
*mark the battery, the switch, the LED* Four parts, one loop.
Marks
A line that names things - “collector on the left, emitter on the right, base
facing you” - asks the eye to find each of them by ear. *mark ...* resolves
its names the way *point at* resolves its target and publishes the world
points in marks; markNames holds the names that resolved, in the same
order. A name that does not resolve is skipped and recorded, and the rest of
the list still marks.
Nothing here draws them. A sequencer has no view, so the points go to whatever
is showing the scene - Clayground.Lab’s MarkLayer is the overlay the labs
use, and a lab’s own is one binding away.
The lifetime is the whole of the rule an author has to hold: a mark set is
raised by its cue, lives for the length of the one line that follows it, and
is gone by the time the next cue starts. stop() and the end of the script
clear it too. A marker says “this one, now”, not “this one, still” - so a
sentence that keeps a mark up needs its own *mark*:
*mark the collector* Collector on the left.
*mark the emitter* Emitter on the right.
*mark the base* And the base, facing you.
Time hints
A spoken run may end with a duration in parentheses:
This is the battery. (2s)
The hint is stripped from the spoken text and becomes that line’s authoritative
duration: the script advances 2 s after the line starts, whether or not the
performer is still talking. Without a hint the line ends when the performer
stops reporting that it is talking, and a backstop timeout
(estimate * 1.5 + 2000 ms, off Speech.estimateDurationMs()) ends it if the
performer never reports anything at all.
The rule is deliberately narrow: the parenthetical counts as a hint only at the
very end of a run and only when whitespace precedes it. This (2s) is the
battery. and Battery(2s) are text.
A whole script
perf.play(
'*point at battery* This is the battery. (2s)\n' +
'*face viewer* *happy* It stores the energy our circuit spends.\n' +
'*pause 500ms* *gesticulate* Watch what happens when I close the switch.\n' +
'*rest* *neutral*')
Ten cues: point, say (hinted at 2000 ms), face, emotion, say, pause, gesticulate, say, rest, emotion.
Performance
| Property | Meaning |
|---|---|
performer |
The character. Duck-typed: each cue calls the method named in the table above if the performer has it, and is skipped and recorded if it does not |
searchRoot |
Node whose children are walked (recursively, by objectName) to resolve a target name to its scenePosition |
resolveTarget |
function(name) returning a vector3d or null; replaces the searchRoot walk |
viewerPosition |
A vector3d or a function() returning one - where “viewer” is |
voiceOf |
function(sayIndex) returning a clip url per spoken line; with a clip and a performer that has tell(), the line is played from the file. sayIndex counts spoken lines from 0 |
spoken |
What a bare line (no clip) does. True (default) goes through say() - the speech engine. False keeps the lab silent: a performer with tell() shows and mouths the line without audio, the professor’s narration mode |
extraVerbs |
Extra directive names the parser accepts |
debug |
Logs [perform] 1234ms cue 3/7: point at battery per cue |
| Method | Meaning |
|---|---|
play(script) |
Parses strictly and plays from the first cue. Returns false and plays nothing when the script has errors |
playFrom(index) |
Plays the last parsed script from a cue - a debugging aid |
stop() |
Disarms every timer, ends speech (stopSpeaking()/quiet()) and drops gestures. Does not emit finished() |
registerVerb(name, handler) |
Teaches the parser a directive and dispatches it to handler(arg). The seam for actions only one character has; a handler that throws is recorded, the script continues |
estimateMs(text) |
What the performer’s speech engine expects the line to take, or 72 ms per character |
Observability
A script is verified by reading state, not by watching it.
| Property / signal | Meaning |
|---|---|
running |
True between play() and the last cue |
done |
True once the last cue has fired |
cueIndex / cueCount |
Which cue is playing, out of how many |
currentCue |
The current cue as a one-liner, e.g. point at battery |
errors |
Parse errors of the last play(), each {at, directive, message} |
skipped |
Cues that could not be carried out, each {cue, reason} - an unresolved target, a missing verb, a handler that threw |
firedLog |
Every cue that fired, each {ms, cue}, ms measured from play() |
marks / markNames |
The world points a *mark ...* cue currently raises, and the names behind them. Empty whenever nothing is marked |
finished() |
Emitted after the last cue |
customCue(verb, arg) |
Emitted for a custom cue with no registered handler |
cueFired(type, arg) |
Emitted for every cue as it fires, so a scene can choreograph around the script |
An unresolved target is never fatal: the cue is skipped, recorded in skipped
and warned about once.
The parser on its own
scripting/performancescript.js is Qt-free - no Qt types, no clock, no
randomness - so scripts can be checked without an engine:
node plugins/clay_character3d/scripting/performancescript.test.js
parse(script, options)->{cues, errors}.options.strict(default true);options.extraVerbsaccepts additional directive names as{type: "custom", verb, arg}cues. Every cue carriesat, its character index in the source.describe(cue)-> the one-linerPerformance.currentCuepublishes.lint(scriptA, scriptB)-> divergences between the directive sequences of two languages of the same script, each{index, a, b, message}.
Cross-language lint
Direction lives inside the translated string, which is what lets a German
script time its cues to German word order - and what lets a translator reorder,
drop or translate a stage direction by accident. lint() compares the two
directive sequences, ignoring the spoken text and the time hints:
const Script = require('.../performancescript.js')
Script.lint(strings.en.introScript, strings.de.introScript)
// [] when the two are in sync
// [{index: 2, a: "point at battery", b: "point at Batterie",
// message: "argument differs: point at battery vs point at Batterie"}]
Speech timing
Speech publishes the numbers a script schedules against, instead of every
caller measuring a speech rate of its own:
estimateDurationMs(text)- how long the engine would take over the text at the current rate, without saying it.durationMs- the current line’s length; stale once the line ends.wordMarks()- the current line’s words as{offset, ms}.
Two engine behaviours a scheduler depends on: an empty or whitespace-only line
reports started() and finished() (asynchronously, never re-entrantly), so a
queue advancing on finished() cannot hang on it; and of several say() calls
in one tick, exactly the last one runs and it is the only one that reports
anything - a line replaced before it began emits neither signal.