← Back to Docs
  • index.html
  • Clayground
  • Clayground.Character3D
  • Speech
  • Clayground 2026.7
  • Speech QML Type

    Voice output with approximate lip-sync for characters. More...

    Import Statement: import Clayground.Character3D

    Detailed Description

    Speech lets a character say things while its mouth moves along. Two kinds of input are supported:

    The mouth shape is published through the continuous properties mouthOpen, mouthWide and mouthRound which Head consumes via its speechSource property. Character wires this up automatically - typically you only call say():

    import Clayground.Character3D
    
    Character {
        id: hero
        Component.onCompleted: hero.say("Hello! Welcome to Clayground.")
    }
    
    // Or play a recorded line:
    // hero.say("dialog/intro.wav")
    
    // Advanced configuration via the speech engine:
    // hero.speech.rate = 0.2
    // hero.speech.volume = 0.8

    Anything scheduling around a line needs this number, and the engine has always known it. See also estimateDurationMs(), which answers the same question without saying anything.

    ConstantDescription
    Speech.EnvelopeLoudness and zero crossings. Says how far the jaw dropped and nothing about the shape of the mouth.
    Speech.SpectralFormant bands: an open vowel opens further than a closed one, a front vowel spreads, a back vowel rounds, a fricative narrows. The default.
    Speech.AlignedThe script's own shapes on the recording's clock. Needs a transcript - see sayAudio() - and falls back to Spectral without one.

    The tiers are not a performance dial: the analysis is a few hundred thousand flops for a whole line, once, before playback starts. What separates them is time-to-first-sound, the samples held while the analysis runs, and - for Aligned - whether anyone wrote the line down.

    No tier leaves the mouth dead. A recording the analyser cannot read falls back to the one below on its own, and Envelope is the floor.

    Only the same as accuracy when the audio could carry it. Asking for Aligned and reading Spectral back is a normal outcome - an unreadable recording or a missing transcript - and not an error.

    Strings ending in a known audio extension (wav/mp3/ogg/m4a/flac) are treated as file paths or URLs, anything else as text to speak. The optional transcript is forwarded to sayAudio() and ignored by the text path.

    The optional transcript is what the recording was recorded saying. Given one, and with accuracy at Speech.Aligned, the mouth takes its shapes from that script and only its timing from the audio - which is the only way an /m/ closes reliably, since a bilabial is voiced and every acoustic measurement of one says "loud" rather than "shut".

    A transcript that does not match the recording is refused rather than used, and the tier below takes the line.

    The same number the mouth is driven by, so anything scheduling around a line stops having to guess a speech rate of its own.

    The character offset in the source text and when that word starts. Empty before the first line. Available for a recorded line too when it was read at Speech.Aligned, since the alignment carries them across.

    See also Character, Head, effectiveAccuracy, and Character::speechAccuracy.