Docs/Voice

Voice

Stream captured audio to an NPC and play back its spoken reply.

Voice runs over the same connection as text. The addon gives you the wire methods — streaming audio chunks in, receiving synthesized speech back — but capturing the player's microphone and deciding when they've stopped talking is up to you, using Godot's own audio APIs.

No built-in microphone capture or voice detection
The addon does not capture your microphone or run voice-activity detection for you. This is a real gap today, not a hidden limitation: you own capture (Godot's AudioEffectCapture on an AudioStreamMicrophone bus is the usual approach) and pass the addon raw PCM bytes.

Sending audio

Stream chunks as you capture them, then send one final chunk with end = true — or let the server auto-finalize once it has buffered about 5MB, whichever comes first:

voice.gd
# pcm_chunk: PackedByteArray of raw audio you've already captured
connection.send_voice_chunk(pcm_chunk, false)

# ...more chunks as they're captured...

connection.send_voice_chunk(final_chunk, true) # end = true
data
Raw PCM bytes for this chunk, as a PackedByteArray. The addon base64-encodes it for the wire.
end
Set true on the final chunk of the utterance.
sender_id
Optional. Overrides the connection's default player id for this utterance.

Once the server finishes transcribing, a transcript signal fires with what it heard, followed by the normal chat signals — voice replies arrive as one complete chat_chunk rather than streamed token by token, since grounding runs before anything is sent back.

Playing the reply

Synthesized speech arrives as raw binary WebSocket frames — no JSON envelope, one frame per sentence, in playback order:

voice.gd
connection.audio.connect(_on_audio)

func _on_audio(bytes: PackedByteArray) -> void:
    # Raw synthesized speech for one sentence. Decode/queue it with
    # AudioStreamGeneratorPlayback in your own playback pipeline.
    play_audio_chunk(bytes)
Microphone permission
Requesting microphone access has its own platform-specific flow (desktop generally works out of the box; mobile and web exports each have their own permission prompt). Request access before the first voice interaction, and offer a text fallback — text and voice can be used freely in the same connection, including in the same conversation.