A text encoder is the part of a voice-based system that represents what was said, reading the transcribed content of speech rather than its acoustic delivery.
Part of Fusion model →The text encoder is the second of the two channels in a typical voice pipeline, working alongside an audio encoder that reads delivery. A fusion model then combines both into a single prediction, so content and delivery both contribute rather than either one alone.
How it shows up
A text encoder reads the same information from a transcript regardless of how it was said, which is exactly why it needs an audio encoder alongside it to capture the rest of the story.
Why it matters
Content alone misses everything about delivery, and delivery alone misses everything about what was said. Reading both is what makes the combination more informative than either.
See what your words alone reveal, and what they miss.
Take the free assessment →