Transcripts · Product guide

How EchoPod word-level transcripts work

EchoPod connects natural-speed audio to the exact words being spoken. When timing data is available, the active word is highlighted so you can listen, read, and recover meaning without leaving the player.

Maintained by EchoPod Product Team

What happens while you listen

  1. Audio and text stay together. Playback position selects the matching transcript segment and available word timing.
  2. A tap opens word help. EchoPod can show meaning, pronunciation, translation, and surrounding context.
  3. AI explains the sentence. Grammar help is requested on demand and is meant to clarify structure and nuance, not replace the original transcript.

A word-timing example

Audio at 12.40–13.85 seconds: “I finally understood the phrase.”

finally is active from 12.62–13.05 seconds. Tap it to open its meaning and pronunciation; select the full sentence to ask why the adverb appears before “understood.” Timing values here are representative—actual boundaries come from each transcription result.

Language coverage

The current EchoPod language settings list nine language codes. Provider support and timing quality can vary by source audio, accent, recording quality, and the selected transcription or translation service.

Listed languages
English, French, Japanese, Chinese, Korean, Spanish, German, Italian, Portuguese
Automatic detection
Available where the selected transcription provider supports it
CJK reading help
Japanese and Chinese can use language-specific annotations
Translation
Requested separately from transcription; results can vary by provider

Limits to expect

Which languages are listed in EchoPod settings?

English, French, Japanese, Chinese, Korean, Spanish, German, Italian, and Portuguese.

Are transcripts always word-level?

EchoPod uses word timing when the transcription result provides it. Some sources may only produce segment-level timing or less precise alignment.

Related guides