How EchoPod word-level transcripts work
EchoPod connects natural-speed audio to the exact words being spoken. When timing data is available, the active word is highlighted so you can listen, read, and recover meaning without leaving the player.
What happens while you listen
- Audio and text stay together. Playback position selects the matching transcript segment and available word timing.
- A tap opens word help. EchoPod can show meaning, pronunciation, translation, and surrounding context.
- AI explains the sentence. Grammar help is requested on demand and is meant to clarify structure and nuance, not replace the original transcript.
A word-timing example
Audio at 12.40–13.85 seconds: “I finally understood the phrase.”
finally is active from 12.62–13.05 seconds. Tap it to open its meaning and pronunciation; select the full sentence to ask why the adverb appears before “understood.” Timing values here are representative—actual boundaries come from each transcription result.
Language coverage
The current EchoPod language settings list nine language codes. Provider support and timing quality can vary by source audio, accent, recording quality, and the selected transcription or translation service.
- Listed languages
- English, French, Japanese, Chinese, Korean, Spanish, German, Italian, Portuguese
- Automatic detection
- Available where the selected transcription provider supports it
- CJK reading help
- Japanese and Chinese can use language-specific annotations
- Translation
- Requested separately from transcription; results can vary by provider
Limits to expect
- Word timing is only as accurate as the transcription result.
- Background noise, overlapping speakers, music, and poor recordings can reduce accuracy.
- Names, specialist vocabulary, and mixed-language speech may need correction.
- AI explanations can be incomplete or wrong; verify important interpretations against trusted references.
Which languages are listed in EchoPod settings?
English, French, Japanese, Chinese, Korean, Spanish, German, Italian, and Portuguese.
Are transcripts always word-level?
EchoPod uses word timing when the transcription result provides it. Some sources may only produce segment-level timing or less precise alignment.