Transcription turns spoken words into text attached to the asset. For anyone holding a large video or audio archive, this is usually the single highest value thing ioMoVo does, because speech is where most of the meaning lives and none of it is in the filename.
Transcribing video
Transcribing audio
Same process for audio only files: interviews, recorded calls, radio, podcasts.
Why it changes how the archive behaves
Before transcription, finding a moment means remembering which file it was in and scrubbing. After it, you search for the words and land on the point in the file where they were said.
The practical effect on a large archive is that footage nobody could locate becomes usable. Most organizations discover they own more than they thought.
Burning captions into a video
Burning captions writes them into the video image itself, so they travel with the file wherever it goes and do not depend on the player supporting a separate caption track. That matters for social platforms, for playback in noisy environments, and for accessibility requirements where you cannot rely on the viewer's setup.
This produces a new video. Your original is untouched. The captioned version is created as a separate file and stored in the directory alongside it, so you end up with both rather than having to choose.
That is worth understanding, because it makes the directory the place where versions of an asset live together. Editors can store a final edited cut there too, so the original, the captioned version and the finished edit all sit in one place rather than scattering across local machines and shared drives.
Give your outputs names that say what they are. A directory holding three versions is useful; a directory holding three files called final is not.
What affects quality
Audio quality. Clear speech transcribes well, heavy background noise and overlapping speakers less so
The model you connected. Transcription needs a speech to text model. If nothing is transcribing, that is the first thing to check
Specialist vocabulary. Names, technical terms and acronyms specific to your organization are where errors cluster. Worth reviewing the transcript before you burn captions from it, since the captioned output is a rendered file rather than something you can edit in place
Related
Translating insights covers producing this in another language. Connecting your own AI models covers attaching a speech to text model.
