The fastest reliable way to caption a prerecorded course video is auto transcription followed by a single focused human edit pass, saved as an .srt or .vtt file. Live sessions and high stakes content need a human or hybrid workflow instead, since accuracy matters more than speed there. Use .srt for broad platform compatibility and .vtt when your video sits on a web player with styling needs. Platforms can host the finished tracks and support multiple languages per course.
TL;DR:
- Use SRT for maximum platform compatibility and VTT when styling, positioning, or metadata are needed, ensuring files are saved in UTF-8 encoding.
- Auto transcription plus one human editing pass offers a quick solution for most prerecorded content, but technical, high-stakes, or accreditation-related videos benefit from full or hybrid human oversight.
- When merging or creating bilingual subtitles, ensure timecodes are aligned and non-overlapping, and keep languages as separate tracks rather than burning both into the video.
- Upload caption files in UTF-8 to avoid encoding issues, and test playback on both desktop and mobile platforms to verify correct display and functionality.
- Captions improve accessibility, support language learners, and enable searchable videos; a thorough editing process enhances these benefits and ensures compliance with accessibility standards.
Table of Contents
- Which subtitle file formats work best for course videos
- Three workflows for adding subtitles: automatic, human or hybrid
- How to generate and edit auto captions step by step
- Briefing translators and checking subtitle translation quality
- Merging subtitle files and choosing bilingual or burnt-in captions
- Uploading captions to your course platform without losing quality
- Meeting WCAG accessibility standards for course captions
- A quick quality check before you publish
- Why captions are worth the extra editing pass
- Publishing captioned courses on BibliOWLteca
- Sources
- FAQ
Which subtitle file formats work best for course videos
Two formats cover nearly every course delivery need. SRT is the plain text format most learning platforms and video players accept without complaint, which makes it the safest default for uploads. WebVTT (.vtt) does the same job but adds support for styling, positioning and metadata that some web based players use. YouTube lists both as its recommended caption formats, with SRT prized for compatibility and VTT for its extra web features. Save either format in UTF-8 encoding, since anything else risks broken accents or missing characters once uploaded.
- SRT: the safest choice for uploading to an LMS, a marketplace or most video hosts.
- VTT: better when your player supports styling, positioning or chapter metadata.
- SBV, TTML or EBU-STL: reserved for broadcast or specialist workflows that most course creators will never need.
Unless a client or broadcaster specifically asks for one of those specialist formats, stick with SRT or VTT and move on.
Three workflows for adding subtitles: automatic, human or hybrid
The right workflow depends on your content, budget and how much accuracy matters.
- Auto transcription plus one edit pass: generate a transcript automatically, then correct it in a single focused session. This is the fastest route for most prerecorded lectures and demos.
- Fully human transcription: a person transcribes and times the captions from scratch, which suits technical content, exams or anything tied to accreditation where errors carry real consequences.
- Hybrid or real time captioning: combines automated output with live human oversight, the standard approach for webinars and events where CART (Communication Access Realtime Translation) or a trained captioner corrects the feed as it happens.
Research comparing automated and edited subtitles found that auto captions can produce learning outcomes similar to edited captions for some learner groups, though quality and the viewer's own proficiency both affect the result. That is a reasonable argument for the auto-then-edit route on lower stakes lessons, but not a licence to skip the edit entirely.
Pro Tip: Budget about as much editing time as the length of the video for a first pass, then adjust once you know your own transcript's typical error rate.
How to generate and edit auto captions step by step
Producing a usable SRT from an automated transcript follows a predictable sequence: export the raw transcript, convert it to SRT if your tool does not do so automatically, then fix timing, punctuation and line breaks before adding speaker labels and non-speech cues such as [music] or [laughter].
Three categories of tool cover this work. In-platform autocaption features built into video hosts or editing software generate a rough first draft. Dedicated transcription services handle longer or more technical recordings with better accuracy. Local or offline converters are useful when you need to reformat a transcript without uploading sensitive material anywhere.
Whichever tool you use, run the same editing checklist:
- Build a short glossary for course-specific terms and abbreviations before you start editing.
- Check timing against the audio, not just the transcript text.
- Fix punctuation and split long lines so no caption runs more than two lines on screen.
Briefing translators and checking subtitle translation quality
Handing a transcript to a translator works best with a clear brief. Include the timecode precision you need, how speakers should be labelled, a glossary of course-specific terms, and instructions for handling non-speech cues consistently across languages.
The workflow itself is straightforward: translate the source SRT, check that timing still matches the translated text length, then run a sample quality check through back-translation or a bilingual reviewer spot-checking a few minutes of dense dialogue. Our course localisation guidance covers this in more depth for creators adapting a course beyond simple translation.
Ask translators to confirm consistency of terminology, a comfortable reading speed, sensible line length and any cultural localisation notes relevant to the target audience.
Pro Tip: Send translators a short style sheet alongside the SRT, covering how you want numbers, names and technical terms handled, so you are not fixing the same inconsistency across every lesson.

Merging subtitle files and choosing bilingual or burnt-in captions
Combining multiple SRT files or building a bilingual track is mostly a matter of care with timecodes.
- Align offsets first: if you are joining two files that were timed against different starting points, adjust one file's timecodes before merging.
- Avoid overlapping timecodes, since most players will only show one caption at a time and silently drop the other.
- Use a plain text editor or a small dedicated utility for concatenation, then scan the merged file for gaps or duplicated lines.
For bilingual delivery, keep the languages as separate selectable tracks rather than burning both into the video. Open, burnt-in captions only make sense for social clips or locked-down devices where the player cannot load a separate track, and they remove the viewer's ability to turn captions off.
Uploading captions to your course platform without losing quality
Most learning platforms and video hosts follow a similar pattern: open the video's caption or subtitle settings, upload your SRT or VTT file, tag it with the correct language code, and publish. The two most common pitfalls are wrong file encoding and a missing or incorrect language tag, both of which can make captions fail silently on some devices.
- Save your file in UTF-8 before uploading, since encoding errors are a frequent cause of caption upload failures.
- Test playback on both desktop and mobile, since caption rendering and track switching sometimes behave differently on smaller screens.
- Check that any custom styling still displays correctly once the file is live, not just in a preview.
Our mobile-friendly course checklist is worth running alongside your caption upload, since a broken caption track on mobile is one of the easier things to miss before publishing. When something does not behave as expected, your platform's own help documentation is usually faster to consult than trial and error.
Meeting WCAG accessibility standards for course captions
Captions are not the same as basic subtitles: proper captions include speaker identification and non-speech cues like [applause] or [phone ringing], both of which matter for viewers who are deaf or hard of hearing rather than simply following along in another language. WCAG guidance recommends closed captions that include these elements, and live events in particular may need CART or a hybrid AI-plus-human approach to reach AA compliance under criterion 1.2.4.
- Include speaker labels whenever more than one person is talking.
- Add non-speech cues for meaningful sounds, not just dialogue.
- Give viewers the option to toggle captions on or off wherever the platform allows it.
National accessibility assessments recommend checking synchronisation tolerance and caption customisation as part of any player review, according to Estonian public sector digital accessibility checks. That kind of check is a useful habit even outside a formal audit.
A quick quality check before you publish
A short checklist catches most caption errors before a learner ever sees them.
- Watch the full video with captions on, checking timing accuracy and that no captions overlap.
- Confirm non-speech cues, speaker labels, correct language tags and consistent encoding are all in place.
- Run a glossary check against your course-specific terms, then test playback on a mobile device.
Pro Tip: Pick the densest two or three minutes of dialogue in the video and play them back at a slower speed. Errors that hide at normal speed tend to surface immediately once you slow down.
Why captions are worth the extra editing pass
Captions do more than tick an accessibility box. They open a course to viewers with hearing loss, support language learners working in a second language, and turn your video into searchable text that helps people find the right lesson later. Research on textual enhancement in captions suggests that highlighting key vocabulary can help higher-proficiency learners, though the effect varies and can add cognitive load for others, so it is worth testing on a sample lesson rather than applying it everywhere by default. Our guide to AI in learning covers how to weigh automated tools like this against the extra editing time they demand. Try the auto-then-edit workflow on one lesson before committing to it across a whole course.
— BibliOWLteca
Publishing captioned courses on BibliOWLteca
Once your captions are edited and saved as SRT or VTT, you need somewhere to host them that does not fight you on languages or formats. Some platforms let creators upload course videos with multiple caption tracks attached, so a single lesson can carry captions in several languages without duplicating the video itself. Course pages can be built for an international audience, with sales processed in multiple currencies for buyers wherever they are.
- Upload SRT or VTT caption files directly alongside your course video.
- Build multilingual course pages that sell to an international audience.
- Sell ebooks, courses, templates and other digital downloads from a storefront.
On pricing, BibliOWLteca charges no published monthly fee for its Creator Plan and instead applies a transaction fee from €0.99 per sale, so cost scales with what you actually earn rather than what you host. Uploading a captioned lesson to some storefronts can be a reasonable way to see how multilingual setups work before committing a full course.

Sources
- Supported subtitle and closed caption files - YouTube Help
- Effects of automated subtitles vs edited subtitles on learning (PMC article)
- The effects of captioning and textual enhancement on L2 viewing comprehension
- Eesti avaliku sektori digiligipääsetavus — video caption checks
FAQ
Should I use SRT or VTT for my course videos?
Use SRT for the widest compatibility across learning platforms and video hosts, since YouTube and most players support it as a basic format. Choose VTT instead when your player supports extra styling or positioning features that SRT cannot carry.
Are automatically generated captions accurate enough for courses?
Auto captions can work well as a starting point, and some research found similar learning outcomes to edited captions for certain learner groups. They still need a human edit pass for punctuation, speaker labels and non-speech cues before publishing, since accuracy and learner proficiency both affect how well auto captions alone perform.
What is the difference between open and closed captions?
Closed captions can be toggled on or off by the viewer and are the format recommended for accessibility, since they let each learner choose. Open captions are burnt directly into the video and cannot be removed, which suits social clips or locked devices but limits flexibility for a full course library.
How do I merge multiple SRT files without breaking the timing?
Align the offsets of each file to a shared starting point before combining them, then check for overlapping timecodes once merged. A plain text editor or a small dedicated utility handles the concatenation itself, but the timecode check is the step that prevents captions from displaying out of sync.
Can BibliOWLteca host courses with multilingual subtitles?
Yes, BibliOWLteca supports uploading caption files in multiple languages alongside a single course video, and course pages can be built for buyers in different currencies. Pricing runs on a transaction fee from €0.99 per sale under the Creator Plan, with no published monthly fee.
