Add subtitles and captions to narrated presentations for accessibility and engagement. Step-by-step guide for self-running decks in multiple languages.
You've built a beautiful deck. You've written a script. You've generated a natural voice-over in your brand's tone. Now you're ready to send it out into the world as a self-running presentation that presents itself, no live speaker required.
But here's the problem: not everyone will watch with sound on.
Someone opens your deck in a noisy office. Someone else is deaf or hard of hearing. A third person is watching in a language they understand, but the audio is still processing in their head slower than they read. A fourth viewer is in a region where bandwidth is spotty, and the audio keeps buffering while the video plays.
Without subtitles and captions, you're losing comprehension, accessibility, and engagement. You're also losing legal compliance in many jurisdictions. When you add synchronized text to your narrated presentation, you're not just being nice to your audience. You're making sure your message lands, no matter how or where it's consumed.
This guide walks you through the why, the how, and the best practices for adding subtitles and captions to narrated presentations, whether you're building investor decks, sales pitches, training materials, or multilingual content.
Before you add subtitles and captions, make sure you have the following in place:
A narrated presentation. You'll need a deck with a voice-over already recorded or generated. If you're using Preso's voice-over feature, your narration is already synced to each slide. If you're using a different tool, your audio should be embedded or linked to your slides.
The script or transcript. You need the exact words that are being spoken. If you generated your narration with AI, you should have access to the underlying script. If you recorded a voice-over yourself, transcribe it now. Accuracy here is critical. Every word in the audio needs to appear in the captions.
A subtitle or caption file format. The most common formats are VTT (WebVTT), SRT (SubRip), and SCC (Scenarist Closed Caption). Your presentation platform will tell you which format it accepts. Microsoft's guide to adding closed captions in PowerPoint covers the technical requirements for PowerPoint specifically.
Timing information. Captions must be synchronized with the audio. You need to know when each phrase starts and stops. If you're generating captions with AI transcription tools, this is usually done automatically. If you're creating them manually, you'll need to time each caption to the nearest millisecond.
An understanding of your audience's needs. Are you serving an international audience? Do you need captions in multiple languages? Are you presenting to people who are deaf or hard of hearing? Are you creating content for a learning management system where captions are required? These questions shape your approach.
Your transcript is the foundation of your captions. Without it, everything else falls apart.
If you used Preso to generate your narration, you already have the script that was read aloud. Export it or copy it into a text editor. Make sure it matches word-for-word what the AI voice is actually saying. If there are contractions, abbreviations, or numbers spoken aloud, they should appear exactly as spoken in the transcript.
If you recorded your own voice-over, transcribe it carefully. Listen to the audio in sections, rewind frequently, and type out every word. Don't paraphrase or clean up the language. Captions should reflect what was actually said, including filler words, pauses, and natural speech patterns, unless you're creating formal closed captions for broadcast (which have different rules).
If you're working with an existing recording, use an AI transcription service. Services like Otter.ai, Rev.com, or Descript can transcribe audio to text with high accuracy. Many of these services also provide speaker identification and timestamp data, which saves you time in the next step.
Once you have the transcript, read it through while listening to the audio. Correct any errors. Check that numbers are spelled correctly, that brand names are capitalized, and that technical terms are accurate. This is your source of truth.
Now you need to synchronize your text with the audio. Each caption needs a start time and an end time, measured in hours, minutes, seconds, and milliseconds.
If you're using a transcription service that provides timestamps, you're ahead. The service has already done the timing work. Export the transcript with timecode data and you're ready to move to the next step.
If you're timing manually, use a caption editor like Subtitle Edit (free, open source) or Aegisub (also free). These tools let you play the audio, highlight sections of text, and set precise start and end times for each caption.
Here's the workflow:
Timing should be tight, but not so tight that captions flash on and off too quickly. Viewers need time to read. A good rule of thumb is that a caption should stay on screen for at least one second, and no more than six seconds, depending on the length of the text.
Not all captions are created equal. The U.S. Section 508 accessibility standard sets requirements for captions in government and federally funded content, but the principles apply universally.
Speaker identification. If multiple people are speaking in your presentation, identify who is speaking. Use the format "SPEAKER NAME: Text of what they said." This is especially important in presentations with interviews, Q&A sections, or multiple narrators.
Sound descriptions. If there's important audio that isn't dialogue, describe it. If there's background music that sets the tone, note it. If there's a sound effect that reinforces a point, caption it. Examples: "[upbeat music plays]," "[keyboard clicking]," "[phone rings]." These are bracketed to distinguish them from dialogue.
Accuracy. Captions must be verbatim or near-verbatim. The only exception is if the speaker is using extremely casual language or heavy accents that would be hard to read if transcribed exactly. In that case, clean it up slightly for readability, but preserve the meaning and tone.
Line length and breaks. Keep caption lines to 32-42 characters maximum. Break longer sentences into logical phrases, not at random points. "The quick brown fox jumps over the lazy dog" should not break as "The quick brown fox jumps over the / lazy dog." Break it as "The quick brown fox / jumps over the lazy dog." This follows natural speech rhythm and makes reading easier.
Timing for comprehension. Harvard's guidance on providing captions recommends that captions stay on screen long enough for viewers to read them comfortably. A typical reading speed is 150-180 words per minute. If a caption has 10 words, it should stay on screen for at least 3-4 seconds.
Color and contrast. If your presentation platform allows it, make sure captions have good contrast against the background. Black text on a white background, or white text on a dark background, is easiest to read. Avoid placing captions over busy images or video backgrounds where they might be hard to see.
Different platforms accept different caption file formats. The most common are:
VTT (WebVTT). This is the modern standard for web video. It's a plain-text format that's easy to edit and widely supported. A VTT file looks like this:
WEBVTT
00:00:00.500 --> 00:00:07.000
Hello, and welcome to our presentation.
00:00:07.500 --> 00:00:15.000
Today we're going to talk about how to build
beautiful presentations with AI.
SRT (SubRip). This is an older format, still widely used. It's similar to VTT but with a slightly different structure. SRT files have a sequence number, timecode, and text:
1
00:00:00,500 --> 00:00:07,000
Hello, and welcome to our presentation.
2
00:00:07,500 --> 00:00:15,000
Today we're going to talk about how to build
beautiful presentations with AI.
SCC (Scenarist Closed Caption). This format is used primarily for broadcast television and includes styling information. It's less common for web presentations, but some platforms still support it.
Check your presentation platform's documentation to see which format it accepts. If you're using PowerPoint, Microsoft's official guide walks you through the process. If you're using Preso, your narrated decks can include synchronized captions that you can customize in the editor.
Most caption editors can export to multiple formats, so create your captions in one format and convert as needed.
Now you're ready to add your caption file to your presentation.
In PowerPoint: Go to the Insert tab, select Audio or Video, and choose your media file. Once the media is inserted, right-click it and select "Add a Caption Track." Choose your VTT or SRT file. PowerPoint will sync the captions to the media. You can adjust the timing and appearance of captions in the Captions pane.
In Google Slides: Google Slides doesn't natively support caption files for embedded audio or video. However, if your video is hosted on YouTube, you can enable YouTube's automatic captions or upload your own caption file to YouTube, then embed the YouTube video in your slide. YouTube will display the captions when the video plays.
In Preso: When you generate a narrated presentation with Preso's voice-over feature, the script is already available in the editor. You can add captions directly in the Preso interface, and they'll be synchronized with the narration. If you're generating decks via Preso's API, you can include caption data in your API request, and the captions will be embedded in the generated presentation.
In other platforms: Check your platform's documentation. Most modern presentation tools have a way to add captions or subtitle files. If not, you may need to export your presentation as a video, add captions using a video editor, and re-export as a presentation file.
After you've added the captions, test them. Play the presentation and verify that the captions appear in sync with the audio, that they're readable, and that they cover all the important content.
If you're presenting to an international audience, you need captions in multiple languages.
The most straightforward approach is to translate your transcript into each target language, then create caption files for each language. When you add captions to your presentation, you can often specify multiple caption tracks, one for each language. Viewers can then choose which language they want to read.
Translation workflow:
Take your original transcript and have it translated by a professional translator or a translation service. Machine translation (Google Translate, DeepL) can be a starting point, but professional translation is more accurate, especially for technical or branded content.
Have the translator review the translation in the context of your presentation. Some phrases might need adjustment to fit the visual content or to maintain the tone of the original.
Create a new caption file for each language, using the translated transcript and the same timing as the original captions.
Add all caption tracks to your presentation, labeled by language.
Alternatively, if you're using Preso's multilingual feature, you can generate the entire presentation in multiple languages at once. Preso writes a narrative for each language, applies your brand styling, and generates captions synchronized with a voice-over in that language. This ensures that the captions, the narration, and the visual design are all coherent and on-brand.
Before you share your presentation, test the captions thoroughly.
Play the presentation in the same environment where your audience will watch it. If it's a self-running presentation that will be emailed or shared as a link, test it in a web browser. If it's a presentation that will be played on a projector in a conference room, test it on that equipment. If it's a training video that will be embedded in a learning management system, test it there.
Check the timing. Do the captions appear and disappear in sync with the audio? Is there any lag or mismatch?
Read the captions for accuracy. Are there any typos, misspellings, or misheard words? Do the captions match the audio exactly?
Verify readability. Can you read the captions comfortably? Is the text large enough? Is there enough contrast? Are the line breaks in logical places?
Test with accessibility tools. If you're creating content for a government agency, a university, or a large corporation, they may have accessibility requirements. Use a tool like WAVE or Axe to check that your presentation meets accessibility standards.
Get feedback from your audience. If possible, show the presentation to a few people from your target audience, including people who are deaf or hard of hearing. Ask them if the captions are clear, if the timing works for them, and if there's anything they'd change.
Make adjustments based on your testing and feedback.
Sync your narration to your slides. The best narrated presentations are designed so that each slide has a specific piece of narration. When you're creating your captions, make sure that each caption corresponds to a single slide or a clear visual moment. If a caption spans multiple slides, adjust the slide timing or the narration so that the visual and the text align.
Use captions to reinforce key points. Captions aren't just for accessibility. They're also a design element. Use them to highlight key phrases, statistics, or calls to action. If a slide says "Three ways to improve your workflow," make sure that's in the caption, so viewers see it in text and hear it in the narration.
Consider the context of your presentation. If your presentation will be watched in a noisy environment (like a trade show or a busy office), captions are even more important. If it's a training presentation that will be watched multiple times, captions help with retention. If it's a multilingual presentation, captions ensure that non-native speakers can follow along.
Plan for caption maintenance. If you edit your presentation after adding captions, you may need to update the captions too. If you change the narration, the captions need to match. If you re-order slides, the timing might change. Build in time for caption updates whenever you revise your presentation.
Use captions to bridge language gaps. If you're presenting in a language that's not your audience's first language, captions help everyone follow along. Even if the narration is in English, captions in the audience's native language can improve comprehension.
In many jurisdictions, adding captions to video content is not optional. It's a legal requirement.
The Web Content Accessibility Guidelines (WCAG) 2.2, published by the W3C, set the standard for accessible web content. WCAG 2.2 requires that all video content with audio must have captions. This applies to presentations that are shared online, embedded in websites, or distributed via email as video files.
In the United States, the Americans with Disabilities Act (ADA) requires that organizations provide equal access to information. For many organizations, this means providing captions for any video or narrated content. The Section 508 standard specifically requires captions for federal content.
In the European Union, the European Accessibility Act requires that digital content, including presentations, be accessible to people with disabilities. This includes captions for audio content.
If you're creating presentations for a government agency, a school, a university, or a large corporation, check their accessibility policy. Many have specific requirements for captions, transcripts, and other accessibility features.
Even if you're not legally required to add captions, doing so is good practice. It makes your content more accessible to everyone, improves engagement, and shows that you care about your audience.
Mismatched timing. The most common mistake is captions that don't sync with the audio. If a caption appears before the speaker says the words, or disappears before they finish, viewers will be confused. Test your timing carefully.
Incomplete captions. Don't caption only the main points. Caption everything that's spoken, including introductions, transitions, and questions. If it's in the audio, it should be in the captions.
Poor readability. Captions that are too small, too long, or placed over busy backgrounds are hard to read. Follow the accessibility guidelines for line length, timing, and contrast.
Ignoring speaker identification. If multiple people are speaking, identify who's speaking. Otherwise, viewers won't know who's talking.
Forgetting about sound descriptions. If there's important audio that's not dialogue (music, sound effects, applause), describe it in brackets. This is especially important for people who are deaf or hard of hearing.
Not testing in the actual platform. Captions that look perfect in your caption editor might not display correctly in your presentation platform. Always test in the actual environment where your audience will watch.
If you're building presentations regularly, especially if you're a sales team creating personalized decks, an agency building client presentations, or an educator creating training materials, you need a workflow that makes captions easy to add and maintain.
When you're using Preso to build your presentations, you can describe your idea in plain English, and Preso designs the deck. If you include narration, Preso's voice-over feature generates a natural-sounding script and audio. The captions are already synchronized, so you don't have to do the timing work manually.
For sales teams creating client-ready decks, adding captions means your prospects can watch a narrated walkthrough even if they're in a noisy environment or prefer to read. For educators and trainers, captions make your training materials accessible to all students and help with retention.
If you're generating presentations programmatically via Preso's API, you can include caption data in your API request. This means your automated presentations are accessible from the moment they're created.
If you're presenting to an international audience, you need more than just translated captions. You need a cohesive experience where the narration, the captions, and the visual design all work together in each language.
Preso's multilingual feature handles this by generating the entire presentation in each target language. The narrative is rewritten for each language, not just translated. The captions are synchronized with a voice-over in that language. The design and branding remain consistent across all languages.
This approach ensures that your message lands the same way in every language, and that viewers can choose to watch with captions in their preferred language.
Once your captions are in place, you need to share your presentation in a way that preserves the captions.
If you're sharing a link to a self-running presentation, the captions will be embedded in the presentation file, so they'll display automatically.
If you're exporting to PowerPoint, Google Slides, or PDF, check that your caption file is included in the export. Some platforms require you to export the captions separately.
When you share presentations securely with Preso, you can control who has access, set expiry dates, and disable downloads if needed. The captions travel with the presentation, so your audience always sees the full experience.
If you're exporting to PDF, note that PDFs don't support audio or video, so captions won't be relevant. But if you're exporting a presentation that will be played as a video, make sure the captions are embedded.
Once you've added captions to your presentations, you might wonder if they're actually making a difference.
Watch your analytics. If your presentation platform provides engagement metrics, check whether viewers are watching longer, replaying sections, or sharing more when captions are present. Some platforms show which slides get the most attention. If captions help viewers stay engaged, you'll see it in the data.
Gather feedback from your audience. Ask viewers if the captions were helpful. If you're presenting to a diverse audience, ask specifically whether captions helped people who are deaf or hard of hearing, or people for whom English is a second language.
For sales presentations, track whether narrated decks with captions lead to more meetings or higher close rates. For training presentations, track whether students with access to captions perform better on assessments.
The investment in captions pays off in accessibility, engagement, and often in better business outcomes.
Subtitles and captions transform a narrated presentation from something that works only in perfect conditions (quiet room, native English speaker, good audio equipment) into something that works for everyone.
When you add captions, you're not just checking an accessibility box. You're ensuring that your message lands, no matter how it's consumed. You're making your presentation more engaging, more professional, and more effective.
The process is straightforward: get an accurate transcript, time it to the audio, follow accessibility standards, choose the right format, implement it in your platform, test it, and refine based on feedback. If you're presenting in multiple languages, create captions in each language and let your audience choose.
If you're building presentations with Preso, the captions are built in. Describe your idea, let Preso design the deck, add narration, and the captions are already synchronized and ready to go. For sales teams, educators, and anyone building presentations at scale, this means you can create accessible, professional presentations without the extra work.
Your next presentation deserves captions. Start with your transcript, sync it to your audio, and share a deck that everyone can follow.
Build your next narrated presentation with Preso. Describe your idea in plain English, and Preso designs a beautiful, accessible deck with synchronized captions and voice-over, ready to present itself. Try Preso today.