whattAI
How-to By whattAI Team ·

How to Make Your First AI Avatar Video with Synthesia in Under 20 Minutes

A step-by-step guide to creating your first AI avatar video with Synthesia , from account setup to finished video , without a camera, microphone, or editing experience.

If the thing standing between you and video content is the idea of setting up lights, looking decent on camera, and recording 12 takes of the same sentence, Synthesia removes most of that friction. You type a script, pick an avatar, and the platform generates a talking-head video from your text. No camera. No microphone. No editing timeline.

This guide walks you through creating your first Synthesia video from scratch, including the planning steps that save you from wasted render credits, the script decisions that determine whether the result looks natural or stilted, and the real limitations you should understand before committing to a paid plan.

Disclosure: This article contains affiliate links. If you sign up to Synthesia through a link on this page, we may earn a commission at no extra cost to you.

What You Need Before You Start

Time: 15-20 minutes for a 1-2 minute video once your script is ready. Longer videos take proportionally more scripting time; the actual platform work stays fast regardless of length.

A script or topic: Synthesia works from text, so you need to know what you want the avatar to say before you open the editor. A rough outline is enough to start, you can refine in the tool, but going in blank wastes time.

A free account: Synthesia’s free (Basic) plan gives you 10 minutes of video per month with no credit card required. That’s enough to complete this guide and evaluate whether the platform fits your needs. Paid plans start at $29/month (Starter) and $89/month (Creator).

No camera, microphone, or software: Everything runs in a browser. You do not need video editing experience, Synthesia’s editor is roughly as complicated as a PowerPoint slide.

Step 1: Plan Your Video Before Opening the Editor

This is the step most first-time users skip, and it’s the one that most affects the result.

Synthesia’s AI avatar reads your script at a fixed speaking pace. Unlike filming yourself, there’s no room for a spontaneous aside, an adjusted emphasis, or a natural pause to let something sink in. The script is the performance, so it pays to write it carefully before you generate anything.

Work through these three questions:

What is this video for, and who will watch it? A 60-second product explainer for a website homepage needs a different structure and tone than a 4-minute internal training walkthrough. The platform supports both, but the scripting approach is different.

What’s the one thing the viewer should walk away knowing or doing? Every scene in a Synthesia video should connect to that outcome. Scenes that don’t connect to it tend to get skipped.

How long does it need to be? Synthesia recommends 2-4 short sentences per scene, with 12-23 scenes as a rough target for a well-paced video. A 1-minute video needs roughly 130-150 words of script; a 3-minute video needs around 400-450 words. Overly long videos with low completion rates are a common first-timer mistake, err shorter than you think you need.

Step 2: Write Your Script Using the FOCA Structure

The quality of your script determines 80% of your output. A well-written prompt generates an avatar that sounds engaged; a rushed one sounds like a terms-and-conditions reading.

Synthesia recommends the FOCA framework for most video types:

ElementWhat it doesExample
FocusA hook in the first 5 seconds”Most people waste their first 30 minutes in Synthesia doing this wrong.”
OutcomeWhat the viewer will learn or be able to do”By the end of this video, you’ll have a finished AI avatar video ready to share.”
ContentYour main message, broken into short scenesStep-by-step scenes, one idea each
ActionA clear CTA at the end”Try it free at the link below.”

A few scripting rules that consistently improve Synthesia output:

  • Write the way you’d say it out loud, not the way you’d write a sentence in an email. Short, direct phrases work better than long compound sentences.
  • Keep each scene to one idea. Cramming three points into one avatar segment usually results in unnatural pacing.
  • Avoid long lists narrated out loud, they are tedious to listen to. If a scene is naturally list-shaped, use on-screen text for the list items and have the avatar introduce the list briefly.
  • Read the script aloud before you generate anything. If it sounds robotic when you read it, it will sound robotic coming from the avatar.

Step 3: Create Your Account and Start a New Video

Go to synthesia.io and create a free account. No credit card required for the Basic plan.

Once you’re in the dashboard, click New video and then AI video generator. You’ll be prompted to either paste a script, upload a file (PDF, PowerPoint, Word, or text), or enter a prompt and let Synthesia generate an outline for you.

Best path for a first video: paste your script. The AI outline generator is useful for longer projects but can produce oddly structured scenes for short, focused videos. Pasting your own script gives you direct control over the result.

After pasting your script, click Generate. Synthesia will produce a scene-by-scene outline with a draft script for each scene. Review this and adjust before moving to the editor, it’s faster to fix structure here than after you’re inside the editor.

Step 4: Choose Your Avatar

In the editor, click Avatar at the top of your screen to open the avatar library.

Synthesia has 240+ stock AI avatars as of mid-2026, organized by gender, age range, and style (casual, business, studio). For a first video, ignore the custom avatar options and pick a stock avatar that fits your use case:

  • For professional or corporate content: David, Emma, or similar business-style presenters
  • For casual, conversational content: look for avatars in the “casual” filter
  • For multilingual content: most stock avatars support all 160+ available languages

A few things that trip up new users on avatar selection:

Don’t change avatars between scenes in your first video. Switching avatars mid-video can feel jarring unless you have a deliberate reason. A single consistent presenter is the safer default.

Preview before committing. Every avatar has a short preview clip. Watch it, then imagine your actual script being delivered in that voice and pace, not just the sample content.

The Express-2 avatars look significantly more natural. These are the newer diffusion-model avatars (you’ll see “Express-2” in the label). They gesture more like a human speaker and are worth the extra few seconds of selection time over the older avatar styles.

Step 5: Select a Voice and Language

After choosing an avatar, select the voice and language from the dropdown near the script panel.

Synthesia automatically matches a voice to the language of your script. If you pasted English text, you’ll be assigned an English voice by default. You can change both the voice and language independently.

A few useful options most first-timers miss:

Voice speed control: You can adjust speaking pace per paragraph or for the whole video. Dense, technical content benefits from a slightly slower pace; short, punchy content often sounds more natural at a slightly faster rate. The default is usually fine, but it’s worth a quick preview.

Multiple languages in one video: You can apply different voices to different sections of your script if you need a bilingual segment. Useful for training content aimed at multilingual teams.

Custom voice (Starter and above): The Starter plan includes a personal avatar, which means you can create a digital twin of your own voice and appearance by filming a short consent video. This is worth exploring after you’ve tried a stock avatar first.

Step 6: Build Your Scenes and Add Supporting Visuals

The editor looks like a simplified presentation tool: scenes run down the left panel, the canvas is in the center, and your avatar and script controls are at the bottom and top.

For each scene, you can:

  • Add on-screen text and shapes to reinforce key points without narrating every word
  • Add motion graphics for transitions, data visualizations, or brand elements
  • Add stock or AI-generated B-roll to run behind or alongside the avatar
  • Add a screen recording, Synthesia has a built-in screen recorder, useful for software demos and walkthroughs
  • Add interactive elements (Creator plan and above), clickable hotspots, quizzes, and branching paths that change what viewers see based on their choices

For a first video, keep it simple: avatar + a few lines of on-screen text per scene + basic branding. Adding too much visual complexity to a first project is how you spend 90 minutes on a 90-second video.

The most important visual rule: anything that is naturally list-shaped should appear as scannable on-screen text, not be read out by the avatar. Narrated bullet lists are one of the most common ways AI avatar videos start to feel tedious.

Step 7: Run Through the Pre-Generate Checklist

Before you click Generate, run through this checklist. Regenerating a video costs credits, so a two-minute review here is worth it.

Script:

  • Does each scene communicate one clear idea?
  • Does the first 5 seconds have a hook that earns continued watching?
  • Have you read the full script out loud at least once?
  • Is there a clear CTA near the end?

Visuals:

  • Are fonts and colors consistent across scenes?
  • Is anything narrated out loud that would be clearer as on-screen text?

Pacing:

  • Have you previewed the video using the Play button?
  • Do transitions between scenes feel natural?

Accessibility:

  • Are fonts large enough to read easily?
  • Have you enabled captions for viewers watching without audio?

Step 8: Generate and Export

Click Generate in the top-right corner. A short video (1-2 minutes) typically renders in a few minutes. Longer videos take proportionally longer.

Once it’s rendered, watch the full video at least once before exporting. Pay attention to:

  • Scenes where the pacing feels rushed or too slow (adjust script length in those scenes)
  • Any word the avatar mispronounces (use the pronunciation editor in the script box to fix individual words)
  • Any visual element that looks misaligned or doesn’t land (fix in the editor and regenerate just that scene)

When you’re satisfied, click Export to download an MP4, generate a shareable link, or embed the video on a webpage. Enterprise plans include SCORM export for LMS integration.

What Synthesia Does Well: and Where It Falls Short

It’s worth being direct about the limitations before you commit to a paid plan.

Synthesia is well-suited for:

  • Internal training, onboarding, and compliance videos where authenticity is less critical than consistency and scalability
  • Product demos and software walkthroughs paired with screen recordings
  • Multilingual content, one script, rendered into 160+ languages with the AI dubbing feature
  • Any situation where a company needs a repeatable, on-brand video process without a production crew

Synthesia is a poor fit for:

  • Consumer-facing brand content where emotional authenticity drives conversion, trained eyes still notice AI delivery, and it can subtly undercut trust in a brand video
  • Content that requires a real person’s specific credibility (a founder message, a thought leadership talk, an investor update)
  • Sales prospecting where a personal, human touch matters, email + real video clip typically outperforms AI avatar content here

Practical pricing reality: The free plan (10 minutes/month) is enough to evaluate the platform but not enough to build a real workflow. The Starter plan at $29/month (billed monthly) or approximately $18/month (billed annually) gives you 10 minutes per month with the ability to download videos and remove the Synthesia watermark. The Creator plan at $89/month (billed monthly) or approximately $64/month annually adds 30 minutes per month, API access, and interactive video features. If you’re doing regular video production for clients or training teams, Creator is the functional plan, Starter is primarily for solo experimentation.

For context on how Synthesia fits into a broader freelance or content production stack, the best AI tools for freelancers in 2026 roundup covers where it sits alongside tools for writing, voice, and editing.

Get Started with Synthesia

Common Mistakes to Avoid

Writing scripts that read well but don’t speak well. The avatar reads exactly what you type. Compound sentences that feel fine on paper become hard to follow when spoken at a fixed pace. Write for speech, not for reading.

Choosing an avatar based on appearance alone. The voice matters as much as the face. Always preview the avatar with audio before committing.

Adding visuals as decoration. Every on-screen element should serve the viewer. If it’s there to “look busy,” cut it. A cleaner video almost always feels more professional than a busy one.

Generating before previewing. The preview button exists for a reason. Use it before spending credits on a generate.

Ignoring the first 5 seconds. Avatar videos have the same drop-off pattern as YouTube videos, viewers decide quickly whether to keep watching. If your opener doesn’t land, most of your viewers won’t see the rest of it.

Using it for everything. Synthesia is a strong tool in a specific range. Knowing when not to use it is part of using it well. If you’re building a video content strategy and thinking about which content types work best for AI tools versus real filming, the blog to YouTube script workflow guide covers the creative decisions around AI-assisted video content more broadly.

Expected Outcome

Following the steps in this guide, a first-time Synthesia user can produce a polished 1-2 minute AI avatar video in 20 minutes or less, including script drafting time. The result is a professional-looking talking-head video with a consistent avatar, on-screen text, and correct audio sync, ready to download or share.

For longer or more complex videos (interactive quizzes, multilingual versions, screen-recording composites), add 15-30 minutes per additional complexity layer. The platform scales well once you’ve worked through a first video and understand how scene structure and script pacing interact.


References

Related articles