You have a headshot. You have a script. You don’t have a camera, a presenter, or an afternoon to spare. D-ID can turn that photo into a talking avatar that delivers your script in about 20 minutes. I’ve used it for client explainers, product demos, and even a birthday message. The first attempt took 22 minutes from login to export. The second took nine. This guide is the process I wish I’d had at the start.
What D-ID actually does
D-ID is a text-to-video and photo-to-video platform. You upload a still image (or choose a stock avatar), type a script, pick a voice, and the system animates the face to match the audio. The mouth, eyes, and slight head movements are generated. It’s not a deepfake tool for impersonation; it’s a production shortcut for explainers, e-learning, and social clips. The company calls the core technology Creative Reality, and it works best for head-and-shoulders framing.
Before you start: three things to prepare
- A clean headshot. Front-facing, even lighting, eyes open, mouth closed, no sunglasses or heavy shadows. A 512×512 pixel crop is the minimum; 1024×1024 is better.
- A script of 80 to 120 words. That produces roughly 30 to 45 seconds of video. Longer scripts work, but the avatar’s subtle movements can start to loop, and you’ll notice it.
- A voice decision. D-ID offers hundreds of stock voices in dozens of languages. You can also upload your own recorded audio if you want your actual voice. Decide before you start typing.
Step 1: Choose a plan without overspending
The free trial gives you a few minutes of video and adds a small watermark. That’s enough to test the lip-sync quality with your own photo. Paid plans start around $5.99 per month for 10 minutes of video, scaling up to $49 for 180 minutes. If you’re comparing tools, a curated list of free AI generators worth using in 2025 includes D-ID alternatives with similar free tiers. And if you’re unsure what free really means, this breakdown of what real free AI tools can do is refreshingly honest about watermarks and export limits.
Step 2: Upload a photo or pick a stock avatar
Click Create Video and choose Photo Avatar. Drag in your image. D-ID will detect the face and show a preview. If the face is off-center or tilted, the animation will look off. Use the crop tool to center the eyes. Stock avatars are fine for internal training, but they have a generic, stock-photo quality. For anything customer-facing, use a real photo. I once used a client’s LinkedIn headshot and the result looked more trustworthy than their previous studio video.
Step 3: Write a script that survives text-to-speech
D-ID’s text-to-speech is good, but it punishes certain habits. Short sentences work best. Avoid abbreviations like e.g. or 24/7 because the voice will read them literally. Spell out numbers under 100: twenty-four seven instead of 24/7. Add commas where you want a pause. For example, Our new roast, the Ethiopian Guji, has notes of blueberry gives the voice a natural beat. Test your first sentence with a 10-second sample before you commit to the full script.
Step 4: Match the voice and language to your audience
Scroll through the voice list. You can filter by language, accent, and gender. For a UK audience, pick a British English voice; for a US audience, pick a General American one. D-ID also lets you adjust speed and pitch with sliders. A voice that’s 10% slower often sounds more authoritative for instructional content. If you want your own voice, record a clean audio file (no background noise, 44.1 kHz WAV) and upload it. The lip-sync will follow your recording.
Step 5: Add motion and background (without making it weird)
D-ID offers gesture options like slight head nods or hand movements, but the avatars are head-and-shoulders. Less is more. A solid color background or a softly blurred office works better than a busy stock video. If you add a logo, place it in a corner and keep it small. The goal is to keep the viewer’s attention on the face.
Step 6: Generate and fix the three most common glitches
- Lip-sync drift on plosives. The mouth can lag on p, b, and m sounds. Fix it by adding a comma before the word or rephrasing to avoid a cluster of plosives. Buy Big Boxes becomes Buy, big boxes.
- Robotic intonation. If the voice sounds flat, shorten the sentence and add a question mark. D-ID’s engine often lifts pitch at question marks. You can also switch to a different voice; some are more expressive than others.
- Unnatural blinking. The avatar may blink too often or not enough. In the settings, there’s a blinking toggle. Turn it off for a more static, professional look, or leave it on for a friendlier feel. Test both with a 5-second clip.
Step 7: Export and edit for social media
D-ID exports an MP4 with the audio embedded. For TikTok, Reels, or Shorts, you’ll want a 9:16 crop. Drop the file into CapCut, Descript, or Premiere and scale it. Add captions because most social video is watched on mute. If you need more control over gestures or full-body avatars, HeyGen is a common alternative, and this explainer of what AI avatars can and can’t do for your video covers the trade-offs without hype. For a deeper look at D-ID’s quirks, there’s a balanced D-ID review that flags the pricing and lip-sync limits.
A real example: 45-second product demo for a coffee roastery
A client sent me a photo of their founder, Maria, standing in front of a brick wall. They wanted a video announcing a new Ethiopian roast. I wrote a 90-word script: Our new Ethiopian Guji has notes of blueberry and dark chocolate. It’s available in 250-gram bags starting today. Order online and we’ll roast it tomorrow morning. I picked a warm female voice, added a blurred brick background, and generated the video in four minutes. The first take had a slight lip-sync lag on blueberry. I added a comma: Our new Ethiopian Guji, has notes of blueberry. The second take was clean. They posted it to Instagram and got 12,000 views in three days. No camera, no studio, no reshoots.
When D-ID is the wrong tool
D-ID shines for talking-head clips under two minutes. It struggles with full-body movement, complex hand gestures, and real-time conversation. If you need a presenter walking through a factory, hire a videographer. If you need a live avatar that answers questions, D-ID’s Agents feature exists but adds latency and cost. And if you need to edit every syllable, a human on camera will always be faster. Knowing the limits saves you from forcing the tool into a job it can’t do.
Your first video: a 15-minute checklist
Open D-ID. Upload a headshot. Paste a 100-word script. Pick a voice. Set the background to a solid color. Generate. Watch it once with the sound off to check for visual glitches. Watch it again with your eyes closed to check pacing. If it passes both, export and crop. If it fails, change one variable at a time. That’s the whole loop. The second video is always faster than the first.

