Back to tools

D-ID

D-ID is an AI talking-avatar and digital human tool that helps creators, marketers, and training teams turn photos plus scripts or voice into presenter-style short videos.

Tool categories
ImageVideo
Tool links

Tool overview

Based on the available evidence, D-ID appears to have real adoption among Chinese users for the specific job of making a photo “speak.” The adoption signal is positive, but it looks more like a mature niche tool than a general AI video platform. Heat proof mainly comes from multiple Zhihu posts and tool-directory mentions. Usability proof is stronger in step-by-step tutorials and one comparative hands-on review, where a user explicitly ranked D-ID above alternatives for photo-to-video quality and ease of use.

Its practical use is fairly narrow and clear: upload a front-facing portrait or choose a preset avatar, enter a script and use built-in TTS, or upload your own audio, then generate a presenter-style video. It is not a text-to-cinematic-video generator for complex scenes or camera motion. A more accurate analogy is a “talking portrait video workstation” or “digital spokesperson maker.” One source mentions a conversational digital-human demo tied to ChatGPT, but the evidence here is much stronger for Studio-based talking-head production than for a full real-time avatar system.

The barrier to entry seems low.

Related social content