For a deeper breakdown of when to choose each mode, see Standard AD vs Extended AD.
- Standard AD
- Extended AD
Input Video
Output AD Video
Output AD Audio
Output AD Text (VTT)
How It Works
Add Your Video
Upload media files or provide a URL. We support all major video formats.
Configure & Generate
Choose configuration options such as audio description (AD) type, language, format, and video category.
Get Descriptions in Text, Audio or Video
Poll for completion or use webhooks.
Quick Links
Quickstart
Get started in 5 minutes
API Reference
Explore all endpoints
Authentication
Secure your requests
Understand the parameters
AD Type
AD Type
- Standard AD — Fits descriptions into existing no-speech gaps. Video runtime stays the same; some elements may go undescribed if there’s no room.
- Extended AD — Briefly pauses the video to make room for fuller descriptions. Runtime increases.
Language
Language
Set
language (BCP-47 code, default en-US) on any generation request to control narration locale. ViddyScribe supports 53 languages including English variants, Spanish, French, German, Hindi, Mandarin, Arabic, and more.See the full list in Languages and Voices.Voice
Voice
Set
voice on audio and video generation requests to pick the narrator (default Achernar). 31 voices are available (15 female, 16 male). Most cover all 53 languages; one extra Robotic voice is English-only.See the full list in Languages and Voices.Custom Instructions
Custom Instructions
Provide custom instructions to guide the AI for specific terminology, style preferences, or focus areas.
Output Format
Output Format
Every generation endpoint returns the text descriptions in the response. Set
format (default vtt) to choose how that text is serialized:json— structured array with timestamps and description textvtt— WebVTT subtitlessrt— SubRip subtitlesedl— Edit Decision List for video editors
API Endpoints
Upload Endpoints
Generation Endpoints
Results Endpoint
Need Help?
Support
Get in touch with our support team
API Reference
Explore all available endpoints

