Overview
AI Talking Photo brings static photos to life by animating faces to speak with realistic lip-sync and natural facial movements. The API analyzes facial features and synchronizes mouth movements, head poses, and expressions with the audio file you provide. Processing: See recent typical API-job times. Jobs run asynchronously, and duration varies with the input, selected settings, and queue load.How It Works
- Provide a photo - Upload an image with a clear face
- Add audio - Provide an audio file (required). To generate speech from text, create the audio first with the AI Voice Generator
- API animates - AI creates realistic lip-sync and facial movements
- Download video - Retrieve your animated talking photo
Use Cases
- Marketing videos - Create spokesperson videos from headshots
- Educational content - Animate historical figures or characters
- Personalized messages - Send video messages from static photos
- Social media - Create engaging content from profile pictures
- Presentations - Add dynamic talking heads to slides
Best Practices
Photo Selection
- Good lighting - Well-lit faces produce better animations
- Front-facing angles - Avoid extreme profile shots
- Clear features - Eyes, nose, and mouth should be unobstructed
- High resolution - At least 512x512 pixels recommended
Audio Guidelines
Generation modes
Setstyle.generation_mode to realistic, the default, to preserve likeness, or prompted to guide the scene with style.prompt. The maximum selected audio duration is 300 seconds for realistic and 45 seconds for prompted.
The older pro, standard, stable, and expressive values are deprecated. Use the current mode names for new integrations; style.intensity is also deprecated.
Code Examples
Basic Talking Photo
start_seconds and end_seconds are required — they control which segment of the audio is
used. Video pricing is duration-based, so shorter segments cost less.Pricing
Talking Photo pricing depends on the output duration,end_seconds minus start_seconds. The create response reports estimated credits_charged; read the completed job for the final amount.
Resolution Limits
Usemax_resolution to constrain the larger output dimension in pixels. The API caps this setting at your plan maximum; supported output sizes also depend on the tool. See Resolution Limits for current limits.
API Reference
AI Talking Photo API Reference
View full API specification
Related Tools
Lip Sync
Sync audio with existing video lip movements
AI Voice Generator
Generate speech audio for your talking photos