Skip to main content

Overview

AI Talking Photo brings static photos to life by animating faces to speak with realistic lip-sync and natural facial movements. The API analyzes facial features and synchronizes mouth movements, head poses, and expressions with the audio file you provide. Processing: See recent typical API-job times. Jobs run asynchronously, and duration varies with the input, selected settings, and queue load.

How It Works

  1. Provide a photo - Upload an image with a clear face
  2. Add audio - Provide an audio file (required). To generate speech from text, create the audio first with the AI Voice Generator
  3. API animates - AI creates realistic lip-sync and facial movements
  4. Download video - Retrieve your animated talking photo

Use Cases

  • Marketing videos - Create spokesperson videos from headshots
  • Educational content - Animate historical figures or characters
  • Personalized messages - Send video messages from static photos
  • Social media - Create engaging content from profile pictures
  • Presentations - Add dynamic talking heads to slides

Best Practices

Photo Selection

Use clear, front-facing photos - Best results come from high-quality headshots with visible facial features.
  • Good lighting - Well-lit faces produce better animations
  • Front-facing angles - Avoid extreme profile shots
  • Clear features - Eyes, nose, and mouth should be unobstructed
  • High resolution - At least 512x512 pixels recommended

Audio Guidelines

Generation modes

Set style.generation_mode to realistic, the default, to preserve likeness, or prompted to guide the scene with style.prompt. The maximum selected audio duration is 300 seconds for realistic and 45 seconds for prompted. The older pro, standard, stable, and expressive values are deprecated. Use the current mode names for new integrations; style.intensity is also deprecated.

Code Examples

Basic Talking Photo

start_seconds and end_seconds are required — they control which segment of the audio is used. Video pricing is duration-based, so shorter segments cost less.

Pricing

Talking Photo pricing depends on the output duration, end_seconds minus start_seconds. The create response reports estimated credits_charged; read the completed job for the final amount.

Resolution Limits

Use max_resolution to constrain the larger output dimension in pixels. The API caps this setting at your plan maximum; supported output sizes also depend on the tool. See Resolution Limits for current limits.
Try this in our Google Colab Cookbook: Run this API with sample code. Just add your API key.

API Reference

AI Talking Photo API Reference

View full API specification

Lip Sync

Sync audio with existing video lip movements

AI Voice Generator

Generate speech audio for your talking photos