Skip to main content

Overview

Text to Video generates complete video content from text descriptions using advanced AI video synthesis. The API creates original video scenes, animations, and visual narratives based on detailed text prompts with customizable styles and durations. Processing: See recent typical API-job times. Jobs run asynchronously, and duration varies with the input, selected settings, and queue load.

How It Works

  1. Write a prompt - Describe the video you want to create
  2. Set duration - Choose how long the video should be
  3. API generates the video - AI creates original video content
  4. Download the result - Retrieve your generated video

Use Cases

  • Social media content - Create engaging videos from ideas
  • Marketing videos - Generate product and promotional content
  • Concept visualization - Bring written ideas to visual life
  • Educational content - Create explainer and demonstration videos
  • Creative projects - Artistic and experimental video creation

Best Practices

Writing Effective Prompts

Be specific and descriptive - Include subject, action, environment, style, and camera motion.
✅ Good prompts:
  • “A majestic lion walking through golden savanna grass at sunset, cinematic slow motion, warm golden lighting”
  • “Underwater scene with colorful tropical fish swimming around a coral reef, crystal clear blue water, nature documentary style”
  • “Futuristic city skyline at night with flying cars and neon lights, cyberpunk aesthetic, sweeping aerial shot”
❌ Avoid:
  • Too vague: “A nice video”
  • No action: “A city” (add what’s happening)
  • Conflicting instructions: “Fast and slow motion”

Prompt Structure

For best results, include these elements:

Duration Guidelines

Model selection

model="default" currently selects kling-3.0 on paid tiers and ltx-2.5 on the free tier. Set a model explicitly when your workflow needs a fixed duration and resolution combination. The examples use ltx-2.5, which supports five-second clips at 480p; the paid default does not support 480p. Other current options include gemini-omni-1.1, minimax-h3, seedance-2.0-mini, seedance-2.5, and veo3.1-lite. Check the API reference for each model’s supported durations, resolutions, and audio settings. Audio is disabled by default and is unavailable on some models.

Code Examples

Basic Text to Video

Pricing

Text to Video pricing depends on the model, resolution, and duration you choose. You pay for the frames that render. The create response estimates credits_charged; read the completed job for the final cost.

Resolution Limits

Supported resolutions are 360p, 480p, 720p, 1080p, and 4k, depending on the model and your subscription tier. For example, kling-3.0 supports 4k. Output defaults to 720p on paid tiers and 480p on the free tier. See Resolution Limits for tier details.
Try this in our Google Colab Cookbook: Run this API with sample code. Just add your API key.

API Reference

Text to Video API Reference

View full API specification

Image to Video

Animate static images into videos

Animation

Create animated videos with motion effects