backblaze-labs/awesome-video-generation
A curated list of AI video generation APIs, SDKs, and tools including text-to-video, video editing, multimodal generation, diffusion models, and generative AI platforms. Covers commercial services, open source models with APIs, and production-ready infrastructure for developers building video applications.
Build with Backblaze B2
SDKs, agent skills, IDE extensions, and reference pipelines from Backblaze Labs. All open source.
Loading star history...
Use Cases & Benefits
- Curates a comprehensive list of AI video generation APIs, SDKs, open source models, and tools for building video applications.
- Provides a centralized resource covering commercial services, open source models, and production-ready infrastructure for easy developer integration.
- Use for discovering text-to-video and image-to-video generation APIs to integrate AI video creation into applications.
- Use for exploring avatar and talking head APIs to build interactive, lip-synced, or real-time streaming avatar video features.
- Use for finding open source video generation models and SDKs to develop custom AI-powered video editing and enhancement workflows.
About awesome-video-generation
Awesome Video Generation

A curated list of AI video generation APIs, SDKs, and production-ready tools. Focused on services developers can integrate today.
Maintained by Backblaze.
Related Lists
- Awesome Image Generation
- Awesome Audio Generation
- Awesome ML Data Pipelines
- Awesome Multimodal Data
- Awesome Agent Infrastructure
- Awesome Physical AI
Contents
- Text-to-Video APIs
- Real-Time and Interactive Video
- Avatar and Talking Head APIs
- Video Style Transfer and Motion
- Video Enhancement and Understanding
- Open Source Models
- SDKs and Developer Tooling
- Infrastructure and Deployment
- Evaluation and Observability
- Templates and Example Projects
Text-to-Video APIs
Commercial text-to-video and image-to-video APIs with developer access.
- Stability AI (SVD) – Image-to-video via Stable Video Diffusion. Hosted API deprecated July 2025; open weights available for self-hosting. Docs
- Fliki – Text-to-video and text-to-speech platform. Enterprise API with 2,500+ voices in 80+ languages. Docs
- Google Veo 2 / Veo 3 – Google's video generation models via Vertex AI and Gemini API. Veo 2 is GA; Veo 3 in paid preview. Docs | SDK: Google Cloud Python
- HappyHorse-1.0 – Native 1080p text-to-video and image-to-video with integrated audio and lip-sync in one pass. Task-based REST with Bearer auth. Docs
- Higgsfield – Cinematic video platform aggregating 15+ premium models (Sora 2, Kling 2.6, Veo 3.1, etc.) with camera simulation, character consistency, and lip-sync. 15M+ users.
- InVideo AI – Turns text prompts into full videos using Sora 2 and Veo 3.1 as underlying models. OpenAI's first official Sora 2 integration partner. 50M+ users.
- Kling AI – Text-to-video and image-to-video from Kuaishou. Up to 30s clips at 1080p/30fps. Async task-based API. Also on fal.ai. Docs
- Krea – Unified API for 10+ video models (Veo 3/3.1, Sora 2, Kling 2.6, Wan 2.5, Hailuo 2.3, Runway Gen-4.5, Ray 2, Seedance Pro). Job-based with webhooks. OpenAPI spec, Python/Node/Go examples. Docs | SDK: Python, Node, Go
- Luma Dream Machine – High-quality text-to-video with character reference and style reference inputs. Ray 3 is the latest model. Docs | SDK: Python, JS
- Magic Hour – Multi-modal AI video generation API. Text-to-video, image-to-video, style transfer, 4K upscaling. Scales to zero when idle. Docs
- MiniMax / Hailuo – Hailuo 2.3 model. Text-to-video and image-to-video up to 1080p, 10s clips. Docs | SDK: Python, Node
- Morph Studio – No-code AI video studio aggregating Wan 2.6, Kling 2.6 Pro, Seedance, Sora 2, Veo 3 into a single canvas with storyboarding and style transfer.
- OpenAI Sora – Text-to-video and image-to-video via the v1/videos endpoint. Sora 2 supports up to 90s at 4K with spatial audio. Docs | SDK: Python, Node
- Pika (v2.2) – Text-to-video and image-to-video with Pikaframes multi-keyframe interpolation. API powered by fal.ai. Docs | SDK: fal Python, fal JS
- Runway (Gen-4) – Text-to-video and image-to-video with Gen-4 Turbo. Async task-based REST API with polling helpers. Docs | SDK: Python, Node
- Seedance 2.0 (ByteDance) – Dual-Branch Diffusion Transformer for simultaneous video + audio generation. Up to 15s at 2K resolution. Available via Dreamina. Docs
- Vidu (Shengshu Technology) – Now on Vidu Q3, the first long-form AI video model with native audio-video generation in a single output. Ranked
- Wan 2.7 (Alibaba) – Commercial successor to open-weight Wan 2.2. Native 1080p, 2–15s clips, first-and-last-frame control, up to 5 reference inputs, instruction-based video editing. Via DashScope and fal.ai at $0.10/sec. Docs
- xAI Aurora / Grok Imagine – Text-to-video and image-to-video using xAI's Aurora autoregressive MoE model. 6–15s clips at 720p with synchronized audio. Docs
Real-Time and Interactive Video
Low-latency video transformation and interactive streaming platforms.
- Krea Realtime 14B – Open-weight 14B autoregressive video model distilled from Wan 2.1 via Self-Forcing. ~11fps on a single B200, ~1s time-to-first-frame. WebSocket streaming server for mid-generation prompt edits. Research/non-commercial license. Docs
- Causal Forcing – Autoregressive diffusion distillation for real-time interactive video generation on a single RTX 4090. Frame-wise and chunk-wise inference; builds on Wan 2.1. HuggingFace weights at zhuhz22/Causal-Forcing. Docs
- CausVid – CVPR 2025. Distills a bidirectional video diffusion transformer into a 4-step autoregressive generator. Enables streaming video generation at 9.4 FPS on a single GPU via KV caching. Supports T2V, I2V, and video-to-video translation. Docs
- Decart (Lucy 2) – Real-time video transformation at 30fps 1080p with near-zero latency. Live-stream style transfer, character swaps, environment transformation, product placement. ~$3/hour. Docs
- HY-WorldPlay (Tencent) – Streaming video diffusion model for real-time interactive world generation at 24fps. Accepts image or text prompt; responds to keyboard/mouse camera inputs. WorldPlay-5B and 8B weights open. Built on HunyuanVideo 1.5. Docs
- PixVerse – Text-to-video and image-to-video platform. PixVerse-R1 adds real-time interactive video at 720p HD with native audio. Docs
Avatar and Talking Head APIs
Avatar video generation, lip-sync, and real-time streaming avatars.
- Anam AI – Real-time interactive avatar API. CARA-3 model, 180ms median latency at 25fps/720x480. Connects to any LLM. JavaScript and Python SDKs. Docs | SDK: JavaScript, Python
- Captions / Mirage – Mirage API generates hyperrealistic talking-head videos from script + image + actor ID. Natural gestures, eye contact, synchronized audio. Docs
- Colossyan – Enterprise avatar video platform. 130+ avatars, 600+ voices, 100+ languages, instant avatar creation from phone recording. Docs
- D-ID – Talking head video generation from text or audio. Express and Premium+ avatars, real-time WebRTC streaming. Docs | SDK: Python
- DeepBrain AI (AI Studios) – AI avatar video platform with REST API. Integrates AWS, Azure, ElevenLabs, IBM Watson, NVIDIA Riva. Docs
- Elai.io – AI video platform with streaming avatar API for interactive e-learning. Turns documents and scripts into avatar-presented videos. Docs
- Hedra – Character-3 omnimodal talking-avatar API (image + text + audio in one pass). Long-form up to 10 min; LiveKit plugin for realtime; Node SDK, REST, Make.com integration. Omnia Fast Alpha adds full scene control. Docs | SDK: Node
- HeyGen – AI avatar video generation and real-time streaming avatars via WebRTC. Template-based workflows. Docs | SDK: JS/TS
- Hour One – AI avatar video generator with 100+ presenters, voice cloning, 100+ languages. API + Zapier integration.
- Simli – Real-time speech-to-video avatar API using 3D Gaussian splatting for full-face animation (not just lip-sync). WebRTC streaming, <200ms latency. Python and JS SDKs. Docs | SDK: Python, JavaScript
- Sync.so – Studio-grade lipsync and visual dubbing API. sync-3 flagship at native 4K with obstruction detection; lipsync-2-pro for diffusion super-res; react-1 for expressive emotion. Batch up to 500 videos per job. From $0.02/sec. Docs | SDK: Python, TypeScript
- Synthesia – Avatar-based video creation from scripts. 140+ languages, custom avatars, template workflows. API in beta. Docs
- SynthLife – Virtual AI influencer creation. Creates AI personas for TikTok, YouTube, Instagram with auto-scheduling and unlimited content generation.
- Tavus – Conversational video AI. Phoenix-4 model does real-time gaussian-diffusion facial synthesis at ~600ms latency. Replica API clones face + voice. Integrates with Pipecat and LiveKit. Docs
Video Style Transfer and Motion
Video-to-video restyling and motion-transfer tools.
- DomoAI – AI video-to-video style remixer. 50+ styles (anime, Ghibli, cinematic). v2.4.1 supports text-to-video, image-to-video, talking avatars, and animation.
- Viggle AI – Motion-transfer video tool that animates static characters to match a motion video or live webcam input. Mix Mode, Live Mode, and VTubing support. 40M+ users.
Video Enhancement and Understanding
Upscaling, restoration, and multimodal understanding APIs for video.
- FlashVSR – CVPR 2026. One-step diffusion framework for streaming 4x video super-resolution at ~17 FPS for 768x1408 on a single A100. Locality-constrained sparse attention and tiny conditional decoder. Apache-2.0; v1.1 weights on HuggingFace. Docs
- REAL Video Enhancer – Cross-platform GUI/CLI for frame interpolation (RIFE), upscaling (Real-ESRGAN, Waifu2x), denoising, and decompression. TensorRT and NCNN backends. Windows/macOS/Linux; v2.4.1 released Jan 2026.
- Topaz Video AI – AI video upscaling, denoising, frame interpolation, and artifact correction. 3M+ users; used by Google, Tesla, NASA. Docs
- Twelve Labs – Video understanding and intelligence API. Multimodal semantic search, analysis, and text generation from video content. Docs
- Video2X – C/C++ framework for AI video super-resolution and frame interpolation. Supports Real-ESRGAN, Real-CUGAN, and RIFE via Vulkan. GUI and CLI. Docker images on GHCR. Docs
Open Source Models
Open-weight video models with hosted inference, Docker, or diffusers support.
- Open-Sora – Open reproduction of Sora-like generation. 2s–15s at 144p–720p. T2V, I2V, V2V.
- Wan 2.1 (Alibaba) – SOTA open T2V-14B model. Also supports I2V, editing, T2I, and V2A. T2V-1.3B runs on consumer GPUs. Also on Replicate and fal.ai. Docs
- Wan 2.2 (Alibaba) – First open-source MoE video diffusion model. +65.6% more image training data and +83.2% more video data vs 2.1. 5B and 14B variants. Docs
- CogVideoX (Zhipu AI / Z.AI) – CogVideoX-5B flagship; supports 10s videos. Commercial product "Ying" available via API. Docs
- AnimateDiff – Plug-and-play animation module for Stable Diffusion models. Merged into HuggingFace diffusers. Docs
- HunyuanVideo (Tencent) – 13B+ param model; v1.5 is 8.3B and runs on consumer GPUs. I2V, Avatar, and Foley variants available. On Replicate and fal.ai. Docs
- LTX-Video / LTX-2 (Lightricks) – First DiT-based real-time video gen model. LTX-2 adds native 4K at 50fps with synchronized audio. ComfyUI nodes available. Docs
- SkyReels (SkyworkAI) – V1: human-centric video fine-tuned on HunyuanVideo. V2: infinite-length video via Autoregressive Diffusion-Forcing. V3: multimodal reaching closed-source SOTA levels. Docs
- MAGI-1 (Sand AI) – 24B param autoregressive denoising model. Generates video chunk-by-chunk (24 frames/chunk). T2V, I2V, V2V with streaming generation. Outperforms Wan 2.1 and HunyuanVideo on benchmarks. Docs
- Step-Video-T2V (StepFun) – 300B parameter text-to-video model, up to 204 frames, bilingual (EN/ZH). Docs
- Pyramid Flow – Efficient autoregressive video generation using pyramidal flow matching. Up to 10s at 768p, 24fps. ICLR 2025. Docs
- LongCat-Video (Meituan) – 13.6B foundation model for text-to-video, image-to-video, and video continuation, tuned for efficient long-form 720p generation. FlashAttention-2 acceleration. Streamlit demo included. Avatar variant released Dec 2025. Docs
- OmniAvatar – Audio-driven full-body avatar video generation with adaptive body animation. Pixel-wise multi-hierarchical audio embedding for diverse scenes. 1.3B and 14B variants (LoRA on Wan 2.1). Local inference via CLI. Docs
- Allegro (Rhymes AI) – 2.8B param VideoDiT. 6s clips at 720p/15fps. Merged into diffusers. Docs
- MOVA (OpenMOSS) – Open foundation model for synchronized video+audio generation in a single inference pass. Asymmetric dual-tower architecture with cross-attention fusion. Multilingual lip-sync and environment-aware SFX. 360p and 720p weights. Diffusers integration planned. Docs
- Duix Avatar – Offline AI avatar video toolkit. Clone appearance and voice to create digital humans, then drive them via text or audio. Fully offline; Windows and Ubuntu support. Free commercial use under community license. Docs
- EchoMimicV2 (Ant Group) – CVPR 2025. Audio-driven semi-body human animation from a single image + audio. Supports English and Mandarin. Accelerated inference (~50s for 120 frames on A100). Gradio UI included. Docs
- LivePortrait (KwaiVGI) – Efficient portrait animation via stitching and retargeting control. 12.8ms/frame on RTX 4090. Humans and Animals modes. Gradio UI and HuggingFace Space. Docs
- Meta Movie Gen – 30B param T2V + 13B audio model. Personalized video from single reference photo, local/global editing, synchronized audio. Research paper public; weights not yet released. Rolling out inside Instagram Reels.
- Mochi 1 (Genmo) – 10B param T2V model with AsymmDiT architecture. 5.4s at 30fps. On fal.ai and Replicate.
- MuseTalk (Tencent Music) – Real-time audio-driven lip-sync in latent space. 30fps+ on a V100. v1.5 improves identity consistency. Supports Chinese, English, and Japanese. Docs
- NVIDIA Cosmos – World foundation model for physical AI (robotics, autonomous vehicles). Cosmos-Predict2.5 generates physics-based video simulations from text/image/video/sensor inputs. Docs
- OmniHuman-1 (ByteDance) – Multimodal human video generation from single image + motion signal (audio, video, or text). Full-body, any aspect ratio. On Replicate.
- Wan-Alpha – CVPR 2026 highlight. Text-to-video model that outputs video with transparent alpha channels via a Shiftable RGB-A Distribution Learner. Enables compositing with arbitrary backgrounds. v1.0 and v2.0 weights on HuggingFace; ComfyUI nodes available. Docs
SDKs and Developer Tooling
Client SDKs and libraries for integrating video generation into apps.
- HuggingFace Diffusers – The canonical PyTorch library for diffusion models including video pipelines. Docs | SDK: Python (pip install diffusers)
- MiniMax MCP Server – Model Context Protocol servers for video gen, TTS, and voice cloning. Docs
- Replicate SDK – Python/JS client for 100+ hosted video models. Async, streaming, webhooks, fine-tuning. Docs | SDK: Python (pip install replicate)
- fal.ai SDK – Serverless AI inference with Python, JS, and Swift SDKs. Hosts Kling, Veo, Pika, Wan, LTX, Luma, and more. Docs | SDK: Python (pip install fal-client), Node (npm install @fal-ai/client)
- Luma AI SDK – Sync and async clients for all Dream Machine generation modes. Docs | SDK: Python (pip install lumaai)
- FastVideo – Unified post-training and inference framework for accelerating video diffusion models. pip install fastvideo. Supports Wan2.1/2.2 distillation; 3x speedup via SageAttention and Teacache. Scales to 64 GPUs. Docs | SDK: Python
- HeyGen Streaming Avatar SDK – TypeScript SDK for real-time WebRTC interactive avatar sessions. SDK: TypeScript (npm install @heygen/streaming-avatar)
- Remotion – React framework for programmatic video generation. Define videos as React components, render server-side via CLI or Lambda (AWS). TypeScript-first; supports any CSS, Canvas, SVG, WebGL. Docs | SDK: TypeScript
- Runway SDK – Official Python and Node.js SDKs with type annotations, async support, and built-in polling. SDK: Python (pip install runwayml), Node (npm install @runwayml/sdk)
- Shotstack – Cloud video editing API. Render data-driven and AI-generated videos at scale via JSON templates. Node, Python, PHP, and Ruby SDKs. Docs | SDK: Node, Python, PHP, Ruby
- Wan2GP – Low-VRAM inference GUI and server for Wan 2.1/2.2, HunyuanVideo, LTX-Video, and Flux. Runs on 6 GB VRAM via int8/fp8/GGUF/NV FP4 quantization. Supports AMD RDNA and older Nvidia GPUs. Gradio web UI with mask editor and motion designer. Custom community license; free for non-commercial use.
Infrastructure and Deployment
GPU platforms, video processing, storage and CDN, and playback libraries.
- HandBrake – Open-source video transcoder wrapping FFmpeg. GUI and CLI. Docs
- hls.js – JavaScript HLS playback via MSE. Used by major streaming platforms.
- Shaka Player – Google's open-source DASH + HLS player.
- Backblaze B2 – S3-compatible object storage at low cost. Free egress via Cloudflare. Docs | B2 integration
- Cloudflare Stream – Video upload, encoding, and CDN delivery billed per minute watched. Live streaming via RTMP/SRT. Docs
- CoreWeave – Kubernetes-native AI cloud with enterprise-scale GPU infrastructure.
- fal.ai – Serverless inference for generative media. 600+ models. Python, JS, Swift SDKs. Docs
- FFmpeg – Industry-standard multimedia processing. Encode, decode, transcode, stream, filter. Docs
- Lambda Labs – On-demand H100/B200 GPUs. SSH and JupyterLab access with REST API for instance management.
- Modal – Serverless Python-first GPU platform. Container spin-up in ~1 second. Docs
- Mux – API-first video infrastructure. Upload, encode, stream (VOD + live), analytics. SDKs for Node, Python, Ruby, Go, and more. Docs
- Pollo AI – Video API aggregator providing access to Kling, Veo 3.1, Runway, Hailuo, Wan 2.6, and Pollo 2.0. Docs
- Replicate – Serverless model hosting. Run open-source video models via REST API. Docs
- RunPod – GPU pods (persistent) and serverless endpoints. REST, GraphQL, and CLI. Docs
- SiliconFlow – Managed inference API for open-source video models including Wan2.1/2.2 T2V and I2V, and HunyuanVideo-HD. OpenAI-compatible REST API at api.siliconflow.cn/v1. English docs available. Docs
- Together AI – Inference API for 200+ open models plus Instant Clusters for self-service GPU clusters.
- WaveSpeedAI – Fast AI inference with no cold starts. 600+ models. 30–50% cheaper than HuggingFace Inference. 99.9% uptime SLA. Docs
Evaluation and Observability
Benchmarks, leaderboards, and perceptual-quality metrics for video.
- VBench / VBench-2.0 – Comprehensive benchmark for video generative models. 16 fine-grained dimensions including subject consistency, motion smoothness, temporal flickering. VBench-2.0 adds Physics and Commonsense evaluation. Docs
- Artificial Analysis Video Arena – Elo-based blind-comparison leaderboard for text-to-video and image-to-video models, with separate tracks for audio-enabled output. Embeddable leaderboard widgets and HuggingFace Space. Docs
Templates and Example Projects
Reference implementations, demos, and starter projects.
- fal.ai Next.js Video Generator – Official Next.js template with queue management and TypeScript. One-click Vercel deploy. Docs
- B2 Video Object Detection with Transformers – Video object detection pipeline using HuggingFace Transformers with Backblaze B2 cloud storage integration. B2 integration
- Google Gemini Streamlit + Cloud Run – Sample app using Gemini multimodal with Streamlit, deployable to Cloud Run.
- HeyGen Streaming Avatar Demo – Next.js/TypeScript starter for real-time WebRTC avatar sessions.
- Stability AI SVD Streamlit Demo – Streamlit demo scripts for Stable Video Diffusion.
Contributing
Contributions are welcome. See CONTRIBUTING.md. One entry per PR — edit entries.yaml only and let the maintainers regenerate README.md.
License
Released under CC0 1.0 Universal. You may copy, modify, and redistribute without attribution.
About Backblaze B2
Backblaze B2 Cloud Storage is S3-compatible object storage designed for AI and media workloads. This list is maintained as part of our work making B2 a convenient storage layer for AI workflows.
Discover Repositories
Search across tracked repositories by name or description