How to use Multimodal AI Video Generator
Create AI videos from a text prompt and an image or video reference in one multimodal workflow.
- Visual reference Use <Picture 1>, <Video 1> and <Audio 1> in your prompt. Click a label to insert it. Video references use visuals only; upload sound separately.
- Video description Describe the scene, motion, and desired look...
- Start processing Run the tool and let it process the file.
- Download the result Done! Your AI-generated video is ready for download.
Why people use this tool
- Prompts and results are not stored.
- Anonymous and private.
- Powered by premium AI models.
- Fast and simple workflow.
FAQ - Multimodal AI Video Generator
- What is multimodal AI video generation and how does the tool work?
- Multimodal AI video generation combines a descriptive text prompt with visual and acoustic references (images, videos, and audio). Instead of processing only plain text or a single still image, the AI model evaluates multiple media inputs simultaneously. This allows you to precisely control characters, camera movement, scene composition, lighting, and mood, producing a coherent, high-resolution video clip.
- How do I use reference tags like <Picture 1>, <Video 1>, and <Audio 1> in the prompt?
- When you upload reference files, they are numbered sequentially. In your prompt, you can type tags such as <Picture 1>, <Video 1>, or <Audio 1>, or insert them by clicking their preview badge. Instruct the AI specifically how to use each reference—for example: "The character from <Picture 1> walks through a futuristic city with the dynamic camera motion from <Video 1>". Video references provide visual guidance and motion; audio tracks guide sound and atmosphere.
- Which file formats, file sizes, and limits are supported for uploads?
- You can upload up to 3 images (JPG, PNG, or WebP up to 10 MB each), 1 video reference (MP4, WebM, or MOV between 2 and 10 seconds, up to 30 MB), and up to 2 audio files (MP3, WAV, M4A, FLAC, or OGG up to 10 MB each). The combined total size of all uploaded reference files cannot exceed 32 MB.
- What is the difference compared to traditional text-to-video or image-to-video?
- In traditional text-to-video, the AI must guess appearance, environment, and motion purely from words, often yielding unpredictable results. Image-to-video usually just animates a single static image. Multimodal generation combines the strengths of both: you can anchor character identity with images, dictate choreographies or camera angles with a reference video, and refine details with text and sound—giving you maximum creative control and visual consistency.
- How are the aspect ratio and resolution of the video determined?
- The tool automatically derives the optimal aspect ratio from your uploaded image or video reference to preserve framing and composition. Alternatively, you can select standard aspect ratios in the settings before generating, such as widescreen (16:9), vertical for Reels, TikTok, and Shorts (9:16), square (1:1), or 4:3.
- How many credits does generation cost and how does billing work?
- The price is calculated server-side: without a video reference, 90 Credits + 5 Credits for every additional started second after 3 seconds, +10 with at least one image, and +5 per standalone audio reference. With a video reference, it is 150 Credits +10 for every additional started second after 3 seconds, +5 with at least one image, and +5 per audio reference. Duration is 3–10 seconds. Example: 10 seconds with a video reference costs 220 Credits. The exact price is shown on the button before starting; no Credits are charged if the request fails or is cancelled.
- How long does generation take and what video format will I receive?
- Because multimodal AI models analyze and synthesize multiple media sources at once, processing typically takes around 2 to 4 minutes. The finished video is delivered as a high-quality MP4 file that you can play directly in the integrated browser player and download to your device without quality loss.
- Can I use the generated AI videos commercially?
- Yes, all videos generated using premium credits can be used freely for commercial projects, social media channels, YouTube, marketing campaigns, and client work in accordance with our terms of service. You retain full usage rights to your generated content.
- Are my data and generated media stored?
- No. To protect your privacy, all uploaded media and prompts are temporarily transmitted for processing, deleted immediately upon completion, and not stored on our servers. No one will ever see them except you. Your chat history and all generated results are stored exclusively locally in your browser on your own device.
Tool currently unavailable
This tool has been switched off temporarily. It should be back shortly — until then, our other tools are ready for you.