Martini Logo

AI Video Models Supported by Martini

We integrate the industry's leading AI video generators including Veo 3.1, Kling 3.0 Pro, Kling O3 Pro, Sora 2 Pro, and more. Martini is a web-based editor that lets you do everything directly in the browser.

Create cinematic AI-generated shots in our cloud-based timeline, then export as XML to import seamlessly into DaVinci Resolve, Adobe Premiere Pro, or any other editing software.

Many models support start frame control for precise matching with existing shots. End frame control is available with Veo 3.1, Kling 3.0 Pro, Kling O1, Hailuo 02, and Seedance 1.5 for even tighter editorial workflows.

Video Generation Models

These AI models generate video from text prompts or still images, creating motion sequences perfect for professional video production. All models are available directly in the Martini web editor.

FLUX.3

Black Forest Labs' video model. Up to twenty seconds in a single generation with native audio, ten keyframes pinned to exact timecodes, and a draft mode you can render up.

  • Start frame
  • End frame
  • Keyframes
  • Audio
  • Extend

Veo 3.1

High-quality video generation with native audio from Google DeepMind. Best for final outputs with start and end frame control.

  • Start frame
  • End frame
  • Audio

Veo 3.1 Fast

Faster variant of Veo 3.1 with native audio. Lower cost with quicker generation times.

  • Start frame
  • End frame
  • Audio

Kling 3.0 Pro

Top-tier video generation with cinematic visuals, fluid motion, native audio, and element support.

  • Start frame
  • End frame
  • Audio
  • Elements
  • Multi-shot

Kling O3 Pro

Video generation with multi-shot support, extended durations up to 15s, and reference-to-video.

  • Start frame
  • End frame
  • Audio
  • Multi-shot

Kling 2.6

Top-tier video generation with cinematic visuals, fluid motion, and native audio support.

  • Start frame
  • Audio

Kling O1

High-quality image-to-video generation with smooth motion and end frame control.

  • Start frame
  • End frame

Sora 2 Pro

OpenAI's state-of-the-art video model with native audio, capable of creating richly detailed, dynamic clips from text or images.

  • Start frame
  • Audio

Sora 2

OpenAI's video model with native audio at 720p. An affordable option for dynamic clips from text or images.

  • Start frame
  • Audio

MiniMax H3

Multimodal 2K video generation with native stereo audio and image, video, or audio references.

  • Start frame
  • End frame
  • Audio
  • Reference media
  • 2K

Hailuo 02

MiniMax's advanced video generation model with 1080p resolution and start/end frame control.

  • Start frame
  • End frame

Seedance 2.5

ByteDance's latest Seedance model, with native 30-second single takes, an Auto duration option, end frame support, and a 50-file reference budget.

  • Start frame
  • End frame
  • Audio
  • Reference images
  • Elements
  • Auto duration
  • 30s

Seedance 2.0

ByteDance's most advanced video generation with cinematic output, native audio, real-world physics, and director-level camera control.

  • Start frame
  • End frame
  • Audio
  • Reference images
  • Elements
  • 4K

Seedance 2.0 Fast

Faster, more affordable Seedance 2.0 with cinematic output, native audio, and reference-to-video, capped at 720p for quicker, cheaper generations.

  • Start frame
  • End frame
  • Audio
  • Reference images
  • Elements

Seedance 2.0 Mini

The quickest, most affordable native Seedance 2.0 tier with audio and reference media, capped at 720p.

  • Start frame
  • End frame
  • Audio
  • Reference images
  • Elements

Seedance 1.5

High-quality video generation from ByteDance with native audio support and flexible durations from 4-12 seconds.

  • Start frame
  • End frame
  • Audio

Grok Imagine

Video generation from xAI with flexible durations from 1-15 seconds and a wide range of aspect ratios.

  • Start frame

Grok Imagine 1.5

Text, image, and reference video generation from xAI with native audio, durations up to 15 seconds, and up to 1080p output.

  • Start frame
  • References
  • Audio
  • 1080p

Pixverse V6

High-quality video generation with native audio, flexible durations up to 15 seconds, and art style presets.

  • Start frame
  • Audio

LTX 2.3

Lightricks' open-source 4K video model with native audio. Supports start/end frame control and durations up to 10 seconds.

  • Start frame
  • End frame
  • Audio

Happy Horse 1.0

Alibaba's #1-ranked video model with native audio, reference-to-video support (up to 9 images), and durations up to 15 seconds.

  • Start frame
  • Audio
  • Reference images

Luma Ray 3.2

Luma's video-to-video model — re-render an existing clip into new cinematic motion from a text prompt, with an optional start frame, at up to 1080p.

  • Video to video
  • Start frame

Gemini Omni Flash

Google's Gemini Omni Flash — fast text-to-video, image-to-video, and reference-to-video with native synchronized audio and role-bound reference images.

  • Start frame
  • Reference images
  • Elements
  • Audio

Still Generation & Image Editing Models

These models specialize in still image generation and editing, perfect for creating reference frames, concept art, or generating motion from static images.

Qwen Image 3.0 Pro

High-fidelity image generation for dense typography, layouts, and precise edits. Add up to three references to preserve identity and details.

  • Reference images
  • Dense typography
  • Precise editing

Grok Imagine Image 2.0

xAI's image model for sharp typography and complex layouts, with photorealistic fidelity across photography, design, and illustration. Add up to three references to edit elements and styles.

  • Reference images
  • Sharp typography
  • Photorealistic fidelity

Reve 2.1

Native 4K image generation with precise layout planning, spatial control, and multilingual text. Add references to edit or remix.

  • Reference images
  • Native 4K
  • Strong text rendering

Nano Banana 2

Fast, affordable image generation with excellent typography and realism. Supports multi-image blending with character consistency.

  • Reference images
  • Web search

Nano Banana 2 Lite

Google's fastest, most affordable image model. Fixed 1K output with reference-image blending and natural-language editing.

  • Reference images

Seedream 5 Lite

ByteDance's fast, affordable image model. High-resolution 2K–4K output with flexible aspect ratios. Add reference images to edit or blend.

  • Reference images
  • 2K–4K output

Seedream 5 Pro

ByteDance's flagship image model. Sharp high-resolution output with strong text rendering and flexible aspect ratios. Add reference images to edit or blend.

  • Reference images
  • Strong text rendering

Nano Banana Pro

State-of-the-art image generation with excellent typography and realism. Supports multi-image blending with character consistency.

  • Reference images
  • Web search

Flux 2 Max

High-quality image generation with flexible aspect ratios and safety controls. Add reference images for creative editing and blending.

  • Reference images

GPT Image 2

OpenAI's GPT Image 2 — photoreal image generation with strong typography and instruction following. Add reference images to edit an existing image.

  • Reference images

Audio Generation Models

Generate music, sound effects, dialogue, and voiceover. Sonilo SFX can also synchronize foley and ambience to an existing shot. Audio outputs remain editable assets in Martini.

MiniMax Music

Generate complete music tracks from a prompt, with optional vocals and custom lyrics.

  • Music
  • Instrumental or vocals
  • Custom lyrics

MiniMax Music 3

Generate a complete song up to five minutes long, holding theme, vocal identity and arrangement steady across the whole track.

  • Music
  • Up to 5 minutes
  • Custom lyrics

Lyria 3

Generate 30-second music sketches with singing vocals and negative-prompt control.

  • Music
  • 30 seconds
  • Vocals

Lyria 3 Pro

Generate complete songs up to about three minutes with vocals, timed lyrics, and full arrangements.

  • Music
  • Up to 3 minutes
  • Timed lyrics

ElevenLabs Music

Generate vocal or instrumental music from three seconds to ten minutes, with optional section planning.

  • Music
  • 3 seconds–10 minutes
  • Section planning

Stable Audio 2.5

Generate stereo music, foley, and ambient sound design up to 190 seconds.

  • Music and SFX
  • Up to 190 seconds
  • Stereo

Seed Audio 1.0

Generate dialogue, sound effects, and music together from one prompt, or create clean speech.

  • Music, SFX, and dialogue
  • Up to 2 minutes
  • Voice presets

ElevenLabs SFX v2

Generate short foley, ambience, impacts, and atmospheric sounds from text.

  • Sound effects
  • 0.5–22 seconds
  • Text to SFX

Beatoven SFX

Generate stereo effects for creatures, vehicles, impacts, science fiction, and abstract textures.

  • Sound effects
  • 1–35 seconds
  • Stereo

CassetteAI SFX

Generate fast sound-effect placeholders and draft variations up to 30 seconds.

  • Sound effects
  • 1–30 seconds
  • Fast generation

Sonilo SFX

Generate standalone sound effects from text, or add a source video for frame-synchronized foley and ambience.

  • Text to sound effects
  • Video synchronization
  • 1–180 seconds
  • AAC, MP3, WAV, FLAC

ElevenLabs TTS v3

Generate expressive dialogue and voiceover with preset voices, language control, and delivery tags.

  • Dialogue and voiceover
  • 21 voices
  • 70+ languages

MiniMax Speech-02 HD

Generate natural narration and character dialogue with broad voice and language support.

  • Dialogue and voiceover
  • 17 voices
  • 40+ languages

Frequently Asked Questions

Yes. Martini is a web-based editor that lets you generate Veo 3 shots directly in the browser. You can then export your project as XML and import it seamlessly into Adobe Premiere Pro, DaVinci Resolve, or any other editing software.

Yes. Kling 3.0 Pro, Kling O3 Pro, and Kling O1 all support start and end frame control, enabling tight in/out precision for editorial workflows. This allows you to match your AI-generated content exactly to your existing shots for seamless continuity.

Martini supports leading AI video generation models including Veo 3.1, Kling 3.0 Pro, Kling O3 Pro, Kling 2.6, Kling O1, Sora 2 Pro, Hailuo 02, Seedance 1.5, and Grok Imagine. For still image generation and editing, we support Nano Banana Pro and Flux 2 Max. All models are available directly in the Martini web editor, allowing you to choose the best model for each shot in a single unified interface.

Yes. Martini supports Nano Banana Pro for high-quality image generation with excellent typography and multi-image blending, and Flux 2 Max for flexible aspect ratios and creative editing with reference images. Generate reference frames directly in your timeline.

For music, Martini supports MiniMax Music, Lyria 3, Lyria 3 Pro, ElevenLabs Music, and Stable Audio 2.5. Sound-effect models include ElevenLabs SFX v2, Beatoven SFX, CassetteAI SFX, Sonilo SFX, and Stable Audio 2.5. Dialogue and voiceover models include ElevenLabs TTS v3 and MiniMax Speech-02 HD. Seed Audio 1.0 can generate music, sound effects, and dialogue together.

Martini works with any professional video editor that supports XML import, including Adobe Premiere Pro and DaVinci Resolve. Simply create in our web editor and export your timeline to XML. This gives you the flexibility of a cloud-based AI studio with the power of your preferred desktop NLE.

Yes. End frame control is available with Veo 3.1, Kling 3.0 Pro, Kling O3 Pro, Kling O1, Hailuo 02, and Seedance 1.5, giving you precise control over both the start and end frames of your AI-generated videos for tighter editorial workflows. This feature is essential for matching existing shots and maintaining continuity in professional productions.

Martini combines a professional timeline interface with AI generation in the browser. Instead of managing scattered video files, you build your sequence in our web editor and export to XML, seamlessly bridging the gap between AI generation and professional post-production.

Yes. We offer a creative partnership program for studios, agencies, and independent creators working on notable projects. Partners receive priority support, early access to new features, collaborative opportunities, and custom integrations. Contact us at koh@c47.studio to learn more about partnership opportunities.

Ready to get started?

For creatives, studios, and agencies who create professional AI videos.