Higgsfield organizes models into image, video, and audio families, each with a different strength. For images, start with Soul, Higgsfield's native model for cinematic and fashion looks, or Nano Banana for speed and reference control. For video, Seedance handles realistic motion and Kling handles longer or multi-shot clips, with Higgsfield DOP adding VFX and cinematic camera control. For voice and audio, use Higgsfield Audio. Choose the family that matches your goal, and Higgsfield uses the latest version automatically.
What model families does Higgsfield offer?
Higgsfield models are organized into three categories: image, video, and audio. Within each category, models are grouped into families, and each family has a distinct strength and use case. You don't need to pick a specific version: start with the family that matches your goal, and Higgsfield uses the latest available version within it.
Which model should I use for image generation?
You want to | Use this model family |
|---|---|
Generate fashion, cinematic, or culture-native visuals. Character consistency | Soul |
Create storyboards or edit across frames | Popcorn |
Generate fast, high-quality images with strong reference control | Nano Banana |
Generate images with accurate text rendering or precise color | GPT Image |
Generate images with intelligent visual reasoning | Seedream |
Generate photorealistic or expressive vector-style images | Recraft |
Generate images with speed-optimized detail | FLUX |
Upscale or enhance an existing image | the Upscale tool |
Soul is Higgsfield's native image model family, built for fashion, aesthetics, and cinematic visuals, with Soul Cinema adding a film-grade look. Its biggest strength is character consistency: train a character once as a Soul ID and reuse the same face across generations. Soul characters can also be carried into some other models by saving them as an Element, so a face you build in Soul isn't locked to Soul alone. If keeping one consistent character matters, Soul is the starting point. See How do I create and use a Soul ID character?
Popcorn is Higgsfield's native model for storyboards and multi-frame work: use it when you need a sequence of connected frames or edits that carry across them.
Nano Banana is built for speed and reference control, with advanced reference handling and swapping. Use it when you need high-quality outputs fast, or when working with product or character references.
GPT Image handles text rendering and precise color better than most generative models in Higgsfield. Use it when your image needs readable text or true-color accuracy.
Seedream is built around visual reasoning, interpreting complex prompts with more contextual understanding. Use it when your prompt is detailed or compositionally complex.
Which model should I use for video generation?
You want to | Use this model family |
|---|---|
Apply VFX and cinematic camera control | Higgsfield DOP |
Generate realistic human motion | Seedance |
Generate stylized, cinematic, or long-form video | Kling |
Generate video with strong first and last frame control | Wan |
Generate video with synchronized audio | Grok Imagine |
Generate video with OpenAI's model | Sora |
Generate video with Google's model | Veo |
Generate high-dynamic video at fast speeds | Minimax Hailuo |
Seedance is built for realistic motion, especially human movement, expressions, and scene dynamics. Use Seedance when realism is the priority.
Kling is built for longer clips, multi-shot continuity, and physically coherent motion, and it can transfer motion from one video to another. Use Kling for cinematic output or when building multi-shot scenes.
Higgsfield DOP is Higgsfield's native cinematography layer: VFX and directed camera moves applied through presets. Use it when the camera work is the point.
Wan gives you precise control over the first and last frame of a generation, useful when you need to continue a scene or match a specific start or end state.
Which model should I use for audio generation?
You want to | Use this model family |
|---|---|
Generate a complete sound scene: voice, music, and effects in one pass | Seed Audio |
Generate expressive AI voice with emotion control | Eleven |
Generate studio-quality text-to-speech | MiniMax Speech |
Generate multilingual voice | Seed Speech |
Generate long-form expressive voice synthesis | VibeVoice |
Seed Audio is an all-in-one audio model: one prompt can produce multi-speaker dialogue, emotional delivery, ambience, music, and effects together as a finished track. It also handles text-to-speech, voice cloning from a short sample, and video dubbing, and it's available through Higgsfield MCP and the Adobe Premiere and DaVinci Resolve plugins.
Audio models in Higgsfield are available through the Voiceover, Change Voice, and Translation features, and the model you choose affects voice quality, expressiveness, and language support. Eleven is the best choice when emotional nuance matters, handling tone, pacing, and expression well. Seed Speech is the multilingual option for content that needs to work across languages. See How do I use lipsync, voiceover, and aspect ratios?
How do models connect to tools?
Models are the engines. Tools like Cinema Studio, Canvas, Marketing Studio, and Supercomputer are where you use them in Higgsfield, and most tools let you choose which model to generate with. For example, Cinema Studio supports multiple video model families, so you can switch between them depending on the scene. For cinematic multi-shot video with full camera control, use Cinema Studio, which is a tool, not a model.
For a full overview of Higgsfield tools and when to use each one, see Which Higgsfield tool should I use?