Voice generation, voice change, and translation live in Audio, while making a character talk (lip sync) lives in Lipsync Studio, which offers several models for image-to-video and video-to-video talking clips. Some video models also generate audio together with video in a single pass. Aspect ratio is set per generation inside each tool: common options are 9:16 for vertical, 16:9 for horizontal, and 1:1 for square, and not every model supports every ratio.
What voice and audio tools are available?
Voice generation lives in Audio (higgsfield.ai → Audio). A mode dial offers three actions:
Mode | What it does |
|---|---|
Voiceover | Generate a voice track and lay it over your video |
Change Voice | Replace the existing voice in a clip with a different one |
Translate | Translate and re-voice your content in another language |
Pick a built-in Voice Preset (a set of ready male and female voices), or create your own: Create custom voice → Add Voice lets you record a sample or upload one (MP3 or WAV), then clones it into a reusable voice you can select for any generation. Only clone a voice that is your own or one you have the person's permission to use.
How do I create a lip-synced (talking) video?
Lip sync lives in Lipsync Studio (higgsfield.ai → Video → Lipsync Studio, "Create Talking Clips"). Upload a character image or video, type what the character should say in Audio text (or generate the audio), pick a model, and generate. Each model offers scene templates (General, Selfie, Podcast, Car Talking, and more) to set the look of the clip.
Model | Input | Notes |
|---|---|---|
Google Veo 3 | Image to video | Cinematic talking video |
Kling 2.6 Lipsync | Image to video | Up to 1080p, with audio |
Wan 2.5 Speak | Image to video | 480p to 1080p, with audio |
Kling Avatars 2.0 | Image to video | Talking avatars, longer clips |
Higgsfield Speak 2.0 | Image to video | Priority-queue speed |
Infinite Talk | Image to video | Long-form talking video |
Kling Lipsync | Video to video | Lip sync on existing footage |
Sync Lipsync 3 | Video to video | Precise lip sync, up to 4K |
What is native audio co-generation?
Native audio co-generation means audio is produced together with video in a single pass, not layered on afterward. The result is video and audio that are physically aligned: lip sync, physics-aware sound, and ambient audio matching the visual environment.
Model | Native audio support |
|---|---|
Kling 3.0 | Sound effects, speech, music synced to on-screen action |
Seedance 2.0 | Voice, lip sync, and ambience as a unified system |
Wan 2.5 | Sound sync with camera movements |
Native audio is not available in Cinema Studio 3.5: use Cinema Studio 3.0 if you need native audio in a cinematic workflow. See How do I use Cinema Studio?
How do aspect ratios work?
Aspect ratio is set per generation inside each tool before confirming, and the available ratios depend on the model.
Ratio | Format | Best for |
|---|---|---|
9:16 | Vertical | TikTok, Reels, YouTube Shorts |
16:9 | Horizontal | YouTube, presentations, film |
1:1 | Square | Instagram feed, profile content |
4:3 | Standard | Classic video formats |
3:4 | Portrait | Pinterest, some mobile formats |
21:9 | Ultrawide | Cinematic widescreen |
Some models support auto aspect ratio, where the system infers the best ratio from your input image.
To change the aspect ratio of an existing video after generation, use the Reframe tool: higgsfield.ai → Edit → Reframe. For images, Expand outpaints the canvas to a new ratio while preserving the original content.
Can I use audio tools through MCP?
Yes. Voiceover, voice change, and video dubbing are available through Higgsfield MCP directly from Claude, and all audio generations through MCP deduct credits at standard rates. See How do I connect Higgsfield to Claude?