// Model benchmarks

AI music models, compared

A working guide for producers deciding which AI tool fits their workflow — what each model is known for, how quality actually gets measured, and where they land on the things that matter.

Content as of August 2026. This space moves fast — model versions and capabilities change often.

// All In's take

Our own comparison — not a published score

These are our editorial judgment calls from hands-on use and public output, not a running lab benchmark. Use them as a starting point for your own testing, not the final word.

Vocal realism

Rated only for models that generate vocals at all — Stable Audio is instrumental-only by design.

SunoStrong
LeVoSolid
Google LyriaSolid
UdioSolid
ElevenLabs MusicSolid
TrebloLimited
ACE-StepLimited

Instrumental quality

Stable AudioStrong
UdioStrong
Google LyriaSolid
LeVoSolid
AIVASolid
ACE-StepSolid
SunoSolid
TrebloSolid

Prompt & instruction following

SunoStrong
Google LyriaStrong
ACE-StepSolid
UdioSolid
Stable AudioSolid
LeVoSolid
TrebloLimited

Workflow / API flexibility

Only rated for tools built around integration or self-hosting — not every model competes on this axis.

MubertStrong
Google LyriaSolid
ACE-StepSolid
Stable AudioSolid

Compositional depth / arrangement

How complex and structurally coherent the underlying composition is — separate from raw audio fidelity.

AIVAStrong
Google LyriaSolid
UdioSolid
SunoSolid
// At a glance

Best for what, exactly

The comparison above in one line each — where to start if you already know what you need.

Best for vocals

01SunoThe benchmark for natural phrasing.
02LeVoBest-in-class vocal accuracy, if you're comfortable running it yourself.
03Google LyriaIncreasingly competitive, not yet frame-accurate on timing.

Best for instrumentals / sync work

01Stable AudioLicensed-safe and precise for timed work.
02UdioRich, editable arrangements.
03Google LyriaStrong structural control.

Best for developers / API

01MubertFastest, easiest integration of any tool here.
02Google LyriaPart of Google's broader API ecosystem.
03ACE-StepSelf-hosted, runs on under 4GB VRAM.

Best for composition / arrangement

01AIVAThe only tool here with real professional composer credentials.
02Google LyriaPrecise verse/chorus/bridge control.
03UdioRegenerate and fix individual sections instead of starting over.
// Commercial models

The commercial leaders

The top platforms have split into distinct niches — your pick depends on whether you need full vocal tracks, licensed-safe instrumentals, or API access.

Vocal & pop

Suno

The most common starting point for producers who want a complete song with vocals fast — strong at natural-sounding vocal phrasing and radio-ready pop structure. Dense, multi-part arrangements can still come out sounding a little tidy compared to a human mix.

Visit official site
Cinematic & remix

Udio

Leans toward atmospheric, orchestral, and electronic textures, with more unpredictable character than Suno's polished pop instinct. Its editing tools let you regenerate or fix specific sections of a track instead of starting over.

Visit official site
Instrumental & commercial-safe

Stable Audio

Instrumental-only, trained on licensed material — the practical choice for background scoring or loops when you need clean commercial rights and don't want to think about where the training data came from. Handles precise, timed generations well, which matters for sync work.

Visit official site
Licensed & clean

ElevenLabs Music

Built around the same fully-licensed-training-data promise ElevenLabs is known for on the voice side — output that's commercially safe first, experimental second.

Visit official site
Enterprise & API

Google Lyria

Positioned more for developers building music features into their own products than for a producer generating a single track by hand — expect API access over a polished consumer app.

Visit official site
Free & style-driven

Treblo (formerly Sonauto)

Built around style tags and community remixing rather than one polished output — generations are free and unlimited, with full commercial rights included by default. Raw audio quality and structural stability trail Suno's, and vocal takes can show artifacting, but it's a strong pick for producers chasing a specific, unusual style fast.

Visit official site
Loop-based & API-first

Mubert

Doesn't synthesize audio from scratch — it arranges millions of pre-recorded, human-made loops and samples into a finished track, trading Suno- or Udio-style novelty for royalty-clean source material. Its real strength is the developer API: teams use it to generate infinite, adaptive background music for games and apps rather than a single standalone song.

Visit official site
AI composer

AIVA

Positions itself as an AI composer rather than a song generator — its strength is producing complex, structured compositions that hold up as a starting point for further arrangement in a DAW, not just a finished single. It renders audio on every plan (MIDI and WAV included on top of that at the highest tier). Ownership depends on the plan: only the top tier grants full copyright ownership of what you create — lower tiers keep it with AIVA.

Visit official site
// Open-source

Open-source alternatives

For producers who'd rather run generation locally, fine-tune on their own material, or skip per-song fees entirely.

Open-source

LeVo

An open-weights model with a reputation in the open-source community for clean raw output quality — a reasonable starting point if you want to fine-tune on your own material rather than prompt a hosted model.

Open-source

ACE-Step

Another open-weights option, generally considered more flexible for downstream tooling and research use than a polished out-of-the-box product — the usual tradeoff open models make against a hosted one.

// How models get benchmarked

"Good" isn't a vibe — here's how it's measured

Beyond subjective listening, researchers lean on a handful of frameworks to compare models on consistent terms.

What is Fréchet Audio Distance (FAD)?

The standard way researchers measure raw audio fidelity without needing a human listener for every comparison. It compares the statistical distribution of a model's output against a large set of real studio recordings — the closer the two distributions, the lower (better) the FAD score.

What is a CLAP score?

CLAP (Contrastive Language-Audio Pretraining) measures how closely generated audio actually matches what you asked for, by comparing the audio and the text prompt in a shared embedding space. A higher CLAP score means tighter prompt adherence — a low score can mean great audio that ignored your prompt.

What about newer frameworks like CMI-RewardBench or SongBench?

Beyond FAD and CLAP, researchers have started building benchmarks aimed specifically at multi-layered instructions (genre, mood, structure and lyrics together) and subjective song-level qualities like memorability and structural coherence. This corner of the field is moving fast and the frameworks themselves are still maturing — treat any single score from them as directional, not definitive.

Made your track with one of these?

All In tags AI provenance honestly and distributes it to 150+ stores — you keep 100% of the royalties.

Start distributing