HomeAIAI ModelsLAION Releases Humaneness Voice Converter, an Open Voice-Conversion Model Built on Chatterbox;...

LAION Releases Humaneness Voice Converter, an Open Voice-Conversion Model Built on Chatterbox; the Scores Are LAION’s Own

LAION has released Humaneness Voice Converter, an open model that re-voices one recording in another person’s voice. It is built on Chatterbox, an open-source speech model from Resemble AI. Hugging Face’s records show the model page was created on Sunday 11 October 2026. The model card gives no date [1][2].

Confirmed: the model card says everything below. Self-reported: every score is LAION’s own result. TSN has not run the model or checked any score.

What voice conversion is

Voice conversion takes a recording of one person speaking and makes it sound like someone else. The words, timing and emotion of the original stay. Text-to-speech is different: it starts from typed text and reads it aloud.

LAION’s card says the model is “an audio-to-audio converter, not a text-to-speech model” [1]. You give it two recordings: the source speech, and a short recording of a “different, consenting” target speaker. It “attempts to say the source content in the target voice while retaining much of the source performance” [1].

In LAION’s tested command the target clip is 5 to 10 seconds long. The card advises a clean clip with one speaker. The output is a 24 kHz audio file [1].

What it is built on

Resemble AI describes Chatterbox as a family of open-source text-to-speech models. Its GitHub repository uses the MIT licence [4]. LAION’s card says its model is “based on Resemble AI’s Chatterbox”, and that LAION developed the voice-conversion design and training [1].

What “reward weighting” means

In its final training stage, LAION trained all 112,620,864 parameters of the model’s flow generator, the part that builds the sound, for 40,000 more rounds. In each round the model made 16 attempts, and attempts that scored better were favoured [1].

“Reward weighting” is the marking scheme. It says how much each goal counts toward one score. LAION weighted three goals [1]:

  • Content Enjoyment, 40%: a score for how pleasant the audio is.
  • Emotion, 30%: how closely the output keeps the source’s two strongest emotions.
  • Speaker similarity, 30%: how close the output is to the target voice.

The scores

LAION tested 40 curated source recordings against two target voices, with one output each. Higher is better, except for emotion error, where lower is better [1].

MeasureUntouched ChatterboxHumaneness Voice Converter
Content Enjoyment6.1696.453
Production Quality7.7668.009
Emotion error (lower is better)0.1130.092
Target speaker cosine0.8210.773

TSN reports these numbers and does not call the model better overall. The card says it is not a blinded listening study [1].

Why the judge matters

Content Enjoyment, Production Quality and emotion are scored by a LAION model, Humaneness Ears Medium. The card says it is “the same judge family used in training” [1].

That is like practising against a marking scheme and then being marked by it. High scores show the model learned what the judge rewards. They do not show listeners agree. The card adds that the test sources were curated for input quality, and that 10 of them overlap the training source list [1].

It also says there is “currently no independent human MOS or WER result” [1]. MOS is a listener rating. WER is the share of words that come out wrong.

Why the speaker-similarity drop matters

Speaker cosine is a number for how close the output voice is to the target’s. It fell from 0.821 to 0.773. The card calls this “a trade-off in target-speaker consistency” [1].

For a voice converter, sounding like the target is the main job. So the model gains polish and emotion, and loses some likeness. The card says the score is only “a proxy”, and a converted voice “is not a verified speaker identity” [1].

LAION also lists an optional best-of-16 mode, which picks the best of 16 takes. Its speaker cosine is 0.795 and its emotion error is 0.072 [1]. It costs more computing.

Licence and responsible use

LAION’s additions are under CC BY 4.0. The bundled Chatterbox parts keep their MIT licence [1][3].

Any voice-conversion tool can be misused. LAION asks users to get consent, to “disclose converted speech”, and not to use the model for “impersonation, fraud, deception or identity spoofing”. It calls these ethical recommendations, not extra legal conditions on the licence [1].

What is not known

  • Independent results. The card lists no human listening or word-error result.
  • Other voices and languages. English and German are in the training and test material. The card makes no multilingual promise.
  • Exact release time. The card gives none.

Sources

  1. LAION, “Humaneness Voice Converter” model card, Hugging Face (primary; read 11 October 2026). https://huggingface.co/laion/Humaneness-Voice-Converter
  2. Hugging Face model record via its API (created and last-modified times; read 11 October 2026). https://huggingface.co/api/models/laion/Humaneness-Voice-Converter
  3. LAION, NOTICE.md in the same repository (attribution and licences). https://huggingface.co/laion/Humaneness-Voice-Converter/raw/main/NOTICE.md
  4. Resemble AI, Chatterbox repository and MIT licence, GitHub (company repository). https://github.com/resemble-ai/chatterbox
  5. ArtRealmAI, “LAION Humaneness Voice Converter: Open Chatterbox Voice Conversion That Keeps the Emotion”, 11 October 2026 (secondary report; also dates the release 11 October). https://artrealmai.com/article/laion-humaneness-voice-converter-open-chatterbox-vc

Related stories

Share this story

Latest stories

More in this category

Latest stories

Free TSN tools: AI funding tracker, DePIN scorecard, AI agent cost calculator and more.