LAION has released Humaneness Voice Converter, an open model that re-voices one recording in another person’s voice. It is built on Chatterbox, an open-source speech model from Resemble AI. Hugging Face’s records show the model page was created on Sunday 11 October 2026. The model card gives no date [1][2].
Confirmed: the model card says everything below. Self-reported: every score is LAION’s own result. TSN has not run the model or checked any score.
What voice conversion is
Voice conversion takes a recording of one person speaking and makes it sound like someone else. The words, timing and emotion of the original stay. Text-to-speech is different: it starts from typed text and reads it aloud.
LAION’s card says the model is “an audio-to-audio converter, not a text-to-speech model” [1]. You give it two recordings: the source speech, and a short recording of a “different, consenting” target speaker. It “attempts to say the source content in the target voice while retaining much of the source performance” [1].
In LAION’s tested command the target clip is 5 to 10 seconds long. The card advises a clean clip with one speaker. The output is a 24 kHz audio file [1].
What it is built on
Resemble AI describes Chatterbox as a family of open-source text-to-speech models. Its GitHub repository uses the MIT licence [4]. LAION’s card says its model is “based on Resemble AI’s Chatterbox”, and that LAION developed the voice-conversion design and training [1].
What “reward weighting” means
In its final training stage, LAION trained all 112,620,864 parameters of the model’s flow generator, the part that builds the sound, for 40,000 more rounds. In each round the model made 16 attempts, and attempts that scored better were favoured [1].
“Reward weighting” is the marking scheme. It says how much each goal counts toward one score. LAION weighted three goals [1]:
- Content Enjoyment, 40%: a score for how pleasant the audio is.
- Emotion, 30%: how closely the output keeps the source’s two strongest emotions.
- Speaker similarity, 30%: how close the output is to the target voice.
The scores
LAION tested 40 curated source recordings against two target voices, with one output each. Higher is better, except for emotion error, where lower is better [1].
| Measure | Untouched Chatterbox | Humaneness Voice Converter |
|---|---|---|
| Content Enjoyment | 6.169 | 6.453 |
| Production Quality | 7.766 | 8.009 |
| Emotion error (lower is better) | 0.113 | 0.092 |
| Target speaker cosine | 0.821 | 0.773 |
TSN reports these numbers and does not call the model better overall. The card says it is not a blinded listening study [1].
Why the judge matters
Content Enjoyment, Production Quality and emotion are scored by a LAION model, Humaneness Ears Medium. The card says it is “the same judge family used in training” [1].
That is like practising against a marking scheme and then being marked by it. High scores show the model learned what the judge rewards. They do not show listeners agree. The card adds that the test sources were curated for input quality, and that 10 of them overlap the training source list [1].
It also says there is “currently no independent human MOS or WER result” [1]. MOS is a listener rating. WER is the share of words that come out wrong.
Why the speaker-similarity drop matters
Speaker cosine is a number for how close the output voice is to the target’s. It fell from 0.821 to 0.773. The card calls this “a trade-off in target-speaker consistency” [1].
For a voice converter, sounding like the target is the main job. So the model gains polish and emotion, and loses some likeness. The card says the score is only “a proxy”, and a converted voice “is not a verified speaker identity” [1].
LAION also lists an optional best-of-16 mode, which picks the best of 16 takes. Its speaker cosine is 0.795 and its emotion error is 0.072 [1]. It costs more computing.
Licence and responsible use
LAION’s additions are under CC BY 4.0. The bundled Chatterbox parts keep their MIT licence [1][3].
Any voice-conversion tool can be misused. LAION asks users to get consent, to “disclose converted speech”, and not to use the model for “impersonation, fraud, deception or identity spoofing”. It calls these ethical recommendations, not extra legal conditions on the licence [1].
What is not known
- Independent results. The card lists no human listening or word-error result.
- Other voices and languages. English and German are in the training and test material. The card makes no multilingual promise.
- Exact release time. The card gives none.
Sources
- LAION, “Humaneness Voice Converter” model card, Hugging Face (primary; read 11 October 2026). https://huggingface.co/laion/Humaneness-Voice-Converter
- Hugging Face model record via its API (created and last-modified times; read 11 October 2026). https://huggingface.co/api/models/laion/Humaneness-Voice-Converter
- LAION, NOTICE.md in the same repository (attribution and licences). https://huggingface.co/laion/Humaneness-Voice-Converter/raw/main/NOTICE.md
- Resemble AI, Chatterbox repository and MIT licence, GitHub (company repository). https://github.com/resemble-ai/chatterbox
- ArtRealmAI, “LAION Humaneness Voice Converter: Open Chatterbox Voice Conversion That Keeps the Emotion”, 11 October 2026 (secondary report; also dates the release 11 October). https://artrealmai.com/article/laion-humaneness-voice-converter-open-chatterbox-vc
Related stories
- AudioShake Launches The Refinery, Speaker-Separated, Quality-Scored Audio for AI Training, by Its Own Account
- The Voice AI Revolution: How Speech Technology Is Reshaping Human-Computer Interaction
- Amazon Nova 2.5 Sonic: AWS Says Its Voice AI Now Thinks and Answers Faster
- Odyssey Opens Its Odyssey-3 World Model to the Public; Its Headline Physics Score Is a Best-of-Eight, and Every Benchmark Is Odyssey’s Own





