HomeAIAI ModelsSperid Labs Releases Iris-3B, an Open Image Model With No VAE; Its...

Sperid Labs Releases Iris-3B, an Open Image Model With No VAE; Its Own Paper Finds No Clear Gain for Depth or Restoration

Sperid Labs has released Iris-3B, an open image model with 3 billion parameters that draws images directly in pixels, with no VAE. Fine-tuned, the same model also estimates depth from one photo and enlarges damaged images four times. The paper is dated 7 October 2026 [1]. The release is not new on 11 October, and the authors call one finding a “negative result” [1].

Confirmed: what was released, the licence and the dates come from the paper, the code and the model pages. Reported: every benchmark number is the authors’ own. TSN has not re-run any of them.

What was released

  • Paper. “Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning”, by Hanqiu Li Cai and Chema Garabito of SperidLabs. It is an arXiv preprint submitted on 7 October 2026 [1].
  • Code. A GitHub repository, created on 5 October and last updated with new code on 8 October. It holds the model, training and fine-tuning code, and inference [2].
  • Weights. On Hugging Face, about 12 GB for the main model. The depth and restoration versions are separate folders, about 12 GB each [2][3].
  • Demo. A Hugging Face Space [2][4].
  • Licence. The README says the code and the weights are under Apache 2.0. Parts from others, such as the Qwen3-VL text encoder, keep their own licences [2].

TSN did not download the weights or try the demo.

The terms, in plain words

  • Diffusion transformer. A diffusion model starts from random noise and cleans it up step by step until an image appears. A transformer is the network design that does the cleaning.
  • VAE and latent space. Most image models first shrink each image with a VAE, a compression tool, into a smaller “latent” version, work on that, and decode it back at the end. Compression is lossy, so detail can be lost.
  • Pixel space. Iris-3B skips the shrinking and works on the pixels themselves, in 16-by-16 patches [1].
  • Monocular depth estimation. Guessing how far each part of a scene is from one photo. Iris-3B’s output is relative, not metric: it says near or far, not metres [2].
  • Image restoration. Cleaning up and enlarging a degraded image. Here “4x” means four times the width and height; the restorer was trained to turn 256-by-256 inputs into 1024-by-1024 outputs [2].

Why dropping the VAE is notable, and what it costs

The authors say latent models make “training and sampling at high resolution affordable”, but at a price: whatever detail the compression discards “lies beyond the generator’s reach” [1]. They expected pixel space to help where fine detail matters [1].

The paper lists costs and gaps. English long-text rendering is Iris-3B’s “clearest gap”. Depth maps show a faint patch grid. The model was “trained with modest compute and data and is not state of the art in raw generation quality” [1]. The paper gives no speed or cost comparison with latent models.

The authors’ own results

Test (self-reported)Iris-3BComparison
OneIGBench, English overall (higher is better)0.540Qwen-Image 20B: 0.539
Depth, mean AbsRel over five benchmarks (lower is better)0.071FLUX.2 Klein, latent: 0.072
DIV2K 4x restoration, LPIPS (lower is better)0.292FLUX.2 Klein, latent: 0.275

On OneIGBench the paper says Iris-3B “matches” Qwen-Image [1]. Qwen-Image is a 20B-parameter model, and its figure is copied from its own paper. On three other tests in the same table Iris-3B scored lower: GenEval 0.798 against 0.87, DPG 86.52 against 88.3, LongText 0.857 against 0.943 [1].

On depth, the paper calls Iris-3B “level” with the latent FLUX.2 Klein: better on two benchmarks, worse on three. It says the differences “should not be read as a win for either” [1]. On restoration, “neither pixel model beats a latent FLUX.2 Klein fine-tune” [1]. The paper adds that each run was single and short, with no significance test, and the Iris-3B restoration run is not matched to the FLUX.2 Klein runs [1].

Their summary: “We find no significant improvement from using a pixel-space generative prior” [1].

What is not known

  • Whether anyone outside SperidLabs has reproduced these numbers.
  • Iris-3B’s speed or running cost against latent models, or its training cost. The paper gives no figure; it says only “modest” compute.
  • Whether the paper has been peer reviewed. It is a preprint.

Sources

  1. Hanqiu Li Cai and Chema Garabito (SperidLabs), “Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning”, arXiv:2610.09450 v1, 7 October 2026 (preprint, authors’ own results; read 11 October 2026). https://arxiv.org/abs/2610.09450
  2. SperidLabs, iris-3b repository and README (primary; created 5 October 2026, last pushed 8 October 2026 per the GitHub API). https://github.com/speridlabs/iris-3b
  3. SperidLabs, Iris-3B model page on Hugging Face (primary; Apache 2.0 listed). https://huggingface.co/speridlabs/iris-3b
  4. SperidLabs, Iris-3B demo Space on Hugging Face (primary; not tried by TSN). https://huggingface.co/spaces/speridlabs/iris-3b
  5. Hugging Face, paper page for arXiv:2610.09450 (shows “Published on Oct 7”). https://huggingface.co/papers/2610.09450

Related stories

Share this story

Latest stories

More in this category

Latest stories

Free TSN tools: AI funding tracker, DePIN scorecard, AI agent cost calculator and more.