All research

2.9 billion pairs in, 105 million out: building a text-to-image model from scratch

Jasper Research published the full recipe for a 4B text-to-image model that renders 1024 by 1024 in under a second. The most interesting part is how much training data got thrown away.

Jasper Research published a technical report in September 2026 covering the construction of a text-to-image model end to end: the variational autoencoder ablations, the denoiser architecture, the dataset, pre-training, post-training, and distillation. The result is a 4 billion parameter model that generates a 1024 by 1024 image in under a second.

Most reports like this spend their pages on architecture. The part of this one worth reading closely is the data, because the headline number is not the model size. It is that 2.9 billion raw image and text pairs were reduced to 104.9 million, and the report explains every stage of the reduction.

The case against more data

The prevailing assumption for years was that scale substitutes for curation. The large open datasets were built on that premise: YFCC100M, LAION-400M and LAION-5B, COYO-700M, all gathering hundreds of millions to billions of web-crawled pairs. They worked, in the sense that they enabled CLIP and Stable Diffusion and much of what followed.

They are also, in the report's own description, largely uncurated, highly redundant, and paired with short, noisy alt-text. Those three properties cause distinct problems. Uncurated means safety and quality vary without bound. Redundant means compute is spent repeatedly on near-identical images. And noisy alt-text means the text side of an image and text pair is often not a description of the image at all, which is a strange foundation for a model whose entire job is to connect the two.

There is a sharper concern underneath. Uncurated web data is a known source of memorization in diffusion models. A model trained on a corpus containing many near-duplicates of the same image is measurably more likely to reproduce it. That is a legal and reputational exposure, not just a quality issue, and it is a direct argument for deduplication as a safety measure rather than an efficiency one.

What MONET actually is

MONET is the dataset built for this, and it has been released openly. The report positions it against a specific gap: at the time, there was no openly available dataset for large-scale text-to-image pre-training that was filtered, deduplicated, and re-captioned by multiple vision language models.

Existing re-captioned datasets were either too small for pre-training, such as ShareGPT4V at 1.2 million images, or relied on a single model to generate captions. That second limitation is the more interesting one. Captioning an entire corpus with one vision language model bakes that model's particular habits into the caption distribution, which shows up later as degraded generation on prompts that fall outside it. Using several models is not about caption quality per image. It is about not inheriting one model's blind spots across a hundred million examples.

The final corpus draws on nine sources, six real and three synthetic:

SourceOriginalFinalCaptions
LAION 2B-en2.1B46.6MAlt-text
COYO747M19.1MAlt-text
Common-Catalog CC-BY14.6M11.2MBLIP-2
Megalith-10M9.6M8.0MNone
Conceptual-12M11.0M6.4MAlt-text
Diffusion-Aesthetic-4K14k12.8kGPT-4o
Z-Image (synthetic)6.2M5.9MSynthetic
FLUX.2-klein-4B (synthetic)3.6M3.5MSynthetic
FLUX.1-schnell (synthetic)4.5M4.4MSynthetic

Two of those rows deserve attention. Megalith-10M arrives with no captions at all, which is fine when you intend to re-caption everything anyway, and it means the images can be selected purely on visual merit. And Diffusion-Aesthetic-4K contributes under thirteen thousand images out of a hundred million, which is a rounding error by count and is included for ultra-high-resolution coverage that nothing else in the pool provides.

The exclusions are also deliberate. DataComp-1B was left out for heavy overlap with LAION and COYO, and the non-English portion of LAION-5B was left out on the grounds that multilingual coverage is more reliably obtained through translation than through noisy alt-text in other languages.

The funnel

The reduction happens in stages, and the first one is blunt. For LAION and COYO, the two largest sources, anything below 512 by 512 pixels is dropped, and anything with an aesthetic score below 5.0 is dropped. Those two filters alone take roughly 2.85 billion images down to about 91 million.

That is worth sitting with. The single largest act of curation in the pipeline is two threshold checks, and they discard around ninety-seven percent of the raw pool before anything sophisticated happens. The reason is visible in the report's breakdown of LAION: roughly three quarters of its images are smaller than 512 pixels. Most of the largest open image dataset is simply too small to train a high-resolution model on.

After merging with the four smaller real sources and removing duplicates within each source by URL and perceptual hash, the pool sits at 121.1 million. From there the remaining stages are multi-classifier safety filtering, two-stage deduplication across sources, domain-based filtering and source governance, multi-model re-captioning, and augmentation with synthetic images from text-to-image models under Apache 2.0 licences. The output is 104.9 million pairs.

Speed as an architectural decision

The other half of the report is about getting the model fast, and the approach is worth separating from the usual claims about efficiency.

Standard diffusion sampling integrates a velocity field over many small steps. Each step is a full forward pass, so the cost of an image is essentially the number of steps multiplied by the size of the model. You can make the model smaller, which costs quality, or take fewer steps, which normally costs coherence.

The distillation approach here targets the second directly by training the model to approximate a flow map, meaning a function that jumps across a large time interval rather than integrating through it. A single evaluation of the full map is one-step generation. The work builds on the Shortcut model framework and on flow map methods, and the report notes it extends an approach Jasper had previously applied to image upscaling.

The framing that matters for anyone evaluating the output: the sub-second figure is not the 4B model running quickly. It is a distilled model that has been taught to take very few steps, which is a different object with a different quality profile. The report is explicit that distillation is a stage applied on top, rather than a property of the base model.

Why publish it

There is a straightforward reading of why a company would document its own training recipe in this much detail, including the ablations that did not work.

Releasing the dataset is the more substantive half. A filtered, deduplicated, multi-VLM re-captioned corpus at this scale is expensive to produce and was not previously available to anyone outside a well-funded lab. Publishing it lowers the barrier for research groups that could not have built it, and it makes results comparable, since two teams can now train on the same data and attribute the difference to their methods rather than to their scraping.

For anyone evaluating image models commercially, the more useful takeaway is narrower. A vendor willing to publish the composition of its training data, its licensing, and its deduplication methodology is a vendor you can actually ask questions about provenance. That is becoming a procurement requirement, and most image models cannot answer it.

Further reading

Run this on your own assets.