All research

Shadows are the hard part of compositing

You cannot hand-annotate a training set for shadows, so Jasper Research rendered one, and built a single-step model that takes light direction, softness, and intensity as explicit inputs.

A ceramic vase with dried branches on a neutral backdrop, with a soft cast shadow grounding it on the surface.

Cut a product out of one photograph and drop it into another and the result almost always looks wrong, even when the cutout is perfect. The edges can be flawless and the color can match, and the object will still appear to hover slightly above the scene rather than sit in it.

The missing piece is usually the shadow. It is the cue that tells a viewer where the light is, how far the object is from the surface, and that the object and the surface occupy the same physical space. Without it, a composite reads as a collage.

This matters commercially in a specific way. In e-commerce, shoppers use photographs to judge whether a product is real and whether it will look right in their own space. A product that floats is a product that looks like a placeholder.

Why you cannot just annotate a shadow dataset

The obvious way to train a shadow model is the way most vision problems get solved: collect images, have people label them, train on the labels. Shadows defeat this almost immediately.

A shadow is not a shape you can trace. It is the joint consequence of the light source position, the light source size, the geometry of the object, and the geometry of the surface receiving it. Asking an annotator to draw a shadow means asking them to solve a physics problem by hand, repeatedly, and to be geometrically correct rather than merely plausible. It is slow, expensive, and the ground truth is only as good as the annotator's intuition about optics.

Worse, the thing you actually want to control is not in the label at all. If you want a model that can make a shadow softer or move the light to the left, you need training data where softness and light position are known quantities. A hand-drawn shadow has no recorded light angle. It is an output with no inputs attached.

Render the data instead

Jasper Research took the other route, described by research scientist Onur Tasar: build the dataset with a rendering engine, where the light parameters are known because you chose them.

The pipeline starts with a curated collection of 3D meshes, made by professional artists under free-use licenses, spanning a deliberate range of shapes, sizes, and materials. Each mesh is placed in a rendering engine with a camera and a light source, and the light source parameters are randomized across:

  • Polar and azimuthal angles, which set the direction the light comes from.
  • Light size, which controls how soft or sharp the shadow edge is.
  • Shadow intensity, which controls how dark the shadow falls.

Iterating over those parameters produces thousands of renders covering a wide span of shadow direction, sharpness, and depth. Each render is saved as a triple: the object, its mask, and the resulting shadow map.

That triple is the whole trick. Because the scene was constructed rather than observed, every training example arrives with its light parameters attached. The model is not being asked to infer what light produced a shadow. It is being taught the forward mapping from light parameters to shadow, which is exactly the mapping a user needs when they want to change one.

Controls that correspond to something

The result is a model that takes an object image plus light parameters and returns a shadow. The parameters are the same ones that were randomized during training, expressed in a spherical coordinate system, so the interface a user touches maps directly onto what the model learned.

In practice this means direction, softness, and intensity behave like three independent dials rather than one prompt that has to be re-rolled until it looks right. Want a soft shadow that blends into the background? Increase the light size. Need a harder edge to emphasize a product silhouette? Decrease it. Need the light to come from the other side to match a campaign's art direction? Change the azimuthal angle.

This is a meaningful difference from a text-prompted approach. A prompt describes an intent and hopes the model agrees. A parameter describes a quantity and the model is obliged. For catalog work, where the requirement is often that four hundred products share one lighting setup, the second is the one that survives contact with production.

Single step, by construction

Speed was a design target rather than an optimization applied afterwards. The model uses rectified flow, a technique that lets it predict the shadow in a single step instead of iterating through the many denoising stages a conventional diffusion process requires.

The reason to care is not the benchmark number, it is what becomes possible at that latency. A model that takes several seconds per image is a batch process, and batch processes get used once and reviewed later. A model that responds immediately can sit behind a slider, which means someone can adjust the light angle and watch the shadow follow. That turns shadow placement from a rendering job into a direct manipulation, and it is the difference between a feature people configure and a feature people actually use.

The benchmark did not exist, so they built one

There is a detail in this work that is easy to skip past and worth underlining. When the team went to evaluate the model, they found there was no existing dataset for the task. Shadow generation with explicit control had no standard way to measure whether a model was any good.

So they built a benchmark and released it. It has three tracks, each testing a different axis of control: shadow softness, horizontal direction, and vertical direction. Rather than asking whether a generated shadow looks plausible, the benchmark asks whether the model actually did what it was told, separately for each control.

That framing is the useful contribution. Plausibility is a low bar and a generative model will usually clear it. Controllability is the property that determines whether the thing can be used in a pipeline, and it is the property nobody had been measuring.

Releasing it as open source is a reasonable trade. A shared benchmark makes competing claims comparable, including claims against Jasper's own model.

Where it goes next

Shadow generation is one component of a larger compositing problem. Placing a subject into a new scene convincingly requires the cutout, the lighting on the subject, and the shadow it casts to agree with each other. Get one wrong and the other two do not save it.

The same reasoning extends past marketing imagery. Anywhere a virtual object has to be placed in a real scene, from augmented reality to film and games, the shadow is what sells the placement. The interesting part of this work is not that a model can draw a shadow. It is that the shadow became a set of parameters you can dial rather than an outcome you have to accept.

Models in this article

Further reading

Run this on your own assets.