Photo-to-Mockup Pipeline

How we turn a product photo into a reusable Kittl mockup, what works today, and the alternatives we tried. State as of 30 September 2026.

Upload a photo of a blank product and get back a mockup in exactly the same format as our catalog mockups. The editor and dashboard render it like any other mockup, with no special code. Apparel comes first, mainly t-shirts worn by a person.

What a Kittl mockup is made of

PieceWhat it does
SceneThe product photo with the print area cut out. The design shows through the hole.
Light layerShadows and folds (multiply, mostly white). Darkens the design where the fabric dips.
Dark layerHighlights (screen, mostly black). Lifts the design where light hits the fabric.
Warp gridA 13×13-point Bézier mesh with curve handles that bends the design to follow the product.
Design positionWhere the design sits and how large it is.

Unlike Photoshop mockups, Kittl has no displacement map. The fine fabric detail comes from the two shading layers. The grid only handles the coarse and medium bend.

The pipeline, step by step

PhotoBlank product, any size
→
1. Cut outFind the garment, cut the print area
→
2. ShadingShadows + highlights from the photo itself
→
3. Warp gridML model predicts how the design bends
→
4. PlacementDesign position + crop
→
Kittl mockup3000×3000, same format as the catalog
StepHowQuality today
1. Cut out SAM (Meta's image segmentation model) finds the garment, and a guided-filter matte cleans the edge. The whole garment is cut out, the same way catalog mockups are built. Good
One known miss: when an arm is pressed against the torso, the shadow gap gets included and the garment colour spills onto the arm.
2. Shading Splits the photo's own brightness into shadows (light layer) and highlights (dark layer). Only a light blur is applied (about 2px at full size), so collar ribbing, care labels and fabric weave stay sharp. Good
Folds survive recolouring, including to dark colours. It assumes a blank garment, so printed or patterned garments aren't handled.
3. Warp grid A model built on Meta's Sapiens-1B predicts the grid points. It was trained on ~825 grids that freelancers made for the catalog. Curve handles are then computed from neighbouring points and clamped so the curves don't overshoot. Note: Sapiens' licence blocks commercial use, see Model options. Decent
On a typical tee it gets close to a freelancer's grid, with median error around 4% of the image. The weak spot: the design sits a bit flat and doesn't bend along medium chest folds or follow a turned torso.
4. Placement Design position, plus a crop window when the photo isn't square (the output space is square). Returned with every result.

Where it runs

How it plugs into the product

Alternatives we tried or considered

ApproachThe ideaWhat happenedStatus
Learned grid from freelancer grids Train a model on the catalog's hand-made grids. Decent, and close to a freelancer on typical tees. It can't beat the freelancers it learned from, and the training data had its curve handles stripped. In use
Geometry from surface normals Estimate the 3D surface pixel by pixel (Marigold, Sapiens normals) and compute the grid from it. No training data needed. It stops over-warping on flat products. But normals only capture local tilt, not body pose, and the folds came out wobbly. Dropped
possible fallback for near-flat products
Train on the rendered result Score the model on how the warped design looks, not on grid coordinates. Better on 4 of 5 image metrics, but it put curve handles on every edge, which made the edges wiggly. The fix didn't pass. Not shipped
Photoshop mockups as answer keys Buy professional Photoshop mockups, read their warp and displacement programmatically (no Photoshop needed), and train our grid to reproduce their final image. On the test template, our reimplementation matches Photoshop's own render almost pixel for pixel. Fitted grids match gentle and moderate warps, but strong warps fail. Raw grid points turned out to be an unstable training target, so the recommended target is now a dense flow field that gets projected onto the grid. Up next
waiting on a go
AI-generated answer keys Have an image-editing model (Nano Banana Pro, GPT Image 2, Qwen-Image-Edit, FLUX.2) put a design on a garment photo, then fit our grid to that result. This would give unlimited variety in poses and products, including real catalog garments. Not started
after the Photoshop track
Displacement map at render time Push pixels Photoshop-style, computed from the photo's brightness. This is how the Custom Mockups extension works. Looks good as a one-off, but the result is a single baked image rather than a reusable mockup, and Kittl's renderer has no displacement support. Ruled out
for the core feature; lives on as the extension
Generate the finished image with AI One prompt produces the whole scene (our AI workflow presets). Results aren't repeatable across a set, the design can drift, and every image costs credits. Good for marketing scenes, not for listing images. Separate surface

Model options

ModelUsed forStatus
Sapiens-1B (Meta, trained on human images)Base of the grid modelIn use Best results so far. Its licence doesn't allow commercial use (CC-BY-NC), so production needs a Meta licence or a swap.
DINOv3 (Meta, released Aug 2025)Replacement for SapiensPlanned Allowed for commercial use. A head-to-head test is planned but hasn't run.
DINOv2-giant, ConvNeXtEarlier bases for the grid modelTried Tested in May. Sapiens came out ahead.
SAM (Meta)Garment masks for the cutoutIn use SAM 3 is the candidate fix for the arm-shadow issue.
MarigoldSurface normalsDropped Used for the geometry approach.
Depth Anything 3DepthParked Only relevant if the geometry approach comes back.

Open before production

  1. Licence: get a Meta licence for Sapiens, or retrain on DINOv3.
  2. Home: move the endpoint and code into Kittl's accounts and GitHub org.
  3. Grid quality: decide the training target for the Photoshop track. We also need more Photoshop mockups with displacement maps from different creators (we have 2 and want 5 or more).
  4. Cold start: keep the endpoint warm, or design the UX so a slow first call is acceptable.
  5. Beyond apparel: flat-lays and non-apparel products need their own grid handling. The shading step already works on them.