Photo-to-Mockup Pipeline
How we turn a product photo into a reusable Kittl mockup, what works today, and the alternatives we tried. State as of 30 September 2026.
Upload a photo of a blank product and get back a mockup in exactly the same format as our catalog mockups. The editor and dashboard render it like any other mockup, with no special code. Apparel comes first, mainly t-shirts worn by a person.
What a Kittl mockup is made of
| Piece | What it does |
|---|---|
| Scene | The product photo with the print area cut out. The design shows through the hole. |
| Light layer | Shadows and folds (multiply, mostly white). Darkens the design where the fabric dips. |
| Dark layer | Highlights (screen, mostly black). Lifts the design where light hits the fabric. |
| Warp grid | A 13×13-point Bézier mesh with curve handles that bends the design to follow the product. |
| Design position | Where the design sits and how large it is. |
Unlike Photoshop mockups, Kittl has no displacement map. The fine fabric detail comes from the two shading layers. The grid only handles the coarse and medium bend.
The pipeline, step by step
PhotoBlank product, any size
→
1. Cut outFind the garment, cut the print area
→
2. ShadingShadows + highlights from the photo itself
→
3. Warp gridML model predicts how the design bends
→
4. PlacementDesign position + crop
→
Kittl mockup3000×3000, same format as the catalog
| Step | How | Quality today |
|---|---|---|
| 1. Cut out | SAM (Meta's image segmentation model) finds the garment, and a guided-filter matte cleans the edge. The whole garment is cut out, the same way catalog mockups are built. | Good One known miss: when an arm is pressed against the torso, the shadow gap gets included and the garment colour spills onto the arm. |
| 2. Shading | Splits the photo's own brightness into shadows (light layer) and highlights (dark layer). Only a light blur is applied (about 2px at full size), so collar ribbing, care labels and fabric weave stay sharp. | Good Folds survive recolouring, including to dark colours. It assumes a blank garment, so printed or patterned garments aren't handled. |
| 3. Warp grid | A model built on Meta's Sapiens-1B predicts the grid points. It was trained on ~825 grids that freelancers made for the catalog. Curve handles are then computed from neighbouring points and clamped so the curves don't overshoot. Note: Sapiens' licence blocks commercial use, see Model options. | Decent On a typical tee it gets close to a freelancer's grid, with median error around 4% of the image. The weak spot: the design sits a bit flat and doesn't bend along medium chest folds or follow a turned torso. |
| 4. Placement | Design position, plus a crop window when the photo isn't square (the output space is square). | Returned with every result. |
Where it runs
- It's a GPU endpoint on Modal (a serverless GPU cloud). One call returns all five pieces.
- About 11 to 15 seconds per photo when warm. A cold start takes 90 to 130 seconds, because the endpoint scales to zero after about 10 minutes idle.
- Gap: it runs under a personal Modal account from a private prototype repo, not the Kittl GitHub org. It needs a proper home before production. Cost per mockup hasn't been measured yet.
How it plugs into the product
- Editor, "Convert to Mockup": select an image on the canvas, convert it, and a mockup board appears next to it. This is a prototype on the current renderer: mono#14323 (draft).
- Storage: mono#15192 saves the result as a workspace-owned custom mockup, so it can be reused and positioned like a catalog mockup. The pipeline itself is outside that PR's scope.
- Dashboard: a My Mockups prototype with own uploads and sets, mono#14525 (draft).
Alternatives we tried or considered
| Approach | The idea | What happened | Status |
|---|---|---|---|
| Learned grid from freelancer grids | Train a model on the catalog's hand-made grids. | Decent, and close to a freelancer on typical tees. It can't beat the freelancers it learned from, and the training data had its curve handles stripped. | In use |
| Geometry from surface normals | Estimate the 3D surface pixel by pixel (Marigold, Sapiens normals) and compute the grid from it. No training data needed. | It stops over-warping on flat products. But normals only capture local tilt, not body pose, and the folds came out wobbly. | Dropped possible fallback for near-flat products |
| Train on the rendered result | Score the model on how the warped design looks, not on grid coordinates. | Better on 4 of 5 image metrics, but it put curve handles on every edge, which made the edges wiggly. The fix didn't pass. | Not shipped |
| Photoshop mockups as answer keys | Buy professional Photoshop mockups, read their warp and displacement programmatically (no Photoshop needed), and train our grid to reproduce their final image. | On the test template, our reimplementation matches Photoshop's own render almost pixel for pixel. Fitted grids match gentle and moderate warps, but strong warps fail. Raw grid points turned out to be an unstable training target, so the recommended target is now a dense flow field that gets projected onto the grid. | Up next waiting on a go |
| AI-generated answer keys | Have an image-editing model (Nano Banana Pro, GPT Image 2, Qwen-Image-Edit, FLUX.2) put a design on a garment photo, then fit our grid to that result. | This would give unlimited variety in poses and products, including real catalog garments. | Not started after the Photoshop track |
| Displacement map at render time | Push pixels Photoshop-style, computed from the photo's brightness. This is how the Custom Mockups extension works. | Looks good as a one-off, but the result is a single baked image rather than a reusable mockup, and Kittl's renderer has no displacement support. | Ruled out for the core feature; lives on as the extension |
| Generate the finished image with AI | One prompt produces the whole scene (our AI workflow presets). | Results aren't repeatable across a set, the design can drift, and every image costs credits. Good for marketing scenes, not for listing images. | Separate surface |
Model options
| Model | Used for | Status |
|---|---|---|
| Sapiens-1B (Meta, trained on human images) | Base of the grid model | In use Best results so far. Its licence doesn't allow commercial use (CC-BY-NC), so production needs a Meta licence or a swap. |
| DINOv3 (Meta, released Aug 2025) | Replacement for Sapiens | Planned Allowed for commercial use. A head-to-head test is planned but hasn't run. |
| DINOv2-giant, ConvNeXt | Earlier bases for the grid model | Tried Tested in May. Sapiens came out ahead. |
| SAM (Meta) | Garment masks for the cutout | In use SAM 3 is the candidate fix for the arm-shadow issue. |
| Marigold | Surface normals | Dropped Used for the geometry approach. |
| Depth Anything 3 | Depth | Parked Only relevant if the geometry approach comes back. |
Open before production
- Licence: get a Meta licence for Sapiens, or retrain on DINOv3.
- Home: move the endpoint and code into Kittl's accounts and GitHub org.
- Grid quality: decide the training target for the Photoshop track. We also need more Photoshop mockups with displacement maps from different creators (we have 2 and want 5 or more).
- Cold start: keep the endpoint warm, or design the UX so a slow first call is acceptable.
- Beyond apparel: flat-lays and non-apparel products need their own grid handling. The shading step already works on them.