What a photo-to-3D pipeline actually does
Generating a 3D model from one product photo is not photogrammetry. Here is what the current generation of models does instead, and where it still falls down.
“Turn a photo into a 3D model” sounds like measurement. It is not. Understanding the difference tells you exactly which products this technology works for and which ones it quietly ruins.
It is not photogrammetry
Photogrammetry reconstructs geometry from many photographs by triangulating features seen from multiple angles. It is measurement: given enough overlapping views, the shape is determined by the data. Its failure modes are honest — too few views, featureless surfaces, bad lighting — and it fails visibly, producing holes and noise.
Single-image 3D generation has one view. There is no triangulation available, because the back of the object was never photographed. What the model produces for the unseen side is not measured. It is inferred from everything the model learned about how objects of that kind tend to be shaped.
This is the whole story in one sentence: the front is reconstruction, the back is a plausible guess.
That is not a criticism. For a great many products the guess is excellent, because chairs, bottles, boxes and lamps are strongly constrained by function and convention. It is a warning about which products to trust it with.
What the pipeline is made of
Current open models converge on a similar decomposition. Two we work with:
Hunyuan3D 2.1 splits the job in two. A 3.3B-parameter shape diffusion transformer produces the geometry, and a separate 2B-parameter model paints PBR materials onto it. Image understanding comes from DINOv2, a self-supervised vision backbone. Splitting shape from appearance matters: the shape model can be trained on untextured geometry, which is far more plentiful than textured geometry.
TRELLIS.2 is a 4B-parameter model that produces PBR natively and supports relighting, under an MIT licence — which is a practical consideration rather than a technical one, but a decisive one if the output has to ship in a commercial product.
The parameter counts tell you the real constraint. These are not models that run on the machine serving your website. They need a datacentre GPU, and generation takes on the order of a minute rather than a page load. Any product that offers photo-to-3D is running an asynchronous job queue behind it, whatever the interface implies.
PBR is where it gets subtle
Producing geometry is the part people expect to be hard. Producing materials is the part that determines whether the result looks like a product or like a video game asset from 2009.
Physically based rendering describes a surface by how it responds to light — base colour, metalness, roughness, normals — rather than by baking a particular lighting situation into the texture. Get it right and the model looks correct under the lighting of whatever room the customer places it in. Get it wrong and it looks like a photograph of a product pasted onto a shape, because that is literally what it is.
The characteristic failure is baked-in lighting. The source photo has highlights and shadows. A naive pipeline treats those as surface colour, so the model carries a permanent highlight on its shoulder that does not move when the light does. In AR, where the model sits in a real room with real lighting, this reads as wrong immediately even to someone who could not name the problem.
The mitigation is a post-processing pass over the material factors before the model ships — correcting metalness and roughness values that the generator was overconfident about. It is unglamorous, it happens after the impressive part, and it does more for perceived quality than another billion parameters would.
Where it works and where it does not
Reliable:
- Convention-constrained shapes. Furniture, containers, appliances, footwear. The model has seen thousands and the guess about the back is well founded.
- Opaque, diffuse surfaces. Fabric, wood, matte plastic, ceramic.
- Products photographed against a clean background, lit evenly, shot from a three-quarter angle that shows two faces at once.
Unreliable:
- Transparency and refraction. Glass and clear plastic. The physics that makes them look right is not recoverable from a single image.
- High-frequency detail that carries meaning. Engraved text, fine jewellery settings, small logos. The model will produce texture that reads as detail without being the actual detail — which is worse than omitting it, because the customer is buying the specifics.
- Genuinely novel shapes. If the object does not resemble a category the model knows, the unseen side is a guess with nothing behind it.
- Anything where the back matters and is unusual. A cable layout, a control panel, a fastening.
How to decide whether to use it
The honest test is not “does the model look good?” It is: would a customer who bought based on this model feel misled when the box arrives?
For a matte ceramic planter, a generated model is a genuinely better product page than a photo, because AR placement answers the question the photo cannot — does this fit on my shelf. For a watch, where the buyer is scrutinising the crown and the bracelet links, a generated model is an invitation to a return.
Photo-to-3D is best understood as a way to reach the long tail. Commission real models for the products that carry the revenue and where detail is the purchase decision. Generate the hundreds of catalogue items that would otherwise never get a 3D model at all, because the alternative to an imperfect model is not a perfect one — it is nothing.
Treat it as experimental, check every output before it goes live, and keep the ability to replace a generated model with a commissioned one for any product that starts selling.