27,000 Visuals In: What AI Photography Still Gets Wrong

In short
- Hands remain the most common failure mode and require frame-by-frame checking.
- Fine repeating detail — embroidery, zari, engraving — degrades into approximation without a very sharp reference.
- Text and logos on products are unreliable and are usually composited from the real photograph.
- Physical plausibility of reflections and shadows is the subtlest failure and the one most viewers feel without identifying.
We have delivered over 27,000 visuals. That volume teaches you where the technology is genuinely excellent and where it quietly fails, and it would be dishonest to write only about the first half.
Hands
Still the most frequent failure, and the most recognisable to viewers. Finger counts drift, joints bend implausibly, a hand holding a product loses its relationship to the object's weight. This matters for us specifically because so much jewellery and accessory imagery involves a hand.
What we do: every frame with a hand is checked individually before it leaves. If it does not survive that check, it is regenerated, not retouched into acceptability.
Fine repeating detail
Intricate embroidery, zari work, filigree, engraved patterns. At a glance they look right; at full resolution the pattern is approximate rather than the actual pattern on your product. For a brand whose value proposition is that craft, approximate is not acceptable.
What we do: insist on very high-resolution references, and for the most detail-critical pieces, composite the real photographed detail into the generated scene rather than generating it.
Text and logos
Any text on a product — a brand mark, an engraved name, a printed slogan — is unreliable. It will look like text and not be text.
What we do: composite from the reference photograph, always. We do not generate text on products.
Closures and hardware
Clasps, zips, buckles, hinges, chain links. These have specific mechanical geometry that generation approximates. A lobster clasp that could not actually open is the kind of detail a jewellery buyer notices immediately, even if they could not articulate why the image feels wrong.
Physical plausibility
The subtlest one. Shadows falling in slightly inconsistent directions. A reflection showing something that could not be where the scene implies. Light that has no plausible source. Most viewers cannot name what is wrong — they simply feel the image is off, and trust drops without them knowing why.
What we do: check shadow direction consistency and reflection logic as an explicit review step, in the same way a retoucher checks a composite.
What this means for choosing a studio
The tools are broadly available. What separates output is whether someone is applying art direction before generation and disciplined review afterwards — the same two things that separate a good photographer from someone who owns a camera.
When you evaluate any AI photography supplier, including us, ask to see work at full resolution rather than as thumbnails, and look at the hands.
Failures that are getting better, and ones that are not
Worth separating, because it affects how you plan.
Improving quickly: overall photorealism, lighting plausibility, fabric behaviour, skin rendering, and scene composition. Work that was obviously synthetic two years ago is now routinely indistinguishable.
Improving slowly: hands and complex articulation, mechanical hardware, fine repeating pattern fidelity.
Not really improving: generating accurate text, and faithfully reproducing a specific real object's fine detail without compositing. These are arguably the wrong tool for the job rather than problems awaiting a fix, which is why our process routes around them rather than waiting.
Why volume produces better output
A studio delivering thousands of images develops something a casual user cannot: a catalogue of known failure modes and the specific directions that avoid them. We know which framings reliably produce bad hands, which staging choices make clasps unreliable, and which surfaces cause reflection logic to break down.
That accumulated knowledge is most of the value. The tools are commodity; knowing where they break is not.
What we will not do
Worth stating, because it is the clearest signal of where the limits actually are:
- We do not generate text or logos on products. Composited from the reference, always.
- We do not generate certification marks, hallmarks or any indication of authentication.
- We do not produce images of identifiable real people without consent.
- We do not take work where the entire value proposition is a specific hand-craft detail that must be documented exactly — we tell the client to photograph it.
- We do not deliver frames that fail the hand check, however tight the deadline.
How to audit any AI imagery, including ours
- View at 100%, never at thumbnail size.
- Look at hands first, then any closure or hinge.
- Trace every shadow back to a single implied light source.
- Check reflections for content that could not be in the scene.
- Compare fine detail against the reference photograph of the real product.
- Look at the set as a whole for colour consistency.
Six checks, a few minutes per set. They catch nearly everything that matters.
Frequently asked
What are the biggest limitations of AI product photography?
Hands, complex closures, fine repeating detail such as embroidery, text and logos on products, and physically consistent reflections. All are checkable, which is why review discipline matters more than model choice.
How do you make sure AI images look real?
Two things: art direction before generation, and hard culling after. Most of what is generated is discarded. Every delivered frame is checked against a specific list — hand anatomy, closure geometry, shadow direction consistency, and faithfulness to the reference product.
Should I be worried about images looking obviously AI-generated?
Viewers do not recognise good AI imagery as AI. They recognise bad AI instantly. The difference is entirely in direction and review discipline, not in the underlying technology, which is why studios producing at volume get better results than tools used casually.
Want Visuals Like This for Your Brand?
Send us 1 photo of your product on WhatsApp and see a campaign-grade editorial sample transformed for free within 24 hours.
Book a Free Sample