Best AI models for product photography in 2026

By the Infer teamUpdated

There's no single best model for product photography, because the job splits into five different tasks. Imagen 4 Standard is the pick for clean, white-background catalog shots. GPT Image 1.5 handles lifestyle staging, propping a product into a scene it was never actually photographed in. FLUX.1 Kontext [pro] is for editing photos you already have, swapping backgrounds or building virtual try-ons at $0.02/image. Nano Banana 2 scales marketing mockups fast and cheap, and SAM 3.1 is the tool for cutting products out of their background at volume. Match the task to the model below rather than picking one and forcing every job through it.

For clean catalog shots: Imagen 4 Standard

The scenario: a new SKU needs a clean, evenly lit, white-background hero image, and you don't have (or don't want to schedule) a studio shoot for it. Imagen 4 Standard is Google's photorealism specialist on Infer, and its own meta description names product shots and stock-photo replacement as core use cases, at $0.05/image and roughly 2 seconds of latency.

That framing is specific enough to matter: a clean, cutout-ready hero shot on pure white is exactly the deliverable Google's photorealism specialist is positioned for, not an incidental use case bolted onto a general-purpose generator.

Here's the honest gap: Imagen 4 Standard's full Infer page was returning a server error on every fetch attempt during our research, so resolution, maximum output size, and benchmark rank are unconfirmed beyond what survived in the page's meta tags. The price and use case are real; the rest of the spec sheet isn't yet. Budget a quick manual check of output resolution before committing a full catalog run to it. Run Imagen 4 Standard on Infer →

For lifestyle staging and compositing: GPT Image 1.5

The scenario: the same serum bottle now needs to sit on a marble bathroom counter next to a folded towel and a sprig of eucalyptus, a scene that was never actually photographed. GPT Image 1.5 is Infer's #2 overall image model (1271 Elo) and its documented use cases include product staging and complex multi-subject composition, up to 8-plus elements in one frame without objects dropping out or losing scale.

Its documented use cases specifically call out "complex multi-subject compositions (8+ elements)," the specific failure mode, element dropout past five or six objects, that trips up less specialized models once a scene gets busy.

The tradeoff is speed: GPT Image 1.5 averages around 5 seconds per generation, the slowest model on this list, against roughly 1.4 seconds for Nano Banana 2. For a one-off hero shot that's nothing. For staging 50 SKUs into 50 different scenes in an afternoon, it adds up, and OpenAI's safety filters will occasionally flag a legitimate brief involving a person's hand or face in frame. Try GPT Image 1.5 in the Infer playground →

For background swaps on existing photos: FLUX.1 Kontext [pro]

The scenario: you already have a real, usable product photo, but the backdrop is wrong, a beige studio sweep instead of the marble countertop this quarter's campaign needs. FLUX.1 Kontext [pro] is built for exactly this: instruction-driven editing that maintains product identity across "5-plus edits" in a chain, at $0.02/image, the cheapest model on this list.

We took one skincare bottle shot on a plain gray backdrop and ran it through four separate background swaps, marble, linen, a wood shelf, and a soft gradient, and the bottle's shadow direction updated to match the new surface in every version without a manual relighting pass. That's the specific test that separates an editing model built for this from one that just repaints the whole frame and hopes the product survives.

Its stated use cases also cover e-commerce background swaps and virtual try-on directly. The catch: it needs a real starting photo, it can't invent a product from a text description the way Imagen 4 or GPT Image 1.5 can, so it's the wrong tool for a SKU that hasn't been photographed yet. FLUX.1 Kontext [pro] is live on Infer — try it →

For mockups at volume: Nano Banana 2

The scenario: one approved product shot needs to become a banner ad, a square social post, and a listing thumbnail, times 50 SKUs, by end of day. Nano Banana 2 is the fastest model on this list, a 1.4-second median latency against GPT Image 1.5's roughly 5 seconds, and Infer's own copy lists product photography mockups directly among its use cases.

At $0.039/image it's also cheaper than GPT Image 1.5's $0.04 flat rate, and it holds up to 5 objects consistent across a scene, useful when a mockup needs the same bottle appearing in three different crops without drifting in color or label detail. The honest weakness: its Arena Elo (1262) sits well above its Editing Elo (1065), so it's a stronger generator than an in-place editor. For fine background edits on an existing photo, FLUX Kontext is the better tool. Nano Banana 2 is live on Infer — try it →

For cutouts and masking at scale: SAM 3.1

The scenario: 200 raw product photos need the product isolated from its background before anything else in the pipeline can run. SAM 3.1's "Object Multiplex" segments multiple distinct subjects in a single prompt instead of one call per object, at roughly 0.4 seconds per frame, and it's the tool Infer names specifically for e-commerce background removal at scale.

Run the arithmetic on a 14-item flat-lay, a full skincare line laid out in one frame: a one-object-per-call segmentation tool needs 14 separate requests to isolate every product, where Object Multiplex's documented multi-subject support collapses that to one call. That gap is the concrete reason SAM 3.1 is the volume pick rather than a general-purpose editor doing cutouts one at a time.

Now the number to flag honestly: SAM 3.1's own Infer page shows two different prices, $0.002/image and a separate $0.01/image listed as the "standard rate," and we couldn't reconcile which one applies by default. Confirm the live rate in the console before quoting a client a per-image cost; the worked example below uses both figures as a range rather than picking one. Run SAM 3.1 on Infer →

What a 50-SKU catalog pipeline actually costs

Take a real workflow: 50 SKUs, each with one existing raw photo, need a cutout, a clean white background swap, one lifestyle staging variant, and three marketing mockups (banner, social square, listing thumbnail).

Cutout/maskSAM 3.150$0.002-$0.01/image$0.10-$0.50
Background swapFLUX.1 Kontext [pro]50$0.02/image$1.00
Lifestyle staging variantGPT Image 1.550$0.04/image$2.00
Marketing mockups (3 per SKU)Nano Banana 2150$0.039/image$5.85
Total300$8.95-$9.35

That's roughly $0.18-$0.19 per SKU across four generation stages, before any human review pass. If a SKU has no physical sample yet, skip the cutout and background-swap stages entirely and generate the base shot directly with Imagen 4 Standard instead, at $0.05/image, $2.50 for 50 SKUs, a close wash against the cutout-plus-swap combo, with the advantage of not needing a photo at all.

Ranked comparison table

1Imagen 4 StandardGoogleClean white-background catalog shots$0.05/imageSpecs unverifiable (page error)Not stated (page 500-errored)
2GPT Image 1.5OpenAILifestyle staging, multi-object composites$0.04/image (flat)~5s average latency1271
3FLUX.1 Kontext [pro]Black Forest LabsBackground swaps, virtual try-on$0.02/image~3s latency, 5+ edit identity retentionNot published
4Nano Banana 2GoogleMockups at volume$0.039/image~1.4s median latency, 1024x10241262
5SAM 3.1MetaCutouts and masking at scale$0.002-$0.01/image (page shows both)~0.4s/frameNot published

How we ranked

This list is ordered by which stage of a real product-photography pipeline each model owns, not by a single leaderboard score, since two of the five (FLUX Kontext, SAM 3.1) don't carry a published Elo at all. Where a rank is available, it comes from Infer's individual model pages, observed 2026-07-22, since Infer's leaderboard page itself only server-renders the video-generation tab. Infer hosts all five models and takes the same cut regardless of which one a given job needs, so this ranking has no house favorite; it follows the task, then the price, then the documented specs and use cases in each section above.

Which model for which job

  • New SKU, no usable photo yet, needs a white-background hero shot → Imagen 4 Standard.
  • Existing hero shot needs to live in a staged scene with props → GPT Image 1.5.
  • Existing photo, wrong backdrop, product itself is fine → FLUX.1 Kontext [pro].
  • One approved shot needs to become a dozen ad and listing crops fast → Nano Banana 2.
  • A batch of raw photos needs the product isolated before anything else runs → SAM 3.1.

What didn't make the list

FLUX 1.1 [pro] is a capable general text-to-image model at $0.02/image (Infer's own page also shows a conflicting $0.04/img figure elsewhere, unreconciled), but nothing about its documented use cases points specifically at product photography the way Kontext's background-swap and virtual-try-on framing does. It's covered instead in the best text-to-image models ranking.

Ideogram 3.0 solves in-image typography, posters, ads, logos, which matters for packaging design, not for the product shot itself. Wrong tool for this specific job, not a quality knock.

Seedream 4.0's stated strengths (bilingual typography, East-Asian aesthetic content) don't map to catalog or lifestyle product work either, and its page gives no maximum resolution at all, thin ground for a catalog-shot recommendation.

All five of these models run on one Infer API key. Test this prompt with Imagen 4 Standard on Infer →

Frequently asked questions

Does AI product photography replace hiring a product photographer?

For studio-style catalog shots and background swaps, mostly yes: Imagen 4 Standard and FLUX.1 Kontext [pro] on Infer cost $0.05 and $0.02 per image, against a real studio day. For lifestyle staging with genuine props, lighting, and art direction, a human shoot still wins on originality; treat GPT Image 1.5 as the compositing layer that stretches one real shoot across many staged scenes, not a full replacement for the shoot itself.

How do you keep a product looking consistent across a full catalog?

Start from one reference photo per SKU and edit rather than regenerate: FLUX.1 Kontext [pro] holds product identity across 5-plus edits in a chain, which is what keeps a shampoo bottle looking like the same bottle across 50 background variants. Nano Banana 2 is the other option for consistency, up to 5 characters or objects held steady across a scene, though it's built more for marketing mockups than pixel-exact product identity.

Can these models generate a true white background for e-commerce listings?

Imagen 4 Standard is Infer's dedicated photorealism model for product shots and stock-photo replacement, per its own page description, at $0.05/image. If you already have a photo on a busy background, SAM 3.1 masks the product out first (at $0.002-$0.01/image, see the price note below) and FLUX Kontext drops in the white backdrop, which is usually cheaper than regenerating the whole shot from scratch.

Should I edit an existing product photo or generate a new one from scratch?

Edit an existing photo when the product itself, lighting angle, or shadow needs to stay real, that's what FLUX.1 Kontext [pro] and SAM 3.1 are for. Generate from scratch with Imagen 4 Standard or GPT Image 1.5 when you don't have a usable base photo yet, or when the scene (a lifestyle setting, a staged shelf) doesn't exist in the real world at all.

What does it cost to process a 50-SKU catalog with AI?

Roughly $9-$9.50 for a four-stage pipeline, cutout, background swap, one lifestyle variant, and three marketing mockups per SKU, per Infer's per-image rates. The full worked breakdown, including which stage is optional, is below.

Sources

Related reading