SAM 3
Open-vocabulary detection. Returns bounding boxes for any prompt.
Meta$0.01 / imagedetectionsam3

Est. cost$0.01For $1 you can run this model 100 times.
Example output

$0.01/ image
$0.01 per image. For $1 you can run this model 100 times.
Pay only for successful generations. No idle, no minimums, no per-seat.
API
Wire it up.
Endpoint
POST https://api.tryinfer.com/v1/inference/sam3/image-to-bboxrequest
curl https://api.tryinfer.com/v1/inference/sam3/image-to-bbox \
-H "Authorization: Bearer $INFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"prompt": "person",
"image_url": "https://example.com/input.jpg"
}
}'
Parameters
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
| prompt | string | Yes | — | — |
| image_url | string (URL) | Yes | — | — |
| threshold | number | — | — | — |
Authenticate with a bearer token. Get an API key →
Further reading
From the Infer content directory.
Best ofBest AI image-editing models in 2026FLUX.1 Kontext [pro] leads AI image editing in 2026 at $0.02/image with 5+ edit identity retention, ahead of GPT Image 1.5, Nano Banana 2, and SAM 3.1 for masking.Best ofBest open-source AI video and image models in 2026Wan 2.2 T2V-A14B leads open-source video on clean Apache 2.0 terms; SAM 3.1 leads segmentation; LTX-2 scores higher but its license is revenue-gated.Best ofBest AI models for product photography in 2026Imagen 4 wins clean catalog shots, GPT Image 1.5 handles lifestyle staging, FLUX Kontext swaps backgrounds, and SAM 3.1 cuts out products at scale in 2026.