Gemma 4 31B
Open-weight 31B dense — thinking, tools, image input, 262K context — self-hosted.
Google (Gemma) — served by Reka$0.09 in · $0.34 out / 1M tokens (≤256K ctx)llmgemmaopen-weightsmultimodal
llmgemmaopen-weightsmultimodal
| Context band | Max tokens | Input / M tok | Output / M tok |
|---|---|---|---|
| ≤256K | 262,144 | $0.09 | $0.34 |
API
Wire it up.
Endpoint
POST https://api.tryinfer.com/v1/chat/completionsrequest
curl https://api.tryinfer.com/v1/chat/completions \
-H "Authorization: Bearer $INFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma4-31b",
"messages": [{ "role": "user", "content": "Explain the difference between sliding-window and global attention in two short paragraphs." }]
}'
OpenAI-compatible — point your SDK at Infer with a bearer token. Get an API key →