Google: Gemini 3.6 Flash

google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hed...

visiontools
MODALITIES
INPUT PRICE
$0.75per 1M
OUTPUT PRICE
$3.75per 1M
CONTEXT
1.05M
RELEASED
Jul 21, 2026
ProviderCacheUptimeChat
$0.75$3.75
Cache read$0.075

Capabilities

Input modalities
audiofileimagetextvideo
Output modalities
text
Features
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
1

Get API Key

Create an API key from the Keys page, then set it as an environment variable:

export ONLIST_API_KEY=sk-...
2

Make your first request

Endpoints

POSThttps://onlist.io/v1/chat/completions

OpenAI Chat Completions format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:google/gemini-3.6-flash
POSThttps://onlist.io/v1/responses

OpenAI Responses format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:google/gemini-3.6-flash
POSThttps://onlist.io/v1/messages

Anthropic Messages format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:google/gemini-3.6-flash

Code samples

curl https://onlist.io/v1/chat/completions \
  -H "Authorization: Bearer $ONLIST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
       "model": "google/gemini-3.6-flash",
       "messages": [
         {
           "role": "user",
           "content": "Explain quantum entanglement in one paragraph."
         }
       ]
     }'

Replace $ONLIST_API_KEY with the API key from your Keys page.

Authentication

All requests must include an Authorization: Bearer <TOKEN> header. Generate keys from the Keys page; keys can be scoped to specific models, groups, IP ranges, and rate limits.

3

Enable streaming

Add "stream": true to receive partial responses as server-sent events in real time.

Streaming example

curl https://onlist.io/v1/chat/completions \
  -H "Authorization: Bearer $ONLIST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
       "model": "google/gemini-3.6-flash",
       "messages": [
         {
           "role": "user",
           "content": "Write a haiku about recursion."
         }
       ],
       "stream": true
     }'

Supported parameters

NameTypeDescription
include_reasoning
max_tokensintegerMaximum number of tokens to generate in the completion.
reasoning
reasoning_effortstringConstrains effort on reasoning. Supported values: "low", "medium", "high".
response_formatobjectSpecifies the output format. Use {"type": "json_object"} for JSON mode.
seedintegerIf specified, the system will attempt deterministic sampling for reproducible results.
stopstring | arrayUp to 4 sequences where the API will stop generating tokens.
structured_outputs
temperaturenumberSampling temperature between 0 and 2. Higher values make output more random.
tool_choicestring | objectControls which tool is called. "auto", "none", "required", or a specific function.
toolsarrayA list of tools the model may call. Currently supports functions.
top_pnumberNucleus sampling. The model considers tokens with top_p probability mass.

These are the request parameters this model accepts. Parameter semantics follow the OpenAI Chat Completions specification.