I

Inception: Mercury 2.5 Preview

inception/mercury-2.5-preview

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, a...

tools
MODALITIES
INPUT PRICE
$0.04per 1M
OUTPUT PRICE
$0.15per 1M
CONTEXT
260K
RELEASED
Sep 1, 2026
ProviderCacheUptimeChat
$0.04$0.15
Cache read$0.004

Capabilities

Input modalities
text
Output modalities
text
Features
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetools
1

Get API Key

Create an API key from the Keys page, then set it as an environment variable:

export ONLIST_API_KEY=sk-...
2

Make your first request

Endpoints

POSThttps://onlist.io/v1/chat/completions

OpenAI Chat Completions format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:inception/mercury-2.5-preview
POSThttps://onlist.io/v1/responses

OpenAI Responses format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:inception/mercury-2.5-preview
POSThttps://onlist.io/v1/messages

Anthropic Messages format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:inception/mercury-2.5-preview

Code samples

curl https://onlist.io/v1/chat/completions \
  -H "Authorization: Bearer $ONLIST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
       "model": "inception/mercury-2.5-preview",
       "messages": [
         {
           "role": "user",
           "content": "Explain quantum entanglement in one paragraph."
         }
       ]
     }'

Replace $ONLIST_API_KEY with the API key from your Keys page.

Authentication

All requests must include an Authorization: Bearer <TOKEN> header. Generate keys from the Keys page; keys can be scoped to specific models, groups, IP ranges, and rate limits.

3

Enable streaming

Add "stream": true to receive partial responses as server-sent events in real time.

Streaming example

curl https://onlist.io/v1/chat/completions \
  -H "Authorization: Bearer $ONLIST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
       "model": "inception/mercury-2.5-preview",
       "messages": [
         {
           "role": "user",
           "content": "Write a haiku about recursion."
         }
       ],
       "stream": true
     }'

Supported parameters

NameTypeDescription
include_reasoning
max_tokensintegerMaximum number of tokens to generate in the completion.
reasoning
reasoning_effortstringConstrains effort on reasoning. Supported values: "low", "medium", "high".
response_formatobjectSpecifies the output format. Use {"type": "json_object"} for JSON mode.
stopstring | arrayUp to 4 sequences where the API will stop generating tokens.
structured_outputs
temperaturenumberSampling temperature between 0 and 2. Higher values make output more random.
tool_choicestring | objectControls which tool is called. "auto", "none", "required", or a specific function.
toolsarrayA list of tools the model may call. Currently supports functions.

These are the request parameters this model accepts. Parameter semantics follow the OpenAI Chat Completions specification.