Xiaomi: MiMo-V2.6-Pro-UltraSpeed

xiaomi/mimo-v2.6-pro-ultraspeed

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while...

visiontools
MODALITIES
INPUT PRICE
$4.35per 1M
OUTPUT PRICE
$8.70per 1M
CONTEXT
1.05M
RELEASED
Sep 22, 2026
ProviderChat
$4.35$8.7
Hit rate
Cache read$0.036
99.0%

Capabilities

Input modalities
audioimagetextvideo
Output modalities
text
Features
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_p
1

Get API Key

Create an API key from the Keys page, then set it as an environment variable:

export ONLIST_API_KEY=sk-...
2

Make your first request

Endpoints

POSThttps://onlist.io/v1/chat/completions

OpenAI Chat Completions format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:xiaomi/mimo-v2.6-pro-ultraspeed
POSThttps://onlist.io/v1/responses

OpenAI Responses format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:xiaomi/mimo-v2.6-pro-ultraspeed
POSThttps://onlist.io/v1/messages

Anthropic Messages format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:xiaomi/mimo-v2.6-pro-ultraspeed

Code samples

curl https://onlist.io/v1/chat/completions \
  -H "Authorization: Bearer $ONLIST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
       "model": "xiaomi/mimo-v2.6-pro-ultraspeed",
       "messages": [
         {
           "role": "user",
           "content": "Explain quantum entanglement in one paragraph."
         }
       ]
     }'

Replace $ONLIST_API_KEY with the API key from your Keys page.

Authentication

All requests must include an Authorization: Bearer <TOKEN> header. Generate keys from the Keys page; keys can be scoped to specific models, groups, IP ranges, and rate limits.

3

Enable streaming

Add "stream": true to receive partial responses as server-sent events in real time.

Streaming example

curl https://onlist.io/v1/chat/completions \
  -H "Authorization: Bearer $ONLIST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
       "model": "xiaomi/mimo-v2.6-pro-ultraspeed",
       "messages": [
         {
           "role": "user",
           "content": "Write a haiku about recursion."
         }
       ],
       "stream": true
     }'

Supported parameters

NameTypeDescription
frequency_penaltynumberPenalizes new tokens based on their existing frequency in the text (-2.0 to 2.0).
include_reasoning
max_tokensintegerMaximum number of tokens to generate in the completion.
presence_penaltynumberPenalizes new tokens based on whether they appear in the text so far (-2.0 to 2.0).
reasoning
response_formatobjectSpecifies the output format. Use {"type": "json_object"} for JSON mode.
stopstring | arrayUp to 4 sequences where the API will stop generating tokens.
structured_outputs
temperaturenumberSampling temperature between 0 and 2. Higher values make output more random.
tool_choicestring | objectControls which tool is called. "auto", "none", "required", or a specific function.
toolsarrayA list of tools the model may call. Currently supports functions.
top_pnumberNucleus sampling. The model considers tokens with top_p probability mass.

These are the request parameters this model accepts. Parameter semantics follow the OpenAI Chat Completions specification.