Qwen: Qwen2.5 VL 72B Instruct

qwen/qwen2.5-vl-72b-instruct

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

vision
MODALITIES
INPUT PRICE
$0.80per 1M
OUTPUT PRICE
$1per 1M
CONTEXT
128K
RELEASED
Feb 1, 2025
ProviderChat
$0.8$1
Hit rate
Cache read$0.4
99.0%

Capabilities

Input modalities
imagetext
Output modalities
text
Features
frequency_penaltylogit_biaslogprobsmax_tokenspresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetop_ktop_logprobstop_p
1

Get API Key

Create an API key from the Keys page, then set it as an environment variable:

export ONLIST_API_KEY=sk-...
2

Make your first request

Endpoints

POSThttps://onlist.io/v1/chat/completions

OpenAI Chat Completions format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:qwen/qwen2.5-vl-72b-instruct
POSThttps://onlist.io/v1/responses

OpenAI Responses format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:qwen/qwen2.5-vl-72b-instruct
POSThttps://onlist.io/v1/messages

Anthropic Messages format

Request Headers
Authorization:Bearer $ONLIST_API_KEY
Content-Type:application/json
Model:qwen/qwen2.5-vl-72b-instruct

Code samples

curl https://onlist.io/v1/chat/completions \
  -H "Authorization: Bearer $ONLIST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
       "model": "qwen/qwen2.5-vl-72b-instruct",
       "messages": [
         {
           "role": "user",
           "content": "Explain quantum entanglement in one paragraph."
         }
       ]
     }'

Replace $ONLIST_API_KEY with the API key from your Keys page.

Authentication

All requests must include an Authorization: Bearer <TOKEN> header. Generate keys from the Keys page; keys can be scoped to specific models, groups, IP ranges, and rate limits.

3

Enable streaming

Add "stream": true to receive partial responses as server-sent events in real time.

Streaming example

curl https://onlist.io/v1/chat/completions \
  -H "Authorization: Bearer $ONLIST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
       "model": "qwen/qwen2.5-vl-72b-instruct",
       "messages": [
         {
           "role": "user",
           "content": "Write a haiku about recursion."
         }
       ],
       "stream": true
     }'

Supported parameters

NameTypeDescription
frequency_penaltynumberPenalizes new tokens based on their existing frequency in the text (-2.0 to 2.0).
logit_bias
logprobsbooleanWhether to return log probabilities of the output tokens.
max_tokensintegerMaximum number of tokens to generate in the completion.
presence_penaltynumberPenalizes new tokens based on whether they appear in the text so far (-2.0 to 2.0).
repetition_penaltynumberPenalizes repeated tokens. Values > 1 discourage repetition.
response_formatobjectSpecifies the output format. Use {"type": "json_object"} for JSON mode.
seedintegerIf specified, the system will attempt deterministic sampling for reproducible results.
stopstring | arrayUp to 4 sequences where the API will stop generating tokens.
structured_outputs
temperaturenumberSampling temperature between 0 and 2. Higher values make output more random.
top_kintegerLimits token selection to the k most likely candidates at each step.
top_logprobsintegerNumber of most likely tokens to return at each position (0-20). Requires logprobs: true.
top_pnumberNucleus sampling. The model considers tokens with top_p probability mass.

These are the request parameters this model accepts. Parameter semantics follow the OpenAI Chat Completions specification.