Skip to main content
Models

QwenQwen3.7-Plus

Direct supply
chatstreamingvisiontoolsthinkingbatchcache

Technical Specifications

Detailed technical parameters and capabilities

Context Window
1,000,000 tokens
Max Output
131,072 tokens

Supported Modes

Chat

Capabilities

Chat

Multi-turn conversational completions

Streaming

Tokens are delivered incrementally as they are generated

Vision

Accepts image input alongside text

Tool Use

Can call functions/tools defined in the request

Thinking

Produces an extended internal reasoning trace before the final answer

Batch

Supports asynchronous batch request processing

Prompt Caching

Reuses previously processed prompt prefixes at a lower rate

Quick Integration

Get started with a single API call

quickstart.py
import openai
 
client = openai.OpenAI(
base_url="https://api.modelsite.ai/v1",
api_key="sk-ms-...",
)
 
response = client.chat.completions.create(
model="qwen3.7-plus",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Recommended for

Ideal for multimodal chat, tool use, and batch reasoning.

Pay-as-you-go
15% off

Cost

No commitments, pay only for what you use

Full Pricing Breakdown

Every billed dimension for Qwen3.7-Plus

Token Pricing

  • Cache Read$0.050 per 1M tokens
    • From 256000 tokens$1.20 per 1M tokens
  • Input$0.251 per 1M tokens
    • From 256000 tokens$6.00 per 1M tokens
  • Output$1.01 per 1M tokens
    • From 256000 tokens$24.00 per 1M tokens
  • Thinking$1.01 per 1M tokens
    • From 256000 tokens$24.00 per 1M tokens

Cache Write

  • Cache Write$0.314 per 1M tokens
    • From 256000 tokens$7.50 per 1M tokens

Other

  • Cache Read (Explicit)$0.025 per 1M tokens
    • From 256000 tokens$0.600 per 1M tokens

Indicative pricing — final cost is confirmed at request time.