All models
DeepSeek

DeepSeek V4 Flash 0731

deepseek/deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.

Pricing

Input$0.14per 1M tokens
Output$0.28per 1M tokens
Cached input$0.028per 1M tokens read from cache

Specifications

Context window1.05Mtokens
Max output66Ktokens
ProviderDeepSeek
CapabilitiesReasoning · Function calling
Input typestext

Use DeepSeek V4 Flash 0731

Point your existing OpenAI SDK at xKiro and pass deepseek/deepseek-v4-flash-0731 as the model. Nothing else in your code changes.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.xkiro.com/v1",
    api_key="YOUR_XKIRO_KEY",
)

response = client.chat.completions.create(
    model="deepseek/deepseek-v4-flash-0731",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

Supported parameters

max_tokens · temperature · top_p · stop · frequency_penalty · presence_penalty · seed · stream · tools · tool_choice · response_format · structured_outputs · reasoning · include_reasoning

Related models

Call DeepSeek V4 Flash 0731 through the same endpoints you already use.

Get an API key