Safeguard

DeepSeek

DeepSeek V4.1 Flash

deepseek/deepseek-flash

Request access

Not served yet. Ask for access and we will tell you when it is live.

DeepSeek's current Flash model (DeepSeek-V4.1-Flash), a 552B-parameter MoE with native image understanding and thinking and non-thinking modes. Weights are published on Hugging Face.

Price

Input
$0.3 per 1M tokens
Output
$1.2 per 1M tokens
Notes
Peak-hour list price (cache miss). Off-peak is half: $0.15 in / $0.60 out. Peak hours 01:00-04:00 and 06:00-10:00 UTC Mon-Fri excluding Chinese public holidays. Cache-hit input $0.006 peak.
Source
DeepSeek pricing , last verified October 8, 2026

No markup on model usage: you pay the provider's list price.

Model

Context
1M (1,000,000 tokens)
Max output
384,000 tokens
Modality
Text and image
Input
text, image
Output
text
Weights
Open weights
Released
September 10, 2026
Tags
open-weights, reasoning, coding, vision, long-context, low-cost

Data policy

Prompts go to
DeepSeek (People's Republic of China)
Retention
Privacy policy: personal data stored in the PRC and kept as long as necessary to provide the service, may be used to train models, with a right to opt out. No API-specific retention period or ZDR offering published
Zero data retention
Not stated by the provider
Region
China (PRC)

Summarised from the provider's published terms. The provider's own policy is what applies.

Call it with any OpenAI SDK

Available when the model is live

from openai import OpenAI

client = OpenAI(
    base_url="https://api.safeguard.sh/v1",
    api_key="SAFEGUARD_API_KEY",
)

response = client.chat.completions.create(
    model="deepseek/deepseek-flash",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

Base URL https://api.safeguard.sh/v1, model deepseek/deepseek-flash.

Self-healing security runs on Safeguard.

Your first fix PR is minutes away.

No sales call required, even your agent can complete the purchase over MCP.