DeepSeek V4 Flash
High-throughput V4 model with the broadest API-format support.
- Model ID
deepseek-v4-flash- Version
DeepSeek-V4-Flash-0731- Thinking
Non-thinking + thinking (default)
- Cached input
- ¥0.02
- Uncached input
- ¥1.00
- Output
- ¥2.00
- Context
- 1M
- Max output
- 384K
- Concurrency
- 2500
- Responses API
- Supported
- API formats
- OpenAI + Anthropic compatible
- Features
- JSON Output, Tool Calls, Anthropic API, prefix completion, FIM (non-thinking only)
