DeepSeek-V4-Flash-0731
A single 512GB machine is a capacity-planning target. Validate the exact artifact, long-context behavior, and customer tasks.
Model card ↗Compare useful work, real capacity, and the cost of keeping it running. Start with the evidence. Adjust the assumptions to your business.
A planning scenario, not a price quote or a measured M5 benchmark.
Token economics do not establish equivalent quality. Compare successful tasks, latency, and human review effort.
Sequential prefill and decode. No batching or prefix caching. Include reasoning tokens in output. The $15,000 hardware cost, 200W power draw, and throughput are planning assumptions from the September 10 feasibility report; $500/month operations is an editable illustration, not a Looski fee. Taxes, financing, model licensing, installation and margin need your own allowances. Cloud cache/batch discounts and tool fees are excluded.
Standard, uncached API prices in USD. Verified September 10, 2026. This is a cost comparison, not a claim that models produce equivalent answers.
| Model | Input / million | Output / million | Blended at 5:1 |
|---|---|---|---|
| Sonnet 5 | $2.00 | $10.00 | $3.33 |
| Opus 5 | $5.00 | $25.00 | $8.33 |
| Fable 5.1 | $10.00 | $50.00 | $16.67 |
Source: Anthropic pricing ↗. Cached inputs and batch discounts can reduce costs; tools can add costs. Refresh rates before a quote.
These independent provider measurements used an M3 Ultra with an 80-core GPU and 512GB memory. They demonstrate local inference on specific model artifacts; they are not measurements of our M5 deployment.
| Model | Format | Decode tok/s | Prefill tok/s | Peak RAM |
|---|---|---|---|---|
| Qwen3.6 35B-A3B | 4-bit MLX | 94.6 | 2,892 | 24.1GB |
| gpt-oss-120B | MXFP4/BF16 MLX | 79.2 | 1,406 | 66.2GB |
| Qwen3-Coder-Next | 4-bit MLX | 77.2 | 2,100 | 47GB |
| Qwen3.5 397B-A17B | 4-bit MLX | 38.1 | 510 | 229.7GB |
| DeepSeek R1-0528 | 4-bit MLX | 20.3 | 207 | 380.7GB |
| Devstral 2 123B | 4-bit MLX · dense | 8.9 | 90 | 72GB |
Source: AI KIZAI measurements and methodology ↗. Short prompts of approximately 2,700–3,100 input tokens and 300 generated tokens, mostly single-run snapshots. Longer context and concurrency change results.
The reference report models 1.5× decode as a bandwidth-sensitive scenario and 3× as conditional compute-sensitive upside. Neither is a measured result or a guaranteed range. Prompt processing and output generation must be tested separately.
Developer benchmarks can identify promising candidates. Local quantization, tool access, context, and your own tasks determine the deployment choice.
| Developer-reported comparison | Benchmark | Open-weight score | Proprietary reference |
|---|---|---|---|
| DeepSeek-V4-Flash-0731 | Terminal Bench 2.1 | 82.7 | Opus 4.8 · 85.0 |
| DeepSeek-V4-Flash-0731 | DeepSWE | 54.4 | Opus 4.8 · 58.0 |
| GLM-5.3 | Terminal Bench 3.0 | 28.3 | Opus 4.8 · 21.1; Fable 5 · 33.7 |
Reported by DeepSeek and Z.ai, checked September 10, 2026. Different benchmark versions are not directly comparable. These results do not transfer automatically to a local quantization or newer proprietary model.
A single 512GB machine is a capacity-planning target. Validate the exact artifact, long-context behavior, and customer tasks.
Model card ↗Suitable quantization may fit one machine with tight headroom; larger precision or workloads may need two. Validate memory and serving support.
Model card ↗Promising memory requirements. The proposed commercial deployment arrangement needs license review before inclusion in an offer.
Model card ↗The reference report targets four machines for native-precision capacity. This is not a validated entry-level configuration or service commitment.
Model card ↗Exact model revision and license. Task success and source fidelity. Time to first token, decode speed, memory pressure, and concurrent-user behavior. Restart recovery and sustained operation. Then total cost, including support, at your actual utilization.
Four machines provide more aggregate capacity, but distributed memory is not one uniformly accessible memory pool. Additional nodes can improve capacity or latency without lowering cost per successful task.
Discuss your deployment →