YOUR ASSUMPTIONS. VISIBLE.

What does your workload cost?

A planning scenario, not a price quote or a measured M5 benchmark.

Hardware & power assumptions
Throughput inputs are for the entire deployment. Adding machines increases cost; it does not automatically multiply speed. Zero resale value assumed.
PER MILLION COMBINED TOKENS
Local scenario$5.07Hardware + power + your operations allowance
Opus 5 · standard API$8.33$5.00 input / $25.00 output per million
Local monthly cost
$945.47
Hardware allocation
$416.67
Electricity
$28.80
Other operations
$500.00
Monthly combined tokens
186.6M

Token economics do not establish equivalent quality. Compare successful tasks, latency, and human review effort.

Sequential prefill and decode. No batching or prefix caching. Include reasoning tokens in output. The $15,000 hardware cost, 200W power draw, and throughput are planning assumptions from the September 10 feasibility report; $500/month operations is an editable illustration, not a Looski fee. Taxes, financing, model licensing, installation and margin need your own allowances. Cloud cache/batch discounts and tool fees are excluded.

THE CLOUD REFERENCE

A familiar point of comparison.

Standard, uncached API prices in USD. Verified September 10, 2026. This is a cost comparison, not a claim that models produce equivalent answers.

ModelInput / millionOutput / millionBlended at 5:1
Sonnet 5$2.00$10.00$3.33
Opus 5$5.00$25.00$8.33
Fable 5.1$10.00$50.00$16.67

Source: Anthropic pricing ↗. Cached inputs and batch discounts can reduce costs; tools can add costs. Refresh rates before a quote.

PUBLISHED MEASUREMENTS · M3 ULTRA

A capable starting point.

These independent provider measurements used an M3 Ultra with an 80-core GPU and 512GB memory. They demonstrate local inference on specific model artifacts; they are not measurements of our M5 deployment.

ModelFormatDecode tok/sPrefill tok/sPeak RAM
Qwen3.6 35B-A3B4-bit MLX94.62,89224.1GB
gpt-oss-120BMXFP4/BF16 MLX79.21,40666.2GB
Qwen3-Coder-Next4-bit MLX77.22,10047GB
Qwen3.5 397B-A17B4-bit MLX38.1510229.7GB
DeepSeek R1-05284-bit MLX20.3207380.7GB
Devstral 2 123B4-bit MLX · dense8.99072GB

Source: AI KIZAI measurements and methodology ↗. Short prompts of approximately 2,700–3,100 input tokens and 300 generated tokens, mostly single-run snapshots. Longer context and concurrency change results.

What about M5 Ultra?

The reference report models 1.5× decode as a bandwidth-sensitive scenario and 3× as conditional compute-sensitive upside. Neither is a measured result or a guaranteed range. Prompt processing and output generation must be tested separately.

MODELS TO EVALUATE

Choose the model for the work.

Developer benchmarks can identify promising candidates. Local quantization, tool access, context, and your own tasks determine the deployment choice.

Developer-reported comparisonBenchmarkOpen-weight scoreProprietary reference
DeepSeek-V4-Flash-0731Terminal Bench 2.182.7Opus 4.8 · 85.0
DeepSeek-V4-Flash-0731DeepSWE54.4Opus 4.8 · 58.0
GLM-5.3Terminal Bench 3.028.3Opus 4.8 · 21.1; Fable 5 · 33.7

Reported by DeepSeek and Z.ai, checked September 10, 2026. Different benchmark versions are not directly comparable. These results do not transfer automatically to a local quantization or newer proprietary model.

Standard deployment candidate

DeepSeek-V4-Flash-0731

A single 512GB machine is a capacity-planning target. Validate the exact artifact, long-context behavior, and customer tasks.

Model card ↗
Higher-capability candidate

GLM-5.3

Suitable quantization may fit one machine with tight headroom; larger precision or workloads may need two. Validate memory and serving support.

Model card ↗
License review required

Qwen3.8-Flash-Next

Promising memory requirements. The proposed commercial deployment arrangement needs license review before inclusion in an offer.

Model card ↗
Larger research deployment

Kimi K3

The reference report targets four machines for native-precision capacity. This is not a validated entry-level configuration or service commitment.

Model card ↗

What we validate before a commitment.

Exact model revision and license. Task success and source fidelity. Time to first token, decode speed, memory pressure, and concurrent-user behavior. Restart recovery and sustained operation. Then total cost, including support, at your actual utilization.

Four machines provide more aggregate capacity, but distributed memory is not one uniformly accessible memory pool. Additional nodes can improve capacity or latency without lowering cost per successful task.

Discuss your deployment →