Article Summary (Model: gpt-5.6-sol)
Subject: Server-Scale AI at Home
The Gist:
Strata is an MIT-licensed inference system for running the 125B-parameter Qwen3.8-Flash-Next MoE model locally on Windows or Linux PCs with supported NVIDIA or AMD GPUs. It combines aggressive quantization with expert placement across VRAM, system RAM, CPU, and SSD, plus speculative decoding. The project reports roughly 100–140 tokens/s on an RTX 3090 and 53–94 tokens/s on an RTX 5070, depending on quantization, while providing chat, vision, and OpenAI-/Anthropic-compatible APIs.
Key Claims/Facts:
- Tiered expert storage: Frequently used experts stay in VRAM; all experts remain in RAM, with CPU execution and an SSD lookup table.
- Speculative decoding: A smaller helper proposes tokens for batch verification, claimed to improve generation speed by 1.6–1.8×.
- Consumer requirements: Minimums are 12 GB VRAM, 32 GB RAM, and about 80 GB disk; 64 GB RAM supports all listed compressed variants.
Discussion Summary (Model: gpt-5.6-sol)
Consensus: Cautiously Optimistic—the speed and accessibility impressed many users, but claims of near-frontier quality drew substantial skepticism.
Top Critiques & Pushback:
Better Alternatives / Prior Art:
Expert Context: