Winery Labs · Qwen3.8-Flash-Next

Cuvée

One 125-billion-parameter blend, four bottles, one gaming PC.

Swift 1.5's short thinking and Huihui's abliteration poured onto Qwen3.8-Flash-Next, merged tensor by tensor in full precision, then bottled for everything from an 8 GB gaming laptop to 96 GB+ machines and served by Winery Strata.

125Bparameters
512experts / layer
10active / token
48layers
3bottles
I · The assemblage

Two fine-tunes that never touch the same grapes twice.

Diffing each fine-tune against the base, tensor by tensor, showed they edit different parts of the network. So nothing has to be averaged away: both deltas are added in full, in FP32, on top of the base.

Cuvée= Base+ (Swift − Base)+ (Huihui − Base)
Base only · Qwen3.8-Flash-Next Swift 1.5 · short thinking Huihui · refusals removed

Hover a donor to see what it changed. Swift and Huihui overlap only on the output projections and the shared expert's down projection, where their deltas simply add. Vision tower and MTP head are the original model's.

II · The cellar

24,576 experts, mapped.

Each row is one of the 48 layers, each column one of its 512 routed experts. Brightness is how often the router picks that expert on general text, from Unsloth's calibration counts. Below, each row is sorted from busiest to rarest: Slim keeps everything left of the gold line. Switch to natural order to see the raw layout.

rarebusy
L0L12L24L36L47
expert 0128256384511
Slim keeps
77%of routed tokens still land on a kept expert in Slim, averaged over layers
72%worst layer
91%best layer
III · The bottles

Pick by RAM, not by ego.

The experts live in system RAM and stream to the GPU, so RAM decides the bottle. Slim, Reserve and Grand keep attention and shared experts in Q6_K/Q8_0 and differ only in how the routed experts are packed. Pocket trims those dense weights to 4-bit as well, so an 8 GB GPU still has room for an expert cache.

64GB
1664128192256
IV · Served by Strata

A 125B model on a single graphics card.

Winery Strata is our fork of Niko1221/Strata, with a wine-red glass UI and Cuvée as the default model. It keeps the busiest experts on the GPU, streams the rest from RAM, and reads the 28.8 GB per-layer-embedding table straight from the SSD.

  1. Clone Winery Stratagit clone https://huggingface.co/WineryLabs/Winery-Strata
  2. Paste one lineIt asks which drive to use (an external SSD is fine), downloads only that bottle, checks the hashes and starts it.
  3. Open the cellarChat in the browser UI, or point any OpenAI client at the local /v1 endpoint.
Get Winery Strata ↗
GPU
Attention, DeltaNet, shared experts, router, hot expertsNVIDIA RTX 20-50 or a recent AMD card · 12 GB+
RAM
The remaining routed experts32 GB Slim · 64 GB Reserve · 96 GB+ Grand
SSD
PLE n-gram table, IQ4_NL28.8 GB, memory-mapped, looked up per token
Still to measure: tokens per second on real hardware, and benchmark scores against the parent models. We'll post them here when we have them, not before.