Two fine-tunes that never touch the same grapes twice.
Diffing each fine-tune against the base, tensor by tensor, showed they edit different parts of the network. So nothing has to be averaged away: both deltas are added in full, in FP32, on top of the base.
Hover a donor to see what it changed. Swift and Huihui overlap only on the output projections and the shared expert's down projection, where their deltas simply add. Vision tower and MTP head are the original model's.
24,576 experts, mapped.
Each row is one of the 48 layers, each column one of its 512 routed experts. Brightness is how often the router picks that expert on general text, from Unsloth's calibration counts. Below, each row is sorted from busiest to rarest: Slim keeps everything left of the gold line. Switch to natural order to see the raw layout.
Pick by RAM, not by ego.
The experts live in system RAM and stream to the GPU, so RAM decides the bottle. Slim, Reserve and Grand keep attention and shared experts in Q6_K/Q8_0 and differ only in how the routed experts are packed. Pocket trims those dense weights to 4-bit as well, so an 8 GB GPU still has room for an expert cache.
A 125B model on a single graphics card.
Winery Strata is our fork of Niko1221/Strata, with a wine-red glass UI and Cuvée as the default model. It keeps the busiest experts on the GPU, streams the rest from RAM, and reads the 28.8 GB per-layer-embedding table straight from the SSD.
- Clone Winery Stratagit clone https://huggingface.co/WineryLabs/Winery-Strata
- Paste one lineIt asks which drive to use (an external SSD is fine), downloads only that bottle, checks the hashes and starts it.
- Open the cellarChat in the browser UI, or point any OpenAI client at the local /v1 endpoint.