The best open-weight models for 8, 16, 32 and 64 GB machines in summer 2026: verified sizes, licenses, speeds and honest limits per tier.
The best open-weight model in autumn 2026 depends on one number: the memory your machine can give it. On 8 GB, run Qwen3.5 4B. On 16 GB, Qwen3.5 9B or Gemma 4 12B. On 32 GB, the sweet spot of the season is GLM-4.7-Flash or Qwen3.5 27B. On 64 GB or a dedicated GPU, Qwen3.5 35B-A3B and Gemma 4 31B do real work, and gpt-oss 120B becomes reachable at the top of the range. Every name, size and license in this article was verified against the official Ollama library and Hugging Face pages at publication. This is the companion piece my local AI suite tutorial promised: that article gets your stack running, this one answers the question that changes every quarter, which weights to actually pull. Model lists rot faster than any other advice in local AI, so this page carries a date in its title and a refresh promise…