</> HitReader
Blog Explore About

Tag: #local llm

Stop Qwen 3.8 from Overthinking: tune reasoning_effort for local runs
machine learning Aug 17, 2026 5 min read

Stop Qwen 3.8 from Overthinking: tune reasoning_effort for local runs

Qwen 3.8 27B’s default `xhigh` reasoning can turn simple prompts into slow, highly elaborated outputs. Learn how thinking mode and `reasoning_effort` interact with context windows so local runs stay responsive.

by ahsan
#llm #local llm #qwen #reasoning #vision-language
Run 80B Qwen on 4.3GB RAM with Expert Streaming on Mac
machine learning Aug 04, 2026 6 min read

Run 80B Qwen on 4.3GB RAM with Expert Streaming on Mac

Swiftlet demonstrates how to run Qwen MoE models locally by keeping only the dense core in RAM and streaming routed experts from SSD on demand. It reports ~4.3GB RAM for a 4-bit 80B model on an M5 Mac, and an on-device 35B experience on iPhone around ~2.5GB RAM.

by ahsan
#apple silicon #local llm #metal #mixture of experts #swift

Categories

  • artificial intelligence
  • machine learning
  • software engineering
  • cybersecurity
  • web development
  • privacy
  • embedded systems
  • hardware
Explore all →

Tags

#open-source #privacy #llm #cybersecurity #artificial intelligence #rust #linux #ai agents #machine learning #ai #security #llm inference
Explore all →

© 2026 HitReader.

About Explore Terms Privacy Facebook