</> HitReader
Blog Explore About

Tag: #metal

Streaming Gemma 4 26B on 2GB RAM: TurboFieldfare on M-series Macs
machine learning Jul 29, 2026 7 min read

Streaming Gemma 4 26B on 2GB RAM: TurboFieldfare on M-series Macs

TurboFieldfare runs Gemma 4 26B-A4B on Apple Silicon with ~2GB RAM by keeping core weights + an FP16 KV cache resident and streaming only the routed MoE experts from SSD. This post explains the memory pieces (MoE routing, KV cache, quantization) and walks through building, installing, and tuning the runtime.

by ahsan
#apple silicon #llm inference #metal #mixture of experts #swift

Categories

  • software engineering
  • machine learning
  • ai & development
  • ai models
  • frontend engineering
  • game development
  • llm engineering
  • music technology
Explore all →

Tags

#open-source #privacy #ai agents #ai security #javascript #llm #llm inference #productivity #software engineering #web development #ai safety #benchmarks
Explore all →

© 2026 HitReader.

About Explore Terms Privacy Facebook