</> HitReader
Blog Explore About

Tag: #moe

Inkling and the new era of open-weights customization
machine learning Jul 16, 2026 7 min read

Inkling and the new era of open-weights customization

Inkling is a multimodal open-weights MoE model (975B total, 41B active) built for customization: up to 1M-token context, controllable thinking effort, and fine-tuning via Tinker (LoRA). The release also demonstrates a self-finetuning loop that defines an objective, trains, evaluates, and swaps weights.

by ahsan
#fine-tuning #inkling #llm systems #moe #tinker
Run GLM-5.2 on a Slow PC with colibrì (Disk-Streamed MoE)
ai infrastructure Jul 10, 2026 7 min read

Run GLM-5.2 on a Slow PC with colibrì (Disk-Streamed MoE)

colibrì runs GLM-5.2 (744B MoE) on consumer hardware by keeping ~9.9GB of dense int4 weights in RAM and streaming routed experts from a ~370GB int4 container on disk. It uses an LRU expert cache, MLA-style compressed KV caching, and native MTP speculative decoding to improve interaction speed once caches are warm.

by ahsan
#cpu inference #glm-5.2 #moe #quantization #systems

Categories

  • machine learning
  • software engineering
  • cybersecurity
  • embedded systems
  • systems programming
  • web development
  • artificial intelligence
  • llm engineering
Explore all →

Tags

#llm #privacy #ai #open-source #rust #cybersecurity #linux #security #ai agents #machine learning #web development #javascript
Explore all →

© 2026 HitReader.

About Explore Terms Privacy Facebook