</> HitReader
Blog Explore About

Tag: #quantization

machine learning Aug 18, 2026 6 min read

Turbovec: Google’s TurboQuant Vector Search in Rust

Turbovec is a Rust vector index built on Google’s TurboQuant. It stores high-dimensional embeddings in 2–4 bits per coordinate, avoids a separate training phase, and uses SIMD-friendly layouts for fast similarity search.

by ahsan
#machine learning #quantization #rust #simd #vector search
Qwen3.8-27B FP8: FP8 Quantization, Thinking Control, and 1M Context in Practice
machine learning Aug 14, 2026 7 min read

Qwen3.8-27B FP8: FP8 Quantization, Thinking Control, and 1M Context in Practice

Qwen3.8-27B-FP8 packages a 27B vision-language model as fine-grained FP8 (block size 128) weights. It adds practical controls like `reasoning_effort` and `preserve_thinking`, plus native 262K context with extensibility toward 1M tokens for long-horizon tasks.

by ahsan
#inference #llm #machine learning #multimodal #quantization
H3-metal: Native MiniMax‑H3 Inference on Apple Silicon (Metal)
machine learning Aug 11, 2026 7 min read

H3-metal: Native MiniMax‑H3 Inference on Apple Silicon (Metal)

h3-metal is a native MiniMax‑H3 inference engine for Apple Silicon built on Metal. It combines BF16/int8 compute paths, unified-memory-aware buffer reuse, and a stateful interactive workflow to generate video and audio locally.

by ahsan
#applesilicon #inference #metal #open-source #quantization
Muse Glimmer 30B: An Open Agentic Coding Model You Can Run Locally
machine learning Aug 10, 2026 7 min read

Muse Glimmer 30B: An Open Agentic Coding Model You Can Run Locally

Muse Glimmer is Meta’s open-weight 30B agentic model released under Apache 2.0, designed to run locally for coding and tool-using workflows. Its local feasibility comes from quantization for memory fit and DFlash-based speculative decoding for faster generation.

by ahsan
#llm #local ai #open-weights #quantization #speculative decoding
Run GLM-5.2 on a Slow PC with colibrì (Disk-Streamed MoE)
ai infrastructure Jul 10, 2026 7 min read

Run GLM-5.2 on a Slow PC with colibrì (Disk-Streamed MoE)

colibrì runs GLM-5.2 (744B MoE) on consumer hardware by keeping ~9.9GB of dense int4 weights in RAM and streaming routed experts from a ~370GB int4 container on disk. It uses an LRU expert cache, MLA-style compressed KV caching, and native MTP speculative decoding to improve interaction speed once caches are warm.

by ahsan
#cpu inference #glm-5.2 #moe #quantization #systems

Categories

  • machine learning
  • software engineering
  • cybersecurity
  • embedded systems
  • systems programming
  • web development
  • artificial intelligence
  • llm engineering
Explore all →

Tags

#llm #privacy #ai #open-source #rust #cybersecurity #linux #security #ai agents #machine learning #web development #javascript
Explore all →

© 2026 HitReader.

About Explore Terms Privacy Facebook