</> HitReader
Blog Explore About

Tag: #qwen

Jeff: The Tiny Local Model That Makes Decisions in ~30 ms
machine learning Sep 29, 2026 5 min read

Jeff: The Tiny Local Model That Makes Decisions in ~30 ms

Jeff turns Qwen3.5 and Gemma 4 into small decision models that classify plain-language situations without generating a paragraph. The 0.8B version reports probabilities in roughly 22–28 ms on listed GPU and Apple silicon hardware, while local fine-tuning can adapt it to specialized tasks.

by ahsan
#edge inference #local ai #model fine-tuning #qwen #zero-shot classification
Qwen3.8-27B Hits ~1,500 Tokens per Second on Cerebras
ai infrastructure Sep 03, 2026 5 min read

Qwen3.8-27B Hits ~1,500 Tokens per Second on Cerebras

Qwen3.8-27B is now listed on Cerebras public endpoints at roughly 1,500 generated tokens per second. This guide explains what the speed means, how reasoning and context limits affect real latency, and how to make a first API call.

by ahsan
#cerebras #developer tools #llm inference #multimodal ai #qwen
Qwen3.8 27B’s AA score of 52: what it means
llm Aug 17, 2026 7 min read

Qwen3.8 27B’s AA score of 52: what it means

Artificial Analysis’ Intelligence Index combines multiple reasoning and coding evaluations into a single 0–100 score. Qwen3.8 27B reportedly scored 52, suggesting strong task-completion behavior for a 27B-size model.

by ahsan
#ai-engineering #benchmarks #llm #open-weights #qwen
Stop Qwen 3.8 from Overthinking: tune reasoning_effort for local runs
machine learning Aug 17, 2026 5 min read

Stop Qwen 3.8 from Overthinking: tune reasoning_effort for local runs

Qwen 3.8 27B’s default `xhigh` reasoning can turn simple prompts into slow, highly elaborated outputs. Learn how thinking mode and `reasoning_effort` interact with context windows so local runs stay responsive.

by ahsan
#llm #local llm #qwen #reasoning #vision-language
Qwen 3.6 27B: The Local Dev Sweet Spot
local ai Jun 30, 2026 5 min read

Qwen 3.6 27B: The Local Dev Sweet Spot

Qwen 3.6 27B stands out as a practical local model: high enough quality for day-to-day development and strong performance with llama.cpp. This guide explains why 27B is the sweet spot and shows how to run it (with MTP) and integrate it into coding tools.

by ahsan
#ai tooling #llama.cpp #llm inference #local llms #qwen

Categories

  • artificial intelligence
  • machine learning
  • software engineering
  • cybersecurity
  • developer tools
  • open source
  • web development
  • embedded systems
Explore all →

Tags

#open-source #artificial intelligence #privacy #ai agents #cybersecurity #llm #machine learning #linux #rust #ai #large language models #android
Explore all →

© 2026 HitReader.

About Explore Terms Privacy Facebook