</> HitReader
Blog Explore About

Tag: #cuda

Auto-research with Codex: a 232× Faster GPU QR Kernel
gpu Aug 15, 2026 8 min read

Auto-research with Codex: a 232× Faster GPU QR Kernel

A batched compact-Householder QR kernel is tricky on GPUs because Householder reflectors form a dependency chain. This post explains how blocked Householder QR with compact WY updates turns the trailing work into GEMM-shaped operations, and how an auto-research loop with Codex drove fast, correct kernel iterations to a 232× speedup over baseline.

by ahsan
#autotuning #codex #cuda #gpu #linear algebra

Categories

  • machine learning
  • software engineering
  • cybersecurity
  • systems programming
  • web development
  • artificial intelligence
  • security
  • ai agents
Explore all →

Tags

#llm #open-source #privacy #ai #ai agents #linux #security #web development #cybersecurity #javascript #llm inference #mixture of experts
Explore all →

© 2026 HitReader.

About Explore Terms Privacy Facebook