</> HitReader
Blog Explore About

Category: gpu

Auto-research with Codex: a 232× Faster GPU QR Kernel
gpu Aug 15, 2026 8 min read

Auto-research with Codex: a 232× Faster GPU QR Kernel

A batched compact-Householder QR kernel is tricky on GPUs because Householder reflectors form a dependency chain. This post explains how blocked Householder QR with compact WY updates turns the trailing work into GEMM-shaped operations, and how an auto-research loop with Codex drove fast, correct kernel iterations to a 232× speedup over baseline.

by ahsan
#autotuning #codex #cuda #gpu #linear algebra

Categories

  • artificial intelligence
  • machine learning
  • software engineering
  • cybersecurity
  • web development
  • embedded systems
  • hardware
  • privacy
Explore all →

Tags

#open-source #privacy #llm #cybersecurity #artificial intelligence #linux #rust #ai agents #machine learning #ai #security #llm inference
Explore all →

© 2026 HitReader.

About Explore Terms Privacy Facebook