- 尝试对Kimi KDA的数学推导&算子实现分析 08-13-2026
- Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G1: 一切的开始 (56→130) 07-30-2026
- Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G2: 从算术指令到 Occupancy 的挣扎 (130→144) 07-30-2026
- Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G4: SM103 + TMA + TCGen05 (227→221) 07-30-2026
- Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G3: Fragment Reshape 与架构级突破 (144→227) 07-30-2026
- Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G6: Cluster 与 2CTA 探索 (727→880) 07-30-2026
- Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G5: Pipeline 重构与双缓冲 (221→727) 07-30-2026
- Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G7: 回归 1CTA (880→987) 07-30-2026
- TopK kernel 优化: from 20GB/s to 2024GB/s on B300 07-06-2026