About This Blog
Made with Hexo
Finder
About Finder
Quit Finder
File
Close Window
Edit
Copy
Select All
View
Toggle Full Screen
Enter Full Screen
Go
Blog
Archives
Readings
Projects
Publications
Links
About
Window
Minimize
Zoom
Help
Blog Help
Hello, jincheng
Wi-Fi
Network
Inicenoj
Status
Connected
Signal
Excellent
Battery
Charge
100%
Source
Power Adapter
Condition
Normal
Notifications
尝试对Kimi KDA的数学推导&算子实现分析
08-13
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G1: 一切的开始 (56→130)
07-30
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G2: 从算术指令到 Occupancy 的挣扎 (130→144)
07-30
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G4: SM103 + TMA + TCGen05 (227→221)
07-30
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G3: Fragment Reshape 与架构级突破 (144→227)
07-30
readings
MoE vs Dense Models in Inference
05-21-2026
1 items
Archives
Readings
Projects
Publications
Links
About
3D Reconstruction
RL Foundations
ML Systems
HPC Games
Basics
Misc
MD
尝试对Kimi KDA的数学推导&算子实现分析
MD
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G1: 一切的开始 (56→130)
MD
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G2: 从算术指令到 Occupancy 的挣扎 (130→144)
MD
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G4: SM103 + TMA + TCGen05 (227→221)
MD
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G3: Fragment Reshape 与架构级突破 (144→227)
MD
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G6: Cluster 与 2CTA 探索 (727→880)
MD
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G5: Pipeline 重构与双缓冲 (221→727)
MD
Flash Attention CUDA Kernel 优化: 从 56 到 986 TFLOPS on B300 — G7: 回归 1CTA (880→987)
MD
TopK kernel 优化: from 20GB/s to 2024GB/s on B300
MD
Paged attention kernel optimization(I)
MD
MoE vs Dense Models in Inference
MD
Streams and Concurrency on CUDA
MD
Foundation of Reinforcement learning(V)
MD
Foundation of Reinforcement learning(IV)
MD
Foundation of Reinforcement learning(III)
MD
Foundation of Reinforcement learning(II)
MD
Foundation of Reinforcement learning(I)
MD
nanoPD:一个 LLM P/D 分离推理引擎的实现笔记
MD
3D Reconstruction Series
MD
学习笔记:Tensor Parallelism(TP)
MD
HPCGames 题解 D E 题
MD
HPCGames 题解 A B C 题
MD
SF3D 论文阅读记录
MD
ViT Transformer 的阅读?(应该算是阅读吧)
MD
回顾一下Transformer
MD
SLAM Former 阅读
MD
重返vggt
MD
论文阅读记录:reloc3r
MD
论文阅读记录:Fast3R
MD
论文阅读记录:MAst3R
MD
VGGT读后有感
MD
为SLAM3R补充实时处理函数方法
MD
SLAM3R读后有感
MD
Celebrate and Introduce My First Page
No Results
Finder
Launchpad
Readings
Projects
Publications
Links
About
Trash