Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
What happens when you run a CUDA kernel? (fergusfinn.com)
294 points by mezark 3 months ago | past | 32 comments
Adaptive speculative decoding: picking draft lengths at runtime (fergusfinn.com)
5 points by hasheddan 3 months ago | past
InfiniBand, RoCE, and All That (fergusfinn.com)
5 points by hasheddan 3 months ago | past
InfiniBand, RoCE, and All That (fergusfinn.com)
4 points by kjeetgill 3 months ago | past
InfiniBand, RoCE, and All That (fergusfinn.com)
5 points by kkm 3 months ago | past
UCCL-EP: DeepEP-style expert parallelism on any NIC, no GPU-initiated comms (fergusfinn.com)
9 points by kkm 3 months ago | past
Anatomy of a high-performance EP kernel (fergusfinn.com)
16 points by kkm 3 months ago | past | 1 comment
The Economics of Speculative Decoding (fergusfinn.com)
30 points by kkm 3 months ago | past | 6 comments
Speculative KV coding: losslessly compressing KV cache by up to ~4× (fergusfinn.com)
155 points by kkm 3 months ago | past | 48 comments
70x faster cold(ish) starts for SGLang (fergusfinn.com)
1 point by kkm 3 months ago | past
Bringing Up DeepSeek-V4-Flash on AMD MI300X (fergusfinn.com)
120 points by kkm 3 months ago | past | 25 comments
Pushing memory bound CUDA kernels past the speed of light with data compression (fergusfinn.com)
2 points by somnial 4 months ago | past
Speculative KV coding: ~4× losslessly compressed KV cache using a small model (fergusfinn.com)
2 points by somnial 4 months ago | past
In search of wasted bits: how much information do LLM weights carry? (fergusfinn.com)
1 point by gmays 4 months ago | past
Redundant Information in LLM Weights (fergusfinn.com)
5 points by mezark 4 months ago | past
Tans: Precomputing RANS (fergusfinn.com)
3 points by mezark 5 months ago | past
Also-RANS: Asymmetric Numeral Systems for Entropy Coding (fergusfinn.com)
25 points by mezark 5 months ago | past
70x faster cold(ish) starts for SGLang (fergusfinn.com)
1 point by somnial 5 months ago | past
70x faster cold(ish) starts for SGLang (fergusfinn.com)
4 points by mezark 5 months ago | past
Parallel Primitives for Multi-Agent Workflows (fergusfinn.com)
1 point by mezark 8 months ago | past
LLM powered data structures: A lock-free binary search tree (fergusfinn.com)
1 point by somnial 8 months ago | past
Parallel Primitives for Multi-Agent Workflows (fergusfinn.com)
1 point by somnial 8 months ago | past
Scheduling in LLM Inference (fergusfinn.com)
1 point by somnial 10 months ago | past
How fast can an LLM go? (fergusfinn.com)
2 points by kkm 10 months ago | past
How fast can an LLM go? (fergusfinn.com)
2 points by gmays 10 months ago | past
How fast can an LLM go? (fergusfinn.com)
2 points by somnial 11 months ago | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: