Hacker Newsnew | past | comments | ask | show | jobs | submit | krackers's favoriteslogin
1.Reasoning LLMs are wandering solution explorers (arxiv.org)
90 points by Surreal4434 11 months ago | 98 comments
2.AI safety is mostly a sex cult in Berkeley (verysane.ai)
121 points by martythemaniak 5 days ago | 30 comments
3.Heretic removes restrictions from language models (heretic-project.org)
279 points by Bluestein 8 days ago | 110 comments
4.Jev's Architecture Unmasked (archerhume.com)
6 points by tosh 10 days ago | discuss
5.DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression (zartbot.github.io)
131 points by mfiguiere 12 days ago | 10 comments
6.A Design Space Exploration of Async/Await (brown.edu)
466 points by wcrichton 20 days ago | 134 comments
7.GPT-6 Astra, looped transformers, and hidden reasoning (sebastianraschka.com)
521 points by ModelForge 20 days ago | 162 comments
8.The browser's main thread is expensive (kciter.so)
436 points by kciter 28 days ago | 151 comments
9.How to build a diffusion language model (kuleshov-group.github.io)
184 points by volodia 30 days ago | 20 comments
10.Hilariously fast volume computation with the divergence theorem (2018) (alyssarosenzweig.ca)
266 points by luu 32 days ago | 68 comments
11.The Conspiracy Against High Temperature Sampling (gist.github.com)
4 points by theanonymousone 69 days ago
12.You only need the frontier model for one single edit (stencil.so)
230 points by jxmorris12 76 days ago | 97 comments
13.A quick look at zero-knowledge proofs (bernsteinbear.com)
89 points by evakhoury 46 days ago | 37 comments
14.The Biology of Claude's Tokenizer: Reverse Engineered (tokencontributions.substack.com)
3 points by goranmoomin 47 days ago
15.How AI text watermarking works (declaude.org)
147 points by padolsey 47 days ago | 102 comments
16.Compression is prediction (ngrok.com)
674 points by nikolay 49 days ago | 301 comments
17.A Tale of Dynamic Programming (2022) (iagoleal.com)
91 points by Brajeshwar 51 days ago | 12 comments
18.An Interesting Fourier Transform – 1/f Noise (2007) (dsprelated.com)
126 points by q7m 53 days ago | 27 comments
19.Honey, I shrunk the embeddings: Matryoshka vs. PCA (dylancastillo.co)
61 points by dcastm 57 days ago | 18 comments
20.Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025) (aleksagordic.com)
151 points by sebg 54 days ago | 10 comments
21.Explorative modeling: Train on the best of K guesses (alexiglad.github.io)
109 points by DSemba 59 days ago | 26 comments
22.The mean means nothing: data visualization to debug a latency problem (fzakaria.com)
116 points by fanf2 62 days ago | 20 comments
23.Kimi K3 Architecture Overview and Notes (sebastianraschka.com)
507 points by ModelForge 63 days ago | 111 comments
24.A walk through of the DeltaNet family of linear attention variants (doubleword.ai)
297 points by AnhTho_FR 63 days ago | 126 comments
25.Jacobian Conjecture for Baby (muchmirul.github.io)
101 points by porphyra 69 days ago | 22 comments
26.The startup's Postgres survival guide (hatchet.run)
523 points by abelanger 69 days ago | 240 comments
27.Coding agents think ahead of time (arxiv.org)
96 points by andre15silva 77 days ago | 78 comments
28.Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't (vettedconsumer.com)
74 points by ermantrout 81 days ago | 58 comments
29.Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train (arxiv.org)
153 points by tcp_handshaker 89 days ago | 43 comments
30.What happens when you run a CUDA kernel? (fergusfinn.com)
294 points by mezark 3 months ago | 32 comments

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: