Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
krackers's favorites
login
submissions
|
comments
1.
Reasoning LLMs are wandering solution explorers
(
arxiv.org
)
90 points
by
Surreal4434
11 months ago
|
98 comments
2.
AI safety is mostly a sex cult in Berkeley
(
verysane.ai
)
121 points
by
martythemaniak
5 days ago
|
30 comments
3.
Heretic removes restrictions from language models
(
heretic-project.org
)
279 points
by
Bluestein
8 days ago
|
110 comments
4.
Jev's Architecture Unmasked
(
archerhume.com
)
6 points
by
tosh
10 days ago
|
discuss
5.
DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression
(
zartbot.github.io
)
131 points
by
mfiguiere
12 days ago
|
10 comments
6.
A Design Space Exploration of Async/Await
(
brown.edu
)
466 points
by
wcrichton
20 days ago
|
134 comments
7.
GPT-6 Astra, looped transformers, and hidden reasoning
(
sebastianraschka.com
)
521 points
by
ModelForge
20 days ago
|
162 comments
8.
The browser's main thread is expensive
(
kciter.so
)
436 points
by
kciter
28 days ago
|
151 comments
9.
How to build a diffusion language model
(
kuleshov-group.github.io
)
184 points
by
volodia
30 days ago
|
20 comments
10.
Hilariously fast volume computation with the divergence theorem (2018)
(
alyssarosenzweig.ca
)
266 points
by
luu
32 days ago
|
68 comments
11.
The Conspiracy Against High Temperature Sampling
(
gist.github.com
)
4 points
by
theanonymousone
69 days ago
12.
You only need the frontier model for one single edit
(
stencil.so
)
230 points
by
jxmorris12
76 days ago
|
97 comments
13.
A quick look at zero-knowledge proofs
(
bernsteinbear.com
)
89 points
by
evakhoury
46 days ago
|
37 comments
14.
The Biology of Claude's Tokenizer: Reverse Engineered
(
tokencontributions.substack.com
)
3 points
by
goranmoomin
47 days ago
15.
How AI text watermarking works
(
declaude.org
)
147 points
by
padolsey
47 days ago
|
102 comments
16.
Compression is prediction
(
ngrok.com
)
674 points
by
nikolay
49 days ago
|
301 comments
17.
A Tale of Dynamic Programming (2022)
(
iagoleal.com
)
91 points
by
Brajeshwar
51 days ago
|
12 comments
18.
An Interesting Fourier Transform – 1/f Noise (2007)
(
dsprelated.com
)
126 points
by
q7m
53 days ago
|
27 comments
19.
Honey, I shrunk the embeddings: Matryoshka vs. PCA
(
dylancastillo.co
)
61 points
by
dcastm
57 days ago
|
18 comments
20.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)
(
aleksagordic.com
)
151 points
by
sebg
54 days ago
|
10 comments
21.
Explorative modeling: Train on the best of K guesses
(
alexiglad.github.io
)
109 points
by
DSemba
59 days ago
|
26 comments
22.
The mean means nothing: data visualization to debug a latency problem
(
fzakaria.com
)
116 points
by
fanf2
62 days ago
|
20 comments
23.
Kimi K3 Architecture Overview and Notes
(
sebastianraschka.com
)
507 points
by
ModelForge
63 days ago
|
111 comments
24.
A walk through of the DeltaNet family of linear attention variants
(
doubleword.ai
)
297 points
by
AnhTho_FR
63 days ago
|
126 comments
25.
Jacobian Conjecture for Baby
(
muchmirul.github.io
)
101 points
by
porphyra
69 days ago
|
22 comments
26.
The startup's Postgres survival guide
(
hatchet.run
)
523 points
by
abelanger
69 days ago
|
240 comments
27.
Coding agents think ahead of time
(
arxiv.org
)
96 points
by
andre15silva
77 days ago
|
78 comments
28.
Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't
(
vettedconsumer.com
)
74 points
by
ermantrout
81 days ago
|
58 comments
29.
Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
(
arxiv.org
)
153 points
by
tcp_handshaker
89 days ago
|
43 comments
30.
What happens when you run a CUDA kernel?
(
fergusfinn.com
)
294 points
by
mezark
3 months ago
|
32 comments
More
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: