That’s a very basic way to keep the LLM inferring past the context window size (...

Ey7NFZ3P0nzAe · 2025-11-09T08:11:03 1762675863

AFAIK nobody does that. They train on much much shorter text but with use tricks in the position encoding steps that can be extrapolated by the LLMs. Lile ROPE and YARN etc.

ErikBjare · 2025-11-09T14:24:50 1762698290

AFAIK (not much) it definitely helps to train on longer sequences even with rope/yarn and is needed if you care about long context performance (and not just the long context capability).