Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There’s another way besides distillation that’s way cheaper: You can have the big model build prescriptive skills that the small model follows.

Take the “train” portion of tasks on some benchmark, have K3 complete it, and then output detailed descriptions of tools used and why, then run the validation tasks with some small model that has access to the skills.

 help



Yes. Using a harness with a strong model to create lots of utilities and tools for yourself is effectively the same thing.

Isn’t that distillation ?

No. Distillation trains on a teach model's logits or output tokens.

How do you have teacher model output anything without it being output tokens (or embeddings / intermediate logits / activations)?



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: