The issue I find with this is that the frontier models still outperform the small finetuned model on its specific task. So much so that the ROI on doing fine tunes is likely negative. I would love to hear some specific example where it did provide value though, if any has any. That would be helpful to start being able to find similar cases.
There are lots of distilled models on huggingface that are much smaller than, say, Opus. They are distilled from Opus or Fable and show clear improvements. I do not have the budget to fine-tune a 30b or more parameter model but from my tiny model experiments, the results are quite clear. Again, I have only a couple small tests.
Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?
I think LLMs just made us think that is the only way. It is not.