Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You can hide a whole lot of essential complexity in a hardware layer.

However, the very next question a researcher will ask once a model fits on one device is “can I make it twice as fast/big if I use two?”



I'd think that such researcher would've already heard of adages like "you can't get nine women to give birth in one month", or "where there's six cooks, there's nothing to eat".

Or more directly, perhaps one should ask such researcher, "if your team was to double in head count, would you do this project twice as fast?".


8 GPUs do a pretty bang-up job of doing 8 months of compute in a month :).

I think the broader point is that the last x0 years of ML research show that more compute is better, both for iteration speed and for resulting performance. Distribution is just the natural outgrowth of that imperative once it reaches the limit of a single device/node. If Cerebras can address models at today's scale on one device, the immediate next step is "what can N of these devices do together to build models at tomorrow's scale"

[0]: https://www.cerebras.net/condor-galaxy-1


Fair enough :).

I still think work on improving single-core/device performance is worthwhile, as distribution will always strictly not-better, and almost always strictly worse, due to coordination costs reducing efficiency. If two Cerberas can be glued together and achieve roughly 2x of their performance, it's still going to be more efficient than achieving equivalent performance from many more regular GPUs. Getting the hardware fast enough so that you need just one device for your problem - that's a special case that will yield extra win.


> Or more directly, perhaps one should ask such researcher, "if your team was to double in head count, would you do this project twice as fast?".

If everybody's job was just to do dot products all day I'd hope the answer would be yes.


Which dot products did you do, I'll do the next one. Oh, that was John's, but he is away on vacation today. Let me take care of John's and tomorrow we have a quick meeting to see which matrix he takes next. Sounds a lot like a bus :)




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: