Apple should get working on a version of the Neural Engine that is useful for these models, and remove the 3GB size limit [1] to take full advantage of the 'unified' memory architecture. Game changer.
Waste of die space currently (on Macbook at least, I'm sure they find uses for it in the iPhone)
It's not a waste on Mac, it will dynamically switch between GPU and NPU whenever CoreML is called. There are a decent amount of applications that use CoreML.
Waste of die space currently (on Macbook at least, I'm sure they find uses for it in the iPhone)
[1] https://github.com/smpanaro/more-ane-transformers/blob/main/...