The interesting part is not whether Apple wins the biggest model race, but whether it changes the economics: If enough AI runs locally, every token avoided is cloud capacity nobody has to build. That is a very different business model from selling ever more cloud compute.
I’m pretty skeptical local models can hold a candle to the cloud-based ones, particularly the ones Apple trains.
Raw capability is only one metric: A local model probably will not beat the best cloud model any time soon, but it does not need to. If it handles 80 to 90% of everyday tasks instantly, privately and at near zero marginal cost, that is a huge win. Reserve the cloud for the genuinely hard requests, not every prompt.
also this is just the beginning. give it 5-10 years and it may be more like 99%.
It seems to have been a plan for a long time, given their huge shift to unified memory architectures across most of their hardware.
They’re pretty much the only vendor where you can cost-effectively deploy a foundational LLM locally.
It would seem so. On the other hand, it is puzzling that they did not also allocate the necessary resources to the development of LLMs. 🤷
And this will have huge ramifications for my field, energy, because a substantial portion of compute power usage will move out of data centers and into the edge (your iPhone)
The decentralised operation of LLMs would also be significantly simpler and cheaper for the use of decentralised renewable energy sources.
At least gpus will finally cease their breakneck price climb, I say, for the 800th time over the last two decades.
I am practicing that strategy now. I recently upgraded to an iPad Pro M5 (the day price increases were announced, jumped on a deal immediately), upgraded several shortcuts with Apple Intelligence, and am refining them to run entirely on-device instead of in PCC (not using ChatGPT at all).
I’m waiting for the next generation of Mac Mini and Mac Studio.




