Qualcomm says its new chip can run a 30B MoE model locally on a phone.

My first reaction was to doubt it. My second was—if it really can be done, then a lot of the judgments in our teams about needing to “tune the API” will have to be revisited. Once we solve the three issues—latency, privacy, and offline capability—the product space for on-device inference is far larger than on the cloud.

Of course, being able to run it and running it fast are two different things. We’ll have to wait for real devices to confirm. But I’m quite optimistic about the direction: models will get squeezed down, chips will get beefed up, and the cloud layer in between will become thinner and thinner.