EXP-001PrototypeAIWebGPU
Edge Inference
An exploration of how far small, quantised models can go when they run next to the user instead of in a data centre.
Placeholder content. Details to be confirmed.
Findings so far
- 01
Quantised models under 500MB run interactively on recent laptops.
- 02
Latency drops sharply once network round-trips disappear.
- 03
Battery cost is the main open question on mobile.