Skip to content
All experiments
EXP-001PrototypeAIWebGPU

Edge Inference

An exploration of how far small, quantised models can go when they run next to the user instead of in a data centre.

Placeholder content. Details to be confirmed.

Findings so far

  1. 01

    Quantised models under 500MB run interactively on recent laptops.

  2. 02

    Latency drops sharply once network round-trips disappear.

  3. 03

    Battery cost is the main open question on mobile.