W3C Freezes WebGPU 2.0 Specifications: Browser-Based AI Inference Speeds Jump 3x
The World Wide Web Consortium (W3C) has finalized the WebGPU 2.0 specifications, standardizing cooperative matrix multiplication and subgroup operations for zero-latency in-browser neural networks.
The World Wide Web Consortium (W3C) GPU for the Web Working Group announced today that the WebGPU 2.0 specification has officially reached Candidate Recommendation status, locking in architectural features designed specifically for browser-native machine learning.
The landmark update introduces standardized subgroup operations and cooperative matrix multiplication primitives across Chrome, Safari, and Firefox. In early benchmarks running quantized 7-billion parameter language models locally via WebLLM, inference throughput increased by an average of 310% compared to WebGPU 1.0.
"WebGPU 2.0 transitions the browser from a simple graphical canvas into a first-class compute runtime for neural architectures," stated Corentin Wallez, co-chair of the W3C GPU Working Group. "Users can now execute advanced diffusion and transformer models without transmitting personal prompts or tokens over the network."
Explore client-side browser capabilities directly on Luminus with our AI Prompt Generator and zero-retention developer utilities.
Subscribe for Updates
Get official press announcements and version releases sent directly to your email.