Apple Unveils M5 Architecture Featuring Dedicated 2-Bit Quantization Matrix Engines
Apple's next-generation M5 silicon introduces dedicated hardware support for INT2 and INT4 weight decompression, allowing edge devices to run 14B parameter models with sub-5-watt power consumption.
During its annual hardware keynote, Apple announced the M5 unified system-on-chip, engineered with a redesigned 32-core Neural Engine specifically tailored for aggressive low-bit integer quantization.
The M5 incorporates dedicated silicon decoders that expand packed 2-bit (INT2) and 3-bit weights into unified memory registers on the fly, bypassing memory bandwidth bottlenecks. This breakthrough enables MacBook Pro and iPad Pro users to run quantized 14-billion parameter open-weights models at speeds exceeding 45 tokens per second on battery power.
This hardware milestone aligns directly with the industry shift toward private, on-device compute that completely bypasses server subscription fees and cloud surveillance.
Subscribe for Updates
Get official press announcements and version releases sent directly to your email.