Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coordinate their execution. We present Cascadia's resident execution of Inkling, a 975B-total/41B-active-parameter model, on eleven Intel Core U...
Tate Berenbaum, Matias Parij, M. Venkatachalam· 0 citations
We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed2551...
Matias Parij, Pawan Paudel, Tate Berenbaum et al.· 1 citation
This paper shows that a handful of AIPCs, working together over an ordinary network, can serve models beyond the capability of any single one, and leverages speculative decoding on stateful OpenVINO models.
Tate Berenbaum, M. Venkatachalam· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.