ShiftWing: an inference engine built from scratch
Why ShiftWing manages graphics memory, system memory and fast storage as one runtime instead of treating local inference as a smaller copy of server inference.
Topic
Building model runtimes around the constraints of hardware people can own and operate.
4 entries
Why ShiftWing manages graphics memory, system memory and fast storage as one runtime instead of treating local inference as a smaller copy of server inference.
Frontier capability matters most when regular people can use it without surrendering their work, context or control to a remote service.
A speculative path was admitted by confidence, verified against the ordinary path and rejected whenever it could not preserve the exact result.
Experiments with direct storage, learned expert locality and predictive prefetch turned three separate resources into one managed path.