ShiftWing: an inference engine built from scratch
Why ShiftWing manages graphics memory, system memory and fast storage as one runtime instead of treating local inference as a smaller copy of server inference.
Topic
Treating graphics memory, system memory and fast storage as one managed resource.
3 entries
Why ShiftWing manages graphics memory, system memory and fast storage as one runtime instead of treating local inference as a smaller copy of server inference.
Low-bit expert sidecars reduced storage pressure, but the quality study made clear that this is a trade to qualify rather than a free saving.
Experiments with direct storage, learned expert locality and predictive prefetch turned three separate resources into one managed path.