Researchers boost LLM efficiency with high-bandwidth flash storage
A team from UC Berkeley and FuriosaAI found that adding high-bandwidth flash (HBF) memory to GPU servers running large language models, like ChatGPT, can speed up processing by 36% to 87% on demanding workloads compared to using only traditional high-bandwidth memory (HBM). Their new system design and scheduling method not only made LLM serving faster but also cut modeled energy use by over 50% in some cases, while tripling the flash memory’s estimated lifespan.
- HBF memory boosts LLM serving speed up to 87%
- Model predicts up to 56% energy savings in some scenarios
- Scheduling method extends flash memory life to 15 years
- More efficient data placement is key for balancing speed and flash health
Sources covering this
More in Chips
Rare AMD Ryzen 9 5900X3D engineering sample revealed
A previously unseen engineering sample of AMD's cancelled Ryzen 9 5900X3D has appeared on a Chinese forum.
Microsoft brings faster game loading to NVIDIA RTX cards
Microsoft's Advanced Shader Delivery, which lets games ship with pre-compiled graphics shaders to reduce load times and cut gameplay…
Researchers detail hardware design model abstractions
A team from Infineon and Technical University of Munich published new research exploring how hardware engineers use different layers of…
AMD Gainsborough chip spotted, Steam Deck 2 speculation grows
A new AMD chip called Gainsborough, surfaced in recent driver code and shipping records, is sparking rumors it could be designed for…