Imagination details GPU strategies for LLM acceleration
Imagination Technologies previewed its work on making large language models (LLMs) like ChatGPT run faster on everyday devices using its PowerVR GPUs. The company explained two core speed metrics—how long until the first word shows up, and how fast each next one appears. They’re focusing on speeding up LLM inference with smarter caching and tailored GPU operations, aiming to bring more practical AI capabilities to phones and edge devices where power and speed matter most.
- Imagination working to speed up LLMs on PowerVR GPUs
- Highlights importance of 'first word' and per-word delays
- LLM inference benefits from GPU-parallel operations
- KV caching skips repetitive calculations for faster AI
- Aims to enable fast, local AI even on modest devices
Sources covering this
More in Chips
Micron warns RAM shortage to worsen through 2028
Micron says the current shortage of computer memory (RAM) will only get worse in 2027 and 2028.
Mod lets AMD GPUs run NVIDIA DLSS 4.5, still experimental
An open-source project called d4r has introduced early support for NVIDIA's latest DLSS upscaling tech on new AMD Radeon RX 9000 series…
Alleged leak details compromises in Apple's A20 chip for iPhone 18
A reported leak suggests that Apple's upcoming iPhone 18 will use a standard A20 chip lacking advanced packaging found in the A20 Pro.
AI drives memory prices higher into late 2026
Strong demand for AI servers is keeping memory chip makers focused on advanced server-grade DRAM, and that's pushing up prices for both…