DeepSeek AI Model Hits 494 Tokens/Sec on 4 DGX Spark Units
A home setup of four NVIDIA DGX Spark units, linked directly without an Ethernet switch, ran the open DeepSeek V4.1-Flash AI model—all 552 billion parameters—at a peak speed of 494 tokens per second. Tech analyst Patrick Moorhead built the DIY rig, which paired the model's efficiency tweaks with 512 GB pooled memory. NVIDIA usually recommends network switching, but the ring linkage here skipped extra hardware costs.
- DeepSeek V4.1-Flash packs 552 billion parameters
- Home rig used 4 NVIDIA DGX Spark units
- Setup delivered 494 tokens/second inference speed
- No Ethernet switch—units linked in a ring topology
- Each unit offers 128 GB unified memory, 31 TFLOPs compute
Sources covering this
In this story
More in AI
OpenAI trials visual ads in ChatGPT's image generator
OpenAI will start testing visual ads in ChatGPT this month, showing paid product images while users generate AI images.
Reflection launches Beam, a large open AI model with cheaper compute
Reflection, a well-funded US AI startup, released Beam, a massive open-weight AI model it claims rivals China’s best open models in…
OpenAI speeds up GPT-6 Astra and Sol models
OpenAI made its GPT-6 Astra and Sol models about 50% faster, now generating around 50 tokens per second instead of 30.
Professor swaps coding tests for AI-guided interviews
A university professor has replaced traditional coding exams with interviews that test how students use AI tools like Gemini in Google…