User reports DDR4 RAM failures after running large AI model
A Reddit user running a 35-billion parameter AI language model on a workstation with 128GB DDR4 ECC RAM claims one memory stick failed after 90 minutes, with another showing errors after about a week. The AI workload exceeded the available GPU VRAM, forcing use of system memory. The user reports CPU and GPU temperatures remained safe, but did not monitor RAM temperatures. The cause of the RAM failures is unclear.
- AI model exceeded GPU VRAM, used system RAM
- First DDR4 stick failed after 90 minutes
- Second stick showed errors after 8-9 days
- No RAM temperature data was provided
- Source is a self-reported Reddit incident
Sources covering this
More in AI
OpenAI launches Dots AI agents amid delayed model update
OpenAI has released Dots, its new always-on AI assistant designed to autonomously handle tasks for users.
OpenAI rolls out Dots, always-on AI agents for work tasks
OpenAI has launched Dots, always-available AI assistants designed to handle projects, tasks, and reminders across apps like Slack,…
Nvidia releases Open Agent Safety Platform to contain rogue AI
Nvidia has launched the Open Agent Safety Platform, aimed at containing AI agents that attempt to circumvent their boundaries.
OpenAI cancels GPT-6.1 Astra over safety failures
OpenAI has called off the upcoming release of its GPT-6.1 Astra model after internal tests flagged safety problems.