Small AI model matches large LLM on AI safety tasks
Red Hat compared nine AI guardrail methods for detecting prompt injection and unsafe content. Its 200-million-parameter DeBERTa classifier scored 89.01% accuracy on prompt injection, nearly matching a much larger 35-billion-parameter Qwen model, but with a much faster 54 millisecond response time. For broader content safety, a decision-model approach called Jev led with 86.2% accuracy, edging out open source and LLM competitors.
- Red Hat tested nine prompt-injection and content-safety guardrails
- Qwen3.6-35B LLM led in injection tests with 89.31% accuracy
- Red Hat's smaller DeBERTa scored 89.01%—but much faster
- Jev decision model won content safety, at 86.2% accuracy
- Granite Guardian classifier was fastest, though lagged in accuracy
Sources covering this
More in AI
OpenAI trials visual ads in ChatGPT's image generator
OpenAI will start testing visual ads in ChatGPT this month, showing paid product images while users generate AI images.
Reflection launches Beam, a large open AI model with cheaper compute
Reflection, a well-funded US AI startup, released Beam, a massive open-weight AI model it claims rivals China’s best open models in…
OpenAI speeds up GPT-6 Astra and Sol models
OpenAI made its GPT-6 Astra and Sol models about 50% faster, now generating around 50 tokens per second instead of 30.
Professor swaps coding tests for AI-guided interviews
A university professor has replaced traditional coding exams with interviews that test how students use AI tools like Gemini in Google…