The AI safety check that runs on a laptop and nearly matched a 35B model
The New Stack Amanda Caswell
Red Hat’s AI Safety team benchmarked nine AI guardrail configurations—comparing a small prompt-injection classifier, LLM judges, and Jev-style decision models—using NVIDIA’s NeMo Guardrails across prompt-injection and content-safety tests. Qwen3.6-35B as an LLM judge scored 89.31% accuracy on prompt injection, nearly matching DeBERTa’s 89.01% while also showing much higher median latency (312.5 ms vs 54.1 ms). Red Hat will ship both classifiers as default guardrails in OpenShift AI 3.6, while concluding that small task-specific classifiers work best for well-defined risks and that policy writing strongly affects results for decision models and other approaches.
Why it matters
Guardrail selection has typically meant choosing between a purpose-built classifier and an LLM acting as a judge. Decision models such The post The AI safety check that runs on a laptop and nearly matched a 35B model appeared first on The New Stack.