Signal Preprint presents 'Grip on LLMs' framework for evaluating government-use AI models
Summary
A preprint presents the Grip on LLMs framework, developed together with a major Dutch municipal organization to evaluate government-use language models. The evaluation suite covers more than 30 multilingual and Dutch-specific models across six dimensions: factuality, honesty, social bias, energy consumption, cost and training-data transparency. It found that no single model performs well across every dimension, and that higher quality tends to come with greater environmental impact and cost. Bias levels, by contrast, were largely unrelated to either quality or cost. The researchers also found that factuality and honesty are independent properties, meaning high factuality does not guarantee high honesty. The team additionally released a public, stakeholder-friendly overview of the models for policymakers and engineers involved in selecting government LLMs.
Classification
Evidence 1
- From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch arXiv (cs.AI) 2026-08-10 accessed 2026-08-12T11:40:08+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-eacccb4375c5
