Quality in figures

How do you know Gem works? And is Gem actually getting better? We measure it, using real residents' questions.

We do not look at a single aspect but at a coherent whole — at what an answer must contain, and equally at what does not belong in it. On this page we collect the measurements of our answer quality. Each measurement compares systems against a fixed set of criteria.

What exactly is quality? Five criteria

  • Correctness Is the answer factually right?
  • Helpfulness Does it actually help the resident along?
  • Simplicity Is it short and understandable?
  • Spelling & grammar Is it written without errors?
  • Tone Does the tone suit a municipality?

These are the five criteria we currently score automatically. Other requirements for a good answer — such as source attribution and safety — are not yet assessed by machine.

Answer quality report

Answer quality per criterion: Gem 1.0 alongside two variants of Gem 2.0

Smoke test · incompletely executed

How do the three variants score on the five general quality criteria that apply to every answer, regardless of the question asked? The same questions were put to Gem 1.0 and to two variants of Gem 2.0.

Scores as percentages · colour = score band · the total is the average of the five criteria, each counting equally.
Criterion Gem 1.0 Gem 2.0 hybrid Gem 2.0 purely generative
Correctness 93.3% 94.3% 93.1%
Helpfulness 62.4% 53.5% 77.8%
Simplicity 65.7% 71.6% 67.8%
Spelling & grammar 85.1% 90.9% 94.3%
Tone 54.9% 56.1% 69.8%
Weighted total 72.3% 73.3% 80.6%

Score bands

  • 90–100%
  • 80–89%
  • 70–79%
  • 60–69%
  • < 60%
What do these criteria mean?

Five criteria that apply to every answer, independent of what was asked.

Correctness
The answer makes no unwarranted assumptions about the asker's situation.
Helpfulness
The answer offers a concrete next step or a way to find out more.
Simplicity
Plain words, short sentences, active voice, logical structure. Made up of four sub-measurements.
Spelling & grammar
Linguistically flawless.
Tone
Professionally friendly, informal second person, empathy where needed.

What does this mean?

Across all three variants, Correctness is the strongest point (93–94%) and Tone the weakest (55–70%): the answers do not consistently use the intended informal second person and lack some empathy. Helpfulness differs most clearly between the variants. Gem 2.0 purely generative scores 78% there, well above Gem 1.0 (62%) and Gem 2.0 hybrid (54%, though that figure rests on a smaller sample).

Across the full set, Gem 2.0 purely generative has the highest weighted total (80.6 out of 100). On the 23 questions that were scored, Gem 2.0 hybrid does better than Gem 1.0, but that is not yet a statement about the full set of 51.

About this measurement

Test set
51 questions from 9 exams. Each score is the average across those questions.
Test setup
Tested at the municipality of Utrecht, assessed by an AI judge.
Gem 1.0
Recognition with non-generative AI and fixed answers based on scripted logic.
Gem 2.0 hybrid
Recognition with generative AI, fixed answers for frequently asked questions and generated answers for less frequently asked ones.
Gem 2.0 purely generative
Recognition with generative AI, and all answering done by generative AI.

Small print: method & caveats

This start-up test was executed incompletely because of limitations in the evaluation system. The figures are therefore provisional.

Gem 2.0 hybrid was scored on 23 of the 51 questions. The remaining 28 (meerinfo, sensitive, simple, tussenvragen, verwijzen, and one question from medewerker) have not yet been assessed for that variant.

The test approach still has to grow from this start-up test into a fuller set of test cases that gives a representative picture of answer quality. To that end we are working on more test capacity for larger test sets, a larger exam with criteria specific to a test case, checking tone against the intended tone, and simplifying the answers. There are still regular long answers with multiple subordinate clauses.

All measurements

Each measurement has its own page and stays there. Only compare measurements that place the same systems side by side.

Why this matters for your municipality

Answer quality is not a technical detail but a matter of reliability and reputation. A single bad answer undermines a resident's trust in your entire digital service.

That is why soundness comes first at Gem: better an honest “let me look that up for you” with a staff member involved than a slick answer that is wrong. You do not get a chatbot that “says something”, but an assistant whose answer quality is defined, measured and traceable.

Generative AI

To answer the difficult, not-yet-recognised questions better as well, we are working on deploying generative AI. That remains a careful judgement: we only use it where it demonstrably improves the answers.

So we test several language models side by side — including GPT-NL, the Dutch language model for the public sector — and compare the quality differences using practical exams: a set of pre-selected questions with matching good answers. Who gets to answer residents in future? We decide that on results, not on gut feeling.

For a long time we checked the criteria largely by hand, supplemented by reports from residents and critical researchers. Valuable, but time-consuming. So we are automating more and more checks with a test system that analyses and scores answers against those same exams. That makes quality measurable, scalable and consistent — without letting go of the human eye.

We publish the outcomes of those exams periodically.