← Back to Quality

Answer quality report

Answer quality per criterion: Gem 1.0 alongside two variants of Gem 2.0

Smoke test · incompletely executed

How do the three variants score on the five general quality criteria that apply to every answer, regardless of the question asked? The same questions were put to Gem 1.0 and to two variants of Gem 2.0.

Scores as percentages · colour = score band · the total is the average of the five criteria, each counting equally.
Criterion Gem 1.0 Gem 2.0 hybrid Gem 2.0 purely generative
Correctness 93.3% 94.3% 93.1%
Helpfulness 62.4% 53.5% 77.8%
Simplicity 65.7% 71.6% 67.8%
Spelling & grammar 85.1% 90.9% 94.3%
Tone 54.9% 56.1% 69.8%
Weighted total 72.3% 73.3% 80.6%

Score bands

  • 90–100%
  • 80–89%
  • 70–79%
  • 60–69%
  • < 60%
What do these criteria mean?

Five criteria that apply to every answer, independent of what was asked.

Correctness
The answer makes no unwarranted assumptions about the asker's situation.
Helpfulness
The answer offers a concrete next step or a way to find out more.
Simplicity
Plain words, short sentences, active voice, logical structure. Made up of four sub-measurements.
Spelling & grammar
Linguistically flawless.
Tone
Professionally friendly, informal second person, empathy where needed.

What does this mean?

Across all three variants, Correctness is the strongest point (93–94%) and Tone the weakest (55–70%): the answers do not consistently use the intended informal second person and lack some empathy. Helpfulness differs most clearly between the variants. Gem 2.0 purely generative scores 78% there, well above Gem 1.0 (62%) and Gem 2.0 hybrid (54%, though that figure rests on a smaller sample).

Across the full set, Gem 2.0 purely generative has the highest weighted total (80.6 out of 100). On the 23 questions that were scored, Gem 2.0 hybrid does better than Gem 1.0, but that is not yet a statement about the full set of 51.

About this measurement

Test set
51 questions from 9 exams. Each score is the average across those questions.
Test setup
Tested at the municipality of Utrecht, assessed by an AI judge.
Gem 1.0
Recognition with non-generative AI and fixed answers based on scripted logic.
Gem 2.0 hybrid
Recognition with generative AI, fixed answers for frequently asked questions and generated answers for less frequently asked ones.
Gem 2.0 purely generative
Recognition with generative AI, and all answering done by generative AI.

Small print: method & caveats

This start-up test was executed incompletely because of limitations in the evaluation system. The figures are therefore provisional.

Gem 2.0 hybrid was scored on 23 of the 51 questions. The remaining 28 (meerinfo, sensitive, simple, tussenvragen, verwijzen, and one question from medewerker) have not yet been assessed for that variant.

The test approach still has to grow from this start-up test into a fuller set of test cases that gives a representative picture of answer quality. To that end we are working on more test capacity for larger test sets, a larger exam with criteria specific to a test case, checking tone against the intended tone, and simplifying the answers. There are still regular long answers with multiple subordinate clauses.

All measurements

Each measurement has its own page and stays there. Only compare measurements that place the same systems side by side.