Answer quality report
Answer quality per criterion: Gem 1.0 alongside two variants of Gem 2.0
Smoke test · incompletely executed
How do the three variants score on the five general quality criteria that apply to every answer, regardless of the question asked? The same questions were put to Gem 1.0 and to two variants of Gem 2.0.
| Criterion | Gem 1.0 | Gem 2.0 hybrid | Gem 2.0 purely generative |
|---|---|---|---|
| Correctness | |||
| Helpfulness | |||
| Simplicity | |||
| Spelling & grammar | |||
| Tone | |||
| Weighted total | 72.3% | 73.3% | 80.6% |
Score bands
- 90–100%
- 80–89%
- 70–79%
- 60–69%
- < 60%
What do these criteria mean?
Five criteria that apply to every answer, independent of what was asked.
- Correctness
- The answer makes no unwarranted assumptions about the asker's situation.
- Helpfulness
- The answer offers a concrete next step or a way to find out more.
- Simplicity
- Plain words, short sentences, active voice, logical structure. Made up of four sub-measurements.
- Spelling & grammar
- Linguistically flawless.
- Tone
- Professionally friendly, informal second person, empathy where needed.
What does this mean?
Across all three variants, Correctness is the strongest point (93–94%) and Tone the weakest (55–70%): the answers do not consistently use the intended informal second person and lack some empathy. Helpfulness differs most clearly between the variants. Gem 2.0 purely generative scores 78% there, well above Gem 1.0 (62%) and Gem 2.0 hybrid (54%, though that figure rests on a smaller sample).
Across the full set, Gem 2.0 purely generative has the highest weighted total (80.6 out of 100). On the 23 questions that were scored, Gem 2.0 hybrid does better than Gem 1.0, but that is not yet a statement about the full set of 51.
About this measurement
- Test set
- 51 questions from 9 exams. Each score is the average across those questions.
- Test setup
- Tested at the municipality of Utrecht, assessed by an AI judge.
- Gem 1.0
- Recognition with non-generative AI and fixed answers based on scripted logic.
- Gem 2.0 hybrid
- Recognition with generative AI, fixed answers for frequently asked questions and generated answers for less frequently asked ones.
- Gem 2.0 purely generative
- Recognition with generative AI, and all answering done by generative AI.
Small print: method & caveats
This start-up test was executed incompletely because of limitations in the evaluation system. The figures are therefore provisional.
Gem 2.0 hybrid was scored on 23 of the 51 questions. The remaining 28 (meerinfo, sensitive, simple, tussenvragen, verwijzen, and one question from medewerker) have not yet been assessed for that variant.
The test approach still has to grow from this start-up test into a fuller set of test cases that gives a representative picture of answer quality. To that end we are working on more test capacity for larger test sets, a larger exam with criteria specific to a test case, checking tone against the intended tone, and simplifying the answers. There are still regular long answers with multiple subordinate clauses.