Answer quality report
Answer quality: Gem 1.0 compared with Gem 2.0
Pilot · subset of questions
We placed the older dialogue agent (Gem 1.0) alongside the new GenAI agent (Gem 2.0): same municipality, same questions, only the engine underneath differs.
| Criterion | Gem 1.0 | Gem 2.0 |
|---|---|---|
| Correctness | ||
| Helpfulness | ||
| Simplicity | ||
| Spelling & grammar | ||
| Tone | ||
| Weighted total | 0.872 | 0.888 |
Score bands
- 0,90–1,00
- 0,80–0,89
- 0,70–0,79
- 0,60–0,69
- < 0,60
What does this mean?
The most important differences between these versions are:
- Gem 2.0 is better at complex questions, for instance those about co-parenting.
- Gem 1.0 is better at noisy questions and informal phrasing (such as “blabla…”, “hi, i want…”).
About this measurement
- Test set
- 100 questions from 12 categories of frequently asked questions. Each score is the average across those questions.
- Criteria
- Criteria and expected answers have not yet been optimised for maximum agreement with human judges. The “correctness” criterion is currently kept general and simple.
Small print: method & caveats
A pilot on a subset: this run is not automatically representative of all topics or all municipalities.