← Back to Quality

Answer quality report

Answer quality: Gem 1.0 compared with Gem 2.0

Pilot · subset of questions

We placed the older dialogue agent (Gem 1.0) alongside the new GenAI agent (Gem 2.0): same municipality, same questions, only the engine underneath differs.

Scores 0.000–1.000 · colour = score band · Correctness has weight 1.0, the other criteria 0.6.
Criterion Gem 1.0 Gem 2.0
Correctness 0.908 0.909
Helpfulness 0.909 0.927
Simplicity 0.714 0.733
Spelling & grammar 0.978 1.000
Tone 0.826 0.860
Weighted total 0.872 0.888

Score bands

  • 0,90–1,00
  • 0,80–0,89
  • 0,70–0,79
  • 0,60–0,69
  • < 0,60

What does this mean?

The most important differences between these versions are:

  • Gem 2.0 is better at complex questions, for instance those about co-parenting.
  • Gem 1.0 is better at noisy questions and informal phrasing (such as “blabla…”, “hi, i want…”).

About this measurement

Test set
100 questions from 12 categories of frequently asked questions. Each score is the average across those questions.
Criteria
Criteria and expected answers have not yet been optimised for maximum agreement with human judges. The “correctness” criterion is currently kept general and simple.

Small print: method & caveats

A pilot on a subset: this run is not automatically representative of all topics or all municipalities.

All measurements

Each measurement has its own page and stays there. Only compare measurements that place the same systems side by side.