Historical Trends
Runs
Queries tested
Overall pass rate
Date range
One LLM pass reading across every run's results and summary — where the system is weak, and why that's likely happening. Regenerated only when new run data shows up.
The system's performance is inconsistent and trending downward, characterized by a failure to prioritize direct answers over generic conversational flows. The chatbot frequently defaults to irrelevant quiz sequences or deflects queries, indicating a significant gap in its ability to retrieve and synthesize specific brand guidelines and product data.
Matched by exact theme text across runs' findings — a rephrased version of the same root cause won't merge with an older one.
| Date | Run | Business | Server | Queries | Pass rate |
|---|---|---|---|---|---|
| 2026-09-04 16:26 | biz27_20260904_162512 | 27 | staging | 54 | 48% |
| 2026-09-04 16:18 | biz21_20260904_161714 | 21 | staging | 54 | 43% |
| 2026-09-02 12:38 | test-furqan_20260902_123716 | 16 | staging | 22 | 73% |
| 2026-09-02 12:20 | barefaced-testing-impaired-barrier_20260902_121212 | 17 | staging | 259 | 59% |
| 2026-09-02 09:01 | rodial-test-report_20260902_085105 | 12 | staging | 372 | 41% |
| 2026-09-01 17:08 | nip-fab-inventory-p1_20260901_170833 | 16 | staging | 11 | 82% |
| 2026-09-01 17:05 | all-qs-all-profiles_20260901_160703 | 16 | staging | 2304 | 65% |
| 2026-09-01 16:37 | biz16_20260901_163736 | 16 | staging | 3 | 0% |
| 2026-09-01 15:42 | nipfab1212_20260901_154119 | 16 | staging | 22 | 64% |