ee
- clear mapping/terminology
- e.g. in eval output at the end “a: ours, b: theirs” -> will also help us to later switch them places without getting confused
- ours -> generated, theirs -> gold, human-> ?? any single terminology is OK as long as it’s consistent
- gold object properties in the same file as well maybe? same as generated/gold text currently there, to show what is actually available to the model instead of matching to other files
general
- hallucination score: more is better
0/geschaftsleute/kompakt
A/ours/generated
“Dieses repräsentative Ein-/Zweifamilienhaus in Ratingen, 40883, bietet Geschäftsleuten eine ideale Kombination aus Wohn- und Nutzungsmöglichkeiten. Mit einer Wohnfläche von 539,47 m² auf einem Grundstück von 3.309,0 m² und einem Kaufpreis von 2.995.000,00 EUR, ist es perfekt für die Nutzung als Büro oder Praxis geeignet. Das 1994 erbaute Haus verfügt über 8 Zimmer, darunter 5 Schlafzimmer und 4 Badezimmer, und ist in einem gepflegten Zustand. Die Energieeffizienzklasse E und die Zentralheizung mit Öl als Energiequelle sorgen für eine zuverlässige Versorgung.\n\nDie Ausstattung umfasst eine Einbauküche, Teppich- und Fliesenböden, sowie eine Einliegerwohnung, die sich ideal für Mitarbeiter oder Gäste eignet. Ein Personenaufzug, eine Sauna, ein Kamin und eine Fußbodenheizung bieten zusätzlichen Komfort. Der Garten, die Terrasse, der Balkon und die Loggia bieten ausreichend Raum für Entspannung und repräsentative Veranstaltungen. Eine Garage und 3 Parkplätze runden das Angebot ab. Die behindertengerechte und seniorengerechte Ausstattung sowie die Erlaubnis für Haustiere machen das Haus zu einer attraktiven Option für Geschäftsleute, die Wert auf Funktionalität und Komfort legen.”
- how do we get b/72 overall assestment when the category scores are 6/6/7/7?
- object description: identical PRC scores, identical hallucination RATE but not SCORE?
- identical metrics breakdown by total # of stuff
- metrics breakdown
-
- leak still present, because ’total agent/gold claims'
-
- both false positives and hallucinated claims present in the dictionary
-
- agent A
- found_issues a lot of stuff
- 0 is not a hallucination
- 2 ideal für geschaftsleute -> again, nota hallucination
- found_issues a lot of stuff
-
- agent b
- why no found issues, but still 7 hallucinated claims?
- agent b
- +comparison summary
- agent B had a lower hallucination score so more accurate description
- a had 5, which is better, but not lower+
- hallucination_score_diff is -1 => should be +1?
- agent B had a lower hallucination score so more accurate description
- object description jq[3] -> why no factual data means hallucination score of 5?
- total groud truth claims -> it’s different because the same keys but different things in the variable list?