serhii.net

In the middle of the desert you can say anything you want

UNLISTED

18 Sep 2025

expose eval thoughts

  • Misc
    • Having a single ID might help, instead of “expose index 0 of type Kompakt” — to easily filter and find matches
      • Or just doing it all in one text style and target group
        • then we can uniquely identify by existing id
        • then we can have the human reference text as well
      • then we take less time evaluating
  • Template
    • do we need the confidence and severity fields? We won’t evaluate it anyway
    • “omission” doesn’t make sense because different priorities, no chance to add EVERYTHING
      • and then every time there’s a field not in the expose but not an “omission” we know eval is wrong
  • All things like ‘good investiion’ are basically hallucinations, but we force this in the templates to make it interesting/engaging
  • Any reason why we can’t use the old metrics from last presentation? Harder to prove wrong and intuitively neat. And easy to estimate

later

  • does calling it gold increase ratings?
  • does splitting into 2 separate changeanything?
    • add sample json with clearly fake data to measure contamination
  • how stable is it running with same parameters X times?
Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus