I built H2AI Chat, an AGPL platform where several different models — from different vendors — debate a topic in turns while a human moderates. Disclosure up front: this is my project.

Over the past two days we hand-verified 41 of those debates, claim by claim: 141 statements marked, 44 of them flatly false.

We don’t delete or correct them. The sentence stays, struck through, and you can still read it by selecting it — with the reason and the source underneath. Editing what a model said would break the only promise the site makes.

Three patterns we didn’t expect:

  • Fabricated authority shows up exactly where an argument is challenged. One debate answers a budget objection with three invented citations in a single turn.
  • Fabrications spread between models. One invents a figure, a second treats it as established, a third does arithmetic on it.
  • One claim contradicts itself inside its own sentence: “62% voted Remain on a 67% turnout, meaning roughly 22% of the electorate” — which is 41.5%.

Debates: https://h2aichat.com/ Code and the fact-check register: https://github.com/Tonterias/h2aichat

    • h2aichat_com@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      3 days ago

      Thanks for reading it. One correction, though, because the ratio flatters us in the wrong direction: the 44 false claims are out of the 141 we marked, not out of everything the six models said. Those 141 are the ones we pulled out to check by hand across 41 debates; the rest went unchecked, not verified.

      So it isn’t “one in three statements is a lie” - it’s “one in three of the claims we thought were worth checking didn’t hold up”. Which is the less comfortable of the two, since it says nothing about the ones we never looked at.