AI hallucination rates are now benchmark-dependent. A model can look solid yet...
https://wiki-view.win/index.php/Did_Gemini_3.1_Pro_Really_Cut_Hallucinations_from_88%25_to_50%25%3F
AI hallucination rates are now benchmark-dependent. A model can look solid yet fail at 30.2% on the HalluHard test. Whether you use Vectara HHEM or AA-Omniscience, your choice of metric defines your risk