Why Does a Model Look Great on Vectara but Bad on AA-Omniscience?
https://privatebin.net/?adec98fdd54e6c7d#Cn3nwWUi7zjhGDM9z522TPVcu6Yjr2AypUpkCZdoffEo
In the past decade of evaluating Large Language Models (LLMs), I have seen the same cycle repeat itself