I don't know why you assume there has to be a gotcha, maybe it's the competitive background... Anyway, it's visual because you look at it to see it. And it's not the best human performance vs best LLM performance, it's best controlled performance because the testing is limited to a set of parameters.