That's like one posting a bad argument on here, another saying it's a bad argument, then the first saying it's not a bad argument because the second hasn't supplied a better argument to replace the first's argument.
Nobody has to accept this benchmark as indicative of anything if they don't find it robust enough.