RT @npew: Existing models failing at your task? Just drop a benchmark and watch the LLM providers trip over themselves to beat it. Problem…