Why tf do most benchmarks only test raw LLMs?
We should also normalize testing them with tools—browsing, running code, using APIs.
AGI is not meant to work in isolation. Give them the tools, then test them.
We should also normalize testing them with tools—browsing, running code, using APIs.
AGI is not meant to work in isolation. Give them the tools, then test them.