Tweet

5.288 tweets12 añosúltimo 17 ago 2026

Una entrada del archivo público.

← volver a todos los años

Why tf do most benchmarks only test raw LLMs?

We should also normalize testing them with tools—browsing, running code, using APIs.

AGI is not meant to work in isolation. Give them the tools, then test them.