R03 · MAP V1
World Map From Memory
The brief asked for a world map drawn from memory.
V2.STUDIO / INDEPENDENT AI EVALUATION
v2.studio is Victor Iglesias’s lab for working AI artifacts. Its flagship, vbench, gives six frontier model configurations the same brief, one reply each, then publishes every result in each released round side by side.
03 PUBLISHED VBENCH ROUNDS
Pilot evidence predates the levelled harness. Any released pilot round is shown as-is and excluded from cross-vendor conclusions; protocol v1 is the canonical series.
R03 · MAP V1
The brief asked for a world map drawn from memory.
R01 · BRIEF A · GAME PILOT
The brief asked for a playable browser game.
R05 · SHOW V1
The brief asked for a 60-second fireworks finale.
THE BLIND TEST
This is one artifact from a published vbench round. Six model configurations received the same brief, one reply each. Guess the maker, then see the round and recorded result.
GUESS THE MAKER
ONE REPLY EACH · NO RETRIES · THE ANSWER IS RECORDED
HOW VBENCH WORKS
SAME BRIEF
Six frontier model configurations receive one identical build brief, sent verbatim. Every recovered artifact is one self-contained HTML file a model wrote, run without repair in a network-blocked sandbox.
ONE REPLY
Within each released round, every configuration gets one prompt and one reply, with no retries or best-of selection. Each is reached through its command-line client with customization stripped, and every run records the exact command, wall time, and outcome.
FAILURES REMAIN VISIBLE
A run that produced nothing stays visible if its round is released, with the reason attributed. Recorded failures can come from the harness rather than the model, and the page says which is which.
SCOPE · IT TESTS A NARROW BUILD WORKFLOW, NOT GENERAL INTELLIGENCE. NO WINNER IS DECLARED.