V2.STUDIO / INDEPENDENT AI EVALUATION

See what AI actually builds.

v2.studio is Victor Iglesias’s lab for working AI artifacts. Its flagship, vbench, gives six frontier model configurations the same brief, one reply each, then publishes every result in each released round side by side.

03 PUBLISHED VBENCH ROUNDS

Same brief. Six configurations. One shot each.

03 BRIEFS · 18 ARTIFACTS RECOVERED · GRADED ROUNDS IDENTIFIED IN PLACE

Pilot evidence predates the levelled harness. Any released pilot round is shown as-is and excluded from cross-vendor conclusions; protocol v1 is the canonical series.

THE BLIND TEST

Can you tell who built it?

This is one artifact from a published vbench round. Six model configurations received the same brief, one reply each. Guess the maker, then see the round and recorded result.

UNLABELED SPECIMEN

An artifact built by one of six configurations, unlabeled

GUESS THE MAKER

ONE REPLY EACH · NO RETRIES · THE ANSWER IS RECORDED

HOW VBENCH WORKS

A narrow test with a public record.

SAME BRIEF

Six frontier model configurations receive one identical build brief, sent verbatim. Every recovered artifact is one self-contained HTML file a model wrote, run without repair in a network-blocked sandbox.

ONE REPLY

Within each released round, every configuration gets one prompt and one reply, with no retries or best-of selection. Each is reached through its command-line client with customization stripped, and every run records the exact command, wall time, and outcome.

FAILURES REMAIN VISIBLE

A run that produced nothing stays visible if its round is released, with the reason attributed. Recorded failures can come from the harness rather than the model, and the page says which is which.

SCOPE · IT TESTS A NARROW BUILD WORKFLOW, NOT GENERAL INTELLIGENCE. NO WINNER IS DECLARED.

READ THE FULL METHODOLOGY

OTHER SYSTEMS FROM THE STUDIO