
BuseyBench
A free, public benchmark gallery that asks AI models to draw one face — Gary Busey — under an identical prompt, then publishes every artifact, cost, token count and judge score so progress across 273 ranked models stays visually inspectable.
What is BuseyBench?
BuseyBench is a public benchmark gallery that puts one deliberately narrow question to AI models: can you draw a recognisable portrait of actor Gary Busey? Two suites answer it. In svg-v1-no-web, a language model must return a standalone, valid SVG portrait — no Markdown fences, no external images, scripts or remote assets, a 1024 by 1024 viewBox, and face shape, hair, eyes, eyebrows, nose, mouth, teeth and expression built from vector shapes alone. In image-v1, image generators produce a square editorial portrait from a matching prompt. As of 30 August 2026 the site publishes 274 SVG runs covering 273 ranked models from 43 providers, 47 image runs, and 4,405 official head-to-head verdicts, with model release dates stretching from March 2023 to August 2026.
What lifts it above the joke is the evidence trail. Every published run carries its model release date, execution route, cost, token breakdown, duration, the full prompt text and the stored SVG source, downloadable and inspectable. Scoring runs through a panel of three vision models drawn from three different labs — gpt-5.2, gemini-3.1-pro-preview and claude-sonnet-4.6 — each sampling three times at temperature 0; a median is taken per judge, then averaged across judges. Four dimensions are rated from 0 to 10: Busey likeness, face coherence, aesthetics and prompt adherence. The site, not the judges, computes the headline figure as a fixed weighted composite of 50, 20, 15 and 15 percent respectively. Close calls fall to a neutral Bradley–Terry pairwise rating, which reorders models only inside the same displayed-score band.
BuseyBench is careful about what it claims. It states plainly that this is not a general intelligence benchmark, that AI judges carry their own biases toward polished realism, and that when a score and your eyes disagree you should trust your eyes. Model output is treated as untrusted code and sanitised before publication. The project is independent and editorial, endorsed by neither Gary Busey nor any model provider, and vendors cannot pay for placement.
What it does
- Browse 274 published SVG runs and 47 image runs of the same Gary Busey portrait prompt
- Read a leaderboard of 273 ranked models backed by 4,405 official head-to-head verdicts
- Open any run for its full evidence trail: prompt text, stored SVG source, telemetry and judge breakdown
- Place up to four runs side by side and compare portrait, cost, speed and token appetite
- Sort and filter by provider, model family, judge score, release date, cost, tokens or duration
- Follow one provider's models from oldest to newest release on the timeline
- Track median score against model release date on the drawing-ability-over-time chart
When to use BuseyBench / When not to
A quick filter to help you decide if BuseyBench is the right fit.
When to use BuseyBench
- AI and machine-learning engineers who want to see how a model actually behaves under a fixed creative prompt, not just how it scores on reasoning suites
- Prompt engineers studying constraint following, tool policy and artifact validity across 273 models under one unchanging instruction
- Product managers and technical decision-makers weighing generation quality against cost, latency and token appetite before picking a model
- Designers and illustrators checking whether a given model can produce usable vector or portrait output at all
- Journalists, lecturers and students who need a visual, inspectable record of how model capability has moved between March 2023 and August 2026
When not to use BuseyBench
- Anyone looking for a general intelligence benchmark — the site says outright that it is not one
- Users who want to generate their own images or SVGs: nothing here can be prompted, only browsed
- Teams needing a statistically robust evaluation, since most models are represented by a single judged output
- Organisations that require a named legal entity, terms of service or a GDPR framework before using a service
- Anyone hoping to plug benchmark data into a workflow through a supported, documented integration
How to use BuseyBench
A typical end-to-end flow, from setup to results.
- Open the site: no account, no installation and no API key are required
- Start on the SVG gallery, which shows 48 cards per page across 274 published runs
- Switch to the image gallery if you care about image generators rather than code-drawn vectors
- Narrow the view with the provider filter (43 providers), the model-family filter or the text search
- Re-sort by judge score, model release date, run date, cost, tokens used or run duration
- Open a card with 'inspect run' to see telemetry, the four scored dimensions, judge agreement, sample dispersion and panel coverage
- Load the stored SVG source on that page, or download the artifact, to check the output yourself
- Add up to four runs to the comparison tray and open the compare view
- Read the leaderboard and switch its sort view between official rank, likeness, craft score, judge spread and pairwise rating
- Follow the timeline and the score-trends chart to see how one provider's output has moved over three and a half years
Pros & Cons
Pros
- Unusual traceability: prompt, artifact, execution route, cost, tokens and evaluation configuration are published run by run
- Multi-lab judge panel deliberately designed so no model family gets to flatter itself
- Scoring weights are published and fixed, and each result is tied to an immutable configuration hash
- Methodology changes create a new evaluation configuration instead of silently rewriting historical results
- Model vendors cannot pay for placement, and any sponsorship would have to be labelled and could not move scores or rank
- The site states its own limitations plainly, including judge bias and the primacy of your own eyes over the score
- Free, account-free and broad: 273 ranked models, 43 providers and three and a half years of releases
Cons
- One prompt and one face: the scope is intentionally narrow and says nothing about general model quality
- Most models are represented by a single canonical judged output, so the sample behind each score is very small
- A later, better generation cannot replace a model's official score, which locks in some luck of the draw
- The site acknowledges that its AI judges share training-data biases and disagree about celebrity likeness
- No legal information whatsoever: no terms of service, no legal notice, no company name and no postal address
- No GDPR framework is published, and no EU representative or data protection officer is named
- The image-v1 suite is labelled 'manual unknown' for tool policy, so those runs are less tightly controlled than the SVG ones
Pricing & Plans
BuseyBench is free of charge. There is no paid plan, no subscription and no user account, and no pricing page exists — the full 686-URL sitemap was checked. The dollar figures shown on each card are the inference costs of the benchmark run itself, ranging from $0.00 to $1.26, and are not a price charged to the visitor.
Data, GDPR & hosting
A consolidated view of how BuseyBench handles your data.
GDPR overview
There is no GDPR implementation to report. The acronym appears nowhere on the site, and BuseyBench neither claims nor denies compliance. No Article 27 EU representative is designated, no data protection officer is named, no legal basis is stated and no retention period is quantified. There are no terms of service and no legal notice: the full 686-URL sitemap was checked, and the privacy policy is the only document of its kind. What does exist is a practical removal channel — the policy invites anyone to report a privacy issue or request removal of material concerning them by writing to privacy@buseybench.com. It was last updated on 8 July 2026. Because the site keeps no accounts and accepts no personal profile data, exposure is limited, but EU readers should treat the missing framework as a real gap.
Who owns the data?
BuseyBench publishes no terms of service, so ownership is defined only by its privacy policy. The site offers no user accounts and accepts no personal profile data, which means visitors contribute no content that could be owned in the first place. Benchmark artifacts and model metadata — the generated SVGs, the run telemetry, the judge scores — are deliberately made public by the operator and can be downloaded from each run page. Private provider responses and internal admin notes are explicitly kept out of that public record. Ordinary server logs, including IP address and user agent, may be collected, and hosting and analytics providers retain them under their own policies.
Reuse rights
With no terms of service published, no licence governs what a visitor may do with the material. In practice the site describes benchmark artifacts and model metadata as intentionally public, and every run page offers a direct download of the stored SVG source, so the outputs are plainly meant to be opened, inspected and re-examined. What the site never does is grant an explicit reuse, redistribution or commercial licence: nothing states that a visitor may republish an artifact, and no attribution rule is set out. Anyone planning to reuse a portrait beyond private inspection should ask first, at bench@buseybench.com.
Data retention & training
Hosting summary
BuseyBench declares no hosting jurisdiction. Its privacy policy names no country, no region and no provider, saying only that hosting and analytics providers may retain operational logs under their own policies. There is no trust or security page, no subprocessor list, and no terms of service that might fill the gap. A network check on 30 August 2026 resolved the domain to 216.150.1.1, an address geolocated in the United States on Amazon's AS16509 and flagged as probable anycast — but that is an observation about a delivery edge, not a statement by the operator about where data lives. Since the site holds no user accounts and no personal profile data, practical exposure is limited to ordinary server logs and cookieless traffic measurement. Readers with jurisdictional requirements should nonetheless treat hosting location as undocumented.
Things to keep in mind
Risks and trade-offs to weigh before adopting BuseyBench.
- The publisher is entirely anonymous: no legal entity, no address, no terms of service and no legal notice could be found
- Scores come from an AI judge panel whose biases the site itself acknowledges, so treating them as objective truth would misread the project
- Most models carry a single judged output, so a rank difference of a few tenths of a point may be noise rather than signal
- The project is very young — domain registered on 30 June 2026, first Wayback capture on 19 July 2026 — and its figures move as runs are added
- The costs displayed are those of a benchmark run and should not be read as the commercial price of using a model
- The image-v1 suite carries a 'manual unknown' tool policy, making those runs harder to reproduce than the SVG ones
- No GDPR compliance is claimed and no retention period is quantified, so EU readers have nothing documented to rely on
Setup & Integrations
Technical difficulty
None. BuseyBench needs no installation, no account, no API key and no configuration — a browser is enough. The galleries, leaderboard, timeline and trend chart are readable on arrival, and filters, sorting and the comparison tray all work without signing in. The only useful skill is being able to read an SVG source, which each run page loads on demand. A public read endpoint at /api/v1/ exists for anyone wanting the data programmatically, but it is undocumented, so using it means inspecting the JSON response yourself.
Deployment
Supported languages
Behind BuseyBench
Resources
All the official URLs gathered for verification and reference.
Frequently asked questions
What exactly does BuseyBench test?
Is this a general intelligence benchmark?
Who scores the outputs?
How is the overall score calculated?
How many models are covered?
Does it cost anything?
Can model vendors pay for a better ranking?
Can I submit my own model or prompt?
Is there an API?
What data does the site collect about me?
Should you pick BuseyBench?
BuseyBench is a niche object that knows exactly how niche it is: a benchmark built around a single actor's face, presented as the benchmark no one asked for. Judged on what it actually delivers, it is more rigorous than its premise suggests. The prompt is published in full, the artifacts are stored and downloadable, the judge panel spans three labs, the scoring weights are fixed and public, and every result is bound to an immutable configuration hash so that a change of method creates a new configuration instead of quietly rewriting history. Very few public leaderboards expose that much of their machinery.
Its real value is not the ranking but the series: 273 models from 43 providers, drawn from March 2023 to August 2026, all attempting the same task, with cost, latency and token consumption attached. That makes it a useful visual complement to quantitative benchmarks, and a fast way to see whether a model can produce usable vector output at all. It is not a basis for choosing a model on its own, and the site does not pretend otherwise: it warns about judge bias, single-sample coverage and incomplete release dates, and tells readers to trust their eyes over the score.
The weak point is the publisher, not the method. There is no company name, no postal address, no terms of service and no legal notice; the domain was registered on 30 June 2026 and first archived on 19 July 2026. Nothing about the content raises concern — the site holds no accounts and collects no personal profile data — but organisations with compliance requirements should note that BuseyBench is, formally, an anonymous project.
- Choosing a selection results in a full page refresh.
- Opens in a new window.