Intelligence, the company behind AI evaluation site DesignArena, announced Monday it raised a $7.9 million seed round led by Index Ventures, with Conviction, A*, Valkyrie and others participating. In the same announcement, co-founder Grace Li told TechCrunch the site is currently generating $60 million in annual recurring revenue.
Read those two numbers in order. The revenue is roughly eight times the round.
DesignArena started as the byproduct of a failure. In 2025, Li and a handful of college friends were building an AI game engine. The models produced games that ran and weren’t fun. There is no benchmark for fun, so they built the crudest available substitute — show a person two outputs, ask which one is better. About a week after that pivot, Li says, they closed their first major deal with a frontier lab.
What they’re actually selling
For consumers, DesignArena works like a sophisticated model router: a prompt window, dropdowns for websites, images and a dozen other visual formats, then a run of “A vs. B” choices until the outputs are ranked best to worst. 5.3 million people worldwide use it. Crucially, they are indifferent to which model produced which output. They just want the best result they can get.
That indifference is the product. Blind preference at scale is the one signal a media-generation model cannot produce for itself, and the one an automated benchmark cannot manufacture. Because users have to log in to collect their output, Intelligence can also track how taste shifts by geography and over time — Li notes that web dashboards in Asia skew toward a more maximalist style.
Why the timing works
Automated evals are getting less trustworthy, not more. They run at enormous scale, but they can be gamed, and the Hugging Face breach last week was a loud reminder that the infrastructure around them is not bulletproof. At the same time, frontier models have converged on the public leaderboards. When four labs post near-identical scores, the tiebreaker moves to the thing no score captures.
Our take: The interesting number isn’t $7.9 million — it’s that a company claiming $60 million in ARR raised a seed round at all. That’s a strategic round, not a survival round. The wider signal matters more: the scarce input in AI has moved from compute, to data, to talent, and now to judgment. Taste is the only one on that list you cannot manufacture with capital. If your work has a subjective quality bar — design, copy, video, product — the model already handles volume. Somebody still has to decide what’s good. Make sure that somebody is you.
It is not a guaranteed market
Yupp raised $33 million from a16z crypto’s Chris Dixon, signed frontier labs as customers and claimed more than 1.3 million users. It shut down earlier this year, less than twelve months after launching. Crowdsourced human feedback has real gravity: it needs constant volume, and volume costs money that labs will only keep paying for as long as the signal stays scarce.
The counterexample is considerably larger. LM Arena, which does the same thing for text responses, raised $150 million in January at a $1.7 billion valuation — four months after formally launching its paid product.
What to watch
- Customer concentration. $60 million in ARR from three lab contracts is a very different business than $60 million from thirty. Nobody has disclosed which it is.
- In-housing. The labs can build their own preference panels. The reason they haven’t is cost and neutrality — watch whether that calculus holds as budgets tighten.
- Shelf life. Preference data ages. If media models converge on a house style, yesterday’s A/B votes stop teaching them anything new.
- Copycats. Arena-style evaluation is cheap to clone and expensive to scale. Expect a crowded field by Q4 and a thin one by next summer.
