As co-founder Grace Li tells it, her company began just weeks before graduation in 2025, when a handful of college friends set out to make their AI game engine work. The models could generate functional games, but none were actually fun—which raised a compelling question: how do you determine if a game will be enjoyable?
They soon concluded that there was no substitute for human judgment, and began brainstorming ways to gather honest, scalable feedback from real users. The result became DesignArena, an AI tool now used by 5.3 million people globally. As it turned out, many AI companies were looking for exactly this kind of scalable user feedback—and plenty were willing to pay for it.
“It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.”
On Monday, the company behind DesignArena—dubbed Intelligence—announced a $7.9 million seed round led by Index Ventures, with participation from Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others.
How DesignArena Works
For non-enterprise users, DesignArena operates much like a sophisticated model router. There’s a ChatGPT-style window for prompts, plus separate dropdowns for websites, images, and a dozen other visual formats. Once you submit a request with your chosen format and style, you’re presented with a series of “A vs. B” comparisons, helping you rank the handful of outputs from best to worst.
While this is a useful service in itself, the platform’s real value lies on the enterprise side. Participating models can treat DesignArena as a constant source of instant feedback for their media-generating models. Users tend to be indifferent about which models they’re ranking—as Li puts it, they simply want the best output—so their rankings provide critical insight into what users truly desire.
Why Human Feedback Matters in 2026
In 2026, as AI models become increasingly sophisticated, automated benchmarks are struggling to capture the subjective quality that users care about—especially in design and creative tasks. This gap is driving a growing demand for human-led evaluation services. For frontier labs, that’s a service worth paying for, Li says. She adds that the site is currently generating $60 million in annual recurring revenue (ARR), solidifying its position as a key source of human-led evaluation data for the AI industry.
Critically, users must log in to receive their output, allowing Intelligence to track how preferences vary across continents and shift over time. (Li notes, for example, that web dashboards in Asia tend to favor a more maximalist design style.) These measures complement automated benchmarks, which can operate at greater scale but are often susceptible to gaming or manipulation—as demonstrated dramatically by the Hugging Face breach last week.
Not Without Challenges
Crowdsourced human feedback is not guaranteed to be a winning market. Less than a year after launching, Yupp shuttered its doors earlier this year despite raising $33 million from a16z crypto’s Chris Dixon. It had attracted some frontier models as customers and claimed over 1.3 million users, but still couldn’t build a sustainable long-term business.
Even so, other startups focused on human evaluation appear to be thriving. LM Arena, which takes a similar approach for language models, has seen steady growth and adoption. The success of these platforms may hinge on their ability to scale while maintaining high-quality feedback loops—something DesignArena seems well-positioned to do.
As the AI industry continues to grapple with the limitations of automated metrics, the demand for human-centered evaluation will likely grow. With its fresh funding and strong traction, Intelligence is betting that “taste” is the next frontier in AI development.
via TechCrunch AI
