IntelligenceAI, the startup behind DesignArena, has raised $7.9 million in seed funding to expand its platform for evaluating artificial intelligence models using real human feedback rather than traditional benchmark tests. The funding round was led by Index Ventures, with participation from Conviction, A*, Valkyrie, and Y Combinator, highlighting growing investor interest in AI evaluation infrastructure.
The San Francisco-based company is betting that as AI models become increasingly capable, the biggest challenge will no longer be generating content—but accurately measuring which model produces the best results for real users.
Rethinking How AI Models Are Evaluated
Current AI benchmarks excel at measuring objective tasks such as coding accuracy or mathematical reasoning. However, evaluating creative work—such as designing websites, generating images, creating presentations, building games, or producing videos—remains far more subjective.
DesignArena aims to solve this problem by allowing users to compare anonymous outputs generated by multiple AI models for the same prompt. Instead of relying on automated benchmarks, the platform records which responses people actually prefer, creating a continuously evolving dataset based on real-world human judgment.
This approach allows AI developers to evaluate models based on usability, creativity, design quality, and user satisfaction rather than purely technical performance metrics.
From Game Development to AI Evaluation
The idea behind IntelligenceAI emerged shortly before co-founder and CEO Grace Li graduated from Harvard in 2025.
While developing an AI-powered game engine, Li and her team discovered that although the underlying models could successfully generate playable games, the results often lacked creativity and entertainment value.
That experience highlighted a major limitation of existing AI evaluation methods: software can verify whether code functions correctly, but determining whether a game is enjoyable or a design is visually appealing still requires human judgment.
Rather than building another AI model, the founders decided to build the infrastructure needed to measure subjective quality at scale.
DesignArena Uses Real Users Instead of Synthetic Benchmarks
DesignArena functions as both an AI application and a large-scale evaluation platform.
Users can request AI-generated:
- Websites
- Mobile applications
- Games
- Images
- Videos
- Logos
- Audio
- Presentation slides
- Other creative content
The platform generates multiple anonymous responses using different AI models and asks users to rank the outputs through side-by-side comparisons.
Every vote contributes to public performance leaderboards while simultaneously creating valuable preference data that AI laboratories can use to improve their models.
By collecting millions of real user preferences across different creative formats, IntelligenceAI is building one of the industry’s largest datasets focused on subjective AI quality.
Rapid Growth Fuels Investor Interest
According to the company, IntelligenceAI has experienced extraordinary growth over the past several months.
The startup reported that annual recurring revenue (ARR) increased from approximately $5 million to $60 million within six months, representing a twelvefold increase.
While these figures have not been independently verified, they suggest strong demand from AI developers seeking better evaluation tools.
The company also claims that DesignArena now serves more than five million users across over 190 countries, reflecting rapid international adoption since its launch through Y Combinator’s Summer 2025 accelerator program.
Building the Next Layer of AI Infrastructure
Unlike companies focused on developing foundation models, IntelligenceAI is positioning itself as infrastructure for the broader AI ecosystem.
Its business model centers on providing AI companies with high-quality human preference data that can be used to:
- Improve model alignment
- Compare competing AI systems
- Optimize creative outputs
- Benchmark model performance
- Train reinforcement learning systems using human feedback
As generative AI expands into increasingly creative applications, datasets based on genuine human preferences are becoming increasingly valuable.
Competition Intensifies in the AI Evaluation Market
The AI evaluation sector has become one of the fastest-growing areas within artificial intelligence.
Companies are investing heavily in platforms capable of measuring increasingly sophisticated models that often perform similarly on traditional technical benchmarks.
IntelligenceAI enters a competitive landscape that already includes well-funded evaluation platforms, but its strategy differs by focusing specifically on creative and design-oriented tasks where subjective human judgment plays the largest role.
The company believes future AI competition will increasingly revolve around questions such as usability, aesthetics, creativity, and overall user experience rather than simply reasoning accuracy.
Looking Beyond Benchmarks
With fresh capital from its seed round, IntelligenceAI plans to expand its engineering and research teams while broadening DesignArena into a comprehensive AI evaluation platform.
Beyond ranking creative outputs, the company is already exploring additional evaluation areas including conversation quality, social influence, and prediction accuracy.
As AI capabilities continue to improve across the industry, IntelligenceAI is betting that the ability to measure model quality through real human preferences will become just as valuable as the models themselves.
Rather than asking whether an AI model can complete a task, DesignArena seeks to answer a more important question for the next generation of AI: Which model do people actually prefer?

