Encyclox

Vals Revolutionizes AI Benchmarking

· curiosity

The Benchmarking Broom: How Vals Is Sweeping Up the Flaws in AI Evaluation

The tech industry’s Achilles’ heel has long been its tendency to fudge numbers. In artificial intelligence, this flaw is particularly egregious. Companies have relied on outdated benchmarking systems that provide a veneer of legitimacy for their models. But in AI, where failure can be catastrophic, such shortcuts won’t suffice.

Enter Vals, a startup founded two years ago by Rayan Krishnan to revolutionize AI evaluation. With Andreessen Horowitz leading its seed round and $40 million in series A funding, Vals is poised to become the gold standard for benchmarking.

Vals differs from its predecessors in one key way: it keeps its test materials secret. This may seem counterintuitive, but it’s a crucial step towards creating a robust evaluation system. By not disclosing its test materials, Vals prevents companies from training their models against them and essentially cheating on the exam.

A More Nuanced Approach

Traditional benchmarking has been woefully inadequate for modern AI. Focusing on abstract evaluations that measure general knowledge, companies have glossed over significant shortfalls in their models’ real-world performance. Vals takes a more nuanced approach, evaluating not just a model’s ability to pass tests but also its capacity to complete complex tasks within specific industries.

Krishnan envisions a system that can accurately predict the real-world impacts of AI models. “Can they do work that produces a product of the same quality as a human within every domain?” he asks. This kind of thinking sets Vals apart from its competitors.

The Dark Side of the Industry

Vals is also evaluating negative outcomes, not just positive ones. Krishnan’s team wants to know: “If these models ran wild in the world, what would be the negative implications?” This question gets to the heart of the industry’s biggest problem.

As AI becomes increasingly integrated into our lives, the stakes are rising. Companies need to demonstrate not just that their models can perform tasks but also that they can do so safely and responsibly. Vals’ evaluation system provides a much-needed check on the industry’s enthusiasm for AI.

Growth Pains

Vals is experiencing growth pains, with revenue eight times what it was last year and a team of 25 employees (up from just eight at the start of the year). But Krishnan remains optimistic about the company’s prospects. As the industry evolves, demand for robust evaluation systems like Vals’ will only continue to grow.

Krishnan sees his company’s system as a key factor in driving AI model usage within companies. “The types of benchmarks and evaluations that we do are going to drive their usage and be central to how these companies submit public filings or discuss prospective investments in AI,” he says.

A New Era for Benchmarking

As the industry hurtles towards a future where AI models are integrated into our daily lives, benchmarking will play a critical role. With Vals at the forefront of this movement, it’s clear that the days of creative accounting and patchwork evaluation systems are numbered.

The wider industry will be scrutinizing Vals’ system as companies begin to publicly file their AI-related investments and regulatory bodies take a closer look. Can the company maintain its commitment to transparency in the face of commercial pressure? Only time will tell, but one thing is certain: with Vals leading the charge, the future of AI benchmarking has never looked brighter.

Reader Views

  • TA
    The Archive Desk · editorial

    While Vals' innovative approach to AI benchmarking is certainly a step in the right direction, one cannot help but wonder about the potential for game-playing in this new system. With test materials kept secret, companies may simply develop models that excel at beating the benchmark without necessarily improving their real-world performance. If not addressed, this could create a paradox where Vals' own evaluation becomes the de facto goal, rather than an honest assessment of AI capabilities.

  • HV
    Henry V. · history buff

    While Vals' innovative approach to benchmarking AI systems is laudable, I worry that its emphasis on secrecy may stifle innovation and transparency within the industry. By not disclosing its test materials, companies are essentially forced to rely on Vals-approved methodologies, rather than being encouraged to push the boundaries of what's possible with their own unique approaches. This could lead to a bottleneck in progress, as the most effective solutions may require collaboration or experimentation outside of Vals' controlled environment.

  • IL
    Iris L. · curator

    Vals is a breath of fresh air in the AI evaluation space, but let's not forget that its closed-source approach raises questions about transparency and replicability. As researchers strive to understand how Vals' secret sauce works, they may inadvertently perpetuate a culture where proprietary methodologies are prioritized over open collaboration. It's crucial for Vals to strike a balance between protecting its intellectual property and contributing to the broader AI research community through transparent sharing of methods and results.

Related articles

More from Encyclox

View as Web Story →