Microsoft's new AI model beats Mythos on security benchmark ...
Relay-Bench, a new AI benchmark posted to arXiv in July 2026, chains problems across seven reasoning domains in a single ...
BIG-bench, the collaborative benchmark suite built by hundreds of researchers, contains a tripwire: a unique “canary” string ...
You're currently following this author! Want to unfollow? Unsubscribe via the link in your email. It's hard to pick the best AI to help you in work and life. What about GPT-4o, 4.5, 4.1, o1, o1-pro, ...
Moonshot AI’s Kimi K3 topped a frontend coding benchmark, beating Claude Fable 5 while adding pressure on US AI leaders.
Want smarter insights in your inbox? Sign up for our weekly newsletters to get only what matters to enterprise AI, data, and security leaders. Subscribe Now A team of Abacus.AI, New York University, ...
An open-source large language model developed in Beijing just shot to the top of the leaderboard for front-end coding tasks.
SAN FRANCISCO--(BUSINESS WIRE)--MLCommons ® and the Autonomous Vehicle Computing Consortium (AVCC) have achieved the first step toward a comprehensive MLPerf ® Automotive Benchmark Suite for AI ...
Researchers studying the emotional impact of tools like ChatGPT propose a new kind of benchmark that measures a model’s emotional and social impact. Researchers at MIT have proposed a new kind of AI ...
Benchmarks can be used to put large language models to the test. Read on for some tips on how to do it right. Today, there is hardly any way around AI. But how do companies decide which large language ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results