How to Benchmark Process

How to build a better AI benchmark

To fix the way we test and measure models, AI is learning tricks from social science. It’s not easy being one of Silicon Valley’s favorite benchmarks. SWE-Bench (pronounced “swee bench”) launched in ...

Morningstar

How to Benchmark Portfolios With Both Public and Private Equities

The challenge of valuing private companies isn’t stopping investment. To improve transparency, Morningstar and PitchBook created the Modern Market 100 Index. Why the Time is Right for a Private Market ...

University of Delaware

Creating Humanity’s Last Exam

The result is Humanity’s Last Exam (HLE). The dramatically titled test is 2,500 questions, crowdsourced from more than 1,000 professors, experts, researchers and graduate students at nearly 500 ...

Forbes

From Strategy To Success: How Process Drives B2B Growth

B2B organizations often find themselves at a crossroads between strategy formulation and strategy execution. The harsh reality is that many fail to transform their well-crafted strategies into ...

Harvard Business Review

How to Marry Process Management and AI

Make sure your people and your technology work well together. by Thomas H. Davenport and Thomas C. Redman When Mars Wrigley decided to digitize its supply chain, it invested in several AI and ...

MIT Technology Review

This benchmark used Reddit’s AITA to test how much AI models suck up to us

The new benchmark, called Elephant, makes it easier to spot when AI models are being overly sycophantic—but there’s no current fix. Back in April, OpenAI announced it was rolling back an update to its ...

Some results have been hidden because they may be inaccessible to you

Show inaccessible results