Skip to content
Today

Topics

Research

Find published briefs and AI signals by topic without losing the daily editorial context.

Current read

Latest in Research

BenchMIRT: What are LLM benchmarks actually measuring? — huggingface_changelog introduced BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions. Across those benchmarks, keeping only 10% of the questions generally preserved nearly the same picture of which models were stronger or weaker o…

1 published update · 1 linked source

Topics

  1. AI decision guides: API costs, local AI and coding agents
  2. Models 2
  3. Products 2
  4. Business 3
  5. Research 1
  6. Policy 2

Research

  1. BenchMIRT: What are LLM benchmarks actually measuring?

    huggingface_changelog introduced BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions. Across those benchmarks, keeping only 10% of the questions generally preserved nearly the same picture of which models were stronger or weaker on the underlying safety or reasoning capability as using the full set. Existing tools already make it possible to trim evaluations in similar ways, and we think the added transparency into what benchmark questions are actually measuring is worth that risk—but it’s a real one.