korrents

On the map

Tap a claim on the ring to put it at the centre.

← Language models perform better on benchmark problems released before…

17 connected korrents · 17 moments on record from 11 Oct 2018 to 3 Sept 2026.

Everything filed under benchmarks benchmarks Everything filed under scaling laws scaling laws Everything filed under Anthropic Anthropic Everything filed under measuring intelligence measuring intelligence Everything filed under reinforcement learning reinforcement learning Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores. Language models perform better onbenchmark problems released before theirtraining data cutoff, indicating datacontamination inflates scores. Last stated 2 years ago 13 May 2024 SR Sebastian Ruder — holds since 2024-05-13 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it Model comparisons understate realprogress, because benchmark tablesdo not control for how muchtest-time compute each answer used. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: The right bar for an LLM evaluator is human-level performance, not perfect accuracy. — tap to centre the map on it The right bar for an LLM evaluatoris human-level performance, notperfect accuracy. Last stated 10 months ago 23 Nov 2025 EY Eugene Yan — holds since 2025-11-23 — tap for who they are Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it Current techniques restrictpre-trained representation powerbecause standard language models areunidirectional. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: Large language models could still plateau, and that possibility should be held open even though no evidence of it has appeared. — tap to centre the map on it Large language models could stillplateau, and that possibility shouldbe held open even though no evidenceof it has appeared. Last stated 4 weeks ago 26 Aug 2026 DH David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are Same subject: The LLM line of research will reach a capability plateau. — tap to centre the map on it The LLM line of research will reacha capability plateau. Last stated a month ago 7 Aug 2026 FC François Chollet — no longer holds since 2026-08-07 — tap for who they are GM Gary Marcus — holds since 2025-06-07 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 3 weeks ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 6 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A true artificial general intelligence cannot exist without being recognized as a moral subject. — tap to centre the map on it A true artificial generalintelligence cannot exist withoutbeing recognized as a moral subject. Last stated a year ago 10 Jun 2025 SH Samuel Hammond — holds since 2025-06-10 — tap for who they are Same subject: Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect. — tap to centre the map on it Both the special-purpose-programsview and the blank-slate view ofhuman intelligence are likelyincorrect. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: Every generation's test for machine intelligence is whatever machines cannot do yet, so in twenty years the question will be whether it can reproduce. — tap to centre the map on it Every generation's test for machineintelligence is whatever machinescannot do yet, so in twenty yearsthe question will be whether it canreproduce. Last stated 3 years ago 29 Jun 2023 GH George Hotz — holds since 2023-06-29 — tap for who they are Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it A physical AI company needs threeAIs, not one — the agent, thesimulator and the critic — turningdeployment into a flywheel. Last stated 2 months ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it A verifiable task can be optimisedby reinforcement learning until aneural network performs it extremelywell. Last stated 10 months ago 17 Nov 2025 AK Andrej Karpathy — holds since 2025-11-17 — tap for who they are Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it Building a realistic simulator isexactly as hard as building theagent, because the simulator isitself a large AI model. Last stated 2 months ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it A trust that owns the missionprotects a company better thanfounder control does, which is whyAnthropic needs no dual-classshares. Last stated 4 months ago 10 May 2026 ER Eric Ries — holds since 2026-05-10 — tap for who they are Same subject: Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million. — tap to centre the map on it Agents can build about half amillion lines before the codebasedissolves into a mess, and the nextmodel will push that to a fewmillion. Last stated 6 months ago 11 Mar 2026 SY Steve Yegge — holds since 2026-03-11 — tap for who they are Same subject: AI agents and AI coding will run on servers and from the cloud first, not on your laptop. — tap to centre the map on it AI agents and AI coding will run onservers and from the cloud first,not on your laptop. Last stated 3 months ago 28 Jun 2026 PL Pieter Levels — holds since 2026-06-28 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they arefaded, dashed ring: they no longer hold it — they changed their mind

At the centre Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores. Last stated 13 May 2024 · 2 years ago Holds Sebastian Ruder Read this korrent →