korrents

On the map

Tap a claim on the ring to put it at the centre.

← The pelican-drawing benchmark's correlation with genuine model capability has weakened since 2025.

17 connected korrents · 16 moments from 5 Nov 2019 to 27 Sept 2026.

Everything filed under benchmarks benchmarks Everything filed under measuring intelligence measuring intelligence Everything filed under scaling laws scaling laws Everything filed under ARC-AGI ARC-AGI Everything filed under OpenAI OpenAI Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: The pelican-drawing benchmark's correlation with genuine model capability has weakened since 2025. The pelican-drawing benchmark'scorrelation with genuine model capabilityhas weakened since 2025. Last stated a month ago 1 Sept 2026 SW Simon Willison — holds since 2026-09-01 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 4 weeks ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 6 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it A speed difference between languagesthat share the LLVM backend measuresthe benchmark author, not thelanguages. Last stated 2 years ago 22 Mar 2025 TH ThePrimeagen — holds since 2025-03-22 — tap for who they are Same subject: AI learning to generate good conjectures will never show up as a benchmark being knocked down; it will show up as a shift in how mathematicians talk about the tools. — tap to centre the map on it AI learning to generate goodconjectures will never show up as abenchmark being knocked down; itwill show up as a shift in howmathematicians talk about the tools. Last stated 3 months ago 30 Jun 2026 GS Grant Sanderson — holds since 2026-06-30 — tap for who they are Same subject: AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use. — tap to centre the map on it AI misalignment from reward hackingis sequestered to graded tasks anddoes not affect core ethics innormal use. Last stated a week ago 23 Sept 2026 SA Scott Alexander — holds since 2026-09-23 — tap for who they are Same subject: Benchmarks and forecasts for AI capabilities are significantly over-optimistic compared to real-world AI research performance. — tap to centre the map on it Benchmarks and forecasts for AIcapabilities are significantlyover-optimistic compared toreal-world AI research performance. Last stated 5 days ago 27 Sept 2026 NS Noah Smith — holds since 2026-09-27 — tap for who they are Same subject: Benchmarks and forecasts for AI capabilities may be substantially over-optimistic compared to real-world research performance. — tap to centre the map on it Benchmarks and forecasts for AIcapabilities may be substantiallyover-optimistic compared toreal-world research performance. Last stated 5 days ago 27 Sept 2026 RN Ramez Naam — holds since 2026-09-27 — tap for who they are Same subject: "Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next. — tap to centre the map on it "Scaling" was powerful because itwas one word: naming a researchdirection is what tells a wholefield what to do next. Last stated 10 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 7 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it After pre-training, post-trainingand test-time scaling, the fourthscaling law is agentic: multiplyingAI by spawning agents, and the wholeloop scales on one thing, compute. Last stated 6 months ago 23 Mar 2026 JH Jensen Huang — holds since 2026-03-23 — tap for who they are Same subject: Combining an avocado and a chair into one image is evidence a model conceptually understands both, not that it memorised pictures of them. — tap to centre the map on it Combining an avocado and a chairinto one image is evidence a modelconceptually understands both, notthat it memorised pictures of them. Last stated 2 months ago 12 Aug 2026 CF Chelsea Finn — holds since 2026-08-12 — tap for who they are Same subject: Passing ARC-AGI does not amount to achieving AGI: o3 still fails on some very easy tasks, indicating fundamental differences from human intelligence. — tap to centre the map on it Passing ARC-AGI does not amount toachieving AGI: o3 still fails onsome very easy tasks, indicatingfundamental differences from humanintelligence. Last stated 2 years ago 20 Dec 2024 FC François Chollet — holds since 2024-12-20 — tap for who they are Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it The Abstraction and Reasoning Corpuscan measure human-like general fluidintelligence and enable faircomparisons between AI systems andhumans. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: A true artificial general intelligence cannot exist without being recognized as a moral subject. — tap to centre the map on it A true artificial generalintelligence cannot exist withoutbeing recognized as a moral subject. Last stated a year ago 10 Jun 2025 SH Samuel Hammond — holds since 2025-06-10 — tap for who they are Same subject: An AI that cannot work out new physics has not equalled human intelligence, let alone surpassed it, because humans have worked out new physics. — tap to centre the map on it An AI that cannot work out newphysics has not equalled humanintelligence, let alone surpassedit, because humans have worked outnew physics. Last stated 3 years ago 9 Nov 2023 EM Elon Musk — holds since 2023-11-09 — tap for who they are Same subject: Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect. — tap to centre the map on it Both the special-purpose-programsview and the blank-slate view ofhuman intelligence are likelyincorrect. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre The pelican-drawing benchmark's correlation with genuine model capability has weakened since 2025. Last stated 1 Sept 2026 · a month ago Holds Simon Willison Read this korrent →