korrents

On the map

Tap a claim on the ring to put it at the centre.

← Judging one model worse than another from a one-shot, probabilistic…

17 connected korrents · 14 moments from 11 Jun 2021 to 27 Sept 2026.

Everything filed under LLMs LLMs Everything filed under benchmarks benchmarks Everything filed under AI writing AI writing Everything filed under open source open source Everything filed under OpenAI OpenAI Everything filed under Anthropic Anthropic Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Judging one model worse than another from a one-shot, probabilistic output on a given day is not a good way to evaluate it. Judging one model worse than another froma one-shot, probabilistic output on agiven day is not a good way to evaluateit. Last stated 5 months ago 19 May 2026 TG Tommy Geoco — holds since 2026-05-19 — tap for who they are Same subject: Trusting model-written code you have not read became defensible only with Opus 4.5, and only for classes of problem you have already watched a model handle. — tap to centre the map on it Trusting model-written code you havenot read became defensible only withOpus 4.5, and only for classes ofproblem you have already watched amodel handle. Last stated 7 months ago 19 Mar 2026 SW Simon Willison — holds since 2026-03-19 — tap for who they are Same subject: Running a model five times and keeping the best answer buys a higher benchmark score without buying a better model. — tap to centre the map on it Running a model five times andkeeping the best answer buys ahigher benchmark score withoutbuying a better model. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: Language models perform better on benchmark problems released before their training data cutoff, indicating data contamination inflates scores. — tap to centre the map on it Language models perform better onbenchmark problems released beforetheir training data cutoff,indicating data contaminationinflates scores. Last stated 2 years ago 13 May 2024 SR Sebastian Ruder — holds since 2024-05-13 — tap for who they are Same subject: Model comparisons understate real progress, because benchmark tables do not control for how much test-time compute each answer used. — tap to centre the map on it Model comparisons understate realprogress, because benchmark tablesdo not control for how muchtest-time compute each answer used. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: Claude Opus 5.5 is a Tier 2 cyber model, despite Anthropic placing it in the lower category. — tap to centre the map on it Claude Opus 5.5 is a Tier 2 cybermodel, despite Anthropic placing itin the lower category. Last stated 2 weeks ago 23 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-23 — tap for who they are Same subject: The release of Opus 4.5 in November 2025 was the dividing line, because it was the first time an agent produced work uncannily close to what its user would have written themselves. — tap to centre the map on it The release of Opus 4.5 in November2025 was the dividing line, becauseit was the first time an agentproduced work uncannily close towhat its user would have writtenthemselves. Last stated a month ago 26 Aug 2026 DH David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are Same subject: Evaluating a model properly would mean delaying its release, and competitive pressure means no lab will. — tap to centre the map on it Evaluating a model properly wouldmean delaying its release, andcompetitive pressure means no labwill. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: Criticism that a simpler, scaled-up architecture underperforms a heavily hand-optimized baseline actually demonstrates why that architecture-scaling research matters. — tap to centre the map on it Criticism that a simpler, scaled-uparchitecture underperforms a heavilyhand-optimized baseline actuallydemonstrates why thatarchitecture-scaling researchmatters. Last stated 5 years ago 11 Jun 2021 GW Gwern — holds since 2021-06-11 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 3 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 3 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: Agents are going to be incorporated into mainstream interfaces — tap to centre the map on it Agents are going to be incorporatedinto mainstream interfaces Last stated 6 months ago 25 Mar 2026 DH David Heinemeier Hansson — holds since 2026-03-25 — tap for who they are Same subject: A CEO who sends out an AI-written strategy memo is modelling that it is fine to outsource thinking and strategy. — tap to centre the map on it A CEO who sends out an AI-writtenstrategy memo is modelling that itis fine to outsource thinking andstrategy. Last stated a week ago 27 Sept 2026 MG Molly Graham — holds since 2026-09-27 — tap for who they are Same subject: A stream of AI-written papers with any error rate at all becomes insufferable, because finding the error costs more than the paper is worth even at ninety-nine percent. — tap to centre the map on it A stream of AI-written papers withany error rate at all becomesinsufferable, because finding theerror costs more than the paper isworth even at ninety-nine percent. Last stated 3 months ago 30 Jun 2026 GS Grant Sanderson — holds since 2026-06-30 — tap for who they are Same subject: AI can now produce in minutes work that used to take weeks to build. — tap to centre the map on it AI can now produce in minutes workthat used to take weeks to build. Last stated 4 months ago 26 May 2026 CB Carlos Alexandro Becker — holds since 2026-05-26 — tap for who they are Same subject: A business making hundreds of millions from an open source project owes it money or contribution, and can be refused its trademarks if it gives neither. — tap to centre the map on it A business making hundreds ofmillions from an open source projectowes it money or contribution, andcan be refused its trademarks if itgives neither. Last stated 2 years ago 26 Sept 2024 MM Matt Mullenweg — holds since 2024-09-26 — tap for who they are Same subject: A codec does not really exist until everyone can decode it. — tap to centre the map on it A codec does not really exist untileveryone can decode it. Last stated 4 months ago 31 May 2026 JK Jean-Baptiste Kempf — holds since 2026-05-31 — tap for who they are Same subject: A company that will not accept a project’s terms is free to go and use a more permissive project instead. — tap to centre the map on it A company that will not accept aproject’s terms is free to go anduse a more permissive projectinstead. Last stated 2 years ago 26 Sept 2024 MM Matt Mullenweg — holds since 2024-09-26 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre Judging one model worse than another from a one-shot, probabilistic output on a given day is not a good way to evaluate it. Last stated 19 May 2026 · 5 months ago Holds Tommy Geoco Read this korrent →