korrents

On the map

Tap a claim on the ring to put it at the centre.

← Studies showing a language model behaving badly prove nothing, because the model was asked to produce that behaviour.

17 connected korrents · 14 moments on record from 5 Nov 2019 to 3 Sept 2026. Nearly all of them are about Anthropic.

Everything filed under LLMs LLMs Everything filed under Google Google Everything filed under measuring intelligence measuring intelligence Everything filed under OpenAI OpenAI Everything filed under benchmarks benchmarks Everything filed under coding agents coding agents Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Studies showing a language model behaving badly prove nothing, because the model was asked to produce that behaviour. Studies showing a language model behavingbadly prove nothing, because the model wasasked to produce that behaviour. Last stated a year ago 2 Sept 2025 BE Benedict Evans — holds since 2025-09-02 — tap for who they are Same subject: A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. — tap to centre the map on it A language model asked to summarizeits own system prompt risks thatprompt's content biasing the summaryit produces. Last stated a week ago 2 Sept 2026 SW Simon Willison — holds since 2026-09-02 — tap for who they are Same subject: Build for the model six months from now, not the model of today. — tap to centre the map on it Build for the model six months fromnow, not the model of today. Last stated 6 months ago 25 Feb 2026 BC Boris Cherny — holds since 2026-02-25 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 6 days ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it A trust that owns the missionprotects a company better thanfounder control does, which is whyAnthropic needs no dual-classshares. Last stated 4 months ago 10 May 2026 ER Eric Ries — holds since 2026-05-10 — tap for who they are Same subject: Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million. — tap to centre the map on it Agents can build about half amillion lines before the codebasedissolves into a mess, and the nextmodel will push that to a fewmillion. Last stated 6 months ago 11 Mar 2026 SY Steve Yegge — holds since 2026-03-11 — tap for who they are Same subject: AI agents and AI coding will run on servers and from the cloud first, not on your laptop. — tap to centre the map on it AI agents and AI coding will run onservers and from the cloud first,not on your laptop. Last stated 2 months ago 28 Jun 2026 PL Pieter Levels — holds since 2026-06-28 — tap for who they are Same subject: An AI should be built as a fiduciary, the equivalent of a lawyer for its user, rather than as an agent pursuing its own notion of the good. — tap to centre the map on it An AI should be built as afiduciary, the equivalent of alawyer for its user, rather than asan agent pursuing its own notion ofthe good. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: Anthropic is a unique instance and not a trend: without it there would be essentially no TPU or Trainium growth at all. — tap to centre the map on it Anthropic is a unique instance andnot a trend: without it there wouldbe essentially no TPU or Trainiumgrowth at all. Last stated 5 months ago 15 Apr 2026 JH Jensen Huang — holds since 2026-04-15 — tap for who they are Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it If intelligence is the process ofacquiring skills, no single taskdemonstrates intelligence unless itis a meta-task of skill-acquisitionacross many tasks. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it Solely measuring skill at any giventask falls short of measuringintelligence, because skill isheavily modulated by prior knowledgeand experience. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it The Abstraction and Reasoning Corpuscan measure human-like general fluidintelligence and enable faircomparisons between AI systems andhumans. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it A technique that lets an AI model'sreasoning shift outside its visibleChain of Thought is dangerous, bothbecause it works and because aleading lab is willing to deploy it. Last stated 6 days ago 3 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 5 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system. — tap to centre the map on it A crewed rocket cannot be made safeby making the booster reliable, sothe only real way to improve safetyis to carry an escape system. Last stated 3 years ago 14 Dec 2023 JB Jeff Bezos — holds since 2023-12-14 — tap for who they are Same subject: A monopolist that can no longer grow by winning new users can only grow by making its product worse for the users it already has. — tap to centre the map on it A monopolist that can no longer growby winning new users can only growby making its product worse for theusers it already has. Last stated 3 years ago 28 Jul 2023 CD Cory Doctorow — holds since 2023-07-28 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Studies showing a language model behaving badly prove nothing, because the model was asked to produce that behaviour. Last stated 2 Sept 2025 · a year ago Holds Benedict Evans Read this korrent →