korrents

On the map

Tap a claim on the ring to put it at the centre.

← Practices that make an LLM's work easier to evaluate also tend to…

17 connected korrents · 16 moments from 18 Mar 2024 to 27 Sept 2026. Nearly all of them are about Anthropic.

Everything filed under LLMs LLMs Everything filed under AI writing AI writing Everything filed under OpenAI OpenAI Everything filed under scaling laws scaling laws Everything filed under benchmarks benchmarks Everything filed under chat interfaces chat interfaces Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Practices that make an LLM's work easier to evaluate also tend to benefit the human developers working alongside it. Practices that make an LLM's work easierto evaluate also tend to benefit the humandevelopers working alongside it. Last stated 8 months ago 27 Jan 2026 SK Steve Klabnik — holds since 2026-01-27 — tap for who they are Same subject: A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. — tap to centre the map on it A language model asked to summarizeits own system prompt risks thatprompt's content biasing the summaryit produces. Last stated 4 weeks ago 2 Sept 2026 SW Simon Willison — holds since 2026-09-02 — tap for who they are Same subject: Being rude to a language model makes it perform worse. — tap to centre the map on it Being rude to a language model makesit perform worse. Last stated 8 months ago 22 Jan 2026 SK Steve Klabnik — holds since 2026-01-22 — tap for who they are Same subject: Build for the model six months from now, not the model of today. — tap to centre the map on it Build for the model six months fromnow, not the model of today. Last stated 7 months ago 25 Feb 2026 BC Boris Cherny — holds since 2026-02-25 — tap for who they are Same subject: Studies showing a language model behaving badly prove nothing, because the model was asked to produce that behaviour. — tap to centre the map on it Studies showing a language modelbehaving badly prove nothing,because the model was asked toproduce that behaviour. Last stated a year ago 2 Sept 2025 BE Benedict Evans — holds since 2025-09-02 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 4 weeks ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A story's quality depends on whether it engages and satisfies the reader, not on the tool used to write it. — tap to centre the map on it A story's quality depends on whetherit engages and satisfies the reader,not on the tool used to write it. Last stated a month ago 30 Aug 2026 EP Eliot Peper — holds since 2026-08-30 — tap for who they are Same subject: A tool for browsing personal notes and data works better with a conversational interface than a traditional file browser. — tap to centre the map on it A tool for browsing personal notesand data works better with aconversational interface than atraditional file browser. Last stated 7 months ago 11 Mar 2026 HR Harper Reed — holds since 2026-03-11 — tap for who they are Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it A trust that owns the missionprotects a company better thanfounder control does, which is whyAnthropic needs no dual-classshares. Last stated 5 months ago 10 May 2026 ER Eric Ries — holds since 2026-05-10 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 3 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it A technique that lets an AI model'sreasoning shift outside its visibleChain of Thought is dangerous, bothbecause it works and because aleading lab is willing to deploy it. Last stated 4 weeks ago 3 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 3 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: "Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next. — tap to centre the map on it "Scaling" was powerful because itwas one word: naming a researchdirection is what tells a wholefield what to do next. Last stated 10 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 7 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: A CEO who sends out an AI-written strategy memo is modelling that it is fine to outsource thinking and strategy. — tap to centre the map on it A CEO who sends out an AI-writtenstrategy memo is modelling that itis fine to outsource thinking andstrategy. Last stated 5 days ago 27 Sept 2026 MG Molly Graham — holds since 2026-09-27 — tap for who they are Same subject: A stream of AI-written papers with any error rate at all becomes insufferable, because finding the error costs more than the paper is worth even at ninety-nine percent. — tap to centre the map on it A stream of AI-written papers withany error rate at all becomesinsufferable, because finding theerror costs more than the paper isworth even at ninety-nine percent. Last stated 3 months ago 30 Jun 2026 GS Grant Sanderson — holds since 2026-06-30 — tap for who they are Same subject: AI can now produce in minutes work that used to take weeks to build. — tap to centre the map on it AI can now produce in minutes workthat used to take weeks to build. Last stated 4 months ago 26 May 2026 CB Carlos Alexandro Becker — holds since 2026-05-26 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre Practices that make an LLM's work easier to evaluate also tend to benefit the human developers working alongside it. Last stated 27 Jan 2026 · 8 months ago Holds Steve Klabnik Read this korrent →