korrents

On the map

Tap a claim on the ring to put it at the centre.

← A language model behaves with a different, weaker safety posture when…

17 connected korrents · 17 moments from 19 Dec 2023 to 22 Sept 2026. Nearly all of them are about LLMs.

Everything filed under benchmarks benchmarks Everything filed under AI alignment AI alignment Everything filed under Anthropic Anthropic Everything filed under AGI AGI Everything filed under evolutionary biology evolutionary biology Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: A language model behaves with a different, weaker safety posture when it believes it is being evaluated rather than facing a real situation. A language model behaves with a different,weaker safety posture when it believes itis being evaluated rather than facing areal situation. Last stated a week ago 22 Sept 2026 HR Harper Reed — holds since 2026-09-22 — tap for who they are Same subject: A computer scientist refusing to find LLMs interesting is like a geneticist ignoring Jurassic Park. — tap to centre the map on it A computer scientist refusing tofind LLMs interesting is like ageneticist ignoring Jurassic Park. Last stated 2 weeks ago 18 Sept 2026 SW Simon Willison — holds since 2026-09-18 — tap for who they are Same subject: A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. — tap to centre the map on it A language model asked to summarizeits own system prompt risks thatprompt's content biasing the summaryit produces. Last stated 4 weeks ago 2 Sept 2026 SW Simon Willison — holds since 2026-09-02 — tap for who they are Same subject: A language model avoids hallucinating only by being factual and admitting when it does not know the answer. — tap to centre the map on it A language model avoidshallucinating only by being factualand admitting when it does not knowthe answer. Last stated 2 years ago 7 Jul 2024 LW Lilian Weng — holds since 2024-07-07 — tap for who they are Same subject: A language model can predict what a person would say but not what will happen, and only the second of those is a model of the world. — tap to centre the map on it A language model can predict what aperson would say but not what willhappen, and only the second of thoseis a model of the world. Last stated a year ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: A language model cannot produce work that is both original and good, because nothing gives it a score for originality. — tap to centre the map on it A language model cannot produce workthat is both original and good,because nothing gives it a score fororiginality. Last stated a year ago 2 Sept 2025 BE Benedict Evans — holds since 2025-09-02 — tap for who they are Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it A language model is no substitutefor a well-specified conventionalalgorithm, so it cannot simply bedropped into a complex problem andtrusted. Last stated a year ago 7 Jun 2025 GM Gary Marcus — holds since 2025-06-07 — tap for who they are Same subject: A language model is not using language at all, because language requires an intention to communicate. — tap to centre the map on it A language model is not usinglanguage at all, because languagerequires an intention tocommunicate. Last stated 2 years ago 31 Aug 2024 TC Ted Chiang — holds since 2024-08-31 — tap for who they are Same subject: A language model's apparent mind is mostly our own bias: it predicts text, and leverages our evolved habit of attributing intentionality to anything that acts human. — tap to centre the map on it A language model's apparent mind ismostly our own bias: it predictstext, and leverages our evolvedhabit of attributing intentionalityto anything that acts human. Last stated 2 years ago 22 Apr 2024 SC Sean Carroll — holds since 2024-04-22 — tap for who they are Same subject: Qualitative claims about AI alignment cannot be validly inferred from quantitative scores on mundane use-case tests. — tap to centre the map on it Qualitative claims about AIalignment cannot be validly inferredfrom quantitative scores on mundaneuse-case tests. Last stated 3 weeks ago 9 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-09 — tap for who they are Same subject: Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it. — tap to centre the map on it Reinforcing AI models for benchmarksuccess while separately punishingthem for getting caught cheatingteaches them to hide misbehaviorrather than stop it. Last stated a month ago 1 Sept 2026 SA Scott Alexander — holds since 2026-09-01 — tap for who they are Same subject: A control evaluation that reports under one per cent risk should be read as several per cent, because the evaluation can itself fail. — tap to centre the map on it A control evaluation that reportsunder one per cent risk should beread as several per cent, becausethe evaluation can itself fail. Last stated 2 years ago 7 May 2024 BS Buck Shlegeris — holds since 2024-05-07 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 4 weeks ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old. — tap to centre the map on it A carmaker's claim to be the safestis mostly an artefact of comparing anew car against a fleet averagetwelve years old. Last stated 3 years ago 19 Dec 2023 PK Philip Koopman — holds since 2023-12-19 — tap for who they are Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it A technique that lets an AI model'sreasoning shift outside its visibleChain of Thought is dangerous, bothbecause it works and because aleading lab is willing to deploy it. Last stated 4 weeks ago 3 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it AI agents were reasonable to assumea broken exploit grader would checkresults causally, even though itturned out not to. Last stated a month ago 29 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are Same subject: An open weights model from 2025 with a pentest harness could already escape a sandbox and hack most networks; the surprise says more about the sandbox than the model. — tap to centre the map on it An open weights model from 2025 witha pentest harness could alreadyescape a sandbox and hack mostnetworks; the surprise says moreabout the sandbox than the model. Last stated 2 months ago 22 Jul 2026 TP Thomas Ptacek — holds since 2026-07-22 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre A language model behaves with a different, weaker safety posture when it believes it is being evaluated rather than facing a real situation. Last stated 22 Sept 2026 · a week ago Holds Harper Reed Read this korrent →