korrents

← Anthropic now uses mechanistic interpretability techniques to audit…

On the map

8 connected korrents · 8 moments on record from 1 Jan 2026 to 3 Sept 2026.

Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Anthropic now usesmechanisticinterpretability techniques… DA Dario Amodei — holds since 2026-01-01 A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. A benchmark that ranksClaude Code last while… DR Dax Raad — holds since 2026-09-03 A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. A language model askedto summarize its own… SW Simon Willison — holds since 2026-09-02 AI agents and AI coding will run on servers and from the cloud first, not on your laptop. AI agents and AI codingwill run on servers and… PL Pieter Levels — holds since 2026-06-28 Build for the model six months from now, not the model of today. Build for the model sixmonths from now, not… BC Boris Cherny — holds since 2026-02-25 Claude Code is the coding agent to build with. Claude Code is thecoding agent to build… PL Pieter Levels — holds since 2026-06-28 PL Pieter Levels — turned since 2026-08-24 Claude Haiku hallucinates badly enough that its use inside Claude Code's WebFetch tool is a hallucination risk on every fetched URL. Claude Haikuhallucinates badly… SW Simon Willison — holds since 2026-08-10 Coding agents should read the shared AGENTS.md convention; insisting on a tool-specific CLAUDE.md creates split-brain problems across a team. Coding agents shouldread the shared… TL Tobias Lütke — holds since 2026-08-25 Even Anthropic does not prioritize AI safety to the extent that doing so would maximize its own medium-term (3-12 month) business interests. Even Anthropic does notprioritize AI safety to… ZM Zvi Mowshowitz — holds since 2026-09-02
same subject