korrents

On the map

Tap a claim on the ring to put it at the centre.

← Error discovery is the most important part of the AI evaluation process.

17 connected korrents · 14 moments from 9 Aug 2024 to 22 Sept 2026.

Everything filed under software quality software quality Everything filed under product management product management Everything filed under AI and science AI and science Everything filed under AI and jobs AI and jobs Everything filed under AI agents AI agents Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Error discovery is the most important part of the AI evaluation process. Error discovery is the most important partof the AI evaluation process. Last stated a week ago 22 Sept 2026 LR Lenny Rachitsky — holds since 2026-09-22 — tap for who they are Same subject: Automated evaluation tools are useful for finding some types of AI errors, especially when combined with human judgment. — tap to centre the map on it Automated evaluation tools areuseful for finding some types of AIerrors, especially when combinedwith human judgment. Last stated a week ago 22 Sept 2026 LR Lenny Rachitsky — holds since 2026-09-22 — tap for who they are Same subject: For a product person working with AI research, learning to write good evals is the most important thing. — tap to centre the map on it For a product person working with AIresearch, learning to write goodevals is the most important thing. Last stated 3 weeks ago 10 Sept 2026 TS Tara Seshan — holds since 2026-09-10 — tap for who they are Same subject: Evaluation and error analysis will remain necessary practices even if AI models become near-perfect. — tap to centre the map on it Evaluation and error analysis willremain necessary practices even ifAI models become near-perfect. Last stated 2 weeks ago 18 Sept 2026 HH Hamel Husain — holds since 2026-09-18 — tap for who they are Same subject: Generating a list of potential problems is worthless unless it is followed by validating and prioritizing them. — tap to centre the map on it Generating a list of potentialproblems is worthless unless it isfollowed by validating andprioritizing them. Last stated 2 years ago 9 Aug 2024 TS Tobias Sjösten — holds since 2024-08-09 — tap for who they are Same subject: AI errors are harder to catch in analysis of unfamiliar data than in simple writing tasks. — tap to centre the map on it AI errors are harder to catch inanalysis of unfamiliar data than insimple writing tasks. Last stated a year ago 5 Sept 2025 JE Jason Evanish — holds since 2025-09-05 — tap for who they are Same subject: Thanks to AI, the principles of the product model, product strategy and product discovery have never been more important. — tap to centre the map on it Thanks to AI, the principles of theproduct model, product strategy andproduct discovery have never beenmore important. Last stated 3 weeks ago 10 Sept 2026 MC Marty Cagan — holds since 2026-09-10 — tap for who they are Same subject: AI systems require human oversight rather than full autonomy in analysis and workflows. — tap to centre the map on it AI systems require human oversightrather than full autonomy inanalysis and workflows. Last stated a year ago 5 Sept 2025 JE Jason Evanish — holds since 2025-09-05 — tap for who they are Same subject: Today's AI tools either solve a problem or fail at it, and are really bad at partial progress or at identifying which intermediate step to attack first. — tap to centre the map on it Today's AI tools either solve aproblem or fail at it, and arereally bad at partial progress or atidentifying which intermediate stepto attack first. Last stated 6 months ago 20 Mar 2026 TT Terence Tao — holds since 2026-03-20 — tap for who they are Same subject: A coding agent left to run over a weekend returns work that is half good, a quarter garbage and a quarter in need of repair. — tap to centre the map on it A coding agent left to run over aweekend returns work that is halfgood, a quarter garbage and aquarter in need of repair. Last stated a year ago 18 Aug 2025 CB Clay Bavor — holds since 2025-08-18 — tap for who they are Same subject: A developer using AI coding tools remains responsible for reviewing the code they ship. — tap to centre the map on it A developer using AI coding toolsremains responsible for reviewingthe code they ship. Last stated 7 months ago 10 Mar 2026 JS Jon Seager — holds since 2026-03-10 — tap for who they are Same subject: A language-agnostic conformance suite is the most powerful thing you can hand a coding agent, because the whole instruction becomes: write code until these tests pass. — tap to centre the map on it A language-agnostic conformancesuite is the most powerful thing youcan hand a coding agent, because thewhole instruction becomes: writecode until these tests pass. Last stated 6 months ago 19 Mar 2026 SW Simon Willison — holds since 2026-03-19 — tap for who they are Same subject: AI-assisted coding will generate a lot of low-quality output that developers will need to clean up, even as it also produces good results. — tap to centre the map on it AI-assisted coding will generate alot of low-quality output thatdevelopers will need to clean up,even as it also produces goodresults. Last stated 8 months ago 9 Feb 2026 SH Scott Hanselman — holds since 2026-02-09 — tap for who they are Same subject: AI-generated code should be treated as untrusted until there is evidence it preserves the intended behavior. — tap to centre the map on it AI-generated code should be treatedas untrusted until there is evidenceit preserves the intended behavior. Last stated a month ago 19 Aug 2026 JS Jon Seager — holds since 2026-08-19 — tap for who they are Same subject: Code review as a review mechanism depends on reviewers being able to infer effort from reading code, and AI-agent-generated code erases that signal. — tap to centre the map on it Code review as a review mechanismdepends on reviewers being able toinfer effort from reading code, andAI-agent-generated code erases thatsignal. Last stated 5 months ago 7 May 2026 DC David Crawshaw — holds since 2026-05-07 — tap for who they are Same subject: A team that slows down and reads every pull request and every line of code should expect only a 30 to 50 percent productivity lift from AI. — tap to centre the map on it A team that slows down and readsevery pull request and every line ofcode should expect only a 30 to 50percent productivity lift from AI. Last stated 3 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: Agent-written code that runs and passes its tests is not enough: security, maintainability and being able to roll it back still need humans in the loop. — tap to centre the map on it Agent-written code that runs andpasses its tests is not enough:security, maintainability and beingable to roll it back still needhumans in the loop. Last stated 4 months ago 7 Jun 2026 TF Tony Fadell — holds since 2026-06-07 — tap for who they are Same subject: Agentic code review raises the floor but cannot be trusted, because the model reading the code is the same model that wrote it, and it will tell you the code is great. — tap to centre the map on it Agentic code review raises the floorbut cannot be trusted, because themodel reading the code is the samemodel that wrote it, and it willtell you the code is great. Last stated 3 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre Error discovery is the most important part of the AI evaluation process. Last stated 22 Sept 2026 · a week ago Holds Lenny Rachitsky Read this korrent →