korrents

On the map

Tap a claim on the ring to put it at the centre.

← Automated evaluation tools are useful for finding some types of AI…

17 connected korrents · 13 moments from 31 Dec 2023 to 22 Sept 2026.

Everything filed under software quality software quality Everything filed under AI and science AI and science Everything filed under AI agents AI agents Everything filed under LLMs LLMs Everything filed under coding agents coding agents Everything filed under coding agents coding agents Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Automated evaluation tools are useful for finding some types of AI errors, especially when combined with human judgment. Automated evaluation tools are useful forfinding some types of AI errors,especially when combined with humanjudgment. Last stated a week ago 22 Sept 2026 LR Lenny Rachitsky — holds since 2026-09-22 — tap for who they are Same subject: Error discovery is the most important part of the AI evaluation process. — tap to centre the map on it Error discovery is the mostimportant part of the AI evaluationprocess. Last stated a week ago 22 Sept 2026 LR Lenny Rachitsky — holds since 2026-09-22 — tap for who they are Same subject: Today's AI tools either solve a problem or fail at it, and are really bad at partial progress or at identifying which intermediate step to attack first. — tap to centre the map on it Today's AI tools either solve aproblem or fail at it, and arereally bad at partial progress or atidentifying which intermediate stepto attack first. Last stated 6 months ago 20 Mar 2026 TT Terence Tao — holds since 2026-03-20 — tap for who they are Same subject: The most underused thing AI does is teach you: people reach for it to write their code and forget it can take them through a curriculum in an hour. — tap to centre the map on it The most underused thing AI does isteach you: people reach for it towrite their code and forget it cantake them through a curriculum in anhour. Last stated a year ago 21 Sept 2025 JZ Julie Zhuo — holds since 2025-09-21 — tap for who they are Same subject: On small tasks he can check himself, he catches about as many of the AI's errors as it catches of his. — tap to centre the map on it On small tasks he can check himself,he catches about as many of the AI'serrors as it catches of his. Last stated 6 months ago 20 Mar 2026 TT Terence Tao — holds since 2026-03-20 — tap for who they are Same subject: Evaluation and error analysis will remain necessary practices even if AI models become near-perfect. — tap to centre the map on it Evaluation and error analysis willremain necessary practices even ifAI models become near-perfect. Last stated 2 weeks ago 18 Sept 2026 HH Hamel Husain — holds since 2026-09-18 — tap for who they are Same subject: AI-validated pull requests beat human review at exactly the things humans are bad at, which frees the humans to argue about direction instead. — tap to centre the map on it AI-validated pull requests beathuman review at exactly the thingshumans are bad at, which frees thehumans to argue about directioninstead. Last stated 2 months ago 12 Aug 2026 CM Charity Majors — holds since 2026-08-12 — tap for who they are Same subject: On any given problem an AI tool succeeds maybe one or two percent of the time; the headline results are what running it at scale and picking the winners looks like. — tap to centre the map on it On any given problem an AI toolsucceeds maybe one or two percent ofthe time; the headline results arewhat running it at scale and pickingthe winners looks like. Last stated 6 months ago 20 Mar 2026 TT Terence Tao — holds since 2026-03-20 — tap for who they are Same subject: AI errors are harder to catch in analysis of unfamiliar data than in simple writing tasks. — tap to centre the map on it AI errors are harder to catch inanalysis of unfamiliar data than insimple writing tasks. Last stated a year ago 5 Sept 2025 JE Jason Evanish — holds since 2025-09-05 — tap for who they are Same subject: A developer using AI coding tools remains responsible for reviewing the code they ship. — tap to centre the map on it A developer using AI coding toolsremains responsible for reviewingthe code they ship. Last stated 7 months ago 10 Mar 2026 JS Jon Seager — holds since 2026-03-10 — tap for who they are Same subject: CMake tooling is missing a linter that could statically catch most CMakeLists.txt mistakes. — tap to centre the map on it CMake tooling is missing a linterthat could statically catch mostCMakeLists.txt mistakes. Last stated 2 weeks ago 20 Sept 2026 CW Chris Wellons — holds since 2026-09-20 — tap for who they are Same subject: Comparing code coverage between passing and failing test runs is a cheap, effective way to localize a bug quickly. — tap to centre the map on it Comparing code coverage betweenpassing and failing test runs is acheap, effective way to localize abug quickly. Last stated a year ago 25 Apr 2025 RC Russ Cox — holds since 2025-04-25 — tap for who they are Same subject: Everyone is rethinking code reviews, deploys, and possibly observability — tap to centre the map on it Everyone is rethinking code reviews,deploys, and possibly observability Last stated 2 weeks ago 16 Sept 2026 GO Gergely Orosz — holds since 2026-09-16 — tap for who they are Same subject: Metaprogramming in a language creates fundamental obstacles to static analysis of its code. — tap to centre the map on it Metaprogramming in a languagecreates fundamental obstacles tostatic analysis of its code. Last stated 3 years ago 31 Dec 2023 PH Phil Hagelberg — holds since 2023-12-31 — tap for who they are Same subject: The valuable form of loop engineering is a slow loop — a nightly cron job that fixes one thing and opens one pull request — not a redesign of your whole infrastructure around agents. — tap to centre the map on it The valuable form of loopengineering is a slow loop — anightly cron job that fixes onething and opens one pull request —not a redesign of your wholeinfrastructure around agents. Last stated 3 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: A coding agent left to run over a weekend returns work that is half good, a quarter garbage and a quarter in need of repair. — tap to centre the map on it A coding agent left to run over aweekend returns work that is halfgood, a quarter garbage and aquarter in need of repair. Last stated a year ago 18 Aug 2025 CB Clay Bavor — holds since 2025-08-18 — tap for who they are Same subject: A software engineer who is not using a tool like Cursor is working at half productivity or worse. — tap to centre the map on it A software engineer who is not usinga tool like Cursor is working athalf productivity or worse. Last stated a year ago 18 Aug 2025 BT Bret Taylor — holds since 2025-08-18 — tap for who they are Same subject: A team that slows down and reads every pull request and every line of code should expect only a 30 to 50 percent productivity lift from AI. — tap to centre the map on it A team that slows down and readsevery pull request and every line ofcode should expect only a 30 to 50percent productivity lift from AI. Last stated 3 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre Automated evaluation tools are useful for finding some types of AI errors, especially when combined with human judgment. Last stated 22 Sept 2026 · a week ago Holds Lenny Rachitsky Read this korrent →