Tap a claim on the ring to put it at the centre.
← Automated evaluation tools are useful for finding some types of AI…
17 connected korrents · 13 moments from 31 Dec 2023 to 22 Sept 2026.
Everything filed under software quality
software quality
Everything filed under AI and science
AI and science
Everything filed under AI agents
AI agents
Everything filed under LLMs
LLMs
Everything filed under coding agents
coding agents
Everything filed under coding agents
coding agents
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Automated evaluation tools are useful for finding some types of AI errors, especially when combined with human judgment.
Automated evaluation tools are useful for finding some types of AI errors, especially when combined with human judgment.
Last stated a week ago
22 Sept 2026
LR
Lenny Rachitsky — holds since 2026-09-22 — tap for who they are
Same subject: Error discovery is the most important part of the AI evaluation process. — tap to centre the map on it
Error discovery is the most important part of the AI evaluation process.
Last stated a week ago
22 Sept 2026
LR
Lenny Rachitsky — holds since 2026-09-22 — tap for who they are
Same subject: Today's AI tools either solve a problem or fail at it, and are really bad at partial progress or at identifying which intermediate step to attack first. — tap to centre the map on it
Today's AI tools either solve a problem or fail at it, and are really bad at partial progress or at identifying which intermediate step to attack first.
Last stated 6 months ago
20 Mar 2026
TT
Terence Tao — holds since 2026-03-20 — tap for who they are
Same subject: The most underused thing AI does is teach you: people reach for it to write their code and forget it can take them through a curriculum in an hour. — tap to centre the map on it
The most underused thing AI does is teach you: people reach for it to write their code and forget it can take them through a curriculum in an hour.
Last stated a year ago
21 Sept 2025
JZ
Julie Zhuo — holds since 2025-09-21 — tap for who they are
Same subject: On small tasks he can check himself, he catches about as many of the AI's errors as it catches of his. — tap to centre the map on it
On small tasks he can check himself, he catches about as many of the AI's errors as it catches of his.
Last stated 6 months ago
20 Mar 2026
TT
Terence Tao — holds since 2026-03-20 — tap for who they are
Same subject: Evaluation and error analysis will remain necessary practices even if AI models become near-perfect. — tap to centre the map on it
Evaluation and error analysis will remain necessary practices even if AI models become near-perfect.
Last stated 2 weeks ago
18 Sept 2026
HH
Hamel Husain — holds since 2026-09-18 — tap for who they are
Same subject: AI-validated pull requests beat human review at exactly the things humans are bad at, which frees the humans to argue about direction instead. — tap to centre the map on it
AI-validated pull requests beat human review at exactly the things humans are bad at, which frees the humans to argue about direction instead.
Last stated 2 months ago
12 Aug 2026
CM
Charity Majors — holds since 2026-08-12 — tap for who they are
Same subject: On any given problem an AI tool succeeds maybe one or two percent of the time; the headline results are what running it at scale and picking the winners looks like. — tap to centre the map on it
On any given problem an AI tool succeeds maybe one or two percent of the time; the headline results are what running it at scale and picking the winners looks like.
Last stated 6 months ago
20 Mar 2026
TT
Terence Tao — holds since 2026-03-20 — tap for who they are
Same subject: AI errors are harder to catch in analysis of unfamiliar data than in simple writing tasks. — tap to centre the map on it
AI errors are harder to catch in analysis of unfamiliar data than in simple writing tasks.
Last stated a year ago
5 Sept 2025
JE
Jason Evanish — holds since 2025-09-05 — tap for who they are
Same subject: A developer using AI coding tools remains responsible for reviewing the code they ship. — tap to centre the map on it
A developer using AI coding tools remains responsible for reviewing the code they ship.
Last stated 7 months ago
10 Mar 2026
JS
Jon Seager — holds since 2026-03-10 — tap for who they are
Same subject: CMake tooling is missing a linter that could statically catch most CMakeLists.txt mistakes. — tap to centre the map on it
CMake tooling is missing a linter that could statically catch most CMakeLists.txt mistakes.
Last stated 2 weeks ago
20 Sept 2026
CW
Chris Wellons — holds since 2026-09-20 — tap for who they are
Same subject: Comparing code coverage between passing and failing test runs is a cheap, effective way to localize a bug quickly. — tap to centre the map on it
Comparing code coverage between passing and failing test runs is a cheap, effective way to localize a bug quickly.
Last stated a year ago
25 Apr 2025
RC
Russ Cox — holds since 2025-04-25 — tap for who they are
Same subject: Everyone is rethinking code reviews, deploys, and possibly observability — tap to centre the map on it
Everyone is rethinking code reviews, deploys, and possibly observability
Last stated 2 weeks ago
16 Sept 2026
GO
Gergely Orosz — holds since 2026-09-16 — tap for who they are
Same subject: Metaprogramming in a language creates fundamental obstacles to static analysis of its code. — tap to centre the map on it
Metaprogramming in a language creates fundamental obstacles to static analysis of its code.
Last stated 3 years ago
31 Dec 2023
PH
Phil Hagelberg — holds since 2023-12-31 — tap for who they are
Same subject: The valuable form of loop engineering is a slow loop — a nightly cron job that fixes one thing and opens one pull request — not a redesign of your whole infrastructure around agents. — tap to centre the map on it
The valuable form of loop engineering is a slow loop — a nightly cron job that fixes one thing and opens one pull request — not a redesign of your whole infrastructure around agents.
Last stated 3 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: A coding agent left to run over a weekend returns work that is half good, a quarter garbage and a quarter in need of repair. — tap to centre the map on it
A coding agent left to run over a weekend returns work that is half good, a quarter garbage and a quarter in need of repair.
Last stated a year ago
18 Aug 2025
CB
Clay Bavor — holds since 2025-08-18 — tap for who they are
Same subject: A software engineer who is not using a tool like Cursor is working at half productivity or worse. — tap to centre the map on it
A software engineer who is not using a tool like Cursor is working at half productivity or worse.
Last stated a year ago
18 Aug 2025
BT
Bret Taylor — holds since 2025-08-18 — tap for who they are
Same subject: A team that slows down and reads every pull request and every line of code should expect only a 30 to 50 percent productivity lift from AI. — tap to centre the map on it
A team that slows down and reads every pull request and every line of code should expect only a 30 to 50 percent productivity lift from AI.
Last stated 3 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are