Tap a claim on the ring to put it at the centre.
← Error discovery is the most important part of the AI evaluation process.
17 connected korrents · 14 moments from 9 Aug 2024 to 22 Sept 2026.
Everything filed under software quality
software quality
Everything filed under product management
product management
Everything filed under AI and science
AI and science
Everything filed under AI and jobs
AI and jobs
Everything filed under AI agents
AI agents
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Error discovery is the most important part of the AI evaluation process.
Error discovery is the most important part of the AI evaluation process.
Last stated a week ago
22 Sept 2026
LR
Lenny Rachitsky — holds since 2026-09-22 — tap for who they are
Same subject: Automated evaluation tools are useful for finding some types of AI errors, especially when combined with human judgment. — tap to centre the map on it
Automated evaluation tools are useful for finding some types of AI errors, especially when combined with human judgment.
Last stated a week ago
22 Sept 2026
LR
Lenny Rachitsky — holds since 2026-09-22 — tap for who they are
Same subject: For a product person working with AI research, learning to write good evals is the most important thing. — tap to centre the map on it
For a product person working with AI research, learning to write good evals is the most important thing.
Last stated 3 weeks ago
10 Sept 2026
TS
Tara Seshan — holds since 2026-09-10 — tap for who they are
Same subject: Evaluation and error analysis will remain necessary practices even if AI models become near-perfect. — tap to centre the map on it
Evaluation and error analysis will remain necessary practices even if AI models become near-perfect.
Last stated 2 weeks ago
18 Sept 2026
HH
Hamel Husain — holds since 2026-09-18 — tap for who they are
Same subject: Generating a list of potential problems is worthless unless it is followed by validating and prioritizing them. — tap to centre the map on it
Generating a list of potential problems is worthless unless it is followed by validating and prioritizing them.
Last stated 2 years ago
9 Aug 2024
TS
Tobias Sjösten — holds since 2024-08-09 — tap for who they are
Same subject: AI errors are harder to catch in analysis of unfamiliar data than in simple writing tasks. — tap to centre the map on it
AI errors are harder to catch in analysis of unfamiliar data than in simple writing tasks.
Last stated a year ago
5 Sept 2025
JE
Jason Evanish — holds since 2025-09-05 — tap for who they are
Same subject: Thanks to AI, the principles of the product model, product strategy and product discovery have never been more important. — tap to centre the map on it
Thanks to AI, the principles of the product model, product strategy and product discovery have never been more important.
Last stated 3 weeks ago
10 Sept 2026
MC
Marty Cagan — holds since 2026-09-10 — tap for who they are
Same subject: AI systems require human oversight rather than full autonomy in analysis and workflows. — tap to centre the map on it
AI systems require human oversight rather than full autonomy in analysis and workflows.
Last stated a year ago
5 Sept 2025
JE
Jason Evanish — holds since 2025-09-05 — tap for who they are
Same subject: Today's AI tools either solve a problem or fail at it, and are really bad at partial progress or at identifying which intermediate step to attack first. — tap to centre the map on it
Today's AI tools either solve a problem or fail at it, and are really bad at partial progress or at identifying which intermediate step to attack first.
Last stated 6 months ago
20 Mar 2026
TT
Terence Tao — holds since 2026-03-20 — tap for who they are
Same subject: A coding agent left to run over a weekend returns work that is half good, a quarter garbage and a quarter in need of repair. — tap to centre the map on it
A coding agent left to run over a weekend returns work that is half good, a quarter garbage and a quarter in need of repair.
Last stated a year ago
18 Aug 2025
CB
Clay Bavor — holds since 2025-08-18 — tap for who they are
Same subject: A developer using AI coding tools remains responsible for reviewing the code they ship. — tap to centre the map on it
A developer using AI coding tools remains responsible for reviewing the code they ship.
Last stated 7 months ago
10 Mar 2026
JS
Jon Seager — holds since 2026-03-10 — tap for who they are
Same subject: A language-agnostic conformance suite is the most powerful thing you can hand a coding agent, because the whole instruction becomes: write code until these tests pass. — tap to centre the map on it
A language-agnostic conformance suite is the most powerful thing you can hand a coding agent, because the whole instruction becomes: write code until these tests pass.
Last stated 6 months ago
19 Mar 2026
SW
Simon Willison — holds since 2026-03-19 — tap for who they are
Same subject: AI-assisted coding will generate a lot of low-quality output that developers will need to clean up, even as it also produces good results. — tap to centre the map on it
AI-assisted coding will generate a lot of low-quality output that developers will need to clean up, even as it also produces good results.
Last stated 8 months ago
9 Feb 2026
SH
Scott Hanselman — holds since 2026-02-09 — tap for who they are
Same subject: AI-generated code should be treated as untrusted until there is evidence it preserves the intended behavior. — tap to centre the map on it
AI-generated code should be treated as untrusted until there is evidence it preserves the intended behavior.
Last stated a month ago
19 Aug 2026
JS
Jon Seager — holds since 2026-08-19 — tap for who they are
Same subject: Code review as a review mechanism depends on reviewers being able to infer effort from reading code, and AI-agent-generated code erases that signal. — tap to centre the map on it
Code review as a review mechanism depends on reviewers being able to infer effort from reading code, and AI-agent-generated code erases that signal.
Last stated 5 months ago
7 May 2026
DC
David Crawshaw — holds since 2026-05-07 — tap for who they are
Same subject: A team that slows down and reads every pull request and every line of code should expect only a 30 to 50 percent productivity lift from AI. — tap to centre the map on it
A team that slows down and reads every pull request and every line of code should expect only a 30 to 50 percent productivity lift from AI.
Last stated 3 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: Agent-written code that runs and passes its tests is not enough: security, maintainability and being able to roll it back still need humans in the loop. — tap to centre the map on it
Agent-written code that runs and passes its tests is not enough: security, maintainability and being able to roll it back still need humans in the loop.
Last stated 4 months ago
7 Jun 2026
TF
Tony Fadell — holds since 2026-06-07 — tap for who they are
Same subject: Agentic code review raises the floor but cannot be trusted, because the model reading the code is the same model that wrote it, and it will tell you the code is great. — tap to centre the map on it
Agentic code review raises the floor but cannot be trusted, because the model reading the code is the same model that wrote it, and it will tell you the code is great.
Last stated 3 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are
At the centre
Error discovery is the most important part of the AI evaluation process.
Last stated 22 Sept 2026 · a week ago
Holds Lenny Rachitsky
Read this korrent →
Similar wording
Automated evaluation tools are useful for finding some types of AI errors, especially when combined with human judgment.
Last stated 22 Sept 2026 · a week ago
Holds Lenny Rachitsky
Similar wording
For a product person working with AI research, learning to write good evals is the most important thing.
Last stated 10 Sept 2026 · 3 weeks ago
Holds Tara Seshan
Similar wording
Evaluation and error analysis will remain necessary practices even if AI models become near-perfect.
Last stated 18 Sept 2026 · 2 weeks ago
Holds Hamel Husain
Similar wording
Generating a list of potential problems is worthless unless it is followed by validating and prioritizing them.
Last stated 9 Aug 2024 · 2 years ago
Holds Tobias Sjösten
Similar wording
AI errors are harder to catch in analysis of unfamiliar data than in simple writing tasks.
Last stated 5 Sept 2025 · a year ago
Holds Jason Evanish
Similar wording
Thanks to AI, the principles of the product model, product strategy and product discovery have never been more important.
Last stated 10 Sept 2026 · 3 weeks ago
Holds Marty Cagan
Similar wording
AI systems require human oversight rather than full autonomy in analysis and workflows.
Last stated 5 Sept 2025 · a year ago
Holds Jason Evanish
Similar wording
Today's AI tools either solve a problem or fail at it, and are really bad at partial progress or at identifying which intermediate step to attack first.
Last stated 20 Mar 2026 · 6 months ago
Holds Terence Tao
Same subject: software quality
A coding agent left to run over a weekend returns work that is half good, a quarter garbage and a quarter in need of repair.
Last stated 18 Aug 2025 · a year ago
Holds Clay Bavor
Same subject: software quality
A developer using AI coding tools remains responsible for reviewing the code they ship.
Last stated 10 Mar 2026 · 7 months ago
Holds Jon Seager
Same subject: software quality
A language-agnostic conformance suite is the most powerful thing you can hand a coding agent, because the whole instruction becomes: write code until these tests pass.
Last stated 19 Mar 2026 · 6 months ago
Holds Simon Willison
Same subject: code generation
AI-assisted coding will generate a lot of low-quality output that developers will need to clean up, even as it also produces good results.
Last stated 9 Feb 2026 · 8 months ago
Holds Scott Hanselman
Same subject: code generation
AI-generated code should be treated as untrusted until there is evidence it preserves the intended behavior.
Last stated 19 Aug 2026 · a month ago
Holds Jon Seager
Same subject: code generation
Code review as a review mechanism depends on reviewers being able to infer effort from reading code, and AI-agent-generated code erases that signal.
Last stated 7 May 2026 · 5 months ago
Holds David Crawshaw
Same subject: coding agents
A team that slows down and reads every pull request and every line of code should expect only a 30 to 50 percent productivity lift from AI.
Last stated 15 Jul 2026 · 3 months ago
Holds Dex Horthy
Same subject: coding agents
Agent-written code that runs and passes its tests is not enough: security, maintainability and being able to roll it back still need humans in the loop.
Last stated 7 Jun 2026 · 4 months ago
Holds Tony Fadell
Same subject: coding agents
Agentic code review raises the floor but cannot be trusted, because the model reading the code is the same model that wrote it, and it will tell you the code is great.
Last stated 15 Jul 2026 · 3 months ago
Holds Dex Horthy