korrents

On the map

Tap a claim on the ring to put it at the centre.

← AI judges used to evaluate AI outputs need ongoing validation and iteration, just like any other AI system.

16 connected korrents · 14 moments on record from 5 Nov 2019 to 15 Sept 2026.

Everything filed under AGI AGI Everything filed under measuring intelligence measuring intelligence Everything filed under benchmarks benchmarks Everything filed under scaling laws scaling laws Everything filed under LLMs LLMs Everything filed under AI agents AI agents Everything filed under coding agents coding agents Everything filed under AI and human skill AI and human skill Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: AI judges used to evaluate AI outputs need ongoing validation and iteration, just like any other AI system. AI judges used to evaluate AI outputs needongoing validation and iteration, justlike any other AI system. Last stated 2 years ago 16 Jan 2025 CH Chip Huyen — holds since 2025-01-16 — tap for who they are Same subject: Progress toward more intelligent artificial systems requires defining and evaluating intelligence so systems can be compared with each other and with humans. — tap to centre the map on it Progress toward more intelligentartificial systems requires definingand evaluating intelligence sosystems can be compared with eachother and with humans. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: Process supervision by an LLM judge cannot be run for long, because a judge with billions of parameters is gameable and RL will find its cracks. — tap to centre the map on it Process supervision by an LLM judgecannot be run for long, because ajudge with billions of parameters isgameable and RL will find itscracks. Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: Reliance on public benchmark results for evaluating AI models will decrease over time. — tap to centre the map on it Reliance on public benchmark resultsfor evaluating AI models willdecrease over time. Last stated 2 years ago 13 May 2024 SR Sebastian Ruder — holds since 2024-05-13 — tap for who they are Same subject: AI agents trained on an expert's published work can only automate a small fraction of that expert's actual job. — tap to centre the map on it AI agents trained on an expert'spublished work can only automate asmall fraction of that expert'sactual job. Last stated 10 months ago 27 Nov 2025 BG Brendan Gregg — holds since 2025-11-27 — tap for who they are Same subject: Review became a worse bottleneck under AI, because companies removed the review automation they had once AI started writing the code. — tap to centre the map on it Review became a worse bottleneckunder AI, because companies removedthe review automation they had onceAI started writing the code. Last stated 6 months ago 22 Mar 2026 NF Nicole Forsgren — holds since 2026-03-22 — tap for who they are Same subject: AI can supply the instruction and the feedback in a practice loop, but it cannot supply the middle step, which is doing the work. — tap to centre the map on it AI can supply the instruction andthe feedback in a practice loop, butit cannot supply the middle step,which is doing the work. Last stated 2 months ago 22 Jul 2026 SY Scott H. Young — holds since 2026-07-22 — tap for who they are Same subject: The two standard criticisms of AI — that it is not deterministic and that it is not creative — cannot both be true. — tap to centre the map on it The two standard criticisms of AI —that it is not deterministic andthat it is not creative — cannotboth be true. Last stated 4 weeks ago 26 Aug 2026 DH David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are Same subject: A true artificial general intelligence cannot exist without being recognized as a moral subject. — tap to centre the map on it A true artificial generalintelligence cannot exist withoutbeing recognized as a moral subject. Last stated a year ago 10 Jun 2025 SH Samuel Hammond — holds since 2025-06-10 — tap for who they are Same subject: A unit of AI inference needs to be defined, for example via a chain of increasingly hard problems where each consecutive pair is solvable by one model. — tap to centre the map on it A unit of AI inference needs to bedefined, for example via a chain ofincreasingly hard problems whereeach consecutive pair is solvable byone model. Last stated 6 days ago 15 Sept 2026 PG Paul Graham — holds since 2026-09-15 — tap for who they are Same subject: Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect. — tap to centre the map on it Both the special-purpose-programsview and the blank-slate view ofhuman intelligence are likelyincorrect. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: "Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next. — tap to centre the map on it "Scaling" was powerful because itwas one word: naming a researchdirection is what tells a wholefield what to do next. Last stated 10 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it A country outside the AI supplychain should just buy the index —which works only in the world whereAI ends up commoditised rather thanconcentrated. Last stated 4 months ago 4 Jun 2026 AI Alex Imas — holds since 2026-06-04 — tap for who they are Same subject: A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean. — tap to centre the map on it A human being is not an AGI: we lacka huge amount of knowledge and relyon continual learning instead, socontinual learning is whatsuperintelligence should mean. Last stated 10 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it A poor country should prioritiseowning a piece of AI over retrainingits workers, but it should not beteverything on that. Last stated 4 months ago 4 Jun 2026 PT Phil Trammell — holds since 2026-06-04 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre AI judges used to evaluate AI outputs need ongoing validation and iteration, just like any other AI system. Last stated 16 Jan 2025 · 2 years ago Holds Chip Huyen Read this korrent →