korrents

On the map

Tap a claim on the ring to put it at the centre.

← When pre-training data is limited, it is better to train a smaller…

17 connected korrents · 16 moments on record from 11 Oct 2018 to 19 Sept 2026.

Everything filed under scaling laws scaling laws Everything filed under benchmarks benchmarks Everything filed under LLMs LLMs Everything filed under reinforcement learning reinforcement learning Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: When pre-training data is limited, it is better to train a smaller model for multiple epochs than to train a larger model on unique data once. When pre-training data is limited, it isbetter to train a smaller model formultiple epochs than to train a largermodel on unique data once. Last stated 3 years ago 1 Dec 2023 SR Sebastian Ruder — holds since 2023-12-01 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it After pre-training, post-trainingand test-time scaling, the fourthscaling law is agentic: multiplyingAI by spawning agents, and the wholeloop scales on one thing, compute. Last stated 6 months ago 23 Mar 2026 JH Jensen Huang — holds since 2026-03-23 — tap for who they are Same subject: AI scaling laws require exponentially more compute and resources to produce only linear gains in model intelligence. — tap to centre the map on it AI scaling laws requireexponentially more compute andresources to produce only lineargains in model intelligence. Last stated 2 days ago 19 Sept 2026 NL Nathan Lambert — holds since 2026-09-19 — tap for who they are Same subject: As models keep scaling up, modularity will become essential to how they are developed. — tap to centre the map on it As models keep scaling up,modularity will become essential tohow they are developed. Last stated 4 years ago 23 Feb 2023 SR Sebastian Ruder — holds since 2023-02-23 — tap for who they are Same subject: Bet on a system that is maximally learned and minimally constrained, and add structure only where it improves the scaling laws. — tap to centre the map on it Bet on a system that is maximallylearned and minimally constrained,and add structure only where itimproves the scaling laws. Last stated 2 months ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: Current AI techniques are four to six orders of magnitude away from optimal in data efficiency and test-time compute efficiency. — tap to centre the map on it Current AI techniques are four tosix orders of magnitude away fromoptimal in data efficiency andtest-time compute efficiency. Last stated a month ago 7 Aug 2026 FC François Chollet — holds since 2026-08-07 — tap for who they are Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it Current techniques restrictpre-trained representation powerbecause standard language models areunidirectional. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 3 weeks ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old. — tap to centre the map on it A carmaker's claim to be the safestis mostly an artefact of comparing anew car against a fleet averagetwelve years old. Last stated 3 years ago 19 Dec 2023 PK Philip Koopman — holds since 2023-12-19 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 6 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it A physical AI company needs threeAIs, not one — the agent, thesimulator and the critic — turningdeployment into a flywheel. Last stated 2 months ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it A verifiable task can be optimisedby reinforcement learning until aneural network performs it extremelywell. Last stated 10 months ago 17 Nov 2025 AK Andrej Karpathy — holds since 2025-11-17 — tap for who they are Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it Building a realistic simulator isexactly as hard as building theagent, because the simulator isitself a large AI model. Last stated 2 months ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: A badly written AI outbound email is evidence of a bad vendor, not of a limit of AI. — tap to centre the map on it A badly written AI outbound email isevidence of a bad vendor, not of alimit of AI. Last stated 9 months ago 1 Jan 2026 JL Jason Lemkin — holds since 2026-01-01 — tap for who they are Same subject: A bigger context window does not give you a smarter model; the intelligence of the model is what decides how much of that window it can actually attend to. — tap to centre the map on it A bigger context window does notgive you a smarter model; theintelligence of the model is whatdecides how much of that window itcan actually attend to. Last stated 2 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: A child who has seen ten cats learns what a machine needs the whole internet of cat photos for, by a learning pathway nobody has solved. — tap to centre the map on it A child who has seen ten cats learnswhat a machine needs the wholeinternet of cat photos for, by alearning pathway nobody has solved. Last stated a month ago 10 Aug 2026 FL Fei-Fei Li — holds since 2026-08-10 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre When pre-training data is limited, it is better to train a smaller model for multiple epochs than to train a larger model on unique data once. Last stated 1 Dec 2023 · 3 years ago Holds Sebastian Ruder Read this korrent →