korrents

On the map

Tap a claim on the ring to put it at the centre.

← Reinforcement learning only works once a model already knows…

17 connected korrents · 16 moments on record from 22 Apr 2024 to 12 Aug 2026. Nearly all of them are about reinforcement learning.

Everything filed under robotics robotics Everything filed under LLMs LLMs Everything filed under scaling laws scaling laws Everything filed under self-driving cars self-driving cars Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Reinforcement learning only works once a model already knows something, which is why robots must be pre-trained by imitation first. Reinforcement learning only works once amodel already knows something, which iswhy robots must be pre-trained byimitation first. Last stated 12 months ago 12 Sept 2025 SL Sergey Levine — holds since 2025-09-12 — tap for who they are Same subject: One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for. — tap to centre the map on it One pre-trained robot model nowmatches or beats the specialiststhat were fine-tuned withreinforcement learning for the verytasks they were built for. Last stated 4 weeks ago 12 Aug 2026 CF Chelsea Finn — holds since 2026-08-12 — tap for who they are Same subject: Reinforcement learning cannot be scaled for robots the way it was for language models, because every attempt spends real robot hours instead of data-centre compute. — tap to centre the map on it Reinforcement learning cannot bescaled for robots the way it was forlanguage models, because everyattempt spends real robot hoursinstead of data-centre compute. Last stated 4 weeks ago 12 Aug 2026 CF Chelsea Finn — holds since 2026-08-12 — tap for who they are Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it A verifiable task can be optimisedby reinforcement learning until aneural network performs it extremelywell. Last stated 10 months ago 17 Nov 2025 AK Andrej Karpathy — holds since 2025-11-17 — tap for who they are Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it Capability does not generalise forfree: a model that will movemountains on an agentic task stilltells the same bad joke it told fiveyears ago. Last stated 6 months ago 20 Mar 2026 AK Andrej Karpathy — holds since 2026-03-20 — tap for who they are Same subject: Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving. — tap to centre the map on it Humans barely use reinforcementlearning for intelligence — what RLthey do use goes into motor tasks,not problem solving. Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: Humans keep their place in AI as judges rather than authors, because telling which of two answers is better is far easier than writing a good one. — tap to centre the map on it Humans keep their place in AI asjudges rather than authors, becausetelling which of two answers isbetter is far easier than writing agood one. Last stated 2 years ago 3 Feb 2025 NL Nathan Lambert — holds since 2025-02-03 — tap for who they are Same subject: Large language models mimic what people say to do rather than work out what to do, which is why they are not about understanding the world. — tap to centre the map on it Large language models mimic whatpeople say to do rather than workout what to do, which is why theyare not about understanding theworld. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments. — tap to centre the map on it Models look far better on evals thanthey are in the world becauseresearchers, inadvertently, takeinspiration from the evals when theybuild RL environments. Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it After pre-training, post-trainingand test-time scaling, the fourthscaling law is agentic: multiplyingAI by spawning agents, and the wholeloop scales on one thing, compute. Last stated 6 months ago 23 Mar 2026 JH Jensen Huang — holds since 2026-03-23 — tap for who they are Same subject: A five-year-old's vision is already good enough to drive a car, on a tiny and strikingly undiverse amount of data. — tap to centre the map on it A five-year-old's vision is alreadygood enough to drive a car, on atiny and strikingly undiverse amountof data. Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A good robotaxi demonstration in one city does not mean the company is ready to scale the service. — tap to centre the map on it A good robotaxi demonstration in onecity does not mean the company isready to scale the service. Last stated a year ago 8 Jul 2025 TL Timothy B. Lee — holds since 2025-07-08 — tap for who they are Same subject: A huge fleet collecting driving data is not a silver bullet for self-driving, because the data arrives unlabelled. — tap to centre the map on it A huge fleet collecting driving datais not a silver bullet forself-driving, because the dataarrives unlabelled. Last stated 2 years ago 21 May 2024 TL Timothy B. Lee — holds since 2024-05-21 — tap for who they are Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it A language model is no substitutefor a well-specified conventionalalgorithm, so it cannot simply bedropped into a complex problem andtrusted. Last stated a year ago 7 Jun 2025 GM Gary Marcus — holds since 2025-06-07 — tap for who they are Same subject: A language model is not using language at all, because language requires an intention to communicate. — tap to centre the map on it A language model is not usinglanguage at all, because languagerequires an intention tocommunicate. Last stated 2 years ago 31 Aug 2024 TC Ted Chiang — holds since 2024-08-31 — tap for who they are Same subject: A language model's apparent mind is mostly our own bias: it predicts text, and leverages our evolved habit of attributing intentionality to anything that acts human. — tap to centre the map on it A language model's apparent mind ismostly our own bias: it predictstext, and leverages our evolvedhabit of attributing intentionalityto anything that acts human. Last stated 2 years ago 22 Apr 2024 SC Sean Carroll — holds since 2024-04-22 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Reinforcement learning only works once a model already knows something, which is why robots must be pre-trained by imitation first. Last stated 12 Sept 2025 · 12 months ago Holds Sergey Levine Read this korrent →