korrents

On the map

Tap a claim on the ring to put it at the centre.

← Learning from a reward ten years away is a solved problem: a value…

17 connected korrents · 15 moments on record from 1 Oct 2023 to 3 Aug 2026.

Everything filed under robotics robotics Everything filed under scaling laws scaling laws Everything filed under self-driving cars self-driving cars Everything filed under reinforcement learning reinforcement learning Everything filed under market efficiency market efficiency Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Learning from a reward ten years away is a solved problem: a value function trained by temporal-difference learning rewards the steps along the way. Learning from a reward ten years away is asolved problem: a value function trainedby temporal-difference learning rewardsthe steps along the way. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: Value functions only make reinforcement learning faster: anything you can do with one you can also do without it, just more slowly. — tap to centre the map on it Value functions only makereinforcement learning faster:anything you can do with one you canalso do without it, just moreslowly. Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: Computer-use agents had to wait for language models: without pre-trained representations the reward is too sparse to ever learn from. — tap to centre the map on it Computer-use agents had to wait forlanguage models: without pre-trainedrepresentations the reward is toosparse to ever learn from. Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: The goal was never good simulation; it was answering counterfactuals, and a value function does that job as well as a simulator does. — tap to centre the map on it The goal was never good simulation;it was answering counterfactuals,and a value function does that jobas well as a simulator does. Last stated 12 months ago 12 Sept 2025 SL Sergey Levine — holds since 2025-09-12 — tap for who they are Same subject: To reward the instinct that a new idea is worth having, training would have to score the smallness of the concepts a solution needs, not just whether it solved the problem. — tap to centre the map on it To reward the instinct that a newidea is worth having, training wouldhave to score the smallness of theconcepts a solution needs, not justwhether it solved the problem. Last stated 2 months ago 30 Jun 2026 GS Grant Sanderson — holds since 2026-06-30 — tap for who they are Same subject: Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving. — tap to centre the map on it Humans barely use reinforcementlearning for intelligence — what RLthey do use goes into motor tasks,not problem solving. Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: Reinforcement learning only works once a model already knows something, which is why robots must be pre-trained by imitation first. — tap to centre the map on it Reinforcement learning only worksonce a model already knowssomething, which is why robots mustbe pre-trained by imitation first. Last stated 12 months ago 12 Sept 2025 SL Sergey Levine — holds since 2025-09-12 — tap for who they are Same subject: Continual learning will probably be solved within a year or two, but the trillions of dollars do not depend on solving it. — tap to centre the map on it Continual learning will probably besolved within a year or two, but thetrillions of dollars do not dependon solving it. Last stated 7 months ago 13 Feb 2026 DA Dario Amodei — holds since 2026-02-13 — tap for who they are Same subject: Even the most advanced AI cannot be followed blindly in investing, where value added is zero-sum and what is widely known is therefore worth little. — tap to centre the map on it Even the most advanced AI cannot befollowed blindly in investing, wherevalue added is zero-sum and what iswidely known is therefore worthlittle. Last stated 3 months ago 10 Jun 2026 RD Ray Dalio — holds since 2026-06-10 — tap for who they are Same subject: A highly dynamic economy that cheaply mass-produces robots may increase rather than decrease the risk of Malthusian resource dilemmas. — tap to centre the map on it A highly dynamic economy thatcheaply mass-produces robots mayincrease rather than decrease therisk of Malthusian resourcedilemmas. Last stated 3 years ago 1 Oct 2023 TC Tyler Cowen — holds since 2023-10-01 — tap for who they are Same subject: A humanoid robot is the wrong tool on a factory floor, because a machine specialised in the task beats it. — tap to centre the map on it A humanoid robot is the wrong toolon a factory floor, because amachine specialised in the taskbeats it. Last stated a year ago 6 Apr 2025 MH Molson Hart — holds since 2025-04-06 — tap for who they are Same subject: A physical AI agent has to be safe on day one, because you cannot ship something merely good enough and let users find the edge cases. — tap to centre the map on it A physical AI agent has to be safeon day one, because you cannot shipsomething merely good enough and letusers find the edge cases. Last stated a month ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it After pre-training, post-trainingand test-time scaling, the fourthscaling law is agentic: multiplyingAI by spawning agents, and the wholeloop scales on one thing, compute. Last stated 6 months ago 23 Mar 2026 JH Jensen Huang — holds since 2026-03-23 — tap for who they are Same subject: A five-year-old's vision is already good enough to drive a car, on a tiny and strikingly undiverse amount of data. — tap to centre the map on it A five-year-old's vision is alreadygood enough to drive a car, on atiny and strikingly undiverse amountof data. Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A good robotaxi demonstration in one city does not mean the company is ready to scale the service. — tap to centre the map on it A good robotaxi demonstration in onecity does not mean the company isready to scale the service. Last stated a year ago 8 Jul 2025 TL Timothy B. Lee — holds since 2025-07-08 — tap for who they are Same subject: A huge fleet collecting driving data is not a silver bullet for self-driving, because the data arrives unlabelled. — tap to centre the map on it A huge fleet collecting driving datais not a silver bullet forself-driving, because the dataarrives unlabelled. Last stated 2 years ago 21 May 2024 TL Timothy B. Lee — holds since 2024-05-21 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Learning from a reward ten years away is a solved problem: a value function trained by temporal-difference learning rewards the steps along the way. Last stated 26 Sept 2025 · 11 months ago Holds Richard Sutton Read this korrent →