korrents

On the map

Tap a claim on the ring to put it at the centre.

← RLHF is not a tax on capability: preference tuning also raises maths…

17 connected korrents · 15 moments on record from 1 Oct 2023 to 3 Aug 2026.

Everything filed under reinforcement learning reinforcement learning Everything filed under LLMs LLMs Everything filed under scaling laws scaling laws Everything filed under robotics robotics Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: RLHF is not a tax on capability: preference tuning also raises maths and code scores, which is why the labs keep reaching for it. RLHF is not a tax on capability:preference tuning also raises mathsand code scores, which is why thelabs keep reaching for it. Last stated 2 years ago 3 Feb 2025 NL Nathan Lambert — holds since 2025-02-03 — tap for who they are Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it A verifiable task can beoptimised by reinforcementlearning until a neuralnetwork performs it extremely… Last stated 10 months ago 17 Nov 2025 AK Andrej Karpathy — holds since 2025-11-17 — tap for who they are Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it Capability does not generalisefor free: a model that willmove mountains on an agentictask still tells the same bad… Last stated 6 months ago 20 Mar 2026 AK Andrej Karpathy — holds since 2026-03-20 — tap for who they are Same subject: Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving. — tap to centre the map on it Humans barely usereinforcement learning forintelligence — what RL they douse goes into motor tasks, not… Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: Humans keep their place in AI as judges rather than authors, because telling which of two answers is better is far easier than writing a good one. — tap to centre the map on it Humans keep their place in AIas judges rather than authors,because telling which of twoanswers is better is far… Last stated 2 years ago 3 Feb 2025 NL Nathan Lambert — holds since 2025-02-03 — tap for who they are Same subject: Large language models mimic what people say to do rather than work out what to do, which is why they are not about understanding the world. — tap to centre the map on it Large language models mimicwhat people say to do ratherthan work out what to do,which is why they are not… Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments. — tap to centre the map on it Models look far better onevals than they are in theworld because researchers,inadvertently, take… Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: Models resemble each other because pre-training is the same everywhere; what differentiates labs now is RL and post-training. — tap to centre the map on it Models resemble each otherbecause pre-training is thesame everywhere; whatdifferentiates labs now is RL… Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: No person learns the way RL does: a human reviews which parts of an attempt were good instead of rewarding every step of a lucky one. — tap to centre the map on it No person learns the way RLdoes: a human reviews whichparts of an attempt were goodinstead of rewarding every… Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a statedbudget, or as a curve againsttest-time compute — never as a… Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research ratherthan on building the nextmodel, because research is… Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: After pre-training, post-training and test-time scaling, the fourth scaling law is agentic: multiplying AI by spawning agents, and the whole loop scales on one thing, compute. — tap to centre the map on it After pre-training,post-training and test-timescaling, the fourth scalinglaw is agentic: multiplying AI… Last stated 6 months ago 23 Mar 2026 JH Jensen Huang — holds since 2026-03-23 — tap for who they are Same subject: A highly dynamic economy that cheaply mass-produces robots may increase rather than decrease the risk of Malthusian resource dilemmas. — tap to centre the map on it A highly dynamic economy thatcheaply mass-produces robotsmay increase rather thandecrease the risk of… Last stated 3 years ago 1 Oct 2023 TC Tyler Cowen — holds since 2023-10-01 — tap for who they are Same subject: A humanoid robot is the wrong tool on a factory floor, because a machine specialised in the task beats it. — tap to centre the map on it A humanoid robot is the wrongtool on a factory floor,because a machine specialisedin the task beats it. Last stated a year ago 6 Apr 2025 MH Molson Hart — holds since 2025-04-06 — tap for who they are Same subject: A physical AI agent has to be safe on day one, because you cannot ship something merely good enough and let users find the edge cases. — tap to centre the map on it A physical AI agent has to besafe on day one, because youcannot ship something merelygood enough and let users find… Last stated a month ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it A language model is nosubstitute for awell-specified conventionalalgorithm, so it cannot simply… Last stated a year ago 7 Jun 2025 GM Gary Marcus — holds since 2025-06-07 — tap for who they are Same subject: A language model is not using language at all, because language requires an intention to communicate. — tap to centre the map on it A language model is not usinglanguage at all, becauselanguage requires an intentionto communicate. Last stated 2 years ago 31 Aug 2024 TC Ted Chiang — holds since 2024-08-31 — tap for who they are Same subject: A language model's apparent mind is mostly our own bias: it predicts text, and leverages our evolved habit of attributing intentionality to anything that acts human. — tap to centre the map on it A language model's apparentmind is mostly our own bias:it predicts text, andleverages our evolved habit of… Last stated 2 years ago 22 Apr 2024 SC Sean Carroll — holds since 2024-04-22 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: how recently it was last stated — full and dark this week, a faint sliver at five yearsa face: someone on record holding the claim — tap it for who they are

At the centre RLHF is not a tax on capability: preference tuning also raises maths and code scores, which is why the labs keep reaching for it. Last stated 3 Feb 2025 · 2 years ago Holds Nathan Lambert Read this korrent →