korrents

On the map

Tap a claim on the ring to put it at the centre.

← It is fundamentally difficult to design a reward function that…

17 connected korrents · 16 moments from 28 Nov 2024 to 21 Sept 2026. Nearly all of them are about reinforcement learning.

Everything filed under AI alignment AI alignment Everything filed under design design Everything filed under LLMs LLMs Everything filed under AGI AGI Everything filed under neural networks neural networks Everything filed under benchmarks benchmarks Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: It is fundamentally difficult to design a reward function that accurately captures the intended goal in reinforcement learning. It is fundamentally difficult to design areward function that accurately capturesthe intended goal in reinforcementlearning. Last stated 2 years ago 28 Nov 2024 LW Lilian Weng — holds since 2024-11-28 — tap for who they are Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it A verifiable task can be optimisedby reinforcement learning until aneural network performs it extremelywell. Last stated 10 months ago 17 Nov 2025 AK Andrej Karpathy — holds since 2025-11-17 — tap for who they are Same subject: Better execution environments are needed for LLMs to properly utilize their ability to evolve systems. — tap to centre the map on it Better execution environments areneeded for LLMs to properly utilizetheir ability to evolve systems. Last stated 2 weeks ago 21 Sept 2026 TL Tobias Lütke — holds since 2026-09-21 — tap for who they are Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it Capability does not generalise forfree: a model that will movemountains on an agentic task stilltells the same bad joke it told fiveyears ago. Last stated 6 months ago 20 Mar 2026 AK Andrej Karpathy — holds since 2026-03-20 — tap for who they are Same subject: Current models are unsafe because they are not smart enough, not because they are too smart. — tap to centre the map on it Current models are unsafe becausethey are not smart enough, notbecause they are too smart. Last stated 3 weeks ago 13 Sept 2026 FC François Chollet — holds since 2026-09-13 — tap for who they are Same subject: Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving. — tap to centre the map on it Humans barely use reinforcementlearning for intelligence — what RLthey do use goes into motor tasks,not problem solving. Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: Humans keep their place in AI as judges rather than authors, because telling which of two answers is better is far easier than writing a good one. — tap to centre the map on it Humans keep their place in AI asjudges rather than authors, becausetelling which of two answers isbetter is far easier than writing agood one. Last stated 2 years ago 3 Feb 2025 NL Nathan Lambert — holds since 2025-02-03 — tap for who they are Same subject: Large language models mimic what people say to do rather than work out what to do, which is why they are not about understanding the world. — tap to centre the map on it Large language models mimic whatpeople say to do rather than workout what to do, which is why theyare not about understanding theworld. Last stated a year ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: Models look far better on evals than they are in the world because researchers, inadvertently, take inspiration from the evals when they build RL environments. — tap to centre the map on it Models look far better on evals thanthey are in the world becauseresearchers, inadvertently, takeinspiration from the evals when theybuild RL environments. Last stated 10 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A blank chat box is lazy: it violates the first rule of a good user experience, that it is obvious what you can do. — tap to centre the map on it A blank chat box is lazy: itviolates the first rule of a gooduser experience, that it is obviouswhat you can do. Last stated a year ago 12 May 2025 JZ Julie Zhuo — holds since 2025-05-12 — tap for who they are Same subject: A board game conquers the world when chance and strategy are in balance in it. — tap to centre the map on it A board game conquers the world whenchance and strategy are in balancein it. Last stated 10 months ago 12 Dec 2025 IF Irving Finkel — holds since 2025-12-12 — tap for who they are Same subject: A browser works better as a single window than as multiple separate windows. — tap to centre the map on it A browser works better as a singlewindow than as multiple separatewindows. Last stated 2 weeks ago 15 Sept 2026 CC Chris Coyier — holds since 2026-09-15 — tap for who they are Same subject: A recurring AI workflow improves more through ongoing user feedback than by being fully specified upfront. — tap to centre the map on it A recurring AI workflow improvesmore through ongoing user feedbackthan by being fully specifiedupfront. Last stated 5 months ago 18 May 2026 JL Jason Liu — holds since 2026-05-18 — tap for who they are Same subject: AI agents should notice useful things the user did not think to ask about, not simply complete the request. — tap to centre the map on it AI agents should notice usefulthings the user did not think to askabout, not simply complete therequest. Last stated 2 weeks ago 21 Sept 2026 LR Lenny Rachitsky — holds since 2026-09-21 — tap for who they are Same subject: AI agents will end traditional UI design and accessibility, because users will stop visiting websites and act only through their agent. — tap to centre the map on it AI agents will end traditional UIdesign and accessibility, becauseusers will stop visiting websitesand act only through their agent. Last stated 2 years ago 21 Feb 2025 JN Jakob Nielsen — holds since 2025-02-21 — tap for who they are Same subject: For now AI is best thought of as average intelligence, which is still very powerful to have across every specialty. — tap to centre the map on it For now AI is best thought of asaverage intelligence, which is stillvery powerful to have across everyspecialty. Last stated 3 weeks ago 10 Sept 2026 EV Elena Verna — holds since 2026-09-10 — tap for who they are Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it A country outside the AI supplychain should just buy the index —which works only in the world whereAI ends up commoditised rather thanconcentrated. Last stated 4 months ago 4 Jun 2026 AI Alex Imas — holds since 2026-06-04 — tap for who they are Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it A poor country should prioritiseowning a piece of AI over retrainingits workers, but it should not beteverything on that. Last stated 4 months ago 4 Jun 2026 PT Phil Trammell — holds since 2026-06-04 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre It is fundamentally difficult to design a reward function that accurately captures the intended goal in reinforcement learning. Last stated 28 Nov 2024 · 2 years ago Holds Lilian Weng Read this korrent →