korrents

On the map

Tap a claim on the ring to put it at the centre.

← A layer should learn a residual with reference to its own input rather…

17 connected korrents · 13 moments on record from 22 Dec 2014 to 3 Sept 2026.

Everything filed under benchmarks benchmarks Everything filed under measuring intelligence measuring intelligence Everything filed under neural networks neural networks Everything filed under reinforcement learning reinforcement learning Everything filed under optimizers optimizers Everything filed under LLMs LLMs Everything filed under mathematics mathematics Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable. A layer should learn a residual withreference to its own input rather than anunreferenced function, which is what makesgreat depth trainable. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases. — tap to centre the map on it Residual networks are easier tooptimize than plain ones and keepgaining accuracy as depth increases. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. — tap to centre the map on it An optimizer that is invariant torescaling the gradients and cheap inmemory is what makes very largemodels practical to train. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it Current techniques restrictpre-trained representation powerbecause standard language models areunidirectional. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: The knowledge a model soaks up in pre-training is holding it back; what we actually want is the intelligence with the knowledge stripped out. — tap to centre the map on it The knowledge a model soaks up inpre-training is holding it back;what we actually want is theintelligence with the knowledgestripped out. Last stated 11 months ago 17 Oct 2025 AK Andrej Karpathy — holds since 2025-10-17 — tap for who they are Same subject: Depth is not free: past a point, deeper neural networks are harder to train rather than simply better. — tap to centre the map on it Depth is not free: past a point,deeper neural networks are harder totrain rather than simply better. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: Machine learning is a shallow field compared with mathematics: even its most important ideas can be explained in a couple of minutes. — tap to centre the map on it Machine learning is a shallow fieldcompared with mathematics: even itsmost important ideas can beexplained in a couple of minutes. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning. — tap to centre the map on it We have almost no automatedtechniques for making a systemtransfer what it learns, and none ofthe few we have are used in moderndeep learning. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. — tap to centre the map on it A deep bidirectional model isstrictly more powerful than aleft-to-right model or a shallowconcatenation of unidirectionalmodels. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it A physical AI company needs threeAIs, not one — the agent, thesimulator and the critic — turningdeployment into a flywheel. Last stated a month ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it Building a realistic simulator isexactly as hard as building theagent, because the simulator isitself a large AI model. Last stated a month ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it Capability does not generalise forfree: a model that will movemountains on an agentic task stilltells the same bad joke it told fiveyears ago. Last stated 6 months ago 20 Mar 2026 AK Andrej Karpathy — holds since 2026-03-20 — tap for who they are Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it If intelligence is the process ofacquiring skills, no single taskdemonstrates intelligence unless itis a meta-task of skill-acquisitionacross many tasks. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence. — tap to centre the map on it Language models are already a formof AGI; what the labs are chasing isa further step, not the arrival ofgeneral intelligence. Last stated 2 years ago 3 Feb 2025 NL Nathan Lambert — holds since 2025-02-03 — tap for who they are Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it Solely measuring skill at any giventask falls short of measuringintelligence, because skill isheavily modulated by prior knowledgeand experience. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated a week ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 5 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it A speed difference between languagesthat share the LLVM backend measuresthe benchmark author, not thelanguages. Last stated a year ago 22 Mar 2025 TH ThePrimeagen — holds since 2025-03-22 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable. Last stated 10 Dec 2015 · 11 years ago Holds Jian SunKaiming He Read this korrent →