korrents

On the map

Tap a claim on the ring to put it at the centre.

← A deep bidirectional model is strictly more powerful than a…

17 connected korrents · 12 moments on record from 10 Jun 2014 to 11 Aug 2026.

Everything filed under scaling laws scaling laws Everything filed under measuring intelligence measuring intelligence Everything filed under reinforcement learning reinforcement learning Everything filed under LLMs LLMs Everything filed under transformers transformers Everything filed under neural networks neural networks Everything filed under mathematics mathematics Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. A deep bidirectional model is strictlymore powerful than a left-to-right modelor a shallow concatenation ofunidirectional models. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it Current techniques restrictpre-trained representation powerbecause standard language models areunidirectional. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: Unidirectional restrictions are sub-optimal for sentence-level tasks and harmful for token-level tasks that need bidirectional context. — tap to centre the map on it Unidirectional restrictions aresub-optimal for sentence-level tasksand harmful for token-level tasksthat need bidirectional context. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it A transduction model can relyentirely on self-attention for inputand output representations withoutsequence-aligned RNNs orconvolution. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more on record Same subject: Depth is not free: past a point, deeper neural networks are harder to train rather than simply better. — tap to centre the map on it Depth is not free: past a point,deeper neural networks are harder totrain rather than simply better. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: Machine learning is a shallow field compared with mathematics: even its most important ideas can be explained in a couple of minutes. — tap to centre the map on it Machine learning is a shallow fieldcompared with mathematics: even itsmost important ideas can beexplained in a couple of minutes. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: What sits in a model's context is far clearer to it than its training data, which is trillions of tokens stirred into a soup of parameters. — tap to centre the map on it What sits in a model's context isfar clearer to it than its trainingdata, which is trillions of tokensstirred into a soup of parameters. Last stated a month ago 30 Jul 2026 JD Jeff Dean — holds since 2026-07-30 — tap for who they are Same subject: Deep generative models have had less impact than discriminative models because of intractable probabilistic computations and difficulty using piecewise linear units. — tap to centre the map on it Deep generative models have had lessimpact than discriminative modelsbecause of intractable probabilisticcomputations and difficulty usingpiecewise linear units. Last stated 12 years ago 10 Jun 2014 YB Yoshua Bengio — holds since 2014-06-10 — tap for who they are Same subject: Large language models start in exactly the wrong place, because they try to get by without a goal and therefore without any sense of better or worse. — tap to centre the map on it Large language models start inexactly the wrong place, becausethey try to get by without a goaland therefore without any sense ofbetter or worse. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget. — tap to centre the map on it A model's capability is now afunction of how much money you spendon it, so asking what a model can domeans nothing until you name abudget. Last stated 2 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it If intelligence is the process ofacquiring skills, no single taskdemonstrates intelligence unless itis a meta-task of skill-acquisitionacross many tasks. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it Solely measuring skill at any giventask falls short of measuringintelligence, because skill isheavily modulated by prior knowledgeand experience. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it The Abstraction and Reasoning Corpuscan measure human-like general fluidintelligence and enable faircomparisons between AI systems andhumans. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it A physical AI company needs threeAIs, not one — the agent, thesimulator and the critic — turningdeployment into a flywheel. Last stated a month ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it A verifiable task can be optimisedby reinforcement learning until aneural network performs it extremelywell. Last stated 10 months ago 17 Nov 2025 AK Andrej Karpathy — holds since 2025-11-17 — tap for who they are Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it Building a realistic simulator isexactly as hard as building theagent, because the simulator isitself a large AI model. Last stated a month ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. Last stated 11 Oct 2018 · 8 years ago Holds Kristina ToutanovaMing-Wei ChangKenton LeeJacob Devlin Read this korrent →