korrents

On the map

Tap a claim on the ring to put it at the centre.

← Looped/recurrent-depth transformer architectures can improve model…

17 connected korrents · 17 moments on record from 10 Dec 2015 to 15 Sept 2026.

Everything filed under scaling laws scaling laws Everything filed under measuring intelligence measuring intelligence Everything filed under reinforcement learning reinforcement learning Everything filed under LLMs LLMs Everything filed under neural networks neural networks Everything filed under transformers transformers Everything filed under recursive self-improvement recursive self-improvement Everything filed under robotics robotics Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough. Looped/recurrent-depth transformerarchitectures can improve model quality ata fixed compute budget when the model islarge enough. Last stated 2 weeks ago 9 Sept 2026 SR Sebastian Raschka — holds since 2026-09-09 — tap for who they are Same subject: Test-time scaling has gained a third axis of latent space reasoning iterations in looped transformers. — tap to centre the map on it Test-time scaling has gained a thirdaxis of latent space reasoningiterations in looped transformers. Last stated 2 weeks ago 5 Sept 2026 FC François Chollet — holds since 2026-09-05 — tap for who they are Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it Scaling up language models greatlyimproves task-agnostic few-shotperformance, sometimes matchingprior fine-tuning approaches. Last stated 6 years ago 28 May 2020 DA Dario Amodei — holds since 2020-05-28 — tap for who they are Same subject: Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases. — tap to centre the map on it Residual networks are easier tooptimize than plain ones and keepgaining accuracy as depth increases. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. — tap to centre the map on it Sequence transduction can use asimple architecture based solely onattention, without recurrence orconvolutions. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more on record Same subject: Larger language models already have enough capacity that extra per-layer embedding tricks add little benefit to them. — tap to centre the map on it Larger language models already haveenough capacity that extra per-layerembedding tricks add little benefitto them. Last stated 4 months ago 16 May 2026 SR Sebastian Raschka — holds since 2026-05-16 — tap for who they are Same subject: As models get better at following instructions, techniques like finetuning and constrained sampling for structured outputs will become less necessary — tap to centre the map on it As models get better at followinginstructions, techniques likefinetuning and constrained samplingfor structured outputs will becomeless necessary Last stated 3 years ago 16 Jan 2024 CH Chip Huyen — holds since 2024-01-16 — tap for who they are Same subject: There is no real impediment to models running their own research loop and improving themselves at a far more rapid rate. — tap to centre the map on it There is no real impediment tomodels running their own researchloop and improving themselves at afar more rapid rate. Last stated 2 months ago 30 Jul 2026 JD Jeff Dean — holds since 2026-07-30 — tap for who they are Same subject: Letting a robot imagine the next image helps, but it is not essential — the model was surprisingly good without it. — tap to centre the map on it Letting a robot imagine the nextimage helps, but it is not essential— the model was surprisingly goodwithout it. Last stated a month ago 12 Aug 2026 CF Chelsea Finn — holds since 2026-08-12 — tap for who they are Same subject: "Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next. — tap to centre the map on it "Scaling" was powerful because itwas one word: naming a researchdirection is what tells a wholefield what to do next. Last stated 10 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it A benchmark result should bereported under a stated budget, oras a curve against test-time compute— never as a single number. Last stated 3 months ago 26 Jun 2026 NB Noam Brown — holds since 2026-06-26 — tap for who they are Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it A lab should spend most of itscompute on research rather than onbuilding the next model, becauseresearch is where the tenfold yearlyefficiency gains come from. Last stated 6 months ago 13 Mar 2026 DP Dylan Patel — holds since 2026-03-13 — tap for who they are Same subject: A true artificial general intelligence cannot exist without being recognized as a moral subject. — tap to centre the map on it A true artificial generalintelligence cannot exist withoutbeing recognized as a moral subject. Last stated a year ago 10 Jun 2025 SH Samuel Hammond — holds since 2025-06-10 — tap for who they are Same subject: A unit of AI inference needs to be defined, for example via a chain of increasingly hard problems where each consecutive pair is solvable by one model. — tap to centre the map on it A unit of AI inference needs to bedefined, for example via a chain ofincreasingly hard problems whereeach consecutive pair is solvable byone model. Last stated 6 days ago 15 Sept 2026 PG Paul Graham — holds since 2026-09-15 — tap for who they are Same subject: Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect. — tap to centre the map on it Both the special-purpose-programsview and the blank-slate view ofhuman intelligence are likelyincorrect. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it A physical AI company needs threeAIs, not one — the agent, thesimulator and the critic — turningdeployment into a flywheel. Last stated 2 months ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it A verifiable task can be optimisedby reinforcement learning until aneural network performs it extremelywell. Last stated 10 months ago 17 Nov 2025 AK Andrej Karpathy — holds since 2025-11-17 — tap for who they are Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it Building a realistic simulator isexactly as hard as building theagent, because the simulator isitself a large AI model. Last stated 2 months ago 3 Aug 2026 DD Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough. Last stated 9 Sept 2026 · 2 weeks ago Holds Sebastian Raschka Read this korrent →