korrents

On the map

Tap a claim on the ring to put it at the centre.

← An optimizer that is invariant to rescaling the gradients and cheap in…

8 connected korrents · 4 moments on record from 22 Dec 2014 to 26 Sept 2025.

Everything filed under optimizers optimizers Everything filed under LLMs LLMs Everything filed under neural networks neural networks Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. An optimizer that is invariant torescaling the gradients and cheap inmemory is what makes very large modelspractical to train. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: An optimizer's hyper-parameters should mean something a practitioner can reason about, and should rarely need tuning. — tap to centre the map on it An optimizer's hyper-parametersshould mean something a practitionercan reason about, and should rarelyneed tuning. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: Gradient descent finds a solution to the problems a model has seen; nothing in the algorithm makes it pick the one that generalises well. — tap to centre the map on it Gradient descent finds a solution tothe problems a model has seen;nothing in the algorithm makes itpick the one that generalises well. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments. — tap to centre the map on it Stochastic optimization can be donewell using only adaptive estimatesof the gradient's lower-ordermoments. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: The same optimizer should handle objectives that move under it and gradients that are noisy or sparse. — tap to centre the map on it The same optimizer should handleobjectives that move under it andgradients that are noisy or sparse. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable. — tap to centre the map on it A layer should learn a residual withreference to its own input ratherthan an unreferenced function, whichis what makes great depth trainable. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases. — tap to centre the map on it Residual networks are easier tooptimize than plain ones and keepgaining accuracy as depth increases. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: Large language models start in exactly the wrong place, because they try to get by without a goal and therefore without any sense of better or worse. — tap to centre the map on it Large language models start inexactly the wrong place, becausethey try to get by without a goaland therefore without any sense ofbetter or worse. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it Scaling up language models greatlyimproves task-agnostic few-shotperformance, sometimes matchingprior fine-tuning approaches. Last stated 6 years ago 28 May 2020 DA Dario Amodei — holds since 2020-05-28 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. Last stated 22 Dec 2014 · 12 years ago Holds Jimmy BaDiederik P. Kingma Read this korrent →