korrents

On the map

Tap a claim on the ring to put it at the centre.

← Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments.

8 connected korrents · 5 moments on record from 22 Dec 2014 to 26 Sept 2025.

Everything filed under optimizers optimizers Everything filed under LLMs LLMs Everything filed under neural networks neural networks Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments. Stochastic optimization can be done wellusing only adaptive estimates of thegradient's lower-order moments. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. — tap to centre the map on it An optimizer that is invariant torescaling the gradients and cheap inmemory is what makes very largemodels practical to train. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: An optimizer's hyper-parameters should mean something a practitioner can reason about, and should rarely need tuning. — tap to centre the map on it An optimizer's hyper-parametersshould mean something a practitionercan reason about, and should rarelyneed tuning. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: Gradient descent finds a solution to the problems a model has seen; nothing in the algorithm makes it pick the one that generalises well. — tap to centre the map on it Gradient descent finds a solution tothe problems a model has seen;nothing in the algorithm makes itpick the one that generalises well. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: The same optimizer should handle objectives that move under it and gradients that are noisy or sparse. — tap to centre the map on it The same optimizer should handleobjectives that move under it andgradients that are noisy or sparse. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases. — tap to centre the map on it Residual networks are easier tooptimize than plain ones and keepgaining accuracy as depth increases. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: Pre-trained representations reduce the need for many heavily-engineered task-specific architectures. — tap to centre the map on it Pre-trained representations reducethe need for many heavily-engineeredtask-specific architectures. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it Current techniques restrictpre-trained representation powerbecause standard language models areunidirectional. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it Scaling up language models greatlyimproves task-agnostic few-shotperformance, sometimes matchingprior fine-tuning approaches. Last stated 6 years ago 28 May 2020 DA Dario Amodei — holds since 2020-05-28 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments. Last stated 22 Dec 2014 · 12 years ago Holds Jimmy BaDiederik P. Kingma Read this korrent →