korrents

On the map

Tap a claim on the ring to put it at the centre.

← An optimizer's hyper-parameters should mean something a practitioner…

5 connected korrents · 3 moments on record from 22 Dec 2014 to 26 Aug 2026.

Everything filed under optimizers optimizers Everything filed under software performance software performance Same subjectSame subjectSame subjectSame subjectSame subject Read this korrent: An optimizer's hyper-parameters should mean something a practitioner can reason about, and should rarely need tuning. An optimizer's hyper-parameters shouldmean something a practitioner can reasonabout, and should rarely need tuning. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. — tap to centre the map on it An optimizer that is invariant torescaling the gradients and cheap inmemory is what makes very largemodels practical to train. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: Gradient descent finds a solution to the problems a model has seen; nothing in the algorithm makes it pick the one that generalises well. — tap to centre the map on it Gradient descent finds a solution tothe problems a model has seen;nothing in the algorithm makes itpick the one that generalises well. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments. — tap to centre the map on it Stochastic optimization can be donewell using only adaptive estimatesof the gradient's lower-ordermoments. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: The same optimizer should handle objectives that move under it and gradients that are noisy or sparse. — tap to centre the map on it The same optimizer should handleobjectives that move under it andgradients that are noisy or sparse. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: Optimization means measuring the gap between your code and what the hardware could theoretically do, not making changes and checking whether the statistics improved. — tap to centre the map on it Optimization means measuring the gapbetween your code and what thehardware could theoretically do, notmaking changes and checking whetherthe statistics improved. Last stated 2 weeks ago 26 Aug 2026 CM Casey Muratori — holds since 2026-08-26 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre An optimizer's hyper-parameters should mean something a practitioner can reason about, and should rarely need tuning. Last stated 22 Dec 2014 · 12 years ago Holds Jimmy BaDiederik P. Kingma Read this korrent →