Tap a claim on the ring to put it at the centre.
← Graph transformer networks allow multimodule document recognition…
13 connected korrents · 11 moments on record from 1 Jan 1998 to 15 Jul 2026.
Everything filed under measuring intelligence
measuring intelligence
Everything filed under pruning
pruning
Everything filed under transformers
transformers
Everything filed under neural networks
neural networks
Everything filed under transformers
transformers
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Graph transformer networks allow multimodule document recognition systems to be trained globally with gradient methods to minimize overall performance.
Graph transformer networks allow multimodule document recognition systems to be trained globally with gradient methods to minimize overall performance.
Last stated 29 years ago
1 Jan 1998
YB
Yoshua Bengio — holds since 1998-01-01 — tap for who they are
Same subject: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. — tap to centre the map on it
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: Work on large language models is not progress toward artificial general intelligence, because the models do not understand what they read. — tap to centre the map on it
Work on large language models is not progress toward artificial general intelligence, because the models do not understand what they read.
Last stated a year ago
30 Jul 2025
AG
Alexey Guzey — holds since 2024-08-09 — tap for who they are
AG
Alexey Guzey — no longer holds since 2025-07-30 — tap for who they are
Same subject: With suitable architecture, gradient-based learning can form decision surfaces that classify high-dimensional patterns such as handwritten characters with minimal preprocessing. — tap to centre the map on it
With suitable architecture, gradient-based learning can form decision surfaces that classify high-dimensional patterns such as handwritten characters with minimal preprocessing.
Last stated 29 years ago
1 Jan 1998
YB
Yoshua Bengio — holds since 1998-01-01 — tap for who they are
Same subject: Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases. — tap to centre the map on it
Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases.
Last stated 11 years ago
10 Dec 2015
JS
Jian Sun — holds since 2015-12-10 — tap for who they are
KH
Kaiming He — holds since 2015-12-10 — tap for who they are
Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it
Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive.
Last stated 2 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: Self-attention can yield more interpretable models. — tap to centre the map on it
Self-attention can yield more interpretable models.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence. — tap to centre the map on it
Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence.
Last stated 2 years ago
3 Feb 2025
NL
Nathan Lambert — holds since 2025-02-03 — tap for who they are
Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it. — tap to centre the map on it
A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it.
Last stated 9 years ago
9 Mar 2018
MC
Michael Carbin — holds since 2018-03-09 — tap for who they are
JF
Jonathan Frankle — holds since 2018-03-09 — tap for who they are
Same subject: A winning ticket wins on its initial weights: the connections it keeps started at values that happen to make training work. — tap to centre the map on it
A winning ticket wins on its initial weights: the connections it keeps started at values that happen to make training work.
Last stated 9 years ago
9 Mar 2018
MC
Michael Carbin — holds since 2018-03-09 — tap for who they are
JF
Jonathan Frankle — holds since 2018-03-09 — tap for who they are
Same subject: Batteries will replace transmission as the cheapest way to keep the lights on, and the grid will shrink rather than grow. — tap to centre the map on it
Batteries will replace transmission as the cheapest way to keep the lights on, and the grid will shrink rather than grow.
Last stated 9 months ago
8 Dec 2025
CH
Casey Handmer — holds since 2023-10-11 — tap for who they are
CH
Casey Handmer — holds since 2025-12-08 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 1998 to today (stretched back to the oldest claim here) — full is today a face: someone on record holding the claim — tap it for who they are faded, dashed ring: they no longer hold it — they changed their mind