Tap a claim on the ring to put it at the centre.
← Research on transformer-style attention mechanisms has not yet reached…
17 connected korrents · 15 moments from 12 Jun 2017 to 18 Sept 2026.
Everything filed under LLMs
LLMs
Everything filed under transformers
transformers
Everything filed under innovation
innovation
Everything filed under OpenAI
OpenAI
Everything filed under behavioural science
behavioural science
Everything filed under benchmarks
benchmarks
Everything filed under AI agents
AI agents
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Research on transformer-style attention mechanisms has not yet reached diminishing returns compared to alternative approaches.
Research on transformer-style attention mechanisms has not yet reached diminishing returns compared to alternative approaches.
Last stated 5 years ago
3 Jun 2021
GW
Gwern — holds since 2021-06-03 — tap for who they are
Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it
Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive.
Last stated 3 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: The current capabilities of existing AI models are barely being used and are often not even well understood. — tap to centre the map on it
The current capabilities of existing AI models are barely being used and are often not even well understood.
Last stated 2 weeks ago
18 Sept 2026
EM
Ethan Mollick — holds since 2026-09-18 — tap for who they are
Same subject: Shrinking attention spans are reshaping real-world business models, not just digital media formats. — tap to centre the map on it
Shrinking attention spans are reshaping real-world business models, not just digital media formats.
Last stated 7 months ago
16 Feb 2026
SW
swyx — holds since 2026-02-16 — tap for who they are
Same subject: Hybrid and sparse-attention model architectures will become more widely adopted as the ecosystem catches up. — tap to centre the map on it
Hybrid and sparse-attention model architectures will become more widely adopted as the ecosystem catches up.
Last stated 3 weeks ago
8 Sept 2026
NL
Nathan Lambert — holds since 2026-09-08 — tap for who they are
Same subject: Computer-use agents had to wait for language models: without pre-trained representations the reward is too sparse to ever learn from. — tap to centre the map on it
Computer-use agents had to wait for language models: without pre-trained representations the reward is too sparse to ever learn from.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: Criticism that a simpler, scaled-up architecture underperforms a heavily hand-optimized baseline actually demonstrates why that architecture-scaling research matters. — tap to centre the map on it
Criticism that a simpler, scaled-up architecture underperforms a heavily hand-optimized baseline actually demonstrates why that architecture-scaling research matters.
Last stated 5 years ago
11 Jun 2021
GW
Gwern — holds since 2021-06-11 — tap for who they are
Same subject: The scarce skill is now context switching across parallel agents, not sustained deep work. — tap to centre the map on it
The scarce skill is now context switching across parallel agents, not sustained deep work.
Last stated 7 months ago
4 Mar 2026
BC
Boris Cherny — holds since 2026-03-04 — tap for who they are
Same subject: We should want more tools and fewer operated machines; the real AI game changers will have little to do with plain content generation. — tap to centre the map on it
We should want more tools and fewer operated machines; the real AI game changers will have little to do with plain content generation.
Last stated 3 years ago
by 1 May 2023
AW
Amelia Wattenberger — holds since 2023-05-01 — tap for who they are
Same subject: "Lots of anonymous people worked it out by trial and error" is never the real story of an invention: on closer investigation, identifiable people always turn out to be responsible. — tap to centre the map on it
"Lots of anonymous people worked it out by trial and error" is never the real story of an invention: on closer investigation, identifiable people always turn out to be responsible.
Last stated 3 years ago
25 Oct 2023
AH
Anton Howes — holds since 2023-10-25 — tap for who they are
Same subject: A business can have market fit, real advantages, and sustained effort behind a hard product and still fail if it doesn't give customers a strong enough reason to switch from what they already use. — tap to centre the map on it
A business can have market fit, real advantages, and sustained effort behind a hard product and still fail if it doesn't give customers a strong enough reason to switch from what they already use.
Last stated 4 months ago
2 Jun 2026
JJ
Justin Jackson — holds since 2026-06-02 — tap for who they are
Same subject: A capability existing is not the same as a capability being absorbed: how fast a technology changes the world is not set by how fast the technology improves. — tap to centre the map on it
A capability existing is not the same as a capability being absorbed: how fast a technology changes the world is not set by how fast the technology improves.
Last stated 2 months ago
21 Jul 2026
RT
Ruxandra Teslo — holds since 2026-07-21 — tap for who they are
Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more
Same subject: Self-attention can yield more interpretable models. — tap to centre the map on it
Self-attention can yield more interpretable models.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more
Same subject: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. — tap to centre the map on it
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more
Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 3 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 4 weeks ago
3 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are
Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it
Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers.
Last stated 3 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are
At the centre
Research on transformer-style attention mechanisms has not yet reached diminishing returns compared to alternative approaches.
Last stated 3 Jun 2021 · 5 years ago
Holds Gwern
Read this korrent →
Similar wording
Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive.
Last stated 15 Jul 2026 · 3 months ago
Holds Dex Horthy
Similar wording
The current capabilities of existing AI models are barely being used and are often not even well understood.
Last stated 18 Sept 2026 · 2 weeks ago
Holds Ethan Mollick
Similar wording
Shrinking attention spans are reshaping real-world business models, not just digital media formats.
Last stated 16 Feb 2026 · 7 months ago
Holds swyx
Similar wording
Hybrid and sparse-attention model architectures will become more widely adopted as the ecosystem catches up.
Last stated 8 Sept 2026 · 3 weeks ago
Holds Nathan Lambert
Similar wording
Computer-use agents had to wait for language models: without pre-trained representations the reward is too sparse to ever learn from.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Similar wording
Criticism that a simpler, scaled-up architecture underperforms a heavily hand-optimized baseline actually demonstrates why that architecture-scaling research matters.
Last stated 11 Jun 2021 · 5 years ago
Holds Gwern
Similar wording
The scarce skill is now context switching across parallel agents, not sustained deep work.
Last stated 4 Mar 2026 · 7 months ago
Holds Boris Cherny
Similar wording
We should want more tools and fewer operated machines; the real AI game changers will have little to do with plain content generation.
Last stated by 1 May 2023 · 3 years ago
Holds Amelia Wattenberger
Same subject: innovation
"Lots of anonymous people worked it out by trial and error" is never the real story of an invention: on closer investigation, identifiable people always turn out to be responsible.
Last stated 25 Oct 2023 · 3 years ago
Holds Anton Howes
Same subject: innovation
A business can have market fit, real advantages, and sustained effort behind a hard product and still fail if it doesn't give customers a strong enough reason to switch from what they already use.
Last stated 2 Jun 2026 · 4 months ago
Holds Justin Jackson
Same subject: innovation
A capability existing is not the same as a capability being absorbed: how fast a technology changes the world is not set by how fast the technology improves.
Last stated 21 Jul 2026 · 2 months ago
Holds Ruxandra Teslo
Same subject: transformers
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 12 Jun 2017 · 9 years ago
Holds Illia Polosukhin Jakob Uszkoreit Niki Parmar Ashish Vaswani Noam Shazeer Aidan N. Gomez Llion Jones Łukasz Kaiser
Same subject: transformers
Self-attention can yield more interpretable models.
Last stated 12 Jun 2017 · 9 years ago
Holds Illia Polosukhin Jakob Uszkoreit Niki Parmar Ashish Vaswani Noam Shazeer Aidan N. Gomez Llion Jones Łukasz Kaiser
Same subject: transformers
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Last stated 12 Jun 2017 · 9 years ago
Holds Illia Polosukhin Jakob Uszkoreit Niki Parmar Ashish Vaswani Noam Shazeer Aidan N. Gomez Llion Jones Łukasz Kaiser
Same subject: OpenAI
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 18 Mar 2024 · 3 years ago
Holds Sam Altman
Same subject: OpenAI
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 3 Sept 2026 · 4 weeks ago
Holds Zvi Mowshowitz
Same subject: OpenAI
Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers.
Last stated 18 Mar 2024 · 3 years ago
Holds Sam Altman