korrents

← People and their mental models

Sebastian Raschka's mental models

3 claims Sebastian Raschka made fit 3 mental models. Most often: Bottlenecks, S-curves and diminishing returns, The map is not the territory. Everything they said here.

Models we see in what they say

Our reading: their claim applies the idea without naming it. The claim is theirs; filing it here is ours.

Bottlenecks

A system moves only as fast as its narrowest point; speed up anything else and nothing changes.

Used by 88 others

As reasoning models and agents retain more tokens for longer, KV-cache size and attention cost become the main constraint on scaling context.

  1. Sebastian Raschka AI research engineer working on large language models As reasoning models and agent workflows keep more tokens around (for longer), KV-cache size, memory traffic, and attention cost quickly become the main constraints, and LLM developers are adding a growing number of architecture tricks to reduce those costs. Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attentionmagazine.sebastianraschka.com · 16 May 2026All korrents from this piece
    As reasoning models and agent workflows keep more tokens around (for longer), KV-cache size, memory traffic, and attention cost quickly become the main constraints, and LLM developers are adding a growing number of architecture tricks to reduce those costs.

S-curves and diminishing returns

Most growth slows as it matures; the question is whether the curve is flattening or something new is restocking it.

Used by 22 others

Pushing an LLM's reasoning budget higher eventually yields diminishing returns and becomes uneconomical.

  1. Sebastian Raschka AI research engineer working on large language models This saturation can be seen more clearly for the GPT 5.6 Sol model, which also shows that increasing reasoning budgets can become uneconomical at some point. Controlling Reasoning Effort in LLMsmagazine.sebastianraschka.com · 18 Jul 2026All korrents from this piece
    This saturation can be seen more clearly for the GPT 5.6 Sol model, which also shows that increasing reasoning budgets can become uneconomical at some point.

The map is not the territory

Every model, word or plan is a simplified map of reality; useful when its structure matches, misleading when mistaken for the thing itself.

Used by 55 others

An AI-detector's score reflects the classifier's training distribution, not a true probability that a given text was written by AI.

  1. Sebastian Raschka AI research engineer working on large language models the score is the classifier's estimated probability for the AI-generated class based on its training distribution. However, we shouldn't interpreted it as a general probability that the text was written by AI. Building an AI Text Detector From Scratchmagazine.sebastianraschka.com · 15 Aug 2026All korrents from this piece
    the score is the classifier's estimated probability for the AI-generated class based on its training distribution. However, we shouldn't interpreted it as a general probability that the text was written by AI.