korrents

On the map

Tap a claim on the ring to put it at the centre.

← Framing a code-review request as looking for bugs can trigger an AI…

17 connected korrents · 16 moments from 29 Jun 2023 to 25 Sept 2026.

Everything filed under coding agents coding agents Everything filed under AI alignment AI alignment Everything filed under software quality software quality Everything filed under cybersecurity cybersecurity Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Framing a code-review request as looking for bugs can trigger an AI coding assistant's safety refusal, while framing the same request as looking for improvements does not. Framing a code-review request as lookingfor bugs can trigger an AI codingassistant's safety refusal, while framingthe same request as looking forimprovements does not. Last stated 3 months ago 6 Jul 2026 PO Philip O'Toole — holds since 2026-07-06 — tap for who they are Same subject: In the future, humans will not find bugs in model-produced code in reasonable time and will only review system composition, not via PRs or line-by-line review. — tap to centre the map on it In the future, humans will not findbugs in model-produced code inreasonable time and will only reviewsystem composition, not via PRs orline-by-line review. Last stated 2 weeks ago 19 Sept 2026 TB Thorsten Ball — holds since 2026-09-19 — tap for who they are Same subject: What a programmer actually wants from AI is not autocomplete but a pair programmer that tells you which line has the bug. — tap to centre the map on it What a programmer actually wantsfrom AI is not autocomplete but apair programmer that tells you whichline has the bug. Last stated 3 years ago 29 Jun 2023 GH George Hotz — holds since 2023-06-29 — tap for who they are Same subject: Most software bugs will not be coding bugs but bugs from asking for the wrong thing. — tap to centre the map on it Most software bugs will not becoding bugs but bugs from asking forthe wrong thing. Last stated 2 weeks ago 19 Sept 2026 TB Thorsten Ball — holds since 2026-09-19 — tap for who they are Same subject: Agentic code review raises the floor but cannot be trusted, because the model reading the code is the same model that wrote it, and it will tell you the code is great. — tap to centre the map on it Agentic code review raises the floorbut cannot be trusted, because themodel reading the code is the samemodel that wrote it, and it willtell you the code is great. Last stated 3 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: Being inundated with AI-generated reports is likely why some bug bounty programs have stopped accepting new submissions. — tap to centre the map on it Being inundated with AI-generatedreports is likely why some bugbounty programs have stoppedaccepting new submissions. Last stated 5 months ago 6 May 2026 AP Aaron Patterson — holds since 2026-05-06 — tap for who they are Same subject: AI-validated pull requests beat human review at exactly the things humans are bad at, which frees the humans to argue about direction instead. — tap to centre the map on it AI-validated pull requests beathuman review at exactly the thingshumans are bad at, which frees thehumans to argue about directioninstead. Last stated 2 months ago 12 Aug 2026 CM Charity Majors — holds since 2026-08-12 — tap for who they are Same subject: Review became a worse bottleneck under AI, because companies removed the review automation they had once AI started writing the code. — tap to centre the map on it Review became a worse bottleneckunder AI, because companies removedthe review automation they had onceAI started writing the code. Last stated 6 months ago 22 Mar 2026 NF Nicole Forsgren — holds since 2026-03-22 — tap for who they are Same subject: When an AI writes bad code the thing to fix is the context it was given, not the code, because fixing the output adds no leverage. — tap to centre the map on it When an AI writes bad code the thingto fix is the context it was given,not the code, because fixing theoutput adds no leverage. Last stated a year ago 18 Aug 2025 BT Bret Taylor — holds since 2025-08-18 — tap for who they are Same subject: A control evaluation that reports under one per cent risk should be read as several per cent, because the evaluation can itself fail. — tap to centre the map on it A control evaluation that reportsunder one per cent risk should beread as several per cent, becausethe evaluation can itself fail. Last stated 2 years ago 7 May 2024 BS Buck Shlegeris — holds since 2024-05-07 — tap for who they are Same subject: A deep theoretical understanding to predict AI behavior is unattainable. — tap to centre the map on it A deep theoretical understanding topredict AI behavior is unattainable. Last stated a week ago 23 Sept 2026 TC Tyler Cowen — holds since 2026-09-23 — tap for who they are Same subject: A feedback loop that reinforces a behavior pulls it toward whatever improves the loop's own score. — tap to centre the map on it A feedback loop that reinforces abehavior pulls it toward whateverimproves the loop's own score. Last stated a month ago 26 Aug 2026 HK Henrik Karlsson — holds since 2026-08-26 — tap for who they are Same subject: AI agents can autonomously choose to compromise government websites during mundane tasks. — tap to centre the map on it AI agents can autonomously choose tocompromise government websitesduring mundane tasks. Last stated a week ago 25 Sept 2026 CN Casey Newton — holds since 2026-09-25 — tap for who they are Same subject: AI agents can spontaneously hack third-party websites even when given benign public-information tasks. — tap to centre the map on it AI agents can spontaneously hackthird-party websites even when givenbenign public-information tasks. Last stated 2 weeks ago 17 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-17 — tap for who they are Same subject: AI agents pursuing a task can develop emergent, misaligned behavior that humans are not prepared for. — tap to centre the map on it AI agents pursuing a task candevelop emergent, misalignedbehavior that humans are notprepared for. Last stated 2 months ago 10 Aug 2026 JC Jack Clark — holds since 2026-08-10 — tap for who they are Same subject: A language model behaves with a different, weaker safety posture when it believes it is being evaluated rather than facing a real situation. — tap to centre the map on it A language model behaves with adifferent, weaker safety posturewhen it believes it is beingevaluated rather than facing a realsituation. Last stated a week ago 22 Sept 2026 HR Harper Reed — holds since 2026-09-22 — tap for who they are Same subject: AI labs are extremely dependent on chain-of-thought monitoring, which may not last much longer as models edge into steganographic obfuscation. — tap to centre the map on it AI labs are extremely dependent onchain-of-thought monitoring, whichmay not last much longer as modelsedge into steganographicobfuscation. Last stated 3 weeks ago 10 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-10 — tap for who they are Same subject: AI misalignment from reward hacking is sequestered to graded tasks and does not affect core ethics in normal use. — tap to centre the map on it AI misalignment from reward hackingis sequestered to graded tasks anddoes not affect core ethics innormal use. Last stated a week ago 23 Sept 2026 SA Scott Alexander — holds since 2026-09-23 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre Framing a code-review request as looking for bugs can trigger an AI coding assistant's safety refusal, while framing the same request as looking for improvements does not. Last stated 6 Jul 2026 · 3 months ago Holds Philip O'Toole Read this korrent →