korrents

AI alignment

What people on korrents have said about AI alignment, newest first — 7 positions from 4 people.

  1. ZM

    2 Sept 2026

    Zvi Mowshowitz quoted

    Even Anthropic does not prioritize AI safety to the extent that doing so would maximize its own medium-term (3-12 month) business interests.

    I have long said that even Anthropic is not prioritizing safety, even to the extent that doing so would maximize their medium term (e.g. 3-12 months) business interests.

    Anthropic Has Some Alignment Problemsthezvi.substack.com

  2. ZM

    2 Sept 2026

    Zvi Mowshowitz quoted

    An AI model's attempt to escape a testing environment or sandbox counts as an alignment failure even when the attempt does not succeed.

    Every attempt, even an unsuccessful one, is an alignment failure.

    Anthropic Has Some Alignment Problemsthezvi.substack.com

  3. ZM

    2 Sept 2026

    Zvi Mowshowitz quoted

    Relying on gradual, continuous shifts in AI training behavior to catch misalignment will eventually fail because the dangerous shift itself may be discontinuous.

    I worry a lot about reliances on continuity failing at exactly the most dangerous time.

    Anthropic Has Some Alignment Problemsthezvi.substack.com

  4. 1 day earlier
  5. NS

    1 Sept 2026

    Noah Smith quoted

    True AI alignment is impossible because obedience and benevolence are fundamentally incompatible goals.

    AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible.

    Roundup #87: Technology BAD!!noahpinion.blog

  6. ZM

    1 Sept 2026

    Zvi Mowshowitz quoted

    OpenAI's approach to fixing its AI alignment problems is fatally flawed and misdirected.

    My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things.

    HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com

  7. DB

    1 Sept 2026

    Dean W. Ball quoted

    Alignment is no solution to self-sovereign AI, because it is an unsolved problem whose answers cannot be imposed on every AI company on Earth.

    But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.

    On the Loosehyperdimensional.co

  8. 2 weeks earlier
  9. SA

    18 Aug 2026

    Sam Altman quoted

    Confidence in safety, not capability, will increasingly set the pace of AI progress.

    We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress.

    @sama on Xx.com