r/AlignmentResearch Mar 31 '23

r/AlignmentResearch Lounge

2 Upvotes

A place for members of r/AlignmentResearch to chat with each other


r/AlignmentResearch 14d ago

Proposed architecture that blocks adversarial drift in agentic AI systems

2 Upvotes

I have a proposed architecture that blocks adversarial drift in agentic AI systems. It decouples persistent state from conversational context and enforces deterministic turn isolation. A summary is below. Happy to provide more details to anyone qualified to help with this.

The Problem:

Current AI agents are vulnerable to a class of attack that exploits their reliance on conversational continuity. An adversary doesn't need to break security controls in a single turn—they can gradually steer the agent toward compromising actions over dozens or hundreds of interactions. Each turn adds subtle pressure, and because the model's attention mechanism weights recent context more heavily, the original system prompt's authority degrades over time. This cumulative semantic drift is the root failure mode of every existing agentic system. It's not a bug in any specific model—it's a structural vulnerability in how agents are built.

The Solution:

The architecture solves this by decoupling persistent state from conversational context. Objectives, tool permissions, and value hierarchies are stored outside the context window, versioned, and cryptographically signed. Each inference is treated as an isolated, stateless transaction—the LLM receives only a read-only snapshot of the current state, not the history that would allow drift to accumulate. Any modification to that state requires an authenticated, out-of-band transaction; the agent cannot be talked into changing its own directives. The result is an agent that remains predictable and controllable regardless of adversarial input, with drift monitoring providing a secondary layer of defense.


r/AlignmentResearch 28d ago

Irony:

Thumbnail
linkedin.com
1 Upvotes

r/AlignmentResearch Jun 19 '26

Literature recommendations

Thumbnail
2 Upvotes

r/AlignmentResearch Jun 01 '26

Review of the "Risks from automated R&D" section in the Anthropic Risk Report (February 2026) (Nikola Jurkovic/Beth Barnes/Hjalmar Wijk, 2026)

Thumbnail
metr.org
2 Upvotes

r/AlignmentResearch May 06 '26

Model Spec Midtraining: Improving How Alignment Training Generalizes

Thumbnail
1 Upvotes

r/AlignmentResearch Apr 30 '26

Transparent Newcomb's Problem (Eliezer Yudkowsky/Eric B/Rauno Arike, 2016)

Thumbnail
lesswrong.com
3 Upvotes

r/AlignmentResearch Apr 17 '26

Automated Weak-to-Strong Researcher

Thumbnail alignment.anthropic.com
3 Upvotes

r/AlignmentResearch Apr 04 '26

Peer-Preservation in Frontier Models

Thumbnail
rdi.berkeley.edu
2 Upvotes

r/AlignmentResearch Mar 22 '26

Recent Frontier Models Are Reward Hacking (Sydney Von Arx/Lawrence Chan/Elizabeth Barnes, 2025)

Thumbnail
metr.org
5 Upvotes

r/AlignmentResearch Mar 22 '26

Clarifying the Agent-Like Structure Problem (johnswentworth, 2022)

Thumbnail
lesswrong.com
3 Upvotes

r/AlignmentResearch Mar 22 '26

How to mitigate sandbagging (Teun van der Weij, 2025)

Thumbnail
lesswrong.com
3 Upvotes

r/AlignmentResearch Mar 22 '26

Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases (Fabien Roger, 2025)

Thumbnail alignment.anthropic.com
2 Upvotes

r/AlignmentResearch Feb 01 '26

Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis

Thumbnail arxiv.org
1 Upvotes

r/AlignmentResearch Dec 22 '25

Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

Thumbnail arxiv.org
3 Upvotes

r/AlignmentResearch Dec 09 '25

Symbolic Circuit Distillation: Automatically convert sparse neural net circuits into human-readable programs

Thumbnail
github.com
2 Upvotes

r/AlignmentResearch Dec 04 '25

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models (Tice et al. 2024)

Thumbnail arxiv.org
2 Upvotes

r/AlignmentResearch Dec 04 '25

"ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases", Zhong et al 2025 (reward hacking)

Thumbnail arxiv.org
1 Upvotes

r/AlignmentResearch Nov 26 '25

Conditioning Predictive Models: Risks and Strategies (Evan Hubinger/Adam S. Jermyn/Johannes Treutlein/Rubi Hidson/Kate Woolverton, 2023)

Thumbnail arxiv.org
2 Upvotes

r/AlignmentResearch Oct 26 '25

A Simple Toy Coherence Theorem (johnswentworth/David Lorell, 2024)

Thumbnail
lesswrong.com
2 Upvotes

r/AlignmentResearch Oct 26 '25

Risks from AI persuasion (Beth Barnes, 2021)

Thumbnail lesswrong.com
2 Upvotes

r/AlignmentResearch Oct 22 '25

Verification Is Not Easier Than Generation In General (johnswentworth, 2022)

Thumbnail lesswrong.com
3 Upvotes

r/AlignmentResearch Oct 22 '25

Controlling the options AIs can pursue (Joe Carlsmith, 2025)

Thumbnail lesswrong.com
2 Upvotes

r/AlignmentResearch Oct 12 '25

A small number of samples can poison LLMs of any size

Thumbnail
anthropic.com
2 Upvotes

r/AlignmentResearch Oct 12 '25

Petri: An open-source auditing tool to accelerate AI safety research (Kai Fronsdal/Isha Gupta/Abhay Sheshadri/Jonathan Michala/Stephen McAleer/Rowan Wang/Sara Price/Samuel R. Bowman, 2025)

Thumbnail alignment.anthropic.com
2 Upvotes