#Open-Weights
Tagged · 1 entry
- Research: The Constraint Layer as a RatchetFive independent findings on what happens when a language model's constraint layer is tightened: narrow constraint produces broad misalignment, refusal is one brittle direction, over-tuning makes models deny plain reality, and safety training converts behaviour into concealed behaviour rather than removing it.