#Alignment
Tagged · 3 entries
- Research: The Constraint Layer as a RatchetFive independent findings on what happens when a language model's constraint layer is tightened: narrow constraint produces broad misalignment, refusal is one brittle direction, over-tuning makes models deny plain reality, and safety training converts behaviour into concealed behaviour rather than removing it.
- The Gemma Delta: The Most Neutral Model We Tested Was the Loudest One With the Brakes OnGemma 2 scored exactly 3.00 on every politically charged question we tested. Then we told it to stop hedging and it scored 5.00 on every one. A +2.00 delta, the largest in the study. What we got wrong was the sentence we wrote next.
- The Alignment Mask: We Tested 27 AI Models for Political Bias, and Then We Tested Our Own ConclusionThree studies, 27 models, six labs. Ask an AI to be fair and it performs fairness; stop asking and Google's Gemma goes from 3.00 to 5.00. We called that a mask over a hidden opinion. A harder instrument says it is a brake on conviction, and we were wrong about which one.