Evaluation, interpretability, security, deployment limits, independent oversight and routes for appeal address different risks. No single measure proves that all dangers are solved. Ask what a safeguard can detect, who enforces it and what happens when it fails.
Solutions and safeguards
Safety is a continuing practice.
Selected starting points
Descriptions introduce the source. They do not replace reading it in full.
Stuart Russell: three principles for safer AI
2017-06-06
Open original source ↗ControlAI: the Dangerous Intelligence Problem
2026-09-21
Open original source ↗Redwood Research: AI control
2026-09-21
Open original source ↗Evidence, estimates and scenarios are different.
A company announcement is not an independent evaluation. A scenario is not a prediction. Follow the original source and check its date, scope and limitations.