AI Risk & Agency Dashboard v2
Illustrative incidents mapped to failure modes and safeguards, with a consciousness/agency research prompt. Data are simplified for demo; replace with audited records before publication.
0
Illustrative incidents
0
Strong evidence records
0
Harm categories represented
0
Failure modes mapped
Incident explorer
Use the filters to inspect current, observable risks rather than treating speculative extinction narratives as established evidence.
| Incident | Harm | Evidence | Failure modes | Action |
|---|
Observed harm profile
Risk-to-safeguard map
Harm
Cyber misuseExploits, phishing, intrusion
Fraud & impersonationScams, synthetic personas
Influence & surveillanceManipulation, monitoring
Failure mode
Oversight bypassExternal controls circumvented
Instrumental deceptionManipulation aids objectives
Evaluation gamingTest/deployment mismatch
Safeguard
Tool-use gatesApproval before consequential actions
Behavioural monitoringDetect abuse patterns and anomalies
Uncertainty-aware refusalPause under high-stakes ambiguity
Research prompt generator
Generate a structured prompt based on the selected filters, then copy it or download it as a plain-text research brief.
Consciousness & agency research prompt
Generate a prompt that explicitly links these harms to missing consciousness-relevant properties and testable architectural hypotheses.
Prototype data are illustrative placeholders. Before publication, replace each record with audited source metadata, direct citations, a verification date, and a transparent evidence-rating protocol. Sources include Anthropic’s Sept 2026 threat report and the International AI Safety Report 2026.
Comments
Post a Comment