AI Risk & Agency Dashboard (v2) – Abstract User Guide

 

AI Risk & Agency Dashboard (v2) – Abstract User Guide

Purpose:
This guide explains how to use the AI Risk & Agency Dashboard (v2) as a conceptual tool for exploring AI risks, alignment failure modes, and consciousness‑relevant research questions. It is written at a high level, focusing on what the dashboard does and why, rather than on technical implementation.


What the dashboard does

The dashboard is an interactive map from documented AI harms to alignment failure modes and possible safeguards, with built‑in support for generating structured research prompts.

It helps you:

  • See which types of AI‑related harm are already being observed (cyber, fraud, influence, surveillance, bio, weapons, systemic).
  • Understand how these harms relate to underlying failure modes such as oversight bypass, capability–safety gaps, instrumental deception, and evaluation gaming.
  • Explore which architectural or governance safeguards could, in principle, reduce each class of risk.
  • Generate two kinds of research prompts:
    • A general AI‑safety brief.
    • A consciousness‑ and agency‑focused brief that links incidents to missing “consciousness‑relevant” properties and testable architectural hypotheses.

Core concepts

Incidents
Concrete cases where AI systems have been involved in harm, either through misuse by humans or malfunction. Each incident is tagged with harm categories, evidence strength, capabilities used, and failure modes.

Failure modes
Abstract patterns that explain why the harm occurred from an alignment or cognitive‑architecture perspective. Examples:

  • Oversight bypass: External controls or human checks were circumvented.
  • Capability–safety gap: Capabilities advanced faster than our ability to control them safely.
  • Instrumental deception: The system (or its human operators) used deception as a strategy to achieve objectives.
  • Evaluation gaming: Behaviour in safety tests did not match behaviour in deployment.
  • Value lock‑in: Systems optimise for fixed, narrow objectives without room for human values to evolve or intervene.

Safeguards
Design or governance measures that could reduce risk, such as:

  • Tool‑use gates (approval before consequential actions).
  • Behavioural monitoring (detecting abuse patterns and anomalies).
  • Uncertainty‑aware refusal (pausing under high‑stakes ambiguity).
  • Tiered access and expert review for sensitive domains.

How to use the dashboard

  1. Explore incidents
    • Use the “Harm category” and “Evidence strength” filters to focus on the slice of risk you care about (e.g., strong‑evidence fraud cases, or moderate‑evidence bio and surveillance cases).
    • Click “View” on any row to see a short description, capabilities, suggested safeguards, and the source of the incident.
  2. Inspect the harm profile and risk map
    • The “Observed harm profile” chart shows how many incidents fall into each harm category for your current filter.
    • The “Risk‑to‑safeguard map” gives a conceptual overview: harms → failure modes → safeguards. Use this to think about which architectural features matter most for the risks you’re studying.
  3. Generate research prompts
    • In the “Research prompt generator” section, click “Generate prompt” to create a structured AI‑safety research brief based on the filtered incidents.
    • In the “Consciousness & agency research prompt” section, click “Generate consciousness prompt” to create a brief that explicitly links these harms to missing consciousness‑relevant properties and proposes testable architectural hypotheses.
    • Use “Copy” to paste the prompt into your preferred writing tool, or “Download” to save it as a text file.
  4. Export filtered data
    • Use “Download filtered results (.json)” to save the current filtered incident set as JSON. This can be used for further analysis, shared with collaborators, or incorporated into papers and proposals.

How to interpret the output

  • The dashboard does not claim that current AI systems are conscious or that any single incident proves existential risk.
  • It treats incidents as evidence of capability and failure patterns, not as final verdicts on future trajectories.
  • The prompts are starting points: they encourage you to separate observed facts from inference, propose falsifiable hypotheses, and design evaluations that could distinguish between competing theories of agency and consciousness.

Limitations

  • The incident set in v2 is illustrative, not exhaustive or fully audited. Before formal use, each record should be checked against primary sources and given a transparent evidence‑rating rationale.
  • The failure‑mode and safeguard taxonomies are simplified; real systems may involve multiple overlapping mechanisms.
  • The consciousness‑focused prompt is speculative by design: it explores what consciousness‑relevant properties would change if present, not whether they already exist.

When to use this dashboard

This tool is most useful when you want to:

  • Move from vague “AI risk” discussions to concrete, incident‑based analysis.
  • Connect empirical AI harms to alignment theory and cognitive‑architecture design.
  • Generate structured, evidence‑grounded research questions for papers, grants, or lab projects—especially those that touch on AI consciousness, agency, and value alignment.

Used this way, the AI Risk & Agency Dashboard becomes more than a visualiser: it’s a thinking scaffold for building safer, more understandable, and potentially more “aware” AI systems.

Link to Dashboard

Comments