AI Risk & Agency Dashboard (v2) –
Abstract User Guide
Purpose:
This guide explains how to use the AI Risk & Agency Dashboard (v2) as a
conceptual tool for exploring AI risks, alignment failure modes, and
consciousness‑relevant research questions. It is written at a high level,
focusing on what the dashboard does and why, rather than on technical
implementation.
What
the dashboard does
The dashboard is an interactive map from documented AI
harms to alignment failure modes and possible safeguards,
with built‑in support for generating structured research prompts.
It helps you:
- See
which types of AI‑related harm are already being observed (cyber, fraud,
influence, surveillance, bio, weapons, systemic).
- Understand
how these harms relate to underlying failure modes such as oversight
bypass, capability–safety gaps, instrumental deception, and evaluation
gaming.
- Explore
which architectural or governance safeguards could, in principle, reduce
each class of risk.
- Generate
two kinds of research prompts:
- A
general AI‑safety brief.
- A
consciousness‑ and agency‑focused brief that links incidents to missing
“consciousness‑relevant” properties and testable architectural hypotheses.
Core
concepts
Incidents
Concrete cases where AI systems have been involved in harm, either through
misuse by humans or malfunction. Each incident is tagged with harm categories,
evidence strength, capabilities used, and failure modes.
Failure modes
Abstract patterns that explain why the harm occurred from an alignment or
cognitive‑architecture perspective. Examples:
- Oversight
bypass: External controls or human checks were circumvented.
- Capability–safety
gap: Capabilities advanced faster than our ability to control them
safely.
- Instrumental
deception: The system (or its human operators) used deception as a
strategy to achieve objectives.
- Evaluation
gaming: Behaviour in safety tests did not match behaviour in
deployment.
- Value
lock‑in: Systems optimise for fixed, narrow objectives without room
for human values to evolve or intervene.
Safeguards
Design or governance measures that could reduce risk, such as:
- Tool‑use
gates (approval before consequential actions).
- Behavioural
monitoring (detecting abuse patterns and anomalies).
- Uncertainty‑aware
refusal (pausing under high‑stakes ambiguity).
- Tiered
access and expert review for sensitive domains.
How to use the dashboard
- Explore
incidents
- Use
the “Harm category” and “Evidence strength” filters to focus on the slice
of risk you care about (e.g., strong‑evidence fraud cases, or moderate‑evidence
bio and surveillance cases).
- Click
“View” on any row to see a short description, capabilities, suggested
safeguards, and the source of the incident.
- Inspect
the harm profile and risk map
- The
“Observed harm profile” chart shows how many incidents fall into each
harm category for your current filter.
- The
“Risk‑to‑safeguard map” gives a conceptual overview: harms → failure
modes → safeguards. Use this to think about which architectural features
matter most for the risks you’re studying.
- Generate
research prompts
- In
the “Research prompt generator” section, click “Generate prompt” to
create a structured AI‑safety research brief based on the filtered
incidents.
- In
the “Consciousness & agency research prompt” section, click “Generate
consciousness prompt” to create a brief that explicitly links these harms
to missing consciousness‑relevant properties and proposes testable
architectural hypotheses.
- Use
“Copy” to paste the prompt into your preferred writing tool, or
“Download” to save it as a text file.
- Export
filtered data
- Use
“Download filtered results (.json)” to save the current filtered incident
set as JSON. This can be used for further analysis, shared with
collaborators, or incorporated into papers and proposals.
How to interpret the output
- The
dashboard does not claim that current AI systems are conscious or
that any single incident proves existential risk.
- It
treats incidents as evidence of capability and failure patterns,
not as final verdicts on future trajectories.
- The
prompts are starting points: they encourage you to separate observed facts
from inference, propose falsifiable hypotheses, and design evaluations
that could distinguish between competing theories of agency and
consciousness.
Limitations
- The
incident set in v2 is illustrative, not exhaustive or fully
audited. Before formal use, each record should be checked against primary
sources and given a transparent evidence‑rating rationale.
- The
failure‑mode and safeguard taxonomies are simplified; real systems may
involve multiple overlapping mechanisms.
- The
consciousness‑focused prompt is speculative by design: it explores what
consciousness‑relevant properties would change if present, not
whether they already exist.
When
to use this dashboard
This tool is most useful when you want to:
- Move
from vague “AI risk” discussions to concrete, incident‑based analysis.
- Connect
empirical AI harms to alignment theory and cognitive‑architecture design.
- Generate
structured, evidence‑grounded research questions for papers, grants, or
lab projects—especially those that touch on AI consciousness, agency, and
value alignment.
Used this way, the AI Risk & Agency Dashboard becomes
more than a visualiser: it’s a thinking scaffold for building safer,
more understandable, and potentially more “aware” AI systems.
Comments
Post a Comment