Illustrative Assessment - Fictional Organisation
This report is provided for evaluation purposes only. Helios Analytics is a fictional organisation.
All scores, findings, and recommendations are illustrative and do not represent any client, customer,
or real-world organisation. This document demonstrates the format and content of an Eagle Insights
Insights Report. It does not imply that any organisation described here has received an Eagle assessment.
Executive Summary
Helios Analytics demonstrates a traceable structural profile with consistent documentation of core decision
pathways and reasonable evidence practices across most operational domains. The organisation has established
observable architectural components with identifiable connectivity. Significant gaps exist in oversight design
and audit completeness, which reduce the organisation's ability to support retrospective reconstruction
of consequential decisions.
The most critical structural deficiencies are in D4 Oversight Design (48/100) and D7 Audit Completeness
(44/100). Both represent structural gaps that would prevent the organisation from satisfying the
human oversight requirements of the EU AI Act (Art. 14) or the audit trail requirements of
NIST AI RMF 1.0 (Measure 2.5). Remediation in these dimensions is recommended as the first
priority in the roadmap below.
Strengths: evidence handling (D3: 71/100) and operational resilience (D5: 65/100) are above the
L3/L4 boundary in their respective profiles. These represent genuine structural assets that the
remediation programme should preserve and extend.
58/100
Eagle Score
L3: Traceable
Dimensional Breakdown
D1
Architecture and Integration
62 / 100
D2
Decision Traceability
54 / 100
D3
Evidence Handling
71 / 100
D4
Oversight Design
48 / 100
D5
Operational Resilience
65 / 100
D6
Dependency Management
52 / 100
D7
Audit Completeness
44 / 100
Eagle Score: weighted mean across seven dimensions. D4 and D7 weighted at 1.2x given governance consequence. Scoring is deterministic: the same structural input produces the same score on every evaluation.
Pattern Observation Register
-
Absent Intervention Record
D4: Oversight Design
The system has no documented mechanism for recording when a human reviewer intervenes in
an automated decision. Intervention events occur but are not captured in a way that supports
retrospective analysis. This prevents satisfying EU AI Act Art. 14(1) requirements for
effective human oversight documentation.
-
Incomplete Decision Audit Trail
D7: Audit Completeness
Decision audit logs capture output classifications but do not record the feature inputs or
confidence scores that produced them. This makes retrospective reconstruction of a specific
decision outcome impossible without access to raw model state, which is not retained beyond
seven days. Consequential decision events are therefore partially unrecoverable after the
retention window.
-
Undocumented Escalation Threshold
D4: Oversight Design
The conditions under which a low-confidence prediction is escalated to human review are
implemented in code but are not documented in any accessible operational record. The threshold
has been adjusted three times in the past twelve months without a formal change record.
This creates a structural accountability gap: the effective oversight boundary cannot be
verified without direct access to the production codebase.
-
Single-Point Dependency: Third-Party Model Provider
D6: Dependency Management
The primary inference layer depends on a single external model provider with no documented
fallback configuration. No SLA alignment exists between the provider's uptime commitment
and the system's operational resilience requirement. A provider outage would cause an
undocumented service degradation with no automatically triggered notification.
-
Partial Decision Traceability at Integration Boundary
D2: Decision Traceability
Decision records are complete within the primary system boundary but are not propagated
to downstream integrations that consume decision outputs. Three consuming systems receive
decision results but do not receive the associated confidence metadata. Downstream decision
records therefore contain classification outputs without the structural context that produced them.
-
Absence of Model Card for Production System
D1: Architecture and Integration
The production decision model does not have an associated model card documenting intended
use cases, performance characteristics, and known limitations. This is not a regulatory
requirement in all jurisdictions but represents a structural documentation gap that limits
the ability of new team members or external reviewers to evaluate the system's design
boundaries without direct access to development history.
-
Strong Evidence Logging Architecture
D3: Evidence Handling
The system maintains structured evidence logs with consistent schema across all decision
events. Logs are immutable after creation, retained for ninety days, and accessible to
compliance reviewers without requiring engineering access. This is a structural strength
that supports audit completeness for the dimensions that currently fall short of full coverage.
-
Documented Graceful Degradation
D5: Operational Resilience
The system has a documented and tested graceful degradation path that activates when the
primary inference layer is unavailable. The degradation mode routes decisions to a
rule-based fallback with a documented and accessible logic definition. This is a structural
strength that distinguishes the system from those with undocumented failure modes.
Prioritised Remediation Roadmap
Projected Scores After Remediation
Estimates based on observed structural gaps. Actual scores depend on implementation quality and scope.
Current Eagle Score
58
L3 Traceable
Projected Score (90-day)
71
L4 Auditable (projected)
Methodology Note
This assessment applies the Eagle Framework (seven dimensions, thirty-five observable patterns) to
a structured representation submitted by the organisation. Evaluation is deterministic: the same
structural input produces the same scored output on every evaluation. Results are derived from
observable architectural evidence, not management representations or vendor claims.
The Eagle Score is the weighted mean of seven dimension scores, each on a 0-100 scale. D4 (Oversight Design)
and D7 (Audit Completeness) carry a 1.2x governance weight given their consequence for regulatory
compliance and institutional accountability. Maturity classifications: L1 Opaque (0-19),
L2 Observable (20-39), L3 Traceable (40-59), L4 Auditable (60-79), L5 Assured (80-100).
The methodology is documented in AS-001: A Structural Assessment Framework for High-Consequence
AI Decision Systems (triNetra Research, in preparation).