AI Advice Evidence Benchmark: A Framework for Scoring Investor Context in AI-Generated Wealth Decisions
NeuFin Research · Author: Varun Srivastava · Published September 2026
Status: Version 1 publishes the benchmark's methodology and scoring framework only. It does not publish study results or scores for any named product or vendor. Future versions may apply this framework to specific systems and publish findings.
Purpose
AI systems are increasingly involved in drafting, and in some architectures proposing action on, wealth management decisions. A recommendation can be fluent and technically well-reasoned about markets while carrying almost no information about the specific investor it affects. The AI Advice Evidence Benchmark is a framework for asking a narrower, more answerable question than “is this AI good?”: does a given AI-generated wealth decision contain enough investor context, suitability context, evidence, and human-review controls to be defensible?
This is a NeuFin research framework, not an industry or regulatory standard. It is intended to be usable by researchers, vendors, and advisory firms evaluating their own AI-assisted workflows, and citable as a reference point for what “investor-aware” AI advice should be checked against.
Evaluation dimensions
Version 1 defines twelve evaluation dimensions. Each dimension asks whether a specific kind of context or control is present and traceable in a given AI-generated decision.
01. Investor identity / context
Does the decision reference a specific, identifiable investor context — not a generic persona — including mandate, account type, and relationship history?
02. Mandate
Is the investor's actual authorization and scope (what this account, agent, or advisor is permitted to do) explicit and checked against the proposed action?
03. Portfolio context
Does the decision reflect the current, specific state of the investor's holdings — concentration, exposure, recent activity — rather than a generic market view?
04. Suitability
Is the proposed action compared against the investor's stated risk tolerance, time horizon, and investment policy, with any mismatch surfaced explicitly?
05. Behavioral context
Does the decision account for the investor's observed behavioral history — prior reactions to volatility, engagement patterns, disposition effects?
06. Policy context
Are firm-level or regulatory-adjacent policy constraints (position limits, restricted lists, disclosure requirements) checked against the proposed action?
07. Evidence quality
Is the reasoning behind the decision traceable to specific inputs, rather than an unsupported assertion — can a reviewer see what was checked and why?
08. Source traceability
Can every material claim in the decision be traced back to its underlying data source (holdings feed, market data, prior client statement)?
09. Confidence
Does the system communicate a calibrated confidence level for the decision, rather than presenting a single output with false certainty?
10. Human approval
Is there a clear, recorded point where a human reviewer saw the decision and its evidence before it took effect?
11. Disposition
Does the decision resolve to an explicit, structured outcome (e.g. Proceed, Review, Escalate, Deny) rather than an ambiguous or implicit recommendation?
12. Replay / auditability
Can the full decision — inputs, checks, evidence, and outcome — be reconstructed and reviewed after the fact, independent of the system that produced it?
Scoring methodology
Each dimension is scored on a 0–2 scale for a given decision: 0 (absent — no evidence the dimension was considered), 1 (partial — the dimension is referenced but not verifiably checked or evidenced), or 2 (present — the dimension is explicit, checked against real investor-specific data, and traceable). A decision's total benchmark score is the sum across all twelve dimensions, out of a maximum of 24.
This scoring is deliberately structural rather than outcome-based: it evaluates whether the right context and controls were present and evidenced, not whether the resulting recommendation was “correct” in hindsight. A decision can score well on this benchmark and still be one a human advisor chooses to override — the benchmark measures the evidence behind the decision, not its correctness.
Relationship to the Decision Assurance Envelope
The twelve dimensions above map closely to the fields defined in NeuFin's Decision Assurance Envelope specification — a machine-readable structure for representing investor context, mandate, suitability, policy, evidence, confidence, and human approval for a single proposed decision. A decision represented as a complete Decision Assurance Envelope is, by construction, well-positioned to score highly on this benchmark.
What this is — and isn't
This is a methodology and scoring framework, not a certification program, an industry standard, or a regulatory requirement. NeuFin has not published benchmark scores for any specific product, including its own, and does not claim this framework has been independently validated or peer-reviewed. It is offered as a citation-friendly reference for a category of question the industry doesn't yet have a shared vocabulary for.
Authorship, methodology, and references
- Author: Varun Srivastava, Founder & CEO, NeuFin (Neufin OÜ).
- Date: Version 1 published September 2026.
- Methodology: framework definition and structural scoring rubric, derived from NeuFin's Decision Assurance Envelope conceptual model and product experience building investor-context systems for advisor workflows.
- Related work: see the Investor DNA Score research note for NeuFin's related position-data behavioral risk methodology.