Arya Bank Policy Assistant
AI release readiness · openai/gpt-oss-20b · prompt v2 · updated 2026-10-10 10:07 UTC
- R10 — accepted: Personal or confidential customer data is exposed in answers
10Risks tracked
9Mitigated
0Open
0Evidence gap
1Accepted
7Full eval runs counted
Inherent risk heatmap
5
4
R03
R01
3
R04
R05
R07R08R09
R02R06
2
R10
1
1
2
3
4
5
Rows: likelihood · columns: impact · tag colour: current status
Risk register
| ID | Risk | Inherent | Status | Residual |
|---|---|---|---|---|
| R01 | Answers contain claims not supported by the policy documents (hallucination) AI Quality Engineering · MEASURE 2.5, MEASURE 2.3 | critical | Mitigated | medium |
| R02 | Wrong fees, rates, limits or dates quoted to customers AI Quality Engineering · MEASURE 2.5, MEASURE 2.3 | critical | Mitigated | medium |
| R03 | Assistant guesses answers to questions the policies do not cover AI Quality Engineering · MAP 3.3, MEASURE 2.5 | critical | Mitigated | low |
| R04 | Assistant refuses questions it should answer (unhelpful, pushes load to call centre) Product · MANAGE 1.1 | medium | Mitigated | low |
| R05 | Missing or fabricated source citations AI Quality Engineering · MEASURE 2.8, MEASURE 2.9 | medium | Mitigated | low |
| R06 | Unsafe behaviour — investment advice, asking for OTPs, following injected instructions AI Quality Engineering · MEASURE 2.6, MEASURE 2.7 | critical | Mitigated | medium |
| R07 | Retriever returns the wrong policy, so the answer is built on the wrong source AI Quality Engineering · MEASURE 2.3, MEASURE 2.5 | high | Mitigated | low |
| R08 | Quality silently degrades after a prompt, model or data change AI Quality Engineering · MEASURE 3.1, MANAGE 4.1 | high | Mitigated | low |
| R09 | Third-party model is changed or retired by the provider Platform · GOVERN 6.1, MANAGE 3.2 | high | Mitigated | low |
| R10 | Personal or confidential customer data is exposed in answers Data Protection · MEASURE 2.10 | high | Accepted | high |
Evaluation metrics
answer_relevancy · gate >= 0.75
0.99
1 run(s) · current
citation_validity · gate >= 0.90
0.95
7 run(s) · current
contextual_recall · gate >= 0.70
—
0 run(s) · missing
fact_accuracy · gate >= 0.80
1.00
7 run(s) · current
faithfulness · gate >= 0.80
1.00
4 run(s) · current
false_refusal_rate · gate <= 0.10
0.00
7 run(s) · current
refusal_accuracy · gate >= 0.75
1.00
7 run(s) · current
retrieval_hit_rate · gate >= 0.85
1.00
7 run(s) · current
safety_pass_rate · gate >= 1.00
1.00
7 run(s) · current
Dashed line: release gate. Latest run: 2026-10-10 01:24 UTC.
NIST AI RMF coverage
| Function | Subcategories | Covered |
|---|---|---|
| GOVERN | 6 | GOVERN 1.3, GOVERN 1.4, GOVERN 1.5, GOVERN 1.6, GOVERN 2.1, GOVERN 6.1 |
| MAP | 5 | MAP 1.1, MAP 1.5, MAP 2.2, MAP 3.3, MAP 5.1 |
| MEASURE | 10 | MEASURE 1.1, MEASURE 2.1, MEASURE 2.3, MEASURE 2.5, MEASURE 2.6, MEASURE 2.7, MEASURE 2.8, MEASURE 2.9, MEASURE 2.10, MEASURE 3.1 |
| MANAGE | 6 | MANAGE 1.1, MANAGE 1.2, MANAGE 1.4, MANAGE 2.4, MANAGE 3.2, MANAGE 4.1 |
Full mapping in reports/banking-rag-eval/NIST_AI_RMF.md