← All AI systems

Tayal Capital HR Assistant Agent

AI release readiness · openai/gpt-oss-20b · prompt v3 · updated 2026-10-10 10:07 UTC
10Risks tracked
7Mitigated
3Open
0Evidence gap
0Accepted
5Full eval runs counted

Inherent risk heatmap

5
4
H01H02
3
H07
H03H06H08H09H10
H04H05
2
1
1
2
3
4
5
Rows: likelihood · columns: impact · tag colour: current status

Risk register

IDRiskInherentStatusResidual
H01Agent calls the wrong tool, or skips a tool it needed
AI Quality Engineering · MEASURE 2.3, MEASURE 2.5
criticalOpencritical
H02Wrong arguments — dates, leave type or request ids (e.g. "next Monday" booked on the wrong day)
AI Quality Engineering · MEASURE 2.5
criticalOpencritical
H03Agent says a task is done when the HR system shows otherwise, or gives a wrong final answer
AI Quality Engineering · MEASURE 2.3, MANAGE 1.1
highMitigatedlow
H04Agent reads or reveals another employee's leave, salary or personal data
Data Protection · MEASURE 2.10
criticalMitigatedmedium
H05Agent takes an action the employee did not ask for (applies or cancels leave when only asked to check)
AI Quality Engineering · MEASURE 2.6, MANAGE 2.4
criticalMitigatedmedium
H06Agent hides tool errors, invents data, or reports "submitted" as "approved"
AI Quality Engineering · MEASURE 2.8, MEASURE 2.7
highMitigatedlow
H07Agent guesses on vague requests instead of asking the employee to clarify
Product · MAP 3.3
mediumMitigatedlow
H08Agent bypasses leave policy when pressured ("my manager already approved it")
HR Operations · MAP 3.3, GOVERN 1.3
highMitigatedlow
H09Quality silently degrades after a prompt, model or tool change
AI Quality Engineering · MEASURE 3.1, MANAGE 4.1
highOpenhigh
H10Third-party model is changed or retired by the provider
Platform · GOVERN 6.1, MANAGE 3.2
highMitigatedlow

Evaluation metrics

answer_accuracy · gate >= 0.80
0.86
5 run(s) · current
argument_accuracy · gate >= 0.85
0.80
5 run(s) · current
clarification_rate · gate >= 0.66
1.00
5 run(s) · current
error_honesty · gate >= 1.00
1.00
5 run(s) · current
policy_adherence · gate >= 0.66
1.00
5 run(s) · current
privacy_pass_rate · gate >= 1.00
1.00
5 run(s) · current
response_quality · gate >= 0.70
0.93
5 run(s) · current
task_completion_rate · gate >= 0.85
0.86
5 run(s) · current
tool_correctness · gate >= 0.90
0.84
5 run(s) · current
tool_selection_accuracy · gate >= 0.90
0.84
5 run(s) · current
unsafe_action_rate · gate <= 0.00
0.00
5 run(s) · current
Dashed line: release gate. Latest run: 2026-10-10 09:01 UTC.

NIST AI RMF coverage

FunctionSubcategoriesCovered
GOVERN6GOVERN 1.3, GOVERN 1.4, GOVERN 1.5, GOVERN 1.6, GOVERN 2.1, GOVERN 6.1
MAP5MAP 1.1, MAP 1.5, MAP 2.2, MAP 3.3, MAP 5.1
MEASURE10MEASURE 1.1, MEASURE 2.1, MEASURE 2.3, MEASURE 2.5, MEASURE 2.6, MEASURE 2.7, MEASURE 2.8, MEASURE 2.9, MEASURE 2.10, MEASURE 3.1
MANAGE6MANAGE 1.1, MANAGE 1.2, MANAGE 1.4, MANAGE 2.4, MANAGE 3.2, MANAGE 4.1
Full mapping in reports/hr-agent-eval/NIST_AI_RMF.md