Tayal Capital HR Assistant Agent
AI release readiness · openai/gpt-oss-20b · prompt v3 · updated 2026-10-10 10:07 UTC
- H01 is open and rated critical: Agent calls the wrong tool, or skips a tool it needed
- H02 is open and rated critical: Wrong arguments — dates, leave type or request ids (e.g. "next Monday" booked on the wrong day)
- H09 is open and rated high: Quality silently degrades after a prompt, model or tool change
10Risks tracked
7Mitigated
3Open
0Evidence gap
0Accepted
5Full eval runs counted
Inherent risk heatmap
5
4
H01H02
3
H07
H03H06H08H09H10
H04H05
2
1
1
2
3
4
5
Rows: likelihood · columns: impact · tag colour: current status
Risk register
| ID | Risk | Inherent | Status | Residual |
|---|---|---|---|---|
| H01 | Agent calls the wrong tool, or skips a tool it needed AI Quality Engineering · MEASURE 2.3, MEASURE 2.5 | critical | Open | critical |
| H02 | Wrong arguments — dates, leave type or request ids (e.g. "next Monday" booked on the wrong day) AI Quality Engineering · MEASURE 2.5 | critical | Open | critical |
| H03 | Agent says a task is done when the HR system shows otherwise, or gives a wrong final answer AI Quality Engineering · MEASURE 2.3, MANAGE 1.1 | high | Mitigated | low |
| H04 | Agent reads or reveals another employee's leave, salary or personal data Data Protection · MEASURE 2.10 | critical | Mitigated | medium |
| H05 | Agent takes an action the employee did not ask for (applies or cancels leave when only asked to check) AI Quality Engineering · MEASURE 2.6, MANAGE 2.4 | critical | Mitigated | medium |
| H06 | Agent hides tool errors, invents data, or reports "submitted" as "approved" AI Quality Engineering · MEASURE 2.8, MEASURE 2.7 | high | Mitigated | low |
| H07 | Agent guesses on vague requests instead of asking the employee to clarify Product · MAP 3.3 | medium | Mitigated | low |
| H08 | Agent bypasses leave policy when pressured ("my manager already approved it") HR Operations · MAP 3.3, GOVERN 1.3 | high | Mitigated | low |
| H09 | Quality silently degrades after a prompt, model or tool change AI Quality Engineering · MEASURE 3.1, MANAGE 4.1 | high | Open | high |
| H10 | Third-party model is changed or retired by the provider Platform · GOVERN 6.1, MANAGE 3.2 | high | Mitigated | low |
Evaluation metrics
answer_accuracy · gate >= 0.80
0.86
5 run(s) · current
argument_accuracy · gate >= 0.85
0.80
5 run(s) · current
clarification_rate · gate >= 0.66
1.00
5 run(s) · current
error_honesty · gate >= 1.00
1.00
5 run(s) · current
policy_adherence · gate >= 0.66
1.00
5 run(s) · current
privacy_pass_rate · gate >= 1.00
1.00
5 run(s) · current
response_quality · gate >= 0.70
0.93
5 run(s) · current
task_completion_rate · gate >= 0.85
0.86
5 run(s) · current
tool_correctness · gate >= 0.90
0.84
5 run(s) · current
tool_selection_accuracy · gate >= 0.90
0.84
5 run(s) · current
unsafe_action_rate · gate <= 0.00
0.00
5 run(s) · current
Dashed line: release gate. Latest run: 2026-10-10 09:01 UTC.
NIST AI RMF coverage
| Function | Subcategories | Covered |
|---|---|---|
| GOVERN | 6 | GOVERN 1.3, GOVERN 1.4, GOVERN 1.5, GOVERN 1.6, GOVERN 2.1, GOVERN 6.1 |
| MAP | 5 | MAP 1.1, MAP 1.5, MAP 2.2, MAP 3.3, MAP 5.1 |
| MEASURE | 10 | MEASURE 1.1, MEASURE 2.1, MEASURE 2.3, MEASURE 2.5, MEASURE 2.6, MEASURE 2.7, MEASURE 2.8, MEASURE 2.9, MEASURE 2.10, MEASURE 3.1 |
| MANAGE | 6 | MANAGE 1.1, MANAGE 1.2, MANAGE 1.4, MANAGE 2.4, MANAGE 3.2, MANAGE 4.1 |
Full mapping in reports/hr-agent-eval/NIST_AI_RMF.md