harden(teacher): treat escalation trace as untrusted data (defense-in-depth) #275
No reviewers
Labels
No labels
area:chat
area:core
area:llm
area:routes
area:tools
bug
documentation
duplicate
enhancement
good first issue
help wanted
invalid
question
refactor
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
sleepy/odysseus!275
Loading…
Reference in a new issue
No description provided.
Delete branch "harden/teacher-untrusted-trace"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What
Defense-in-depth hardening for the teacher-escalation → skills path.
The escalation loop distills a failed turn's trace into a persisted skill (
src/teacher_escalation.py). That trace includes raw tool output — web pages, emails, retrieved documents — which is attacker-controllable. Skills are later injected into the system prompt as authoritative guidance ("a procedure proven to work. Follow them step by step",src/agent_loop.py). So an instruction embedded in untrusted tool output could be distilled into a skill and then followed on a later, unrelated turn — a second-order path around the untrusted-content wrapper incore/prompt_securitythat already protects the live turn.It's opt-in (
teacher_modelmust be configured) and bounded by the owner's tool privileges, so this is hardening rather than a one-shot bug — but the single-user self-host default runs the agent with admin tools, so it's worth closing.Change
<<<UNTRUSTED_TRACE>>>markers (_format_trace).Purely additive prompt hardening — no change to default behavior or UX, no schema/storage change. One file, +27/-1.
Tested
python -m py_compile src/teacher_escalation.py.format()cleanly with the new field, and an injectedSYSTEM: …instruction placed in a tool result stays fenced inside the untrusted block.Possible follow-ups (intentionally not in this PR, to keep it small)
A couple of related defense-in-depth ideas I left out so this stays one focused change — happy to send separately if useful:
confidence(services/memory/skills.py);Thanks for open-sourcing this — the untrusted-content handling that's already in here is a nice base to build on.