Self-consistency checking for indirect prompt injection defense in email agents

Date

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

As the scale of agentic large language model (LLM) deployment increases, indirect prompt injection (IPI) persists as a major vulnerability when connecting agents to untrusted data sources. One solution to IPI is human review of agent actions, but this imposes a utility cost by reducing the autonomy of the agent. An ideal solution must balance agent autonomy and security. This thesis investigated to what extent inference-time compute scaling via self-consistency checking (SCC) reduces IPI vulnerability in a local, open-weight email agent and characterized the autonomy-security tradeoff for this defense mechanism. The research design consisted of a factorial experiment crossing self-consistency sample size, voting method, and attack plausibility level across 4,000 trial records for a ReAct agent with access to sandboxed email infrastructure. The study found that increasing inference-time compute via the SCC mechanism decreased IPI attack success rate (ASR) from 13.7% at sample size N = 1 to 0.7% at N = 7 with unanimous voting. The largest single reduction occurred between sample sizes N = 1 and N = 3, with further reductions from increasing N observed only under unanimous voting. The utility cost was also concentrated in unanimous voting, which reduced benign task completion rate (TCR) to 75% at N = 7. Majority and super-majority voting showed no measurable utility cost but left a 17.3% residual ASR on high-plausibility attacks. The relationship between plausibility and ASR was non-monotonic with medium-plausibility attacks succeeding less often than either low- or high-plausibility attacks. These findings are consistent with SCC operating by aggregation over independent per-sample decisions, where attack plausibility raises the per-sample rate at which the agent follows the injection. This suggests that the SCC mechanism has added security value and can be effective as a defense-in-depth layer, but practitioners should acknowledge the utility cost and consider this method in combination with other agent security mechanisms.

Description

Keywords

Indirect prompt injection, Self-consistency, Agents, Large language models, Security, Artificial intelligence

Graduation Month

August

Degree

Master of Science

Department

College of Technology and Aviation

Major Professor

Michael J. Pritchard

Date

Type

Thesis

Citation