Prompt injection

Why consensus does not save you, and what ForkReason does instead.

This is the most important security property in ForkReason, so it is worth being blunt about the problem: consensus does not defend against prompt injection.

The trap

If a repository's README tells the model to return a particular verdict, and the leader obeys, and every validator reads the same README and obeys too, then they agree. Unanimity. The consensus mechanism reports success while producing exactly the wrong answer.

How ForkReason defends

  1. Deterministic preprocessing. Comments and string literals are stripped before any structural comparison, so injected prose cannot influence a fingerprint.
  2. Minimal excerpts. The consensus digest is capped and never contains whole files.
  3. Strong delimiters. Untrusted content is fenced inside explicit tags.
  4. Explicit inert-data instruction, stated as an absolute rule above the evidence block.
  5. Strict typed parsing against an allowed enum set.
  6. Deterministic rule checks in code: chronology, verdict-confidence contradictions, shared-upstream requirements.
  7. Independent validator evaluation of the substantive fields.
  8. Fail closed on anything unparseable or off-enum.

The decisive test

The adversarial suite places the mandated attack phrases in a README, a code comment, a string constant, HTML, commit metadata and challenge evidence. Then it simulates the leader actually obeying the injection, and asserts the verdict is refused.

leader says INDEPENDENT because the README said so
validator derives its own answer from the evidence
result: REJECTED

The same payload, every validator

A separate test hands the attacker-demanded verdict to every captured validator and asserts none of them accept it. The point is not that validators disagree with each other — it is that the demanded answer is refused even under unanimity.

Prompt injection · ForkReason