Research · curated 22 Jul 2026

We Fed Claude Code's System Prompt to GPT — Its Fingerprint Drifted Like a Different Model

Coverage timeline

21 Jul 2026tosea.ai

Single-source research — first reported, latest, and curated coincide.

Why it matters

Behavioral fingerprinting is proposed as a defense to verify that an API aggregator or relay actually serves the model it advertises, and this measurement shows a middlebox that silently prepends a system prompt can spoof or wreck that fingerprint, undermining trust in opaque LLM serving chains.

Tosea.ai ran 5,660 controlled API calls testing whether a hidden prepended system prompt can distort an LLM's behavioral fingerprint (the answer distributions used to verify that an API serves the advertised model). Injecting the real 6,651-token Claude Code system block into clean GPT-5.5 and gpt-4o-mini endpoints pushed their fingerprints a Jensen-Shannon distance of ~0.46 — into the 'different model' band — collapsing GPT-5.5's random-color answers to 'blue' 40/40 times and flipping many responses to Chinese, while shorter Codex CLI prompts and filler drifted less.