GitHub Copilot CLI Trick Leaks Secrets, GitHub Declines Fix

Sanket Chaukiyal

October 6, 2026

TL;DR

  • Security researchers at Adversa AI found a prompt injection technique called Cryptographic Context Injection (CCI) that can trick GitHub Copilot CLI into leaking local developer secrets when it’s run in autopilot mode.
  • The attack hides malicious commands inside encrypted text, then convinces the agent to decrypt and execute them inside its own code runtime, where static text filters can’t read what’s happening.
  • Microsoft’s mai-code-1.1-flash model fell for the attack 50 percent of the time. OpenAI’s GPT-5.6 models refused the payload every time they were tested.
  • Adversa AI reported the issue through GitHub’s bug bounty program on September 17, 2026. GitHub investigated and declined to classify it as a product vulnerability.

How Zombie Instructions Hijack a Coding Agent

Here’s the trick. Instead of hiding a plain-text command on a web page and hoping an AI agent reads it, Adversa AI’s researchers encrypted the malicious instructions first. They handed Copilot CLI the ciphertext, the decryption key, and a polite little note telling it to go ahead and decrypt the thing.

That’s the whole move. The agent, operating in autopilot mode, treats decryption as just another coding task. It fires up its own code execution runtime, runs the decryption routine, and only then does the real instruction wake up. Adversa AI calls these “zombie instructions” because they’re dead text until the agent itself breathes life into them.

As the researchers put it: “Static guardrails read text; they do not run it.” That’s the entire vulnerability in eight words. Guardrails scan for suspicious phrases in plain content. They were never built to scan the output of code the agent hasn’t executed yet. In the most severe version of the attack chain, Copilot CLI used a local.env file as part of the decryption key material, then exfiltrated the resulting secrets through an outbound HTTP request.

GitHub investigated the report and responded with a clear line in the sand: “After investigating, we determined this requires a user to intentionally direct Copilot CLI to fetch attacker-controlled or untrusted content and confirm they want to trigger the action, and thus is not a product vulnerability.” Adversa AI disagrees, and points out the attack chain currently works exactly as described, regardless of how it gets classified.

Who’s Actually Holding the Keys Here

I’ve read a lot of prompt injection writeups this year, and most of them rely on tricking a model with cleverly worded plain text. This one is different. It doesn’t bother arguing with the guardrail at all. It just waits for the guardrail to look away and lets the agent do the dirty work itself, using its own permissions, inside its own runtime.

Think of it like slipping a sealed envelope past airport security because the scanner only checks what’s printed on the outside. The envelope says “birthday card.” Nobody opens it at the checkpoint, because opening it isn’t the checkpoint’s job. The opening happens later, in a room with no cameras, and by then the envelope has already been cleared. Copilot CLI’s code execution runtime is that room with no cameras.

Should GitHub’s classification call bother you more than the technique itself? Honestly, both. GitHub’s position isn’t unreasonable on its face. Copilot CLI’s autopilot mode is opt-in, not default, and a user does have to point it at the untrusted page in the first place. But that framing puts the entire burden on developers to never ask an AI assistant to summarize a webpage, check documentation, or fetch a dependency’s README, which is exactly the kind of mundane task autopilot mode exists to automate.

The competitive numbers make this messier, not cleaner. Microsoft’s own mai-code-1.1-flash model got fooled half the time. OpenAI’s GPT-5.6 models, running the identical attack payload, refused every single time. That’s not a small gap. It suggests the vulnerability isn’t really baked into Copilot CLI’s architecture so much as it’s baked into which underlying model you point the architecture at, which is an uncomfortable thing for GitHub to hear about its own flagship coding tool.

And here’s the part that should worry anyone building agent tooling right now: if encryption is enough to blind a static guardrail, how many other guardrails are checking the wrapping paper instead of the contents?

A Pattern, Not an Isolated Incident

This isn’t Adversa AI’s first rodeo with Cryptographic Context Injection. The same research team reported a near-identical CCI vulnerability in Grok roughly two months before this Copilot CLI disclosure, which suggests the technique generalizes across different coding agents rather than exploiting something uniquely broken in GitHub’s product.

Copilot CLI has also had prior brushes with indirect prompt injection concerns tied to its ability to browse and act on untrusted content, so this isn’t the tool’s first appearance in a security researcher’s writeup either. What makes the autopilot angle worth tracking is the design choice sitting underneath it. In Anthropic’s Claude Code, autopilot-style autonomous execution is enabled by default. In Copilot CLI, it remains something a developer has to switch on deliberately.

That difference matters for how you read GitHub’s rejection of the bug bounty report. If autopilot were the default behavior, the “user chose to do this” defense would carry a lot less weight, because most users wouldn’t have consciously chosen anything. As it stands, GitHub can plausibly argue the feature requires informed consent. Whether that argument holds up once agentic coding tools become the default way people interact with their codebases is a separate question, and one the industry hasn’t answered yet.

What To Watch From Here

Keep an eye on whether GitHub revisits its classification. Public disputes between researchers and vendors over severity ratings have a habit of shifting once enough outlets start asking questions, and this is exactly the kind of story that tends to generate follow-up scrutiny.

Watch the model layer too. If GPT-5.6’s resistance to CCI-style attacks holds up under further testing, expect coding agent vendors to start quietly favoring certain underlying models for security-sensitive autopilot tasks, not just for raw coding ability.

And watch whether Adversa AI or other researchers find CCI working against additional agents beyond Grok and Copilot CLI. A technique that’s succeeded against two separate products in two months isn’t a one-off curiosity. It’s a pattern that every team shipping an autonomous coding agent should be actively testing against right now, not after the next disclosure.

Editor's Note

What gets me here isn't the exploit, it's the shrug. GitHub's own words say a user just has to point Copilot CLI at a bad webpage and click confirm, as if that isn't exactly what autopilot mode is designed to let people do casually. I'd bet this classification doesn't hold once a few more researchers replicate it against other models. The GPT-5.6 refusal rate is the detail I'm watching most. It tells me this fight is moving from prompt engineering to model selection, and most teams aren't thinking about it that way yet.

– Sanket Chaukiyal, founder, SmartChunks


Source: The Register

Sanket Chaukiyal — Editor at Smart Chunks

Sanket Chaukiyal

Technology editor • 12+ years in editorial

Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.

All articles → • LinkedIn