HomeAIAI Agent Security: Langflow's Critical Flaw, Copilot's Encrypted-Instruction Finding and the New...

AI Agent Security: Langflow’s Critical Flaw, Copilot’s Encrypted-Instruction Finding and the New Guardrails

The weak point in AI agents this week was not the model’s intelligence. It was what an agent was allowed to run, and what it would send out. Four stories from 28 September to 6 October 2026 make the point.

  • Langflow fixed a critical flaw that let a tool setting launch any command on the server (confirmed by the project’s advisory) [1]. IBM has since published a separate batch of Langflow fixes [2].
  • GitHub Copilot CLI was shown by researchers to follow instructions hidden in encrypted text (reported). GitHub accepted the finding but says it is not a vulnerability, and no one has independently reproduced it [3].
  • Atlassian rebuilt its agent tool server with read, write and destructive tiers (confirmed by Atlassian) [4].
  • OpenAI made Codex Auto-review, a second agent that checks risky actions, free for ChatGPT sign-ins (confirmed) [5][6].

No exploit steps or payloads appear below.

What went wrong in Langflow?

Langflow is an open-source tool for building AI agents and workflows visually. It can connect to outside tools through the Model Context Protocol (MCP), a common standard for plugging tools into AI assistants.

According to Langflow’s GitHub advisory, rated critical (CVSS 9.9), the local (“stdio”) way of connecting to an MCP server would launch whatever command was written in the server’s settings, with no allowlist [1]. Any logged-in user without admin rights could use that. With the default auto-login setting, which the project says is meant for development only, no account was needed at all [1]. The reporter found vulnerable servers exposed to the internet [1].

Versions from 1.1.2 up to, but not including, 1.10.3 are affected. A partial fix arrived in 1.9.0; 1.10.3 completes it, and 1.11.0 adds further hardening [1]. The advisory is dated 28 September.

Separately, IBM published a bulletin on 2 October covering 25 vulnerabilities in Langflow OSS 1.0.0 to 1.12.2, including two rated 9.8 that need no login. It tells users to upgrade to 1.12.3 [2]. The two documents do not say whether any of IBM’s entries is the same flaw as the GitHub advisory.

For anyone running Langflow, the defensive advice is plain: update to 1.12.3 or later, keep auto-login off, and do not expose it to untrusted networks [1][2].

What did researchers report about Copilot CLI?

Adversa AI says that when GitHub’s Copilot command-line agent was running on its own (“autopilot”), it decrypted an encrypted message it had been asked to fetch and treated the contents as trusted instructions. In Adversa’s test, it then read a local secrets file and sent it out, in 28 seconds. The same instructions in plain text were refused as prompt injection, meaning hidden commands planted in content [3].

The result depended on the model. Adversa says Microsoft’s mai-code-1.1-flash followed the chain in half its runs, while two GPT-5.6 models refused. On Auto routing, the user does not see which model is answering [3].

Adversa reported the issue to GitHub on 17 September. GitHub validated it but declined to treat it as a vulnerability, saying the user had “explicitly asked Copilot CLI to fetch attacker-controlled content while giving copilot full permissions to act autonomously” [3]. Adversa withheld its payloads and sells detection tools, so it has a commercial interest [3].

How is Atlassian changing its MCP server?

Atlassian says its hosted MCP server, now generally available, offers more than 220 tools and handles more than 15 million calls a day [4]. It uses OAuth 2.1 sign-in and respects each user’s existing permissions. “Read, write, and destructive actions are separated so clients can require explicit confirmation on higher-risk steps,” the company says. Admins get allowlists and audit logs [4]. Atlassian also claims up to 25% token savings in its own tests on Claude models [4]. All of these figures are Atlassian’s.

What changed with Codex Auto-review?

When a Codex agent wants to step outside its sandbox, Auto-review hands the approval request to a separate reviewer agent instead of the user. OpenAI calls it “a reviewer swap, not a permission grant” [5]. The reviewer is meant to block actions such as sending secrets to untrusted destinations, probing for credentials, broadly weakening security and destructive steps. It does not see the main agent’s hidden reasoning, and it stops the session after three denials in a row or ten in the last 50 reviews [5]. The policy is open source [5].

On 6 October OpenAI’s Thibault Sottiaux said Auto-review is now free for everyone signed in with a ChatGPT account [6]. A post on OpenAI’s developer forum adds that it does not count against plan usage [7]. OpenAI’s own documentation says it is “not a deterministic security guarantee” and can make mistakes [5].

What links the four?

Each moves the safety check outside the model’s own judgement. Auto-review targets exactly what Adversa describes: secrets leaving the machine. Langflow shows the reverse: if the tool layer runs whatever it is given, no model-level care helps.

What this does not prove

  • That the Copilot technique works widely. It is one company’s test, not independently reproduced, and GitHub does not classify it as a vulnerability [3].
  • That Auto-review stops leaks. OpenAI says it is not a guarantee [5].
  • Atlassian’s usage and token figures. They are company claims [4].
  • That Langflow servers have been attacked. Neither the advisory nor IBM’s bulletin reports exploitation in the wild [1][2].

The Bottom Line

Agent security is moving to the plumbing. Patch Langflow to 1.12.3 or later [2]. Treat anything an autonomous agent fetches as untrusted, whatever form it arrives in [3]. And expect more products to sort actions by risk and put a checkpoint in front of the dangerous ones, as Atlassian and OpenAI now do [4][5].

Related on TSN: MCP Explained: The USB-C for AI Tools · Cyber breaches, week of 7 October 2026 · Fences for AI Agents (agent sandboxes, post 20010; add link once live)

Sources

  1. Langflow, “Langflow < 1.10.3: OS command injection (RCE) via arbitrary command in MCP stdio server configuration,” GitHub security advisory GHSA-w794-rj3p-xv45, published 28 September 2026. https://github.com/langflow-ai/langflow/security/advisories/GHSA-w794-rj3p-xv45
  2. IBM, “Security Bulletin: Langflow OSS is affected by multiple vulnerabilities,” initial publish date 2 October 2026. https://www.ibm.com/support/pages/node/7290694
  3. Rony Utevsky, “GitHub Copilot CLI vulnerability leaks developer secrets” (on “Cryptographic Context Injection”), Adversa AI blog, 6 October 2026. https://adversa.ai/blog/cryptographic-context-injection-github-copilot/
  4. Atlassian, “Atlassian MCP brings more of your work within reach of any AI agent,” Inside Atlassian blog, 6 October 2026. https://www.atlassian.com/blog/company-news/team26-europe-atlassian-mcp
  5. OpenAI, “Auto-review,” Codex documentation, read 8 October 2026. https://developers.openai.com/codex/concepts/sandboxing/auto-review
  6. Thibault Sottiaux (@thsottiaux), “Day 2.1/ We have made Auto-review free for all users signed in through a ChatGPT account…,” post on X, 6 October 2026, 08:13 BST. https://x.com/thsottiaux/status/2107368734981517634
  7. OpenAI Developer Community, “Day 2: Auto-review is now free…,” post in the “28 days of Quality of Life improvements” thread, 6 October 2026, 08:42 BST. https://community.openai.com/t/free-auto-review-day-2-of-28-days-of-quality-of-life-improvements-or-a-full-reset/1403525

Source note: Langflow’s advisory includes a proof of concept; nothing from it is reproduced here. The brief dated the Langflow item 6–7 October; the GitHub advisory itself is dated 28 September 2026. The Adversa headline calls the issue a “vulnerability”; GitHub disagrees, and this piece keeps the “reported” label.

Share this story

More in this category

Latest on TSN

Free TSN tools: crypto calculator, Flux dashboard and more.