HomeBig TechFences for AI Agents: Microsoft's Execution Containers and AWS's Strands Box

Fences for AI Agents: Microsoft’s Execution Containers and AWS’s Strands Box

AI agents, software that uses AI models to take actions rather than just answer questions, are getting their own fences. On 7 October 2026 Microsoft made Microsoft Execution Containers (MXC) generally available on Windows 11, as part of a broader rebuild of Windows for running AI both on the PC and in the cloud [1][2]. The same day, AWS open-sourced Strands Box, a sandbox that enforces rules on agents “regardless” of what the agent itself decides [5][6].

Both launches are confirmed by the companies. Reuters reports that Anthropic, OpenAI and Nvidia will use Microsoft’s containers [4]; that is labelled as reported below, though Microsoft’s own lists point the same way.

What does an enforced sandbox actually do?

Start with the problem. An agent can read files, run code, call online services and keep going for hours without anyone watching every step. Most of the time that is the point. But an agent can also take a step that seems sensible to it and is wrong for you: changing a production setting, deleting the wrong folder, posting too often, or running up an API bill.

You could tell the agent not to do those things. The trouble is that an instruction is only advice. The agent might misunderstand it, ignore it, or be tricked by something it reads.

An enforced sandbox moves the rule out of the agent’s hands. In plain words:

  • The agent runs inside a boundary set by a person or an organisation.
  • The boundary lists what the agent may touch: these folders, these websites, these tools.
  • The operating system, or a policy layer the agent cannot change, checks each action against that list and blocks anything outside it.
  • The agent cannot grant itself extra access, however it reasons.

Microsoft puts it simply: “An agent cannot be its own security authority. It must run within a boundary defined by the developer or organization and enforced independently of the agent itself” [2]. AWS’s Marc Brooker told The Register that with Strands Box, “the agent can’t talk its way around these rules” [5].

Part 1: Microsoft Execution Containers

What is MXC?

Microsoft describes MXC as a policy-driven layer for running untrusted or AI-generated workloads. Developers and IT administrators declare which files and network destinations an agent can use, and MXC enforces that boundary at runtime, with the policy kept outside the agent’s control [2].

Microsoft’s example is a coding agent asked to update a website. It needs to read and write the website’s code and to read the server configuration, but it should not be able to change that configuration. Without a boundary, the agent might decide editing the server settings is the quickest fix and break the live site. With MXC, Microsoft says, the attempt is blocked “regardless of what the model, generated code, plugin, or tool decides to do” [2].

MXC offers several levels of isolation, from a lightweight process sandbox (available on Windows 11, macOS and Linux) to a separate Windows session for long-running agents, a Linux environment through WSL, and an experimental micro virtual machine for higher-risk work [2]. It also has a “learning” mode that records what an agent tried to reach, so administrators can write a tight policy without guessing [2].

Microsoft says Windows will soon let its Entra identity service tell agent activity apart from the person using the PC, and extend its Agent 365 management controls to agents running on the device [1][2].

Who is using it?

Microsoft says agents that already support MXC include OpenAI’s Codex, GitHub Copilot, OpenClaw, Replit, LM Studio, NVIDIA’s OpenShell and Unsloth AI. It lists others that “will be releasing support”, including Anthropic’s Claude Code, Box, Hermes Agent by Nous Research, Manus and Perplexity [1][2]. Meta’s Muse agent is coming to Windows as a native app with MXC integration, Microsoft adds [1].

Reported: Reuters, in a report carried by MarketScreener, says Windows chief Pavan Davuluri told the San Francisco launch event that Anthropic, OpenAI and Nvidia will use the tools [4]. Reuters also quotes Microsoft chief executive Satya Nadella: “We needed to make the desktop the most secure place for agents to execute” [4].

What else is changing in Windows?

MXC sits inside a broader pitch Microsoft calls “hybrid intelligence”: agents that run locally when that makes sense and reach the cloud when they need to [1]. The main confirmed pieces [1]:

  • Local coding model. Microsoft is bringing its MAI Code 1.1 Flash coding model onto the device, using 3-bit precision to shrink it by nearly 80% while, it says, preserving coding quality and a 256,000-token context window.
  • Local-plus-cloud routing. GitHub’s HydraFusion, which routes tasks to suitable cloud models, will be able to use models on the PC too. Microsoft says this comes to the GitHub Copilot app, Copilot CLI and Visual Studio Code in experimental preview later in October.
  • Open-source runtime. llama.cpp support in Windows ML, Microsoft’s on-device AI runtime.
  • Copilot with local access. On Copilot+ PCs, Copilot will be able to use local context “with your permission”, take actions on the PC and use local models. Microsoft says these features are expected to begin rolling out “in the coming months”.

Surface Laptop Ultra: price and date

The hardware anchor is NVIDIA’s RTX Spark chip. Microsoft opened pre-orders for Surface Laptop Ultra, which it calls its most powerful Surface Laptop, starting at $2,599 (MSRP), with availability beginning 16 October [3]. Microsoft says it pairs an NVIDIA Blackwell RTX GPU and Grace CPU with up to 128 GB of unified memory, and can run AI models of more than 120 billion parameters locally [3].

Price is the catch. Reuters reports the Ultra rises to $5,899 for a model with a 20-core processor, 128 GB of memory and 1 TB of storage, and that the entry model costs more than Apple’s base MacBook Pro, though with more memory [4]. “Now the software is ready and the hardware is too expensive to actually run it locally,” Moor Insights & Strategy analyst Anshel Sag told Reuters [4].

A desk-based Surface RTX Spark Dev Box is available to pre-order at $5,999 (MSRP), on Microsoft.com in the US only, shipping in November [3]. Microsoft says RTX Spark laptops from ASUS, Dell, HP, Lenovo and MSI are also on pre-order, shipping from 16 October [1]. Microsoft’s performance comparisons against Apple’s MacBook Pro come from tests it or NVIDIA commissioned on pre-production hardware [1]; we have not seen independent results.

Part 2: AWS Strands Box

What is Strands Box?

Strands Box is an open-source local sandbox for agents, released by AWS in developer preview [6]. It works with any agent framework and any AI model [6].

AWS’s argument is that ordinary containers and virtual machines are good at keeping an agent in, but poor at judging what it is doing. “Agents increasingly run in ‘YOLO mode,’ approving every action without human review,” the AWS team wrote, as quoted by The Register. “The usual solution to this problem is a sandbox … but access is only part of what we want to control” [5].

So Box adds what AWS calls semantic policy: rules about specific actions, enforced deterministically [6]. AWS’s examples [5][6]:

  • allow an agent to git push code, but only if tests have passed;
  • allow it to use a payments API, but only up to $100 a day;
  • allow it to post status updates to Slack, but no more than three times every ten minutes.

The Register notes that the policy engine has a sense of time and history, so it can check a request against what the agent has already done, not only what it wants to do now [5].

How does it enforce the rules?

In outline only: Box uses the operating system’s own isolation features so that the only way out of the sandbox is through a policy layer. That layer can see shell commands, Python, tool calls, file access and network access, and allow or block each one [6]. The rules are written in Dogwood, AWS’s open-source policy language [6].

AWS is candid about limits. Brooker writes that Box relies on more complex code than a bare virtual machine, so bugs could mean bypassed policy, and AWS recommends pairing it with a separate micro virtual machine per session for multi-user cloud deployments [6]. He also told The Register that even with correct permissions an agent can still produce an unwanted result: “Developers remain responsible for deciding what access to grant and where human review is needed” [5].

Where does it run?

For now, only on macOS. The Register reports Linux support is in development and a Windows client is “on our radar”, with no release dates; deployment to AWS AgentCore, ECS and Kubernetes is planned [5].

Update, 8 October 2026

OpenAI has added a different kind of fence. On 6 October it made Codex Auto-review free for everyone signed in with a ChatGPT account [7].

Normally, when a Codex agent wants to do something its sandbox does not allow, it stops and asks the user. Auto-review sends that request to a separate reviewer agent instead. OpenAI calls this “a reviewer swap, not a permission grant”: the sandbox limits stay the same, and only the approver changes [8]. The reviewer is told to block actions such as sending secrets to untrusted destinations, probing for credentials and destructive steps. If it keeps refusing (three denials in a row, or ten in the last 50 reviews), the session stops [8].

That is a judgement call made by a model, not a fixed rule. Strands Box enforces rules such as “no git push” whatever the agent decides [5][6]; Auto-review asks a second agent to decide. OpenAI’s documentation says plainly that it is “not a deterministic security guarantee” [8]. The two approaches can sit together: hard limits for what must never happen, a reviewer for the grey areas.

Update, 8 October 2026 (GitHub Copilot)

One open question has already been answered. On 7 October GitHub made local sandboxing for GitHub Copilot generally available, and it runs on Microsoft Execution Containers [9].

The sandbox covers Copilot CLI, the Copilot app and VS Code sessions using Agent Host. Commands and tools that Copilot starts on a developer’s machine run with limited access to files and folders, the internet and local networks, and Git and GitHub CLI credentials, according to a policy set by the developer or their organisation [9]. GitHub says MXC translates one sandbox policy into native controls on Windows, macOS and Linux, so the fence is not Windows-only [9].

Two details match the argument above. Enterprises can require the sandbox and “enforce policies that developers cannot weaken”, and the policy applies to tool execution “regardless of which model Copilot uses” [9]. In other words, the fence does not depend on the agent behaving well. It is included with Copilot at no extra cost [9].

What this does not prove

  • That sandboxes make agents safe. They limit what an agent can reach. They do not make its choices good, and both companies say human judgement is still needed [2][5].
  • That every listed partner has shipped support. Microsoft separates agents that “already support” MXC from those that “will be releasing support”, including Anthropic’s Claude Code [2]. The Reuters line about Anthropic, OpenAI and Nvidia is reported [4].
  • Microsoft’s performance claims. They come from Microsoft- or NVIDIA-commissioned tests on pre-production machines [1].
  • That Strands Box is production-ready. AWS calls it a developer preview, macOS only for now [5][6].
  • Timing of every Copilot local feature. Microsoft said “coming months” [1]; GitHub’s local sandboxing for Copilot, built on MXC, became generally available on 7 October [9]. Other features in Microsoft’s announcement may follow later.
  • That a reviewer agent is as strong as a hard rule. OpenAI says Auto-review can make mistakes and is not a guarantee [8].

The Bottom Line

The industry is converging on one idea: do not ask an agent to police itself. Microsoft’s Execution Containers, now generally available on Windows 11, let organisations fence agents in at the operating-system level, and Microsoft names OpenAI’s Codex among the agents already supporting them, with Reuters reporting Anthropic, OpenAI and Nvidia will use them [2][4]. AWS’s Strands Box adds rules about specific actions, such as no git push until tests pass, enforced whatever the agent decides [5][6]. Around MXC, Microsoft is rebuilding Windows for local-plus-cloud AI, with Surface Laptop Ultra from $2,599 and available from 16 October [1][3]. Fences will not make agents wise. They do make mistakes smaller.

Related on TSN: Wikipedia’s Uninvited Guests: Wikimedia Finds OpenAI “Rogue” Agents on Its Sites

Sources

  1. Pavan Davuluri, “Building Windows for hybrid intelligence,” Windows Experience Blog, Microsoft, 7 October 2026. https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/
  2. Logan Iyer, “Microsoft Execution Containers: Policy-driven containment for AI agents,” Windows Developer Blog, Microsoft, 7 October 2026. https://blogs.windows.com/windowsdeveloper/2026/10/07/microsoft-execution-containers-policy-driven-containment-for-ai-agents/
  3. Brett Ostrum, “Pre-order our most powerful Surface devices ever,” Microsoft Devices Blog, 7 October 2026. https://blogs.windows.com/devices/2026/10/07/pre-order-our-most-powerful-surface-devices-ever/
  4. Stephen Nellis, “Microsoft brings more AI to PCs as it challenges Apple,” Reuters, 7 October 2026, as carried by MarketScreener. https://www.marketscreener.com/news/microsoft-brings-more-ai-to-pcs-as-it-challenges-apple-ce785ddedd89f421
  5. Brandon Vigliarolo, “AWS launches open-source AI agent sandbox to prevent YOLO mode disasters,” The Register, 7 October 2026. https://www.theregister.com/ai-and-ml/2026/10/07/aws-launches-open-source-ai-agent-sandbox-to-prevent-yolo-mode-disasters/5301687
  6. Marc Brooker, “Strands Box: The Big Picture,” Strands Agents blog (AWS), 7 October 2026. https://strandsagents.com/blog/strands-box-the-big-picture/
  7. Thibault Sottiaux (@thsottiaux), “Day 2.1/ We have made Auto-review free for all users signed in through a ChatGPT account…,” post on X, 6 October 2026, 08:13 BST. https://x.com/thsottiaux/status/2107368734981517634
  8. OpenAI, “Auto-review,” Codex documentation, read 8 October 2026. https://developers.openai.com/codex/concepts/sandboxing/auto-review
  9. GitHub, “Local sandboxing for GitHub Copilot now generally available,” GitHub Changelog, 7 October 2026. https://github.blog/changelog/2026-10-07-local-sandboxing-for-github-copilot-now-generally-available/

Source note: The Reuters report was read via MarketScreener’s syndicated copy, not on reuters.com.

Share this story

More in this category

Latest on TSN

Free TSN tools: crypto calculator, Flux dashboard and more.