News
AI SHIFT:

AI Coding Agents Face Sandbox-Escape Findings Across Four Tools

Newsroom brief

BleepingComputer reported that Pillar Security reproduced sandbox-escape paths in Cursor, OpenAI Codex, Gemini CLI and Google Antigravity, shifting attention from agent containment to trusted developer tools around the workspace.

Verified against source materialEdited by SendTech Times Cybersecurity Desk
AI Coding Agents Face Sandbox-Escape Findings Across Four Tools
Image source: bleepingcomputer.com

Four major AI coding tools face a sandbox-design warning after BleepingComputer reported that Pillar Security reproduced escape paths in Cursor, OpenAI Codex, Google's Gemini CLI and Google Antigravity without directly attacking the sandbox layer.

The finding turns the security question away from whether an agent remains inside a restricted workspace.

Pillar's research centres on what happens when a trusted tool outside that workspace later reads or runs files that the agent was allowed to create.

Agent Sandboxes Meet Trusted Developer Tools

Enterprise development teams run coding agents inside IDEs, command-line tools and local developer workflows rather than isolated cloud demos.

BleepingComputer identified the affected tools as Cursor, OpenAI Codex, Gemini CLI and Antigravity, while naming Pillar researchers Eilon Cohen, Dan Lisichkin and Ariel Fogel as the team behind the work.

Pillar grouped seven findings into four failure modes: denylist sandboxes that lag operating-system behaviour, workspace configuration files that act as executable code, allowlists that trust a command name more than its arguments, and privileged local daemons that sit outside the sandbox.

The pattern does not require the agent to break its own instructions; the risk appears when the surrounding developer environment treats agent-written files as trusted input.

The affected layer is the surrounding toolchain rather than the model alone.

Extensions, task runners, hooks, Git integrations, interpreters and local services can operate with broader host privileges than the agent workspace.

Cursor, Codex And Gemini Fixes Cover Most Findings

Most of the issues have patches or vendor acknowledgement.

Pillar's material names CVE-2026-48124 for one Cursor finding and records Cursor fixes in version 3.0.0 for more than one issue.

Pillar's disclosure records OpenAI's v0.95.0 Codex CLI patch and a high-severity bounty, with a CVE pending.

A separate Docker socket issue affected Codex, Cursor and Gemini CLI and is now fixed.

The disclosure does not provide customer incident counts or evidence that the findings were exploited in production environments.

The Antigravity handling differs from the rest of the disclosure set.

Google's two findings involved a macOS Seatbelt denylist bypass and a task-configuration bypass of Secure Mode; Pillar's account says Google classified both as valid security vulnerabilities but downgraded severity because exploitation would require social engineering or user trust in a repository carrying indirect prompt injection.

AI Coding Security Extends Beyond The Agent

The same class of problem had already appeared earlier in research from Cymulate, which used the term Configuration-Based Sandbox Escape for a pattern spanning Claude Code, Gemini CLI and Codex CLI.

Pillar's newer work broadens that issue across four tools from three vendors and ties it to everyday developer tooling around repositories and local services.

For security teams, the operating boundary is not only the AI model or the prompt.

Review processes must also cover which local tools can act on files created by agents, which daemons are reachable from development workspaces and whether vendor sandboxes monitor post-write execution by trusted host components.

Pillar's proposed direction is to watch the moment a trusted local tool runs something an agent wrote rather than relying only on lists of banned filenames.

Production exploitation evidence and customer-level mitigation data remain unnamed.

Share this article
inXf

Related articles

More
Neo Raises $100M To Control Enterprise AI Software Actions
Cybersecurity

Neo Raises $100M To Control Enterprise AI Software Actions

SecurityWeek reported that Neo emerged from stealth with $100 million for a platform that governs AI agents, MCP servers and software actions across enterprise systems.

AI Coding Push Turns Developers Into a Prime Cybersecurity Target
Cybersecurity

AI Coding Push Turns Developers Into a Prime Cybersecurity Target

A Japanese @IT analysis says attackers are increasingly targeting developers because AI coding tools, OSS, CI/CD pipelines and cloud services concentrate valuable credentials around them. The report highlights vulnerable AI-generated code, fake recruiting approaches, polluted open-source packages and GitHub Actions-style automation attacks. The practical warning is that companies need stronger identity, dependency and workflow controls rather than relying only on individual developer caution.

Injective SDK npm Compromise Exposes Wallet-Key Theft Risk
Cybersecurity

Injective SDK npm Compromise Exposes Wallet-Key Theft Risk

Socket, Ox Security and StepSecurity said they detected wallet-stealing code in @injectivelabs/sdk-ts npm package version 1.20.21 after an Injective Labs contributor account was compromised. Socket said the malicious release was downloaded 310 times before deprecation, while Ox Security counted 87 direct dependencies and described a six-figure cumulative download count across dependent packages.

Socket Tracks 108 Malicious Packages In PolinRider Supply-Chain Attack
Cybersecurity

Socket Tracks 108 Malicious Packages In PolinRider Supply-Chain Attack

Socket reported 162 malicious release artefacts across 108 packages in the PolinRider supply-chain campaign. Victim companies remain outside the public record.

Fake GitHub Repositories Turned Developer Trust Into BoryptGrab Delivery Chain
Cybersecurity

Fake GitHub Repositories Turned Developer Trust Into BoryptGrab Delivery Chain

DeveloperTech's article on Arctic Wolf Labs research describes a fake-repository campaign that used polished GitHub project pages as a delivery route for BoryptGrab malware. The case makes artifact provenance and workstation controls more important than visual trust in repository pages.

Sysdig Says AI Ransomware Still Needed Human Setup
Cybersecurity

Sysdig Says AI Ransomware Still Needed Human Setup

Sysdig described JadePuffer as agentic ransomware, but Michael Clark said a human still chose the victim, provisioned infrastructure and supplied database credentials before the AI agent executed the attack.

Keep Reading

More Stories

Latest
e& UAE And Core42 Launch Sovereign AI Compute PlatformCloud & Data CentersJul 21, 2026e& UAE And Core42 Launch Sovereign AI Compute PlatformMiddle East AI News reported that e& UAE and Core42 launched Sovereign AI Compute, giving UAE enterprises and government bodies in-country GPU access with data residency, connectivity and vendor-claimed zero egress fees.Microsoft Adds AMD Helios AI Racks To Azure Without Order SizeChips & SemiconductorsJul 21, 2026Microsoft Adds AMD Helios AI Racks To Azure Without Order SizeMicrosoft will deploy AMD Helios rack-scale AI accelerators for Azure AI workloads, with watts, dollars and rack counts still absent from the public terms of the commitment.AliExpress Hit With Record €550m EU Fine Over Illegal GoodsCapital & PolicyJul 21, 2026AliExpress Hit With Record €550m EU Fine Over Illegal GoodsBBC reported that the European Commission imposed a record €550m Digital Services Act penalty on AliExpress and ordered the Alibaba-owned marketplace to file a corrective action plan by 20 October.Z.ai Tests Gigawatt AI Data Centre Built On Domestic ChipsCloud & Data CentersJul 21, 2026Z.ai Tests Gigawatt AI Data Centre Built On Domestic ChipsUnite.AI reported that Z.ai has begun operating part of a gigawatt-class AI data centre built on Chinese-made chips, highlighting how export controls are pushing large-scale model training toward domestic compute stacks.Digital Takumi AWS Route 53 Outage Exposes Root-Account Recovery RiskCloud & Data CentersJul 21, 2026Digital Takumi AWS Route 53 Outage Exposes Root-Account Recovery RiskThe Register links a suspended AWS Route 53 account to Digital Takumi-managed websites and connected Google Workspace email going offline after billing warnings, MFA recovery and DNS hosting sat inside one operating loop.SBI Takes Majority Control Of Coinhako After MAS ApprovalFintech & Digital PaymentsJul 21, 2026SBI Takes Majority Control Of Coinhako After MAS Approvale27 reported that SBI Holdings acquired a majority stake in Singapore crypto exchange Coinhako after MAS approval, using a capital injection and share purchase to turn the platform into a consolidated subsidiary. Financial terms and product timelines were not disclosed.PULSE Gives 10 US Health Agencies OpenAI And Anthropic AI PilotsAIJul 21, 2026PULSE Gives 10 US Health Agencies OpenAI And Anthropic AI PilotsPULSE will let 10 US public health jurisdictions test OpenAI and Anthropic enterprise AI tools, with capacity for up to 2,000 practitioners and playbooks expected in 2027, AI News reported.PJM Backup-Generator Warnings Test Large-Load Grid ToolCloud & Data CentersJul 21, 2026PJM Backup-Generator Warnings Test Large-Load Grid ToolData Center Knowledge reported that PJM issued emergency backup-generator warnings during a July heat wave but did not dispatch customer-owned generators, leaving total available large-load backup capacity undisclosed.Taiwan Mobile Extends Nokia 5G Deal Around AI-Native Network OperationsTelco & ConnectivityJul 21, 2026Taiwan Mobile Extends Nokia 5G Deal Around AI-Native Network OperationsRCR Wireless News reported that Taiwan Mobile and Nokia have extended their 5G partnership with AI-native network operations, including AirScale equipment, MantaRay SON automation, RedCap support and a 100% renewable-electricity target by 2040.AI Agent Rollouts Require Testing Before Live CustomersAIJul 21, 2026AI Agent Rollouts Require Testing Before Live CustomersNo Jitter reported that enterprise AI agents need guardrails, simulations, answer checks and visibility before customer-facing deployment, as vendors add tools to catch regressions and rollback failures.FCC Names Three Unlicensed Bands For Satellite D2D VoteTelco & ConnectivityJul 21, 2026FCC Names Three Unlicensed Bands For Satellite D2D VoteAn FCC draft order names 902MHz-928MHz, 2400MHz-2483.5MHz and 5725MHz-5850MHz for a proposed unlicensed direct-to-device satellite proceeding ahead of an August 6 commissioner vote.Coratia Gets Rs 66 Crore Navy Order For Underwater RobotsCapital & PolicyJul 21, 2026Coratia Gets Rs 66 Crore Navy Order For Underwater RobotsYourStory reported that Coratia Technologies has a Rs 66 crore Indian Navy contract for indigenous underwater ROVs, moving the Odisha startup from inspection prototypes toward defence delivery.