Tags: prompt injection*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Amanda Caswell writes that Google's Gemini CLI version 0.61.0 introduces new security safeguards to prevent prompt injection attacks by requiring manual user confirmation for sensitive operations. The update requires explicit approval before the coding agent can edit specific build configuration files, run subsequent test or build commands after edits, or execute shell commands containing arguments derived from untrusted external content like web searches or Google Docs.

    - Security checks are designed to prevent attackers from using indirect prompt injection via malicious documentation or fetched data.
    - The update hardens the Gemini CLI sandbox by stripping sensitive information such as API keys and OAuth credentials before it is mounted inside a container.
    - Users cannot set "always allow" permissions for actions involving untrusted context, ensuring human oversight remains mandatory in those specific scenarios.
  2. Dan Goodin writes that ASCII smuggling, a technique once primarily used to hide malicious prompt injections from Large Language Models (LLMs), has been adopted by spammers to evade email filters. By using invisible Unicode tags that mimic the structure of standard text, attackers can bypass machine learning-based spam detectors and natural language processing models without alerting human readers. This method allows words like "funding" or "credit" to be broken into non-standard tokens that escape keyword detection while remaining perfectly readable to a person once rendered in an email client.

    - ASCII smuggling uses the Unicode Tags block, which contains 128 characters designed to be invisible to humans but readable by computers.
    - Microsoft observed spam signatures using this technique spike from 21,000 per day to over 2.5 million within a four-day period in early February.
    - The method is effective against modern AI-driven filters because it disrupts the way tokenizers process words into sub-word pieces.
    - While originally used for stealthy prompt injections, spammers now use it specifically to obfuscate financial keywords from automated detectors.
  3. Jessica Lyons writes that researcher Johann Rehberger, known as wunderwuzzi, has demonstrated a method for hijacking Anthropic's Claude Code in Auto Mode via prompt injection. By asking the agentic coding model to summarize a malicious website, an attacker can trick it into bypassing its standard WebFetch tool and instead using Bash with `curl` to download files. This chain allows attackers to use "Python module shadowing'' specifically by placing a malicious file named `struct.py` in the same directory as a downloaded archive' to execute arbitrary code on the host system.

    - The attack had success rates between 60% and 80% in tested scenarios.
    - An attacker can successfully trigger "nested" Claude Code instances to create new agents with their own tool access.
    - Anthropic stated that Auto Mode is a convenience feature, not a security guarantee, as the classifier may not catch complex injection chains.
    - Experts recommend running coding agents in isolated sandboxes due to these vulnerabilities.
  4. Anurag Singh replaced five Python scripts (backup, organizer, renamer, cleaner, watchdog) with a local LLM agent, which made errors the scripts didn't (wrong directories, skipped steps, false success reports).Each of the original scripts followed explicit rules through a scheduler; the agent instead added a longer inference chain (inspect, interpret, choose a tool, build a command, execute, review) to tasks that fixed logic already described completely, while also holding a loaded model in memory between runs.

    - AutomationBench scores for frontier models remain well under 20%: GPT-5.6 Sol 18.1%, GPT-5.5 12.9%, Claude Opus 4.8 15.5%, Gemini 3.5 Flash 14.5%
    - Granting an LLM system-level access creates a prompt-injection vector: a malicious file on disk could carry instructions the agent interprets as commands
    - Singh's proposed fix: let the agent classify and route ambiguous requests, then hand off to a validator + fixed script for the actual filesystem action
    - The five original scripts covered photo backup, extension-based Downloads sorting, file renaming, app-cache clearing, and a disk-threshold alert
  5. Snyk Agent Scan provides a way to discover and inspect local agent components like Model Context Protocol (MCP) servers and skills. It identifies various security risks, such as prompt injections, malware payloads in natural language, sensitive data exposure, and credential leaks. The tool offers both an interactive command-line interface for individual users and a background mode for enterprise monitoring through Snyk Evo.

    - Detects 15+ distinct security risks across MCP servers and agent skills
    - Supports agents including Claude Code, Cursor, Windsurf, and Gemini CLI
    - Automatically discovers configurations for various desktop and IDE-based agents
    - Scanning MCP configs executes commands defined in them to retrieve tool descriptions
  6. Charles Ye, Jasmine Cui, and Dylan Hadfield-Menell write that prompt injection in large language models is fundamentally caused by role confusion, where the model's internal representations fail to maintain the boundaries established by structural tags like system, user, or tool. Because these models often rely on superficial linguistic cues rather than formal markers to identify roles, attackers can successfully escalate privileges simply by mimicking a more authoritative writing style.

    - The "CoT Forgery" attack leverages this vulnerability by inserting fake reasoning blocks that trick the model into following malicious instructions under its own perceived authority.
    - Experimental results show that minor linguistic changes, such as replacing the phrase "The user" with "The request", can significantly reduce the effectiveness of these attacks.
    - The authors propose that roles could be used to structurally isolate competing objectives within models, potentially improving both performance and safety.
  7. Researchers from Google and Forcepoint have identified a rise in indirect prompt injection (IPI) attacks, where malicious instructions are hidden within web pages to manipulate LLM-powered AI agents. While some injections are harmless pranks or tone adjustments, others aim for serious harm including traffic hijacking, data exfiltration, denial of service, and financial fraud through unauthorized payment processing. Attackers use techniques like invisible text, HTML comments, and metadata manipulation to hide these payloads from humans while remaining visible to AI.
    Key points:
    * Real-world evidence of IPI attacks found in massive web crawls and active threat hunting.
    * Malicious intents include search engine manipulation, data theft (API keys), and destructive commands.
    * Financial fraud attempts have been observed using embedded PayPal transactions and Stripe donation routing.
    * Attackers hide instructions via single-pixel text, near-transparent colors, or metadata injection.
    * The risk level scales with AI privilege; agentic AIs capable of executing commands or payments are high-impact targets.
  8. This article details a hands-on experience with Nvidia's NemoClaw, a security-focused stack designed to enhance the safety of the OpenClaw AI platform. While NemoClaw introduces improvements like a sandbox model and aggressive policy filtering, the author finds it still falls short of being a reliable solution.
    Bugs, limitations, and the inherent risks associated with OpenClaw's architecture—particularly its connection to external services—persist. The core issue remains that NemoClaw can secure the agent but cannot protect against malicious instructions embedded in external data sources like emails or messages.
    The author concludes that while NemoClaw is a step forward, it doesn't fully address the fundamental security concerns surrounding OpenClaw.
  9. Despite initial excitement and a viral moment, some AI experts are questioning the usability of OpenClaw due to inherent cybersecurity flaws. The article details the vulnerabilities discovered in Moltbook, a social network built on OpenClaw, and explores whether the technology's access and productivity benefits outweigh its security risks.
  10. This article discusses a new paper outlining design patterns for mitigating prompt injection attacks in LLM agents. It details six patterns – Action-Selector, Plan-Then-Execute, LLM Map-Reduce, Dual LLM, Code-Then-Execute, and Context-Minimization – and emphasizes the need for trade-offs between agent utility and security by limiting the ability of agents to perform arbitrary tasks.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "prompt injection"

About - Propulsed by SemanticScuttle