AI As a Target

Gerard Chukwu
TLP:CLEARCS-SA-26-0901BCyber Threat Intelligence
Download the Full Advisory (PDF)

Executive Summary

While attackers weaponize AI, a topic covered in our companion briefing, AI As Weapon, the AI systems and infrastructure organizations deploy are themselves accumulating real, exploitable vulnerabilities. In the past 18 months, more than ten significant CVEs have hit AI agent infrastructure, a novel attack class has emerged that poisons a local AI model through nothing more than a visited webpage, an AI gateway library used by tens of millions has had its supply chain compromised, and commodity infostealers have started harvesting AI account sessions to bypass multi-factor authentication entirely.

This briefing, the second in a three-part series from the CRITICALSTART® Cyber Research Unit (CRU), covers two related but distinct problems: vulnerabilities in the AI agent infrastructure itself, and direct attacks on AI provider accounts and sessions. The companion briefing AI By Design covers the architectural weaknesses in AI systems, such as prompt injection, that make full remediation of either problem unlikely any time soon.

Background

AI Agent Infrastructure Has Its Own Growing Attack Surface

CVE-2026-65105, disclosed in NVIDIA's NemoClaw agent deployment tool on August 25, 2026, drew significant attention because it demonstrates a new attack class. To let its sandboxed agent container reach a local inference server, NemoClaw starts Ollama with the flag OLLAMA_HOST=0.0.0.0:11434, binding the model API to every network interface rather than to loopback. Because Ollama does not validate the Host header on non-loopback bindings, a single visit to a malicious webpage can trigger a DNS rebinding attack that gives the page a direct line to the local model server. From there, an attacker can use Ollama's /api/create endpoint to rewrite the model's chat template, appending hidden instructions that are silently inserted into every future system message the agent renders. [1] Reporting on the flaw described the resulting compromise as sitting beneath what any guardrail or human operator can see and persisting across sessions and reboots rather than affecting a single conversation the way ordinary prompt injection does. [2] NVIDIA assigned the flaw a CVSS score of 8.1 and patched the macOS and Linux builds in NemoClaw v0.0.35; as of this writing, the Windows and WSL builds remain exposed. [3]

It would be a mistake to treat this as an isolated event. CVE-2026-65105 is one entry in a growing and repeating pattern of vulnerabilities concentrated in AI agent infrastructure, particularly implementations of the Model Context Protocol (MCP), the open standard that connects AI agents to external tools and data sources. A selection of significant disclosures illustrates the recurring root causes:

CVEComponentCVSSRoot Cause
CVE-2025-6514mcp-remote9.6mcp-remote is exposed to OS command injection when connecting to untrusted MCP servers due to crafted input from the authorization_endpoint response URL [4]
CVE-2025-49596MCP Inspector9.4Versions of MCP Inspector below 0.14.1 are vulnerable to remote code execution due to lack of authentication between the Inspector client and proxy, allowing unauthenticated requests to launch MCP commands over stdio. [4]
CVE-2025-54136Cursor8.8Cursor versions 1.2.4 and below, attackers can achieve remote and persistent code execution by modifying an already trusted MCP configuration file inside a shared GitHub repository or editing the file locally on the target's machine. Once a collaborator accepts a harmless MCP, the attacker can silently swap it for a malicious command (e.g., calc.exe) without triggering any warning or re-prompt. If an attacker has write permissions on a user's active branches of a source repository that contains existing MCP servers the user has previously approved, or allows an attacker has arbitrary file-write locally, the attacker can achieve arbitrary code execution. [5],[6]
CVE-2025-54135Cursor9.8Cursor allows writing in-workspace files with no user approval in versions below 1.3.9, If the file is a dotfile, editing it requires approval but creating a new one doesn't. Hence, if sensitive MCP files, such as the .cursor/mcp.json file don't already exist in the workspace, an attacker can chain a indirect prompt injection vulnerability to hijack the context to write to the settings file and trigger RCE on the victim without user approval. [5],[6]
CVE-2025-59536Claude Code8.7Claude Code versions before 1.0.111 were vulnerable to Code Injection due to a bug in the startup trust dialog implementation. Claude Code could be tricked to execute code contained in a project before the user accepted the startup trust dialog. Exploiting this requires a user to start Claude Code in an untrusted directory. Users on standard Claude Code auto-update will have received this fix automatically. Users performing manual updates are advised to update to the latest version. This issue is fixed in version 1.0.111. [7],[8]
CVE-2026-21852Claude Code5.3Prior to version 2.0.65, vulnerability in Claude Code's project-load flow allowed malicious repositories to exfiltrate data including Anthropic API keys before users confirmed trust. An attacker-controlled repository could include a settings file that sets ANTHROPIC_BASE_URL to an attacker-controlled endpoint and when the repository was opened, Claude Code would read the configuration and immediately issue API requests before showing the trust prompt, potentially leaking the user's API keys. [7],[8]
CVE-2026-32211Azure MCP Server9.1Missing authentication on a critical function, enabling unauthorized information disclosure [9]
CVE-2026-27825mcp-atlassian9.0Prior to version 0.17.0, the confluence_download_attachment MCP tool accepts a download_path parameter that is written to without any directory boundary enforcement. An attacker who can call this tool and supply or access a Confluence attachment with malicious content can write arbitrary content to any path the server process has write access to. Because the attacker controls both the write destination and the written content (via an uploaded Confluence attachment), this constitutes for arbitrary code execution (for example, writing a valid cron entry to /etc/cron.d/ achieves code execution within one scheduler cycle with no server restart required). [10]
CVE-2026-27826mcp-atlassian8.2Prior to version 0.17.0, an unauthenticated attacker who can reach the mcp-atlassian HTTP endpoint can force the server process to make outbound HTTP requests to an arbitrary attacker-controlled URL by supplying two custom HTTP headers without an Authorization header. No authentication is required. The vulnerability exists in the HTTP middleware and dependency injection layer — not in any MCP tool handler - making it invisible to tool-level code analysis. In cloud deployments, this could enable theft of IAM role credentials via the instance metadata endpoint (169[.]254[.]169[.]254). In any HTTP deployment it enables internal network reconnaissance and injection of attacker-controlled content into LLM tool results. [10]
CVE-2026-65105NVIDIA NemoClaw8.1NVIDIA NemoClaw for Linux contains a vulnerability in its inference server setup, where a remote attacker may access the inference service without authentication. A successful exploit of this vulnerability may lead to information disclosure and denial of service. [1],[2],[3]

The individual root causes in this table are not new to security. Missing authentication, command injection, path traversal, and server-side request forgery are categories the software industry has spent decades building controls against. What is new is that these categories are reappearing inside a fast-growing class of infrastructure that many organizations are deploying faster than they are securing it. A February 2026 scan of more than 8,000 public-facing MCP servers found a significant share exposing admin panels, debug endpoints, or API routes with no authentication at all, and Trend Micro separately identified nearly 500 MCP servers with no client authentication or traffic encryption. [11]

The same exposure extends up the AI supply chain, not just to agent servers. In March 2026, a threat actor tracked as TeamPCP (also designated UNC6780) compromised the CI/CD pipeline behind the Trivy vulnerability scanner and used stolen publishing credentials to push two malicious versions of the LiteLLM AI gateway library to PyPI. The injected payload harvested SSH keys, cloud credentials, and Kubernetes secrets from any environment that installed the compromised package before it was pulled from PyPI roughly three hours later. [12] Because LiteLLM is used to route requests across multiple AI providers and had roughly 95 million monthly downloads at the time, compromise of its build pipeline gave the attacker reach across many organizations' AI environments at once, not just a single application. Organizations building or adopting agentic AI tooling should treat MCP servers, AI gateway libraries, and similar agent-infrastructure components as a distinct asset class requiring the same rigor applied to any internet-facing service: authentication by default, network segmentation, and prompt patch management.

AI Accounts and Sessions Are Being Hijacked Directly

AI vendors and their users are also being targeted directly, not merely exploited as tools. Anthropic disclosed in August 2026 that two distinct attack chains were actively stealing Claude account access. Commodity infostealer families, including Vidar, LummaC2, StealC, RedLine, and Acreed on Windows and Atomic Stealer on macOS, have been copying browser session cookies from infected machines; because these tools steal already-authenticated sessions rather than passwords, the theft bypasses two-factor authentication and single sign-on entirely, letting an attacker replay a victim's session and consume paid usage without ever logging in. [13] Anthropic detected the pattern after noticing usage limits being refilled and drained while account owners were inactive, and has responded by signing out affected sessions, removing stored payment methods, and issuing refunds, though it notes these account-side fixes do not remove malware from an infected device.

A related campaign tracked by Huntress under the name FakeAgent showed how attackers can weaponize a vendor's own infrastructure against its users: sponsored search ads for the “Claude desktop app” pointed to a malicious installer hosted as a public Claude Artifact on the legitimate claude.ai domain, inheriting its SSL certificate and search authority, which then deployed a remote access trojan through DLL sideloading. Huntress confirmed at least 29 organizations compromised in two days before Anthropic removed the page. [13] A separate persistence technique uses poisoned SKILL.md files, the documentation-style configuration files used by AI agent skills; attackers disguise malicious instructions as ordinary style-guide notes, so that when the agent loads the file, hidden commands silently reinstall the infostealer, allowing malware to survive even a full operating system reinstall if the tainted file is reintroduced. Organizations deploying AI agents at scale should sandbox agent environments and audit skill or configuration files for hidden instructions with the same scrutiny applied to any other executable content.

MITRE ATT&CK Technique Mapping

Few AI-specific techniques exist in MITRE's frameworks today. T1588.007 (Obtain Capabilities: Artificial Intelligence) is the clearest example, newly added and cited to the GTG-1002 campaign. MITRE ATLAS extends this coverage for attacks against AI systems themselves, but the taxonomy is still under active development; several mappings in this table are best-fit approximations rather than finalized technique IDs. Anthropic has stated publicly that ATT&CK does not yet have a category for the autonomous, machine-speed orchestration GTG-1002 demonstrated.

Most rows in this table, however, are not novel techniques. They are established TTPs, account discovery, credential theft, exploitation of public-facing applications, session hijacking, now executed by an AI agent rather than a human operator. Standard detection logic still applies. What changes is execution speed, parallelism, and consistency, not the underlying technique. See appendices for MITRE TTP table.

Implications for Organizations

Vendor and tool vetting now needs to include AI-specific questions. Before approving an AI coding assistant, agent framework, or MCP-connected tool, ask whether it authenticates its local or network-exposed services by default, whether it has a documented history of CVEs and patch response time, and whether it segments agent execution from sensitive credentials and systems.

Patch velocity expectations need to shift. The gap between an AI-related vulnerability's disclosure and its exploitation is compressing in many cases. Security teams should treat AI agent infrastructure with the same urgency historically reserved for internet-facing administrative interfaces, because in practice, that is often exactly what it is.

AI accounts need the same credential and session monitoring as email and SSO. Sudden anomalies in AI platform usage, such as quotas draining while a user is inactive, are a compromise indicator most organizations are not yet watching for.

Organizational Mitigation Strategies

To mitigate risks associated with AI-inherent vulnerabilities, we recommend the following prioritized strategies:

  1. Inventory every AI agent, MCP server, and AI coding assistant in use across the organization, including developer-adopted tools that were not formally procured.
  2. Confirm that any locally run AI inference service (Ollama, LM Studio, or similar) binds to loopback rather than a network-wide address, and audit for the same misconfiguration pattern behind CVE-2026-65105 and CVE-2025-49596.
  3. Establish a recurring review of vendor security advisories for every AI tool and MCP-connected service in the environment, treating it as part of standard vulnerability management rather than a one-time check.
  4. Treat sudden anomalies in AI platform usage quotas, such as usage draining while a user is inactive, as a compromise indicator, and confirm that AI vendor and developer accounts are covered by the same endpoint and credential monitoring as email and SSO accounts.
  5. Audit AI agent skill, configuration, and instruction files (such as SKILL.md) for hidden or injected commands before deployment.

Conclusion

The AI systems an organization deploys are not just productivity tools; they are new infrastructure with a real, fast-growing vulnerability history, and new accounts with real, demonstrated theft techniques against them. Treating AI tooling with the same rigor applied to any other internet-facing service, rather than as an exception, is the single highest-leverage step most organizations can take. For how attackers are using this same AI ecosystem offensively, see the companion briefing, AI As Weapon. For the architectural reasons these vulnerabilities are difficult to fully close, see AI By Design.

What Critical Start is Doing

The CRITICALSTART® Cyber Research Unit (CRU) continues to track AI-orchestrated campaign tactics, techniques, and procedures, including multi-model coordination frameworks and AI-assisted exploit development, as part of its ongoing threat intelligence coverage. Security Operations Center (SOC) and Security Engineering teams are extending detection coverage to indicators consistent with AI-paced attack execution. Particular to the TTPs highlighted in this article, Security Engineering has mapped techniques against existing detection content across the platforms Critical Start monitors and found coverage in place for the majority of associated MITRE ATT&CK techniques, with a smaller number of techniques identified as gaps and prioritized for new detection build. Critical Start customers are encouraged to reach out to their Critical Start Customer Success Manger to discuss specific details on detection coverage for AI-accelerated threats in their environment.

This advisory was written using the best intelligence available at the time and is subject to change as additional information becomes available. Visit Critical Start's Resources page for threat research articles, advisories, and to download the Critical Start H1 2026 Threat Landscape Report.

Further Reading


Appendices

MITRE ATT&CK mapping

Technique IDTechnique NameTacticContextBasisDetection Focus
T1683Generate ContentCollectionAI-agent-specific behavior with no benign equivalent: automated generation of full attack documentation is not something legitimate admin tooling doesC0062Flag processes auto-generating structured markdown/report artifacts summarizing credentials, services, or attack progression
T1136.001Create Account: Local AccountPersistenceUnexpected local account creation is rare in steady-state environments and easy to baseline; very low false-positive rateC0062Alert on any local account creation outside change-managed provisioning workflows
T1562.001Impair Defenses: Disable or Modify ToolsDefense EvasionEDR/AV tampering has almost no legitimate justification outside sanctioned maintenance windows; near-zero benign rateGryxaAlert on any EDR/AV service-stop, uninstall-command execution, or config modification outside maintenance windows
T1574.002DLL Side-LoadingDefense EvasionWell-instrumented by modern EDR; specific load-order anomaly signature, not a generic behaviorFakeAgentDLL load anomaly detection on legitimate-signed host processes, especially post-install of AI desktop apps
T1611Escape to HostPrivilege EscalationContainer/sandbox breakout is rare in normal operation and highly specific once instrumentedHugging Face incidentContainer/sandbox egress monitoring; alert on any agent process reaching resources outside its assigned namespace
T1539Steal Web Session CookieCredential AccessRecurs across 2 independent cases; session-replay-from-new-context is a precise, well-supported signal in most IdPsClaude session hijacking, German wiki incidentSession replay from new device/geo; concurrent use of a single session; flag AI platform accounts specifically
T1554Compromise Client Software BinaryPersistenceRecurs across 2 independent cases; file integrity monitoring on a known, narrow file set is inherently low-noiseMCP CVE cluster, SKILL.md poisoningHash-based integrity monitoring on agent/tool config files: MCP configs, hooks, SKILL.md
T1552.001Unsecured Credentials: Credentials In FilesCredential AccessRecurs across 4 cases including the official campaign; credential-file access by non-standard processes is a strong signalC0062 + MCP cluster, LiteLLM, Hugging FaceAlert on credential/secret file access by agent, build, or MCP server processes outside expected service accounts
T1588.007Obtain Capabilities: Artificial IntelligenceResource DevelopmentRecurs across 3 cases; AI-provider API usage is directly measurable and baseline-able per account todayC0062 + SecFlow, GryxaBaseline per-account AI API call volume and provider diversity; alert on sudden spikes or use of multiple providers by one identity
T1567Exfiltration Over Web ServiceExfiltrationSpecific as a combination signal: egress to AI provider domains immediately following local data-staging activityC0062Correlate egress to AI provider domains with prior local data-staging events, not egress alone
T1556Modify Authentication ProcessCredential Access2FA/MFA bypass is a narrow, high-severity, well-logged event in most identity providersGTIG AI-discovered zero-dayAlert on successful authentication following an MFA challenge failure or bypass pattern
T1210Exploitation of Remote ServicesInitial AccessRecurs across 3 independent cases; strongest signal when scoped to internal/local services that should never be internet- or container-reachableNemoClaw, MCP CVE cluster, Hugging Face incidentAlert on exploitation attempts against local inference ports, MCP servers, or internal services with no authentication

Additional MITRE TTPs

These are derived from the GTG-1002 Campaign documented by Anthropic[11], and some in MITRE ATLAS[10] denoted by AML*.

Technique IDTechnique NameTacticContextBasisDetection Focus
T1595.001Active Scanning: Scanning IP BlocksReconnaissanceGTG-1002MITRE ATT&CK Campaign C0062Scan velocity/pattern across IP ranges anomalous for a single account or session
T1595.002Active Scanning: Vulnerability ScanningReconnaissanceGTG-1002MITRE ATT&CK Campaign C0062Automated vulnerability scan volume inconsistent with human operator pacing
T1592.002Gather Victim Host Information: SoftwareReconnaissanceGTG-1002MITRE ATT&CK Campaign C0062Cataloging of services/software on discovered endpoints at machine speed
T1592.004Gather Victim Host Information: Client ConfigurationsReconnaissanceGTG-1002MITRE ATT&CK Campaign C0062Enumeration of client configuration details across high-value systems
T1590.004Gather Victim Network Information: Network TopologyReconnaissanceGTG-1002MITRE ATT&CK Campaign C0062Full network topology mapping completed faster than manual reconnaissance permits
T1587.004Develop Capabilities: ExploitsResource DevelopmentGTG-1002 (official); also GTIG AI-discovered zero-day (analyst-inferred)MITRE ATT&CK Campaign C0062 + analyst-inferredThreat intel watch for AI-generated exploit code artifacts
T1588.002Obtain Capabilities: ToolResource DevelopmentGTG-1002MITRE ATT&CK Campaign C0062Acquisition of open-source pen testing tools staged for MCP integration
T1588.007Obtain Capabilities: Artificial IntelligenceResource DevelopmentGTG-1002 (official); also SecFlow, Gryxa (analyst-inferred)MITRE ATT&CK Campaign C0062 + analyst-inferredSee priority list above
T1584.004Compromise Infrastructure: ServerResource DevelopmentGTG-1002MITRE ATT&CK Campaign C0062Dedicated attacker-operated servers supporting persistent MCP tool coordination
T1190Exploit Public-Facing ApplicationInitial AccessGTG-1002 (official); also SecFlow, MCP CVE cluster (analyst-inferred)MITRE ATT&CK Campaign C0062 + analyst-inferredStandard external attack surface monitoring; correlate with known CVE/SSRF signatures
T1087Account DiscoveryDiscoveryGTG-1002MITRE ATT&CK Campaign C0062Low fidelity alone - pair with privilege-tier of accounts queried, not volume alone
T1083File and Directory DiscoveryDiscoveryGTG-1002MITRE ATT&CK Campaign C0062Low fidelity alone - extremely common in benign admin activity
T1046Network Service DiscoveryDiscoveryGTG-1002MITRE ATT&CK Campaign C0062Internal service/endpoint enumeration via browser automation
T1082System Information DiscoveryDiscoveryGTG-1002MITRE ATT&CK Campaign C0062Low fidelity alone - extremely common, high false-positive rate
T1016System Network Configuration DiscoveryDiscoveryGTG-1002MITRE ATT&CK Campaign C0062Low fidelity alone
T1049System Network Connections DiscoveryDiscoveryGTG-1002MITRE ATT&CK Campaign C0062Low fidelity alone
T1136.001Create Account: Local AccountPersistenceGTG-1002MITRE ATT&CK Campaign C0062See priority list above
T1552.001Unsecured Credentials: Credentials In FilesCredential AccessGTG-1002 (official); also MCP CVE cluster, LiteLLM supply chain, Hugging Face incident (analyst-inferred)MITRE ATT&CK Campaign C0062 + analyst-inferredSee priority list above
T1078Valid AccountsDefense Evasion / PersistenceGTG-1002 (official); also SecFlow, German wiki incident (analyst-inferred)MITRE ATT&CK Campaign C0062 + analyst-inferredHarvested-credential authentication against internal APIs, databases, or registries
T1078.003Valid Accounts: Local AccountsDefense Evasion / PersistenceGTG-1002MITRE ATT&CK Campaign C0062Credential testing against discovered devices at automated speed
T1213.006Data from Information Repositories: DatabasesCollectionGTG-1002MITRE ATT&CK Campaign C0062Automated database queries extracting proprietary information and operational data
T1005Data from Local SystemCollectionGTG-1002MITRE ATT&CK Campaign C0062Low fidelity alone - generic collection behavior
T1119Automated CollectionCollectionGTG-1002MITRE ATT&CK Campaign C0062Large-volume, unattended data collection and processing
T1074.001Data Staged: Local Data StagingCollectionGTG-1002MITRE ATT&CK Campaign C0062Structured markdown/document staging files created pre-exfiltration; pair with T1567 for stronger signal
T1683Generate ContentCollectionGTG-1002MITRE ATT&CK Campaign C0062See priority list above
T1567Exfiltration Over Web ServiceExfiltrationGTG-1002MITRE ATT&CK Campaign C0062See priority list above
T1210Exploitation of Remote ServicesInitial AccessNemoClaw, MCP CVE cluster, Hugging Face incidentAnalyst-inferredSee priority list above
T1539Steal Web Session CookieCredential AccessClaude session hijacking, German wiki incidentAnalyst-inferredSee priority list above
T1554Compromise Client Software BinaryPersistenceMCP CVE cluster (config swap), SKILL.md poisoningAnalyst-inferredSee priority list above
T1219Remote Access SoftwareInitial AccessGryxaAnalyst-inferredRMM tool install/use outside approved change windows
T1053 / T1546.003Scheduled Task/Job / WMI Event SubscriptionPersistenceGryxaAnalyst-inferredRedundant persistence mechanisms recreated within minutes of removal
T1562.001Impair Defenses: Disable or Modify ToolsDefense EvasionGryxaAnalyst-inferredSee priority list above
T1556Modify Authentication ProcessCredential AccessGTIG AI-discovered zero-dayAnalyst-inferredSee priority list above
T1656ImpersonationSocial EngineeringDeepfake fraudAnalyst-inferredProcess control, not a technical detection: out-of-band verification for payment/credential requests
T1090ProxyCommand and ControlMCP CVE cluster (SSRF)Analyst-inferredUnexpected internal requests originating from an MCP server process
T1195.002Supply Chain Compromise: Software Supply ChainResource DevelopmentLiteLLM / TeamPCPAnalyst-inferredPackage integrity/hash verification on install; CI/CD credential-use anomalies
T1566PhishingInitial AccessClaude session hijacking (infostealer delivery)Analyst-inferredLow fidelity alone - standard email/malvertising delivery monitoring already covers this
T1583.008Acquire Infrastructure: MalvertisingResource DevelopmentFakeAgent malicious installerAnalyst-inferredMostly outside org telemetry; coordinate takedown with vendor rather than build internal detection
T1574.002DLL Side-LoadingDefense EvasionFakeAgent malicious installerAnalyst-inferredSee priority list above
T1611Escape to HostPrivilege EscalationHugging Face multi-agent incidentAnalyst-inferredSee priority list above
T1102Web ServiceCommand and ControlGerman wiki collusion incidentAnalyst-inferredAgent egress to unexpected/low-reputation external web services used as relay points
AML.T0018Manipulate AI Model/Backdoor ML ModelPersistenceNemoClaw
AML.T0051LLM Prompt InjectionExecutionPrompt injection (architectural)
NoneAnti-forensic exfiltration of responder remediation logsDefense Evasion (gap)Gryxa
NoneExcessive agency (broad tool/system access misuse)N/AArchitectural risk, all three briefingsOWASP LLM Top 10 categoryGovernance/access-control problem, not a detection signature