Sep 15, 2026

Attackers Are Going Dark on AI Providers. Enterprises Need Their Own AI Visibility

4

As threat actors move cyber operations to open-weight models, provider visibility will recede. Enterprises need their own evidence of how humans and AI agents use trusted access.

Anthropic’s latest threat intelligence report offers an unusually detailed view of how cybercriminals are using AI. It may also represent one of the last periods in which a model provider can observe this much of an attacker’s operation from the inside.

The report describes actors using Claude to direct reconnaissance, exploitation and data exfiltration, sometimes through autonomous or multi-agent workflows. One Russian-speaking group targeted AI companies, stole production API keys and pursued access to a pre-release Claude model. The campaign was classified as financially motivated, and there is no public evidence attributing it to a nation-state. Still, the interest in unreleased frontier capability deserves attention from defenders and intelligence agencies. Advanced models are becoming strategic assets, while the credentials that unlock them are becoming high-value targets.

The deeper security issue is where these operations go next.

The visibility window is narrowing

Commercial AI providers can identify abuse because they retain visibility into activity occurring on their infrastructure. That allowed Anthropic to connect internal prompts with subsequent Claude actions and reverse engineer how operators pursued their objectives. This kind of evidence is exceptionally valuable because it reveals the relationship between human instruction and the work delegated to the model.

But that visibility becomes far harder once an attacker downloads an open-weight model, removing the provider from the operational chain. In a separate statement, Anthropic CEO Dario Amodei warned that open-weight models are difficult to monitor and that safeguards cannot be updated or the model withdrawn once its weights are released. The threat report shows limits of provider intervention. In one surveillance case, account enforcement disrupted the actor’s use of Claude but did not affect the locally deployed platform.

Following intent beyond the prompt

AI security cannot depend on a provider seeing the malicious prompt. Even when prompts are visible, one interaction rarely establishes intent on its own. Threat actors can distribute an objective across different sessions or delegate seemingly routine tasks to multiple agents.

Open-weight models make this a longitudinal behavioral-risk problem. Individual actions may appear benign, but their cumulative sequence can reveal reconnaissance, expanding capability, adaptation, evasion and movement towards a harmful outcome. Intent emerges through the order of events:

  • Which data was accessed?
  • Which processes followed?
  • Where were credentials used?
  • How did the workflow adapt?
  • Was the outcome consistent with its authorized purpose?

Detecting that trajectory requires visibility across the full chain of activity, not just the individual event. As offensive activity moves towards locally hosted models and private agent frameworks, organizations will need to preserve the connection between the person setting an objective, the agent acting under delegated authority and the resulting activity across the enterprise.

Without that chain, an investigation may show a legitimate identity accessing a permitted resource while missing the malicious operation assembled around it. Authentication can establish who or what possessed the credential, but it cannot explain whether the activity that followed remained trustworthy.

Balancing AI visibility with privacy

There is an unavoidable privacy tension. Anthropic’s report demonstrates the defensive value of detailed model telemetry, yet enterprises will have legitimate concerns about a third party retaining prompts that may contain credentials or commercially sensitive information. The report also alleges that some third-party AI services silently routed customer requests to Claude, exposing users’ information to a provider they did not know was involved.

Enterprises should not respond by collecting every prompt without regard for privacy. They need proportionate visibility that establishes accountability while limiting unnecessary access to sensitive content. Often, the clearest signal is what happens next: sensitive files are opened, credentials are used, new processes are launched or data is sent somewhere it should not go. Connecting those actions to the AI workflow can reveal emerging intent without requiring indiscriminate access to prompt content. Security teams need to know who set an AI workflow in motion and what it ultimately did inside the enterprise.

Frontier model providers are giving defenders valuable threat intelligence today. We should use it while recognizing its limits. As threat actors gain access to capable open-weight models, more of their activity will move beyond provider oversight. The priority now is for enterprises to build their own line of sight into how humans and AI agents use trusted access in their environment, regardless of where the underlying model runs.

FAQ

Anthropic’s threat intelligence provides valuable insight into how threat actors use Claude and multi-agent workflows in cyber operations. It also shows why defenders need their own behavioral evidence connecting human instructions with subsequent AI agent activity, especially as more operations move beyond provider-hosted models.

Open-weight models can be deployed outside provider infrastructure, where activity is no longer visible to the original provider. Provider threat intelligence remains critical, but enterprises need their own AI visibility because reduced observed misuse may reflect activity moving elsewhere, not disappearing.

Enterprises need AI visibility that connects who initiated a workflow with what the AI agent ultimately did. Because intent may emerge across a sequence of actions, behavioral evidence can support investigations without requiring security teams to collect every prompt or unnecessarily expose sensitive content.

Subscribe today to stay informed and get regular updates from DTEX