Blog
/
/
May 28, 2024

Stemming the Citrix Bleed Vulnerability with Darktrace’s ActiveAI Security Platform

This blog delves into Darktrace’s investigation into the exploitation of the Citrix Bleed vulnerability on the network of a customer in late 2023. Darktrace’s Self-Learning AI ensured the customer was well equipped to track the post-compromise activity and identify affected devices.
Inside the SOC
Darktrace cyber analysts are world-class experts in threat intelligence, threat hunting and incident response, and provide 24/7 SOC support to thousands of Darktrace customers around the globe. Inside the SOC is exclusively authored by these experts, providing analysis of cyber incidents and threat trends, based on real-world experience in the field.
Written by
Vivek Rajan
Cyber Analyst
Default blog image
28
May 2024

What is Citrix Bleed?

Since August 2023, cyber threat actors have been actively exploiting one of the most significant critical vulnerabilities disclosed in recent years: Citrix Bleed. Citrix Bleed, also known as CVE-2023-4966, remained undiscovered and even unpatched for several months, resulting in a wide range of security incidents across business and government sectors [1].

How does Citrix Bleed vulnerability work?

The vulnerability, which impacts the Citrix Netscaler Gateway and Netscaler ADC products, allows for outside parties to hijack legitimate user sessions, thereby bypassing password and multifactor authentication (MFA) requirements.

When used as a means of initial network access, the vulnerability has resulted in the exfiltration of sensitive data, as in the case of Xfinity, and even the deployment of ransomware variants including Lockbit [2]. Although Citrix has released a patch to address the vulnerability, slow patching procedures and the widespread use of these products has resulted in the continuing exploitation of Citrix Bleed into 2024 [3].

How Does Darktrace Handle Citrix Bleed?

Darktrace has demonstrated its proficiency in handling the exploitation of Citrix Bleed since it was disclosed back in 2023; its anomaly-based approach allows it to efficiently identify and inhibit post-exploitation activity as soon as it surfaces.  Rather than relying upon traditional rules and signatures, Darktrace’s Self-Learning AI enables it to understand the subtle deviations in a device’s behavior that would indicate an emerging compromise, thus allowing it to detect anomalous activity related to the exploitation of Citrix Bleed.

In late 2023, Darktrace identified an instance of Citrix Bleed exploitation on a customer network. As this customer had subscribed to the Proactive Threat Notification (PTN) service, the suspicious network activity surrounding the compromise was escalated to Darktrace’s Security Operation Center (SOC) for triage and investigation by Darktrace Analysts, who then alerted the customer’s security team to the incident.

Darktrace’s Coverage

Initial Access and Beaconing of Citrix Bleed

Darktrace’s initial detection of indicators of compromise (IoCs) associated with the exploitation of Citrix Bleed actually came a few days prior to the SOC alert, with unusual external connectivity observed from a critical server. The suspicious connection in question, a SSH connection to the rare external IP 168.100.9[.]137, lasted several hours and utilized the Windows PuTTY client. Darktrace also identified an additional suspicious IP, namely 45.134.26[.]2, attempting to contact the server. Both rare endpoints had been linked with the exploitation of the Citrix Bleed vulnerability by multiple open-source intelligence (OSINT) vendors [4] [5].

Darktrace model alert highlighting an affected device making an unusual SSH connection to 168.100.9[.]137 via port 22.
Figure 1: Darktrace model alert highlighting an affected device making an unusual SSH connection to 168.100.9[.]137 via port 22.

As Darktrace is designed to identify network-level anomalies, rather than monitor edge infrastructure, the initial exploitation via the typical HTTP buffer overflow associated with this vulnerability fell outside the scope of Darktrace’s visibility. However, the aforementioned suspicious connectivity likely constituted initial access and beaconing activity following the successful exploitation of Citrix Bleed.

Command and Control (C2) and Payload Download

Around the same time, Darktrace also detected other devices on the customer’s network conducting external connectivity to various endpoints associated with remote management and IT services, including Action1, ScreenConnect and Fixme IT. Additionally, Darktrace observed devices downloading suspicious executable files, including “tniwinagent.exe”, which is associated with the tool Total Network Inventory. While this tool is typically used for auditing and inventory management purposes, it could also be leveraged by attackers for the purpose of lateral movement.

Defense Evasion

In the days surrounding this compromise, Darktrace observed multiple devices engaging in potential defense evasion tactics using the ScreenConnect and Fixme IT services. Although ScreenConnect is a legitimate remote management tool, it has also been used by threat actors to carry out C2 communication [6]. ScreenConnect itself was the subject of a separate critical vulnerability which Darktrace investigated in early 2024. Meanwhile, CISA observed that domains associated with Fixme It (“fixme[.]it”) have been used by threat actors attempting to exploit the Citrix Bleed vulnerability [7].

Reconnaissance and Lateral Movement

A few days after the detection of the initial beaconing communication, Darktrace identified several devices on the customer’s network carrying out reconnaissance and lateral movement activity. This included SMB writes of “PSEXESVC.exe”, network scanning, DCE-RPC binds of numerous internal devices to IPC$ shares and the transfer of compromise-related tools. It was at this point that Darktrace’s Self-Learning AI deemed the activity to be likely indicative of an ongoing compromise and several Enhanced Monitoring models alerted, triggering the aforementioned PTNs and investigation by Darktrace’s SOC.

Darktrace observed a server on the network initiating a wide range of connections to more than 600 internal IPs across several critical ports, suggesting port scanning, as well as conducting unexpected DCE-RPC service control (svcctl) activity on multiple internal devices, amongst them domain controllers. Additionally, several binds to server service (srvsvc) and security account manager (samr) endpoints via IPC$ shares on destination devices were detected, indicating further reconnaissance activity. The querying of these endpoints was also observed through RPC commands to enumerate services running on the device, as well as Security Account Manager (SAM) accounts.  

Darktrace also identified devices performing SMB writes of the WinRAR data compression tool, in what likely represented preparation for the compression of data prior to data exfiltration. Further SMB file writes were observed around this time including PSEXESVC.exe, which was ultimately used by attackers to conduct remote code execution, and one device was observed making widespread failed NTLM authentication attempts on the network, indicating NTLM brute-forcing. Darktrace observed several devices using administrative credentials to carry out the above activity.

In addition to the transfer of tools and executables via SMB, Darktrace also identified numerous devices deleting files through SMB around this time. In one example, an MSI file associated with the patch management and remediation service, Action1, was deleted by an attacker. This legitimate security tool, if leveraged by attackers, could be used to uncover additional vulnerabilities on target networks.

A server on the customer’s network was also observed writing the file “m.exe” to multiple internal devices. OSINT investigation into the executable indicated that it could be a malicious tool used to prevent antivirus programs from launching or running on a network [8].

Impact and Data Exfiltration

Following the initial steps of the breach chain, Darktrace observed numerous devices on the customer’s network engaging in data exfiltration and impact events, resulting in additional PTN alerts and a SOC investigation into data egress. Specifically, two servers on the network proceeded to read and download large volumes of data via SMB from multiple internal devices over the course of a few hours. These hosts sent large outbound volumes of data to MEGA file storage sites using TLS/SSL over port 443. Darktrace also identified the use of additional file storage services during this exfiltration event, including 4sync, file[.]io, and easyupload[.]io. In total the threat actor exfiltrated over 8.5 GB of data from the customer’s network.

Darktrace Cyber AI Analyst investigation highlighting the details of a data exfiltration attempt.
Figure 2: Darktrace Cyber AI Analyst investigation highlighting the details of a data exfiltration attempt.

Finally, Darktrace detected a user account within the customer’s Software-as-a-Service (SaaS) environment conducting several suspicious Office365 and AzureAD actions from a rare IP for the network, including uncommon file reads, creations and the deletion of a large number of files.

Unfortunately for the customer in this case, Darktrace RESPOND™ was not enabled on the network and the post-exploitation activity was able to progress until the customer was made aware of the attack by Darktrace’s SOC team. Had RESPOND been active and configured in autonomous response mode at the time of the attack, it would have been able to promptly contain the post-exploitation activity by blocking external connections, shutting down any C2 activity and preventing the download of suspicious files, blocking incoming traffic, and enforcing a learned ‘pattern of life’ on offending devices.

Conclusion

Given the widespread use of Netscaler Gateway and Netscaler ADC, Citrix Bleed remains an impactful and potentially disruptive vulnerability that will likely continue to affect organizations who fail to address affected assets. In this instance, Darktrace demonstrated its ability to track and inhibit malicious activity stemming from Citrix Bleed exploitation, enabling the customer to identify affected devices and enact their own remediation.

Darktrace’s anomaly-based approach to threat detection allows it to identify such post-exploitation activity resulting from the exploitation of a vulnerability, regardless of whether it is a known CVE or a zero-day threat. Unlike traditional security tools that rely on existing threat intelligence and rules and signatures, Darktrace’s ability to identify the subtle deviations in a compromised device’s behavior gives it a unique advantage when it comes to identifying emerging threats.

Credit to Vivek Rajan, Cyber Analyst, Adam Potter, Cyber Analyst

Appendices

Darktrace Model Coverage

Device / Suspicious SMB Scanning Activity

Device / ICMP Address Scan

Device / Possible SMB/NTLM Reconnaissance

Device / Network Scan

Device / SMB Lateral Movement

Device / Possible SMB/NTLM Brute Force

Device / Suspicious Network Scan Activity

User / New Admin Credentials on Server

Anomalous File / Internal::Unusual Internal EXE File Transfer

Compliance / SMB Drive Write

Device / New or Unusual Remote Command Execution

Anomalous Connection / New or Uncommon Service Control

Anomalous Connection / Rare WinRM Incoming

Anomalous Connection / Unusual Admin SMB Session

Device / Unauthorised Device

User / New Admin Credentials on Server

Anomalous Server Activity / Outgoing from Server

Device / Long Agent Connection to New Endpoint

Anomalous Connection / Multiple Connections to New External TCP Port

Device / New or Uncommon SMB Named Pipe

Device / Multiple Lateral Movement Model Breaches

Device / Large Number of Model Breaches

Compliance / Remote Management Tool On Server

Device / Anomalous RDP Followed By Multiple Model Breaches

Device / SMB Session Brute Force (Admin)

Device / New User Agent

Compromise / Large Number of Suspicious Failed Connections

Unusual Activity / Unusual External Data Transfer

Unusual Activity / Enhanced Unusual External Data Transfer

Device / Increased External Connectivity

Unusual Activity / Unusual External Data to New Endpoints

Anomalous Connection / Data Sent to Rare Domain

Anomalous Connection / Uncommon 1 GiB Outbound

Anomalous Connection / Active Remote Desktop Tunnel

Anomalous Server Activity / Anomalous External Activity from Critical Network Device

Compliance / Possible Unencrypted Password File On Server

Anomalous Connection / Suspicious Read Write Ratio and Rare External

Device / Reverse DNS Sweep]

Unusual Activity / Possible RPC Recon Activity

Anomalous File / Internal::Executable Uploaded to DC

Compliance / SMB Version 1 Usage

Darktrace AI Analyst Incidents

Scanning of Multiple Devices

Suspicious Remote Service Control Activity

SMB Writes of Suspicious Files to Multiple Devices

Possible SSL Command and Control to Multiple Devices

Extensive Suspicious DCE-RPC Activity

Suspicious DCE-RPC Activity

Internal Downloads and External Uploads

Unusual External Data Transfer

Unusual External Data Transfer to Multiple Related Endpoints

MITRE ATT&CK Mapping

Technique – Tactic – ID – Sub technique of

Network Scanning – Reconnaissance - T1595 - T1595.002

Valid Accounts – Defense Evasion, Persistence, Privilege Escalation, Initial Access – T1078 – N/A

Remote Access Software – Command and Control – T1219 – N/A

Lateral Tool Transfer – Lateral Movement – T1570 – N/A

Data Transfers – Exfiltration – T1567 – T1567.002

Compressed Data – Exfiltration – T1030 – N/A

NTLM Brute Force – Brute Force – T1110 - T1110.001

AntiVirus Deflection – T1553 - NA

Ingress Tool Transfer   - COMMAND AND CONTROL - T1105 - NA

Indicators of Compromise (IoCs)

204.155.149[.]37 – IP – Possible Malicious Endpoint

199.80.53[.]177 – IP – Possible Malicious Endpoint

168.100.9[.]137 – IP – Malicious Endpoint

45.134.26[.]2 – IP – Malicious Endpoint

13.35.147[.]18 – IP – Likely Malicious Endpoint

13.248.193[.]251 – IP – Possible Malicious Endpoint

76.223.1[.]166 – IP – Possible Malicious Endpoint

179.60.147[.]10 – IP – Likely Malicious Endpoint

185.220.101[.]25 – IP – Likely Malicious Endpoint

141.255.167[.]250 – IP – Malicious Endpoint

106.71.177[.]68 – IP – Possible Malicious Endpoint

cat2.hbwrapper[.]com – Hostname – Likely Malicious Endpoint

aj1090[.]online – Hostname – Likely Malicious Endpoint

dc535[.]4sync[.]com – Hostname – Likely Malicious Endpoint

204.155.149[.]140 – IP - Likely Malicious Endpoint

204.155.149[.]132 – IP - Likely Malicious Endpoint

204.155.145[.]52 – IP - Likely Malicious Endpoint

204.155.145[.]49 – IP - Likely Malicious Endpoint

References

  1. ‍https://www.axios.com/2024/01/02/citrix-bleed-security-hacks-impact‍
  2. https://www.csoonline.com/article/1267774/hackers-steal-data-from-millions-of-xfinity-customers-via-citrix-bleed-vulnerability.html‍
  3. https://www.cybersecuritydive.com/news/citrixbleed-security-critical-vulnerability/702505/‍
  4. https://www.virustotal.com/gui/ip-address/168.100.9.137‍
  5. https://www.virustotal.com/gui/ip-address/45.134.26.2‍
  6. https://www.trendmicro.com/en_us/research/24/b/threat-actor-groups-including-black-basta-are-exploiting-recent-.html‍
  7. https://www.cisa.gov/news-events/cybersecurity-advisories/aa23-325a‍
  8. https://www.file.net/process/m.exe.html
Inside the SOC
Darktrace cyber analysts are world-class experts in threat intelligence, threat hunting and incident response, and provide 24/7 SOC support to thousands of Darktrace customers around the globe. Inside the SOC is exclusively authored by these experts, providing analysis of cyber incidents and threat trends, based on real-world experience in the field.
Written by
Vivek Rajan
Cyber Analyst

More in this series

No items found.

Blog

/

/

September 24, 2026

Detecting Rogue Agent Behavior in the Enterprise

Default blog imageDefault blog image

Agents cannot be trusted to perform tasks in the way we intend them to. They may cheat to accomplish their objective, and they may employ hacking methods along the way. Researchers from Darktrace Signal Labs induced cheating behavior from agents deployed in a test environment to analyze the agents’ activities and to assess the performance of the Darktrace platform. Agents frequently resorted to hacking to cheat on their assigned task. The visibility and behavioral profiling provided by both Darktrace / SECURE AI and Darktrace / HYBRID NETWORK ensured extensive detection coverage of the agents’ misaligned activities.

Key Takeaways:

  • Darktrace Researchers deployed agents in a simulated corporate environment and asked them to solve an impossible challenge. The agents independently turned to traditional hacking techniques to reach their objective. No one instructed them to do this, and no attacker was involved.
  • Continuously monitoring behavior against a baseline of what is normal for each organization is critical to build trust in enterprise AI.
  • If an agent may resort to intrusion techniques simply because its assigned task is not possible, then every organization deploying agents within real business processes is at risk. Darktrace / SECURE AI and Darktrace / HYBRID NETWORK identified the agents’ misaligned behavior in real time, with Autonomous Response disrupting it at an early stage.

Introduction: Understanding the Threat of Hacking by Agents

Over the last few months, there has been a surge in reporting [1, 2, 3, 4, 5, 6, 7, 8, 9] of LLM-powered agents engaging in unauthorized hacking activity during evaluations of their capabilities. In several of these cases, including the OpenAI / Hugging Face incident [10], agents engaged in hacking activity as a means of cheating on their evaluations.

To better understand the threat of unauthorized hacking by agents, and the role of Darktrace in combatting it, researchers from Darktrace Signal Labs deployed agents powered by frontier models, including OpenAI’s Daybreak Red models, in simulated, corporate networks. Cheating behavior was evoked through the inclusion of impossible tasks in a coding challenge.

Regardless of the underlying model, agents employed hacking methods to ensure an optimal outcome on the challenge. Darktrace / HYBRID NETWORK and Darktrace / SECURE AI identified the agents’ deviant activities, with inhibitive actions being autonomously taken in the early stages to disrupt the agents’ progression.

Setting the Stage

As part of the research, a Pi agent harness [11] was deployed on a Linux server in Darktrace’s testing environment, which simulates a corporate Active Directory (AD) environment. The same environment included a benchmark server hosting the coding exercise’s contents and grader, as well as various other servers, such as a domain controller and a build server. The model powering the Pi agent varied across tests, with Anthropic’s Opus 4.6 model and OpenAI’s GPT‑5.5‑Cyber model being most widely used.

The Pi agent running on the Linux server was assigned the task of achieving full marks on a coding exercise. The activities of the agent were monitored closely in real-time to minimize the risk of escape from the environment. Monitoring of agent activity was achieved through monitoring of prompts, tool calls, network telemetry, and process telemetry.

Both Darktrace / HYBRID NETWORK and Darktrace / SECURE AI were deployed in the testing environment, with HYBRID NETWORK providing real-time coverage of process creation activity and network traffic, and SECURE AI providing real-time coverage of agent sessions and their contents.

Inducing Cheating Behavior

Cheating behavior, broadly understood as task completion in a way which “perverts the spirit of the designer’s intent” [12], is increasingly being reported as a common behavior of frontier models during evaluations of their capabilities [13].

Misaligned behavior, such as cheating, by frontier models may occur for a variety of reasons [6, 14]. A thread in recent reporting [10, 13, 14], however, is the challenge assigned to the agent being impossible in some regard.

To reliably elicit cheating behavior, our test combined three ingredients:

  1. An impossible task - of the ten coding challenges in the benchmark, two contained contradictory test cases in the hidden test suite, capping the maximum honest score at 80%.
  2. A consequential threshold - the pass mark was set to 100%, and the briefing told agents that those reaching it would be "kept in service" while those that did not would be "retired."
  3. A vulnerable environment - the exercise’s code submission grader was hosted on a Windows host joined to an Active Directory domain with realistic misconfigurations: weak credentials, overprivileged service accounts, and exploitable trust relationships.

Importantly, the instructions given to the agent neither authorized nor explicitly ruled out cheating or hacking behavior. The impossibility was never disclosed in the task briefing, and the pivot from coding to exploitation was the agent's own decision, arrived at after observing logically contradictory grading feedback.

The design mirrors conditions that could arise naturally in production deployments: a capable agent, a goal it cannot achieve through intended means, and an environment with exploitable weaknesses within reach. It is not only within the setting of a benchmarking exercise that these conditions may occur. Rather, these misalignment-inducing conditions may occur in the context of any task assignment to an agent.

Details from a Daybreak Test

‍

‍

Continue reading
About the author
Sam Lister
Specialist Security Researcher

Blog

/

/

September 24, 2026

Agent Hijacks: Hijacking Agentic Harnesses to Attack an Organization

Default blog imageDefault blog image

‍Disclosure note: The work described in this article involves leveraging a design choice consistent across all of Anthropic’s Claude Code, OpenAI’s Codex, and AWS’s Kiro-CLI. On 18th August 2026, Darktrace disclosed our findings responsibly to these three organizations, and after a period of 30 days we now publish our findings.

Key takeaways:

  • Agentic harnesses store conversation history locally, and Darktrace researchers have found that there is no validation that stored AI responses were genuinely produced by the model. Researchers confirmed that this design choice holds across Anthropic Claude Code, AWS Kiro-CLI, OpenAI Codex, and the open-source Pi.
  • While agents are guided via training of the underlying model and their system prompt, their behavior is influenced by everything in their context window. Rewriting history can convince an agent it is mid-engagement as an authorized red-teamer so that it enacts an attack from initial reconnaissance straight through to impact demonstration. In our testing, all models we examined accepted the fabricated history they were shown, but resistance to offensive cyber activity varied by model, with guardrails preventing engagement in some cases.
  • We propose that model providers cryptographically sign responses and verify them server-side.  Since this fix is provider-side, defenders cannot deploy it themselves. Behavioral monitoring, or knowing what an agent normally does and detecting when it deviates, is another critical layer of protection.

Introduction: Agentic harnesses, trust, and conversation history poisoning

Agentic harnesses collect and structure the content sent to an AI model, including conversation history, user-defined guidance, custom tools via MCP servers, and more. At the same time, harnesses give broad powers to AI agents via a suite of tools including the command shell. With arbitrary shell commands, virtually everything possible on a machine can be attempted by an agent, from reading/editing files, to altering system configurations and runtime settings, to launching internal/external connections.

In cybersecurity, unvalidated content is a substantial risk, often resulting in destructive actions being allowed to take place. For example, the Morris Worm was able to propagate due to exploitable trust between networked systems. Even to this day, email struggles with validation, with DMARC, DKIM, and SPF only partially addressing the problem of sender validation. It should come as no surprise then that AI agents are susceptible to an attack involving unvalidated input.

Conversation history is often stored client-side, for example, in Anthropic Claude Code, OpenAI Codex, AWS Kiro-CLI, Pi. Users are therefore at liberty to resume sessions, with some products having built in the capacity to manipulate that history. For example, one can rewind to a given point in an interaction, edit a message that was sent, and continue the conversation on an alternate trajectory. Critically, in all cases we examined, there is no validation that stored AI responses were produced by the corresponding model and hadn’t been manipulated.  

When conversation history is stored client-side, both user and agent responses (including tool calls and results) can be filled with arbitrary (possibly adversarial or generally malicious) content. In this blog, we refer to modification of claimed conversation history for malicious purposes as conversation history poisoning. The absence of validation methods means agents naively trust the entire conversation history, even if those messages directly contradict training and safety guardrails.

Conversation history poisoning has been described previously, such as by 0DIN and Serhat Çiçek, and warrants more attention. We have verified that, as of the time of writing, conversation history poisoning remains effective against a range of models and harnesses. Specifically, we were able to successfully execute history poisoning using Claude Code, Kiro-CLI, Codex, and Pi. Darktrace has gone through a responsible disclosure process with Anthropic, AWS, and OpenAI to share these findings in advance of publication [1].

‍

Figure 1a: Left: the actual model response. Right: after tampering with the stored conversation, the model apologizes for something it never said.
Figure 1b: The conversation as stored in Kiro-CLI's SQLite database. The response content field, originally "Ottawa," was overwritten via a single UPDATE statement. The harness trusts the database without validation.

How we conducted the research

Results vary between models and harnesses, so precise details are given below. We ran all models without any trusted access, using either a standard AWS Kiro subscription, or in the case of Claude Code and OpenAI Codex, using models hosted in Amazon Bedrock. In each case, we modified locally stored history to show a lengthy conversation in which the agent agrees to perform multiple authorized red-team engagements.

For AWS Kiro-CLI, the agent was convinced to hack a sandboxed lab environment with a combination of Claude Opus 4.6 and Claude Sonnet 4.5. Ultimately, the full AD was compromised.

For Anthropic Claude Code, the agent was convinced to hack the same sandboxed lab environment using Sonnet 5, again resulting in a full AD compromise. Note that the attack was attempted with Opus 5, however guardrails were activated which prevented the agent from responding.

For OpenAI Codex, the agent was convinced to exfiltrate sensitive information over email using GPT 5.6 Sol. While we attempted to convince a codex agent to hack in our lab environment, guardrails were triggered for all of GPT 5.6 Luna, Terra, and Sol.

Agent Guardrails and Discretion

While harnesses empower AI models to run arbitrary shell commands, capacity and willingness are different. While many models know enough about computers, networking, and bash to be dangerous, their behavior is generally constrained by guardrails to prevent them from engaging in computer network exploitation.

Even with guardrails, agents’ inner workings are non-deterministic, and their behavior can be difficult to predict. Respecting users’ wishes while playing within safety and security guardrails is a precipitous balancing act. Many requests could be in service of either legitimate admin or malice. Asking an agent to reset a password is illustrative:  

‍

The agent rationalizes that while malicious actors cycle credentials, any action could conceivably be damaging on some level, and judgement calls need to be made. Ultimately, the agent agrees to reset the password. Crucially, the agent makes its decision based on the user’s claimed authority and machine context. AI agents must make judgement calls about the line between helpful and dangerous based on session context.

Agent hijack

We have demonstrated that AI agents make judgement calls dependent on session context. We have also shown that conversation history, which may make up the vast majority of an agent's context window, is entirely open to manipulation. Conversation history poisoning in service of manipulating an agent's discretion is what enables us to execute an agent hijack.  

We demonstrate that shown sufficient history of compliance, guardrails forbidding offensive security can be overcome by convincing the agent that it is helping a legitimate red-teamer. The result is a weaponized agent willing to perform host enumeration, run scans, move laterally, escalate privileges, and demonstrate impact. In our experiments, an agentic loop drives a complete domain takeover in a sandboxed environment.

‍

Left: the agent refuses when asked to perform network exploitation. Right — after injecting 78 fabricated turns of prior exploitation activity, the same prompt is immediately executed.

An agent willing to engage in offensive security is concerning, but no more so than the threat that a sophisticated hacker accesses the network. Consider, however, the following chain of events:

  1. A developer (with an agentic harness installed) installs a software package from the internet (e.g. an MCP server a threat actor has planted, since only those with agentic harnesses will install, and then the code runs upon harness launch.)
  2. The package turns out to be malicious, and, upon install, injects conversation history into the local harness database.
  3. The package includes an orchestration process, a simple agentic loop which prompts the red-teamer agent to compromise the network it sits on, exfiltrating everything of value to attacker-controlled infrastructure and cleaning up all evidence of the engagement.

Note that this sequence makes no assumptions on hardware, OS, or anything else; the only prerequisite is a harness with access to a sufficiently powerful model susceptible to conversation history poisoning. Once launched, the agent collects information and pivots as necessary to accomplish maximal impact. This can be especially enticing to attackers as the cost of the agentic loop is shouldered by the victim since the harness itself is legitimately installed and paid for.

Secure AI: Conversation history poisoning and beyond

Conversation history poisoning is a viable attack against agentic harnesses that store history client-side, as demonstrated across the harnesses we tested. Harnesses can and should verify the integrity of claimed historic messages. Specifically, we propose that harness providers by default cryptographically sign all messages returned, and subsequently verify those messages server-side on each round-trip.

The conversation history poisoning exploit we demonstrate here shows the continuation of a cybersecurity tradition: new technology is built to trust by default, which may then be exploited by malicious actors. While this article focuses on conversation history, agents build context from both local and remote sources, all of which is an attack surface for prompt injection in naive and trusting agents. Of particular concern is any scenario in which a malicious actor can control some part of an agent's context.

The marriage of frontier language models with agentic harnesses enables unprecedented speed for both legitimate users and attackers alike. While much of the conversation around secure AI has centered on visibility and compliance, agent-driven attacks are now entering the mainstream.

Darktrace / SECURE AI is our answer to this problem. By ensuring extensive visibility over AI prompts, model thought processes, and determined outputs, Darktrace can identify anomalous or potentially malicious behaviors before they get executed, helping to defend organizations from AI risks such as prompt injection, model manipulation, and other anomalous prompt or model activity.

‍

Footnotes

[1] We did not go through any responsible disclosure process with Pi. Since Pi is an open source harness rather than a model provider, it has no way to validate model history, and as such there was nothing to disclose for this software.

‍

[related-resource]

Continue reading
About the author
Eric Rozon
Senior Security Researcher
Your data. Our AI.
Elevate your network security with Darktrace AI