# ColossalX CyberTwins: attack the twin, not production.

> ColossalX CyberTwins attacks a twin of your AI agents instead of production. It captures a snapshot of your agents, their guardrails and the discovered estate, stands it up as shadow agents and runs real attacks against them. Breaches become findings, a guardrail change is tried on the twin first, and it reaches production only if it held.

Attack a snapshot of your agents and their guardrails instead of production, then promote a guardrail change only if it held on the twin.

Canonical page: https://colossalx.tech/platform/cybertwins · Last reviewed: 6 Oct 2026

*Illustration:* Attack the twin, promote what held: campaign to shadow agents (attack); campaign refused before production (never attacked); shadow agents found calling guardrail change (breach filed).

## The threat and the control

- **The threat:** A guardrail change meant to stop tool misuse goes straight to production, and breaks something else.
- **The control:** ColossalX tries the change on shadow agents first, re-attacks them, and refuses promotion if the breach still lands.

## How it works: From a breach on the twin to a deliberate promotion.

A support agent breaks under tool misuse on the twin. The fix is tried there, re-attacked, and promoted only after it held.

### Workflow: a guardrail change, tried on the twin first (illustrative)

1. **Replica captured.** Agents and their guardrails are captured as a snapshot, not a copy.
   Replica: support agents · configuration; Kept apart: identity, findings, scores; Shared: the model, at real cost
2. **Twin attacked.** Shadow agents are attacked; you confirm the real model spend first.
   Sends real prompts, spends budget (done: confirmed) | support-bot · shadow agent · Breach | Finding: tool misuse · ASI02
3. **Fix tried on twin.** The proposed guardrail change is applied to the shadow agents only.
   Rule: refund tool · require approval; Applied to: twin only | support-bot · shadow agent · Breach -> support-bot · shadow agent · Held
4. **Promoted, or refused.** It reaches production only after it held; revert stays one step away.
   Guardrail owner: Promoted to production after the twin held | Promoted · Revert kept | x, tested

## What you see: Campaigns that say what they did, and did not.

Red-team campaigns against named agents through your gateway, with findings by severity and a record of each phase attempted and each phase never attempted.

1. **Capture a replica.** A configuration snapshot of your agents, guardrails and the discovered estate. A replica is a snapshot, not a running system, and its coverage gaps are listed. Governed agents from the registry are projected into the twin, observed agent-to-tool paths come from the MCP firewall, and AWS discovery adds cloud resources.
2. **Stand up shadow agents.** Shadow agents get their own identities and the captured guardrail profile. The panel states what is kept apart from production, identity, findings and scores, and guardrail binding, and what is not: the model itself, shared provider limits and any system the model reaches.
3. **Attack and file.** Breaches are filed as findings, ticketed, and closed on a clean re-attack. You confirm that a run sends real adversarial prompts to the real model and spends its budget. Campaigns record each phase: reconnaissance and exploitation are attempted, while lateral movement, exfiltration and persistence are named as never attempted.
4. **Try, then promote.** Tried on the twin, promoted only if it held. A change that still breaks on the twin cannot be promoted. Promotion keeps the rules it replaced, so revert puts them back, and promote and revert both need permission to manage guardrails. Tearing the twin down retires the shadow agents and keeps the findings.

*Illustration:* An illustrative ColossalX CyberTwins campaign: production agents untouched and not a target, the campaign run against shadow agents in the twin, one breach filed, and a guardrail change that held on the twin and can now be promoted.

## How we know

- The panel says what the twin keeps apart from production and what it shares, before you attack.
- A change that still breaks on the twin cannot be promoted.
- A stopped campaign still shows what it sent before it stopped.
- Narratives drop and count any reference the campaign's findings cannot support.

## Where a breach on the twin goes next.

A breach on the twin travels the same path as one found in testing.

- **Findings and tickets.** Breaches join the same ticket and risk path as code scans, closed by a clean re-attack.
- **Guardrail profiles.** A promoted change lands in the production guardrail profile, with the replaced rules kept.
- **Attack paths.** Your scan findings are chained into routes an attacker could walk, with the evidence behind each hop.
- **Board narratives.** A completed campaign is written up for the board, the CISO, the engineers or the regulator.

Where an x ends up: x, tested.

## Specs: delivery and data

- **Delivery:** SaaS, from one login.
- **Isolation:** Each customer runs in an isolated workspace with its own database.
- **Certifications:** None held. Frameworks are mapped to and assessed against.

## Frameworks

- Covered in probes and scans OWASP LLM Top 10 (2025): Campaign techniques from the LLM Top 10.
- Assessed per agent against OWASP Top 10 for Agentic Applications: Breaches tagged by agentic risk.
- Coverage measured from runs MITRE ATLAS: Observed paths tagged with ATLAS techniques.

## What it does not do

- The twin is a configuration snapshot of your agents, not a replica of your systems or network.
- Shadow agents call the same real model, at real cost and visible in the vendor's logs.
- Infrastructure discovery for the twin is AWS only; other clouds are shown as not available yet.
- Campaigns do not attempt lateral movement, exfiltration or persistence, and say so phase by phase.

*Illustration:* Twin · what is kept apart: Agent identity Kept apart; Findings, scores Kept apart; Guardrail binding Kept apart; The model Shared, at real cost; Provider limits Shared. Stated before each attack.

## Questions

### What is a digital twin in cyber security?

A digital twin in cyber security is a model of a system that can be examined or attacked instead of the system itself. In ColossalX CyberTwins the twin is narrow on purpose: a configuration snapshot of your agents, their guardrails and the discovered estate, stood up as shadow agents, not a copy of your network.

### How does ColossalX CyberTwins attack AI agents?

It sends real adversarial prompts to shadow agents that carry your captured guardrails, one run per agent, through your gateway. Campaigns can also target named registered agents and record each phase: what was attempted, what worked, what the gateway stopped and by which control, and the phases never attempted.

### Can we test a guardrail change before it reaches production?

Yes. A proposed change names the rule the gateway would apply differently and the probes behind it. You try it on the twin's shadow agents only, re-attack them, and promote it to production only if it held. Promotion keeps the rules it replaced, so a revert restores them exactly.

### What are attack paths and blast radius?

An attack path is a route an attacker could walk, such as a committed credential, then a vulnerable dependency, then reachable personal data. Blast radius is everything affected if one part falls. ColossalX derives paths from your own scan findings, says each step is available rather than taken, and highlights blast radius on the twin graph.

### Who are the twin's narratives written for?

A completed campaign can be written up for the board, the CISO, the engineers or the regulator, from its real findings and attack paths, by the AI model your workspace chose. Each reference is checked, and references the findings do not support are dropped and counted, so a narrative is only as good as its evidence.

## Related

- [Red-teaming and validation](https://colossalx.tech/platform/red-teaming)
- [Runtime guardrails](https://colossalx.tech/platform/runtime-guardrails)
- [Exposure management](https://colossalx.tech/platform/exposure-management)

---

ColossalX is an AI security and governance platform from Quantexra Labs LLP, delivered as SaaS. Book a walkthrough: https://colossalx.tech/demo · client.success@quantexra.tech
