ProveColossalX CyberTwinsx, tested

ColossalX CyberTwins: attack the twin, not production.

Attack a snapshot of your agents and their guardrails instead of production, then promote a guardrail change only if it held on the twin.

How it works

Specs

ColossalX CyberTwins, in detail

Delivery and data

Delivery
SaaS, from one login.
Isolation
Each customer runs in an isolated workspace with its own database.
Certifications
None held. Frameworks are mapped to and assessed against.

Frameworks

OWASP LLM Top 10 (2025)
Covered in probes and scans: Campaign techniques from the LLM Top 10.
OWASP Top 10 for Agentic Applications
Assessed per agent against: Breaches tagged by agentic risk.
MITRE ATLAS
Coverage measured from runs: Observed paths tagged with ATLAS techniques.
See the frameworks

Last reviewed 6 Oct 2026

Attack the twin, promote what heldIllustrative

Attack the twin, promote what held: campaign to shadow agents (attack); campaign refused before production (never attacked); shadow agents found calling guardrail change (breach filed).

In short

ColossalX CyberTwins attacks a twin of your AI agents instead of production. It captures a snapshot of your agents, their guardrails and the discovered estate, stands it up as shadow agents and runs real attacks against them. Breaches become findings, a guardrail change is tried on the twin first, and it reaches production only if it held.

A guardrail change meant to stop tool misuse goes straight to production, and breaks something else.

ColossalX tries the change on shadow agents first, re-attacks them, and refuses promotion if the breach still lands.

How it works

From a breach on the twin to a deliberate promotion.

A support agent breaks under tool misuse on the twin. The fix is tried there, re-attacked, and promoted only after it held.

Workflow · a guardrail change, tried on the twin firstIllustrative

01 Replica captured

Agents and their guardrails are captured as a snapshot, not a copy.

02 Twin attacked

Shadow agents are attacked; you confirm the real model spend first.

03 Fix tried on twin

The proposed guardrail change is applied to the shadow agents only.

04 Promoted, or refused

It reaches production only after it held; revert stays one step away.

What you see

Campaigns that say what they did, and did not.

Red-team campaigns against named agents through your gateway, with findings by severity and a record of each phase attempted and each phase never attempted.

  1. Capture a replica
  2. Stand up shadow agents
  3. Attack and file
  4. Try, then promote
Read the detail, step by step4
  1. Capture a replica. A configuration snapshot of your agents, guardrails and the discovered estate. A replica is a snapshot, not a running system, and its coverage gaps are listed. Governed agents from the registry are projected into the twin, observed agent-to-tool paths come from the MCP firewall, and AWS discovery adds cloud resources.
  2. Stand up shadow agents. Shadow agents get their own identities and the captured guardrail profile. The panel states what is kept apart from production, identity, findings and scores, and guardrail binding, and what is not: the model itself, shared provider limits and any system the model reaches.
  3. Attack and file. Breaches are filed as findings, ticketed, and closed on a clean re-attack. You confirm that a run sends real adversarial prompts to the real model and spends its budget. Campaigns record each phase: reconnaissance and exploitation are attempted, while lateral movement, exfiltration and persistence are named as never attempted.
  4. Try, then promote. Tried on the twin, promoted only if it held. A change that still breaks on the twin cannot be promoted. Promotion keeps the rules it replaced, so revert puts them back, and promote and revert both need permission to manage guardrails. Tearing the twin down retires the shadow agents and keeps the findings.
Twin campaignIllustrative

An illustrative ColossalX CyberTwins campaign: production agents untouched and not a target, the campaign run against shadow agents in the twin, one breach filed, and a guardrail change that held on the twin and can now be promoted.

How it connectsx, tested

Where a breach on the twin goes next.

A breach on the twin travels the same path as one found in testing.

  1. Breaches join the same ticket and risk path as code scans, closed by a clean re-attack.

  2. A promoted change lands in the production guardrail profile, with the replaced rules kept.

  3. Your scan findings are chained into routes an attacker could walk, with the evidence behind each hop.

  4. A completed campaign is written up for the board, the CISO, the engineers or the regulator.

Honest by design

What it does, and what it does not.

Twin · what is kept apartIllustrative

Twin · what is kept apart: Agent identity Kept apart; Findings, scores Kept apart; Guardrail binding Kept apart; The model Shared, at real cost; Provider limits Shared. Stated before each attack.

What it does not do

x, not measured

The twin is a configuration snapshot of your agents, not a replica of your systems or network.

All 4 limits
  • Shadow agents call the same real model, at real cost and visible in the vendor's logs.
  • Infrastructure discovery for the twin is AWS only; other clouds are shown as not available yet.
  • Campaigns do not attempt lateral movement, exfiltration or persistence, and say so phase by phase.

How we know

  • The panel says what the twin keeps apart from production and what it shares, before you attack.
  • A change that still breaks on the twin cannot be promoted.
  • A stopped campaign still shows what it sent before it stopped.
  • Narratives drop and count any reference the campaign's findings cannot support.

Questions

Questions buyers ask

What is a digital twin in cyber security?

A digital twin in cyber security is a model of a system that can be examined or attacked instead of the system itself. In ColossalX CyberTwins the twin is narrow on purpose: a configuration snapshot of your agents, their guardrails and the discovered estate, stood up as shadow agents, not a copy of your network.

How does ColossalX CyberTwins attack AI agents?

It sends real adversarial prompts to shadow agents that carry your captured guardrails, one run per agent, through your gateway. Campaigns can also target named registered agents and record each phase: what was attempted, what worked, what the gateway stopped and by which control, and the phases never attempted.

Can we test a guardrail change before it reaches production?

Yes. A proposed change names the rule the gateway would apply differently and the probes behind it. You try it on the twin's shadow agents only, re-attack them, and promote it to production only if it held. Promotion keeps the rules it replaced, so a revert restores them exactly.

What are attack paths and blast radius?

An attack path is a route an attacker could walk, such as a committed credential, then a vulnerable dependency, then reachable personal data. Blast radius is everything affected if one part falls. ColossalX derives paths from your own scan findings, says each step is available rather than taken, and highlights blast radius on the twin graph.

Who are the twin's narratives written for?

A completed campaign can be written up for the board, the CISO, the engineers or the regulator, from its real findings and attack paths, by the AI model your workspace chose. Each reference is checked, and references the findings do not support are dropped and counted, so a narrative is only as good as its evidence.

Related

Next step

Know your x, tested.

See a guardrail change tried on a twin of your own agents before it reaches production.

  1. 01Tell us what you run
  2. 02See the four verbs on it
  3. 03Decide where to start