# Data lineage: which personal data reached which model.

> AI data lineage in ColossalX shows which agent sent which kinds of personal data to which model, drawn from real gateway traffic rather than a questionnaire. Each flow separates what was sent, what was held back before it left and what came back in answers, and impact analysis shows what a change to one source would touch.

Lineage drawn from real AI traffic: each agent, each model, and the personal data sent, held back or returned.

Canonical page: https://colossalx.tech/platform/data-lineage · Last reviewed: 6 Oct 2026

*Illustration:* One flow, end to end: client records to kyc-agent (read); kyc-agent to gateway (name, PAN); gateway with data redacted, to provider A (PAN held back); gateway to lineage (captured).

## The threat and the control

- **The threat:** A data protection officer is asked which models saw customer identifiers, and has only a questionnaire.
- **The control:** ColossalX draws each flow from gateway traffic, with the personal data sent, held back and returned.

## How it works: From captured traffic to an answer for the auditor.

One question from a data protection officer, which models saw card numbers, answered from captured flows rather than memory or a spreadsheet, with the source of each flow named.

### Workflow: which models saw card numbers, answered (illustrative)

1. **Asked.** A data protection officer asks which models received card numbers this quarter.
   Data protection officer: Which models saw card numbers this quarter? | Card data · All agents
2. **Captured.** The latest capture draws each caller-to-model flow from gateway traffic.
   kyc-agent · to provider A · Captured | claims-bot · to provider B · Captured | Source: gateway traffic; Last capture: today, 08:50
3. **Separated.** Each flow shows card data sent, held back before leaving, or returned.
   kyc-agent to provider A (done: held back); claims-bot to provider B (failed: sent); Answers from provider B (done: none returned)
4. **Answered.** One flow sent card data; impact analysis shows what it touches.
   Flow: claims-bot to provider B; Source: claims archive; Next: switch on the card preset | Impact analysis | x, found

## What you see: The lineage graph, captured from the gateway.

Sources, agents, models and stores on one graph, with personal-data paths marked and the time of the last capture on the view.

1. **Captured from traffic.** Caller-to-model flows are drawn from real gateway traffic, and refreshed daily. Capture runs on demand and each day. Each flow counts the requests served and refused. A flow no longer seen is marked inactive, never deleted.
2. **Sent, held, returned.** Each flow separates personal data sent, held back, and returned in answers. Kinds of personal data come from the gateway's detectors: 6 presets, such as India DPDP, GDPR, HIPAA and PCI DSS, plus identifiers you define.
3. **Impact analysis.** Pick a source and everything upstream and downstream of it stays lit. Click a node for its classification, region, personal-data fields and retention. Impact analysis dims everything a change to that node would not touch.
4. **Removed, never erased.** Removing a flow flags it with who and why; the record stays. A captured flow can be annotated. A flow drawn by hand can be edited. Removal is a flag with a person and a reason, so the history survives an audit.

*Illustration:* An illustrative data lineage map drawn from gateway traffic: client KYC records and a card ledger read by two agents, through the gateway to two models. The PAN is withheld at the gateway, the name is sent on, and the flow that crosses a region boundary is marked.

## How we know

- Flows are drawn from real gateway traffic, not from a questionnaire or a declaration.
- Each flow separates personal data sent, held back before it left, and returned in answers.
- A removed flow is flagged with who and why, and a flow no longer seen is marked inactive.
- The same flows light the data lens on the AI Agent Map.

## Where a mapped flow goes next.

A flow on the graph is a fact about your data. Each one feeds the controls and the records that depend on it.

- **Runtime guardrails.** Personal data is held back with 6 presets and your own identifiers, on prompts and answers.
- **Runtime consent.** A person who withdraws consent has their next AI request refused.
- **Compliance.** Captured flows become evidence for privacy assessments and records of processing.
- **The agent map.** The data lens on the map shows which agents touch which data.

Where an x ends up: x, found.

## Specs: delivery and data

- **Delivery:** SaaS, from one login.
- **Isolation:** Each customer runs in an isolated workspace with its own database.
- **Certifications:** None held. Frameworks are mapped to and assessed against.

## Frameworks

- Mapped to India DPDP: Personal-data flows, as evidence.
- Mapped to GDPR: Records of processing, from real traffic.

## What it does not do

- Captured flows cover traffic through the ColossalX gateway; other flows are drawn by hand.
- Cross-border flows are drawn only between registered assets with different storage locations.
- Personal-data counts come from the detectors you switch on: presets and your own identifiers.
- Content-scan violations are not yet linked to the lineage view.

*Illustration:* What lineage can see: kyc-agent to gateway (captured); gateway with data redacted, to provider A (PAN held back); batch script to drawn by hand.

## Questions

### What is AI data lineage?

AI data lineage is the record of where data goes when it enters an AI system: from a source, through the agents that read it, to the models that receive it and the answers that come back. Done well, it is drawn from what actually happened, so it can answer which models saw which kinds of personal data.

### How does ColossalX build lineage?

From real traffic. ColossalX captures caller-to-model flows from the requests that pass through its gateway, on demand and each day, and counts the requests served and refused on each flow. Flows that do not pass through the gateway can be added by hand, and each one says how it got there.

### Can we see what was sent, held back and returned?

Yes. Each captured flow separates the kinds of personal data sent to the model, the kinds found and held back before they left, by redaction or refusal, and the kinds returned in answers or redacted from them. That is the difference between knowing a flow exists and knowing what it carried.

### How does AI data lineage support DPDP and GDPR assessments?

It gives an assessment facts instead of assumptions: which agents send which kinds of personal data to which models, and what was held back. ColossalX maps its controls to India DPDP and GDPR and keeps captured flows as evidence. It supports your privacy assessments and records of processing; it does not decide them for you.

### What is impact analysis for AI data flows?

Impact analysis shows what a change would touch. Pick a source, an agent or a model on the lineage graph and everything upstream and downstream of it stays lit while the rest dims, so you can see which agents and models depend on a dataset before you move, reclassify or retire it.

## Related

- [Runtime guardrails](https://colossalx.tech/platform/runtime-guardrails)
- [Runtime consent](https://colossalx.tech/platform/runtime-consent)
- [Shadow AI](https://colossalx.tech/platform/shadow-ai)

---

ColossalX is an AI security and governance platform from Quantexra Labs LLP, delivered as SaaS. Book a walkthrough: https://colossalx.tech/demo · client.success@quantexra.tech
