Legal Register
Part II — Target Operating Model & Core Workflows
Part II — Target Operating Model & Core Workflows
Chapter 07 · 3,042 words
15 min read

Chapter 7 — Discovery, Processing Inventory and Data-Flow Mapping

1. The enterprise that does not know what it holds

The inventory begins with a question that looks easier than it is: what personal data does this organisation process, where, for whom, and why? A declared database list cannot answer for an unregistered marketing export, a supplier’s retained copy, or a digitised document archive. These are plausible discovery risks, not measured prevalence claims. The fictional Company used below has declared and unverified systems; the distinction is the starting evidence, not an embarrassment to hide.

Chapter 8 reviews purpose and authority for each activity. Chapter 14 assesses qualified retention and disposal when purpose ends. Chapter 16 will scope a breach to the affected principals. Chapter 4 will answer a principal who asks what is held and with whom it was shared. Every one of those controls presupposes the answer to the simple question — and an enterprise that cannot answer it cannot do any of them. The failure mode has a specific shape: each downstream control fails silently, in exactly the places the inventory missed.

The inventory is an author-recommended dependency map, not an independently prescribed statutory register. Start broad enough to discover relevant operations, then record actual scope and exemption decisions. Breadth of discovery does not make every discovered record legally in scope.


2. What the law actually demands, by implication

No provision of the Act says “maintain an inventory.” The demand is distributed across the obligations, and reading it that way shows why the inventory cannot be thin:

Section 2(x) defines processing as wholly or partly automated operations on digital personal data, including collection, storage, use, sharing and erasure. This supplies the lifecycle vocabulary, not a blanket statutory requirement to catalogue every physical object (ACT:122–126).[1]

Section 3(a) covers processing within India where data is collected digitally or subsequently digitised. Never-digitised paper is not brought in by that limb alone. Section 3(b) covers offshore processing connected with activity related to offering goods or services to Data Principals within India’s territory: location and offering facts matter, not Indian nationality. Section 3(c)‘s personal/domestic and qualifying public-data exclusions need their actual facts; online discoverability alone is insufficient (ACT:135–157).[1]

Section 8(3) requires completeness, accuracy and consistency where data is likely to be used for a decision affecting the principal OR disclosed to another fiduciary. Lineage helps fulfil both limbs (ACT:338–342).[1]

Section 8(7) links erasure to withdrawal or reasonable purpose-end, whichever is earlier, unless retention is necessary for compliance with law; its processor limb makes undiscovered copies operationally important. Read the Rules’ separate retention layer before issuing disposal instructions (ACT:351–385; RULES:1101–1106,1142–1166).[1][3]

Section 11 requires an eligible access response to include a summary of data and activities, recipient identities and descriptions of shared data. Its prior-consent framing includes Section 7(a); Section 11(2) limits specified sharing information on qualifying facts, not the entire inventory or every access answer (ACT:441–462).[1]


3. The tension: breadth of discovery versus cost and completeness

The inventory’s central tension is between breadth and feasibility. The following failure modes are author analysis; no client programme or completed processor audit is claimed.

The breadth pull. The recommended discovery sweep is broad: the paper archive, the shadow SaaS, the vendor copies, the logs, the derived features. The broader the target set, the more honest the inventory — and the more the effort costs, in time, in tooling, in the organisational friction of asking forty teams what data they hold and being believed by none of them.

The feasibility pull. So programmes narrow. They inventory the declared systems, scan the obvious databases, declare victory, and file the artefact. The narrowed inventory is complete enough to look convincing and incomplete enough to be dangerous — because every system it missed becomes a permanent blind spot that downstream controls will faithfully never reach.

The failure modes, named:

  • Declared-only. Owner nominations are a useful seed, not a completeness certificate. Compare declarations with procurement, identity-provider and network records. A discrepancy is a finding to investigate, not proof that a named supplier is necessarily a processor or unlawful.
  • Assumed-complete scanning. The mirror error: a PII scanner runs, “finds” data everywhere, and the programme treats the scan as the inventory. But a scanner without measured precision and recall is a claim, not a measurement — the Chapter 7 validation discipline exists because a scanner that misses Indian identity formats silently produces an inventory with confidence it did not earn.

The insight. The reconciliation is triangulation with measured validation, and an honest unknown set. Declared sources, structural sources, content scans, runtime observation, and SaaS hunting each find what the others miss; the inventory records how each dataset was found and grades its confidence; and — the discipline most programmes skip — the residual “unknown” set is named, owned, and scheduled for closure rather than silently absorbed. A visible, owned gap is more actionable than an invisible one; neither a gap register nor a scan certifies legal compliance.


4. The inventory row: what each dataset carries

An inventory is only as useful as its rows are rich, and the row design determines what downstream controls can even ask. The minimum workable schema, drawn from what the later chapters actually consume:

FieldWhy the later control needs it
Stable dataset / flow / system IDsidentify the same object across rights, withdrawal, breach and later migrations
Format and itemised categoriesdistinguish records, logs, digitised images and potentially identifiable derivatives
Scope / exemption decision and factsterritorial offering facts, Section 3 exclusion or precise Section 17 branch, evidence and reviewer; unknown is not exempt
Owner / source of truth / verification timewho can resolve a stale or conflicting lineage assertion
Purpose links, each with ground and statusmany purposes per dataset; rejected training must not inherit loan consent
Collection point and subject provenancedirect provision by the person vs someone else’s contact details
Retention decisions and permitted usesseparate trigger, minimum, hold and disposal; not one universal min/max
Recipient activity / role / flow / contractprocessor instructions vs another fiduciary’s decisions; purpose-specific disclosure descriptions
Storage, support and backup geographytransfer review input, including discovered but unapproved locations
Discovery evidence, sample denominator and unknown ownermeasured sample results separated from the unmeasured estate

The last row is the one that makes the inventory an instrument rather than a list, and it is where the worked example below earns its keep.


5. The discovery techniques, and what each finds

Six techniques, each with a distinct yield; the honest inventory runs several and records which found what.

Declared / self-reported. Business owners nominate their systems and data. This provides a fast initial map whose coverage must be tested.

Structural discovery. Schemas, ERDs, data dictionaries, ETL/ELT lineage, API contracts. Yields: the shape of the known estate with high fidelity, including flows between systems the declared pass garbled. Misses: data outside the documented structures.

Content / PII scanning. Classifiers and recognizers (Chapter 28’s primitives, Chapter 29’s Presidio) over structured and unstructured stores. Yields: data where nobody declared it — the log directory, the shared drive, the forgotten export. The measurement discipline: precision and recall measured on the entity’s own labelled sample, per class, documented per iteration. Misses: classes the recognizers do not know — Indian identity formats, entity-specific codes — which is why custom recognizers (Ch.29) are the second half of the technique, not an optional extra.

Traffic / runtime discovery. Observability on collection endpoints, form submits, API ingestion, third-party SDKs. Yields: what is actually flowing in — including the adtech or analytics SDK nobody connected to “processing.” Misses: data at rest that never moves.

Shadow-IT / SaaS hunting. SSO app inventories, browser extension and domain analysis, expense and procurement records. Yields: the unsanctioned estate — the marketing tool, the survey service, the AI note-taker someone subscribed to. Misses: tools outside every register (personal accounts on corporate data — a policy problem the inventory surfaces rather than solves).

Paper-to-digital discovery. Document archives, scanning pipelines, OCR stores. Yields: the digitised estate Section 3(a) explicitly captures — frequently the oldest and most sensitive data the organisation holds. Misses: paper never digitised (out of scope — but worth recording as out-of-scope by design, not by omission).


6. Data-flow mapping: the inventory’s second half

The inventory says what is believed to exist; the flow map says what happens to it. Mapping is the book’s recommended way to connect the statutory lifecycle to executable instructions. For each high-value dataset, trace:

collection (form / API / third party)
  → validation and storage (source of truth)
  → use (queries, models, derived features — Ch.21's registry consumes this edge)
  → sharing (internal roles + processors/fiduciaries — the s.11(1)(b) answer)
  → retention (purpose-completion / legal hold — Ch.14's triggers)
  → erasure (the deletion receipt's targets, including s.8(7)(b) propagation)
  → backups and restore (the orphan test's blast radius)

Each edge carries three attributes the later chapters will demand: the purpose it serves, the ground that authorises it, and the third-party role and applicable contract/disclosure authority. Section 8(2) requires a valid contract for the specified processor engagement, not every relationship with any third party (ACT:335–337).[1] A map without those attributes is a diagram; with them, it is the input to the purpose matrix (Ch.8), the transfer map (Ch.18), and the breach scoping (Ch.16).


7. The control-test-evidence set

The inventory is itself a control — the one that fails first when neglected — and it takes the Chapter 22 discipline like any other:

  • Recall test: a known set of PII is seeded in a sandbox spanning formats the estate uses; the inventory and its scanners must find it. The measured miss rate is recorded, not hidden.
  • Coverage test: every business-owned system is represented in the inventory — the declared estate is reconciled against the org chart and the SSO register, and each gap is a named finding.
  • Accuracy / drift test: a re-scan after a fixed interval detects new endpoints, shadow SaaS, and schema changes — the inventory’s freshness is measured, because a stale inventory is the close-and-decay failure of Chapter 36 in miniature.
  • Purpose-resolution test: every released processing activity resolves to an approved authority. Unresolved discoveries remain visible in the inventory but outside the allow list; do not delete the unknown set to make a launch report green.

8. A populated discovery validation pack

INV-001 below is synthetic CASE-001 data, owned by DataOps with Architecture as reviewer, version Q03-1, last illustrative verification 1 June 2027. A verification date here is an authored scenario field, not a real inspection. The complete reusable fragment is out/remediation/Q03/inventory.json. The source/version context is the frozen canonical register and dossier contract.

Dataset / flowLocation, purpose and role factsScope / evidence statusDecision and unknown owner
DS-001 / FLOW-001SYS-001 to SYS-002; Company determines lending purpose; PUR-001 consentdigital processing in India; no asserted exemption; inventory and CONSENT-001 stipulatedallow only authorised lending scope; DataOps owns reconciliation
DS-003 / FLOW-002SYS-002 to ENT-004’s SYS-003; optional PUR-002; processor acts on instructionsIndian customer journey stipulated; CONSENT-002 separate; processor internal copies not inspectedconditional before withdrawal; Supplier Manager owns copy-list gap
DS-006 / FLOW-003warehouse SYS-004 to training SYS-005; proposed PUR-003identifiable training inputs; no grant; no research exemption establishedDEC-001 stop, ML owner; existence in inventory is not permission
DS-009 / FLOW-004SYS-002 to offshore SYS-008 backupoffering-linked offshore facts; exact transfer/sector analysis unresolvedDEC-004 isolate/reroute; Architecture owns missing decision
DS-007, DS-008 / FLOW-007SYS-006 to digitised SYS-007 archiveemployee PUR-006 distinct from applicant PUR-014; Section 7(i) applicant boundary unresolvedrestrict applicant reuse; HR/Legal own determination

DS-001 has several purpose links; a single ground field would lose the distinction between lending, optional marketing and the requested receipt PUR-007. ENT-002’s independently determined insurance activity is not labelled a processor merely because FLOW-006 originates at the Company. The table records role facts as scenario assumptions; the operator must replace them with contract and decision-right evidence in a real engagement.

The labelled sample and what its arithmetic can establish

The pack contains twelve explicitly synthetic labelled tokens, three per class: email, phone, identity-like code and internal customer code. Each class has two known positives and one negative. A stipulated baseline classifier output finds both emails, one phone, one identity-like code and neither internal code; it falsely labels the negative identity-like code. A proposed recognizer revision still misses one internal code. These are authored prediction arrays, not the output of Presidio or a client scan. out/remediation/Q03/checks.py computes their confusion matrices; it does not benchmark a vendor.

ClassBaseline TP / FP / FN / TNBaseline precision / recallRevised TP / FP / FN / TN
Email2 / 0 / 0 / 11 / 12 / 0 / 0 / 1
Phone1 / 0 / 1 / 11 / 0.52 / 0 / 0 / 1
Identity-like1 / 1 / 1 / 00.5 / 0.52 / 0 / 0 / 1
Internal code0 / 0 / 2 / 1undefined / 01 / 0 / 1 / 1

Baseline micro precision is 0.8 and recall 0.5; revised recall is 0.875, not completeness. Undefined precision is not zero error. With only two positives per class, the point values are not estimates of estate-wide detection quality and no confidence interval is claimed. Different prevalence and undocumented file formats would change performance. A production study should sample actual formats and adjudicate labels independently, rather than tuning and reporting on the same tiny fixtures.

Coverage has a separate denominator: eight reviewed systems out of the fourteen declared SYS-001–SYS-014 systems, not eight out of an unknowable universe. SYS-003, SYS-005, SYS-008, SYS-009, SYS-013 and SYS-014 remain unverified in this initial discovery snapshot. Later event examples stipulate their behaviour for a teaching exercise; they do not retroactively constitute inspection evidence. Confidence rubric: A requires reconciled owner, structural and content/flow evidence; B has owner and structural evidence but incomplete copy checks; C is declaration only; U is unverified. The synthetic rows are B or U, never A by assertion.

The unknown register assigns the ENT-004 copy-list to Supplier Manager and offshore backup facts to Architecture, each with an illustrative review date of 8 June 2027 and escalation to the DPO if absent. No unresolved flow may enter an approved launch list merely because a date is assigned. This is the key difference between an honest gap register and a register that launders uncertainty into approval.

Exercise and failure branch

Add a second purpose to DS-001, omit its authority, and ask whether the existing loan consent unlocks it. Expected answer: no; the missing purpose link is a separate decision. Next remove SYS-003 from the destination list while FLOW-002 still names it. The reference check must reject the pack, not quietly improve its coverage percentage. A passing local consistency check only establishes agreement among these authored rows; it does not prove the absence of shadow systems.


9. What remains for the reader and the reviewer

The open items, each <residual>:

  1. Confirm offering, location and role facts for the entity’s offshore relationships; a contract does not itself prove complete discovery.
  2. Apply Rule 8(1)/(2) only to the Third Schedule classes: threshold-qualified e-commerce/social-media entities at two crore and gaming intermediaries at fifty lakh registered users in India, with the three-year period and stated account/token-purpose exclusions; include the at-least-48-hour warning. These are available source facts, not unread residuals (RULES:1142–1152,1598–1680).[3]
  3. Resolve record-specific retention overlap, reset dates and holds under documented legal review. Rule 8(3)‘s one-year minimum and Rule 6’s security retention remain separate inputs even when a Third Schedule class does not apply (RULES:1101–1106,1153–1166).[3]

The question that hands the book its next chapter

The inventory is now a scoped, versioned set of claims with visible uncertainty. The next question is why each processing activity is authorised. Chapter 8 gives each purpose a fact-specific ground decision; it does not turn discovery into permission.


Evidence and reusable artifacts

The primary-text line keys ACT, COMM, RULES and CORR resolve to the retained files below. Line numbers count physical newlines, not PDF form feeds. The canonical provision register supplies actor, trigger, conditions, exceptions and effective dates; chapter recommendations and synthetic examples are not statutory forms. The Q03 source manifest preserves URL, retained retrieval metadata and recalculated hashes.

Completed chapter fragments, before/after evidence and actual local checks: out/remediation/Q03/. Blank operating templates remain under research/operations/templates/; The reconciled integrated dossier is the reader working copy; these chapter fragments preserve the earlier bounded examples and their run evidence.

Sources

[1] https://www.meity.gov.in/static/uploads/2024/06/2bf1f0e9f04e6fb4f8fef35e82c42aa5.pdf — ACT [2] https://www.meity.gov.in/static/uploads/2025/11/c56ceae6c383460ca69577428d36828b.pdf — COMM [3] https://www.meity.gov.in/static/uploads/2025/11/53450e6e5dc0bfa85ebd78686cadad39.pdf — RULES [4] https://www.meity.gov.in/static/uploads/2025/12/3c7ebbae0e5456f493f486e6845df86b.pdf — CORR


Contents · Reader guide and citation conventions · Artifact index