Data Security in the Age of AI
Why traditional masking approaches start to fall short when AI begins reasoning over sensitive data, why the real risk often lies in data pipelines rather than databases, and what it takes to maintain accurate record matching when sensitive values are no longer directly visible. Drawing on insights from a closed-door roundtable convened by Skyflow and facilitated by AIM Research, the discussion brought together senior leaders from data, privacy, technology, and risk functions across India’s banking, financial services, and insurance sector to explore the challenges of protecting sensitive data while enabling AI-driven use cases.
For years, data security in financial services has been a question of storage: encrypt the disk, lock down the schema, control who can query the table. AI has quietly made that question beside the point. The sensitive data that matters most today isn't sitting still - it's moving through prompts, agent memory, retrieval calls, model outputs, and the dozens of tool invocations an AI agent makes to complete a single task. None of that activity looks like a database, and almost none of the tooling built for databases was designed to watch it. This is the shift that ran through the entire roundtable: security has to move from where data is stored to where data is actually used - at runtime, inside the pipeline, in the moment an AI system reads, reasons over, or generates something containing personal information. Leaders described a growing gap between institutions that still think of data protection as a storage problem and the reality of agentic AI, which touches sensitive information continuously, contextually, and often unpredictably. A specific and recurring frustration was masking. Most institutions already mask sensitive fields somewhere in their stack, and most leaders admitted it gives them false comfort. Masking is a static, cosmetic control - built for screens and reports, not for AI pipelines that need to reason over unstructured documents, chat transcripts, audio, and scanned forms in real time. When masked data hits an AI workflow, one of two things happens: the model gets garbage and the workflow breaks, or someone quietly routes the raw, unmasked value around the mask so the automation keeps working. Both outcomes defeat the purpose of masking. The second one is how PII actually ends up spewing through AI pipelines - not through a breach, but through a bypass nobody documented. A second, less obvious theme was referential integrity - not in the database sense of foreign keys, but in the practical sense of whether two systems can still recognise the same customer, transaction, or account once their sensitive identifiers have been protected. Institutions increasingly found that once they masked or hashed data differently across systems and partners, they lost the ability to match records at all. Protection and usability were pulling in opposite directions. The roundtable converged on tokenisation with runtime resolution as the way out of this bind: replace sensitive values with consistent, governed tokens that preserve the ability to match and reconcile data across systems, and resolve them to the real value only at the moment of use, only for the party authorised to see them, and only for as long as the purpose requires. Compliance, including India's DPDP Act, was treated as a natural by-product of getting this architecture right - not the reason for building it.
Category: Whitepaper | Published: 2026-08-24