What is PII data discovery?
PII data discovery is the automated process of scanning your databases, documents, cloud storage, APIs and AI systems to locate every instance of personally identifiable information your organisation holds. It combines data discovery (finding where data lives), data classification (labelling each finding by type and sensitivity) and data mapping (tracing how that data flows between systems). The output is a complete, continuously updated inventory of personal data that supports DPDP Act compliance, breach response and data subject rights fulfilment.
What is PII classification and how does it work?
PII classification is the automated labelling of every discovered personal data field by its type (Aadhaar, PAN, mobile number), sensitivity level, applicable regulation and business criticality. Our platform classifies each finding in real time using AI models rather than manual tagging, so your data catalog stays current as new data appears. Classification determines what protection each field needs: an Aadhaar number in a public-facing system demands different handling from a name in an internal HR file.
Does the platform store the PII it discovers?
No. The platform uses a scan-in-place architecture. Data is analysed at source and never copied to our servers. The platform stores only metadata about each PII finding: the location (system name, table, column), the PII type detected, the risk score and the timestamp. No actual personal data, no Aadhaar numbers, PAN details or customer records, ever leaves your environment.
Can it scan unstructured documents like PDFs and scanned images?
Yes. The platform uses OCR-powered scanning to detect PII in PDFs, scanned images, Word documents, spreadsheets, email attachments and shared drives, including KYC documents written in Hindi, Tamil, Bengali, Telugu and 7 other Indian scripts. This covers Aadhaar cards, PAN cards and other identity documents uploaded by users that would be completely invisible to a database-only scanner.
Does it support monitoring of AI workflows and LLM usage?
Yes. The platform monitors LLM prompt-response pairs sent to OpenAI, Gemini, Claude and Azure OpenAI APIs, ML training datasets, vector databases and RAG pipelines, and AI-generated outputs, flagging personal data that appears in any of these AI surfaces. This is critical under the DPDP Act, where sending Aadhaar or PAN data to a third-party LLM provider without consent and a valid purpose creates a direct regulatory violation.
How accurate is detection? What about false positives?
The platform achieves 98.7 percent detection accuracy with under 2 percent false positives across all data source types. This is achieved through a multi-layer approach combining Named Entity Recognition (NER), regex validation with format-specific rules and contextual analysis, rather than the single-layer regex matching that most sensitive data discovery tools use. India-specific identifiers (Aadhaar, PAN, GSTIN) include checksum and format validation to eliminate structural false positives.
Is the platform suitable for DPDP Act compliance?
The platform was built specifically for DPDP Act 2023 compliance. It detects all India-specific personal data identifiers natively, maps every finding to specific DPDP Act obligations (Sections 4, 8 and 9), auto-generates RoPA documentation, links the PII inventory to consent records, and produces audit-ready reports for the Data Protection Board. It also covers GDPR, HIPAA, PCI DSS, ISO 27001 and RBI and SEBI obligations in the same scan.
Can we deploy it on-premise or in our own cloud?
The platform supports cloud-hosted, on-premise and hybrid deployment models. For organisations with strict data residency requirements, particularly BFSI institutions subject to RBI data localisation guidelines, the scan engine can be deployed entirely within your own cloud environment or on-premise infrastructure. Data never leaves your perimeter in any deployment model, as the scan-in-place architecture analyses data at source regardless of where the engine runs.
How long does the initial scan take?
Most organisations complete their first full PII discovery scan in under 4 hours for standard environments of up to 50 connected data sources. Large enterprise environments with 100 plus data sources and petabyte-scale storage typically complete initial scans within 24 to 48 hours. Scan speed does not affect data source performance. Scans are designed to consume under 5 percent of available database or storage I/O capacity to avoid any impact on production systems.