Documentation Index
Fetch the complete documentation index at: https://docs.getcollate.io/llms.txt
Use this file to discover all available pages before exploring further.
Overview
Auto-Classification is a Collate workflow that automatically detects and tags sensitive data — such as PII — across your database columns. It removes the need for manual tagging by scanning both column names and sample data during ingestion, then applying or suggesting tags likePII.Sensitive and PII.NonSensitive.
How It Works
Auto-Classification uses two complementary detection approaches:-
Column Name Scanner: Validates column names against a set of regex rules that identify common sensitive patterns — email addresses, names, SSNs, bank account numbers, and similar fields.
For example, columns
emailandfull_nameare auto-tagged asPII.Sensitivebased on their column names.
-
Entity Recognition: If sample data ingestion is enabled, scans the actual row values using an NLP-based entity recognition engine. This catches sensitive data even when the column name is generic or ambiguous. The
confidenceparameter (0–100, default80) controls the minimum score required to tag a column asPII.Sensitive. If a column already has aPIItag, it is skipped during execution. For example, the columnI_FORMULATIONis also tagged asPII.Sensitive, even though its name gives no indication of sensitive content. Inspecting the Sample Data tab reveals that the actual row values contain sensitive information, which the entity recognition engine detected. This shows that auto-classification works beyond column names and relies on the data itself when sample ingestion is enabled.

Glossary Term Associated Tags
Separate from the auto-classification workflow, Collate can derive classification tags from glossary terms. If a glossary term has associated classification tags, applying that glossary term to an asset also applies the associated tags as derived tags. For example, if the glossary termAccount has PII.Sensitive associated with it, adding the Account glossary term to a table or column also adds PII.Sensitive. This behavior is configured on glossary terms; it is not generic classification-tag-to-classification-tag mapping.
Set Up Auto-Classification
Workflow
Add an Auto Classification Agent to a database service directly from the Collate UI.
External Workflow
Run the Auto Classification Workflow externally using a YAML pipeline configuration.
Auto PII Tagging
Understand the tagging logic and troubleshoot common issues like SSL certificate errors.
Custom Recognizers
Define custom rules to detect and tag sensitive data using regex patterns, exact terms, or pre-built detectors.
Tag Feedback and Approvals
Report false positives on auto-applied tags and manage approval workflows to continuously improve classification accuracy.
Sample Data
Store sample data collected during auto-classification to an S3 bucket in Parquet format.