699 raw rows compared across exact-row, features-only, and ID-only duplicate definitions
The cleaned dataset removes only exact duplicate rows, reducing the raw file from 699 rows to 691. That approach is conservative: it preserves records when the ID or any clinical feature differs.
Two alternative duplicate definitions behave very differently. If duplicates are defined by identical clinical features while ignoring ID, 236 rows become redundant and only 463 rows survive. If duplicates are defined by ID alone, 54 rows are redundant and 645 survive.
Repeated IDs are not automatically interchangeable records: 39 of the 46 repeated IDs have more than one distinct feature profile, and 4 repeated IDs span both class labels. That means the ID-only rule can collapse clinically distinct entries.
Each bar uses the same 699-row baseline so the effect of each deduplication rule is directly comparable.
Bars count repeated IDs by frequency; the overlaid line shows how many extra rows those repeats contribute under the ID-only rule.
Bars show unique feature profiles per ID, while the line shows total rows for that ID. When the bar equals the line, every reuse introduces a different feature pattern.
| ID | Rows | Distinct Profiles | Class Values |
|---|
| ID | Rows in Group | Class |
|---|