Poeta, E., Pastor, E., Cerquitelli, T., & Baralis, E.
PKDD/ECML Workshops
2026
Modern vision models often achieve strong average predictive performance but can exhibit substantial disparities across demographic and contextual groups, visual attributes, or environmental contexts that are hidden by aggregate metrics. Existing auditing approaches typically rely on predefined demographic annotations or manually specified groups, limiting their ability to uncover previously unknown sources of bias. Consequently, important model failures may remain concealed within latent populations not captured by available metadata.
We propose TRACE, an auditing framework that combines mechanistic interpretability with subgroup discovery, without access to the metadata. Starting from frozen CLIP representations, TRACE trains a Sparse Autoencoder (SAE) to uncover latent visual concepts, assigns semantic names through CLIP text–image alignment, and represents each image through its most active concepts. These concept signatures are then analyzed by subgroup discovery to identify concept-defined populations exhibiting anomalous predictive behavior. By linking subgroup performance disparities to interpretable concept combinations, TRACE provides a practical framework for auditing predictive disparities in vision models without demographic annotations. More broadly, our work demonstrates how mechanistically discovered features can serve as an interpretable basis for subgroup auditing and reliability analysis.
[Poeta et al., 2026] Poeta, E., Pastor, E., Cerquitelli, T., & Baralis, E. (in press). TRACE: Transparent representation auditing via concept extraction. PKDD/ECML Workshops 2026, Naples, Italy, September 7–11, 2026. Accepted for publication.


Lascia un commento