Redescription Mining Framework for Post-Hoc Explanation of Deep Learning Models
An architecture-independent method for interpreting a trained deep network by mining statistically significant redescriptions of its neuron activations, linking neurons to target labels or descriptive attributes and relating layers within a model or across different models. It supports multi-label and multi-target settings and can reproduce pedagogical and decompositional rule-extraction behaviours, offering interpretability information distinct from mainstream explainable-AI techniques.
D
Data
Text
A redescription mining framework for post-hoc explaining and relating deep learning models Matej
Primary AI research proposing a framework that uses redescription mining to post-hoc explain and relate deep learning models. It performs cohort analysis of an arbitrary trained network by finding st…