Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

Publication
Non-Proceedings Track