SONAR is an evaluation toolkit for multilingual ASR that goes beyond WER/CER. It combines semantic similarity, the Poseidon Score, and analysis across dialect, demographic, and metadata-based failure modes. 🌍
Our goal is to make it easier for everyone to understand why an ASR model fails, not just how often. 🔍 You can plug in your own models + audio, extend it to new languages and datasets, or contribute directly. 🛠️
MIT licensed. Would love feedback from the HF community! 🤗