Corpora

mGeNTE

mGeNTE (Multilingual Gender-Neutral Translation Evaluation) is a natural, multilingual corpus designed to benchmark gender-neutral language and automatic translation.mGente is built upon European Parliament speech data extracted...

Read More

MOSEL

The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled...

Read More

Speech-MASSIVE

Spoken Language Understanding (SLU) involves interpreting spoken input using Natural Language Processing (NLP). Voice assistants like Alexa and Siri are real-world examples of SLU applications. The core tasks in SLU include...

Read More

INES

The INclusive Evaluation Suite (INES) is a test set designed to assess MT systems ability to produce gender-inclusive translations for the German→English language pair. By design, each German source sentence in INES includes an...

Read More

GeNTE

GeNTE (Gender-Neutral Translation Evaluation) is a natural, bilingual corpus designed to benchmark the ability of machine translation systems to generate gender-neutral translations. Built from European Parliament speeches,...

Read More
Loading

🚀 New tech report out! Meet FAMA, our open-science speech foundation model family for both ASR and ST in 🇬🇧 English and 🇮🇹 Italian.

The models are live and ready to try on @huggingface 👇
🔗

#ASR #ST #OpenScience #MultilingualAI

🚀 New shared task at #WMT2025 (co-located with @emnlpmeeting ): Model Compression for Machine Translation!
Can you shrink an LLM and keep translation quality high?🔧
Submit by July 3 and push the limits of efficient NLP!
👉 https://www2.statmt.org/wmt25/model-compression.html #NLP #ML #LLM #ModelCompression

More great news! 🎉
Our paper “Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution” was accepted at #Interspeech2025!

Interested in interpretability for speech models? Preprint coming soon!

✍🏼 @mgaido91, @negri_teo, M.Cettolo, @luisabentivogli

Load More