mGeNTE
mGeNTE (Multilingual Gender-Neutral Translation Evaluation) is a natural, multilingual corpus designed to benchmark gender-neutral language and automatic translation.mGente is built upon European Parliament speech data extracted...
Read Moreby Beatrice Savoldi | Jan 13, 2025 | Corpora | 0
mGeNTE (Multilingual Gender-Neutral Translation Evaluation) is a natural, multilingual corpus designed to benchmark gender-neutral language and automatic translation.mGente is built upon European Parliament speech data extracted...
Read Moreby Beomseok Lee | Aug 21, 2024 | Corpora | 0
Spoken Language Understanding (SLU) involves interpreting spoken input using Natural Language Processing (NLP). Voice assistants like Alexa and Siri are real-world examples of SLU applications. The core tasks in SLU include...
Read Moreby Mauro Cettolo | Apr 30, 2024 | Corpora | 0
Ready-to-use version for MT research purposes of the multilingual transcriptions of TED talks
Read Moreby Dennis Fucci | Oct 20, 2023 | Corpora | 0
Text corpora for Spanish, French, and Italian containing gendered words referring to the first-person speaker
Read Moreby Beatrice Savoldi | Oct 19, 2023 | Corpora | 1
The INclusive Evaluation Suite (INES) is a test set designed to assess MT systems ability to produce gender-inclusive translations for the German→English language pair. By design, each German source sentence in INES includes an...
Read Moreby Beatrice Savoldi | Oct 9, 2023 | Corpora | 0
GeNTE (Gender-Neutral Translation Evaluation) is a natural, bilingual corpus designed to benchmark the ability of machine translation systems to generate gender-neutral translations. Built from European Parliament speeches,...
Read Moreby Marco Gaido | Jul 7, 2023 | Corpora | 0
EC Short Clips is a test set dedicated to evaluate automatic subtitling systems.
Read Moreby Marco Gaido | Jul 7, 2023 | Corpora | 0
EuroParl Interviews is a test set dedicated to evaluate automatic subtitling systems.
Read Moreby Matteo Negri | Jun 1, 2023 | Corpora | 0
Multilingual benchmark built from European Parliament speeches and annotated with Named Entities and Terminology
Read Moreby Mauro Cettolo | May 30, 2023 | Corpora | 0
Annotation of dubbing segments based on the Heroes corpus
Read Moreby Beatrice Savoldi | May 30, 2023 | Corpora | 0
This multilingual dataset was created within the TOSCA-MP project as ground truth data for the evaluation of automatic transcription and spoken language translation technologies.
Read More
🚀 New tech report out! Meet FAMA, our open-science speech foundation model family for both ASR and ST in 🇬🇧 English and 🇮🇹 Italian.
The models are live and ready to try on @huggingface 👇
🔗
#ASR #ST #OpenScience #MultilingualAI
Our pick of the week by @lina_conti: "Languages in Multilingual Speech Foundation Models Align Both Phonetically and Semantically" by @soheunshim, Domenico De Cristofaro, Chengzhi Martin Hu, Alessandro Vietti, and @barbara_plank (2025).
#speech #SFM #multilingual #speechtech
Pick of the week @fbk_mt: https://arxiv.org/abs/2505.19606 by @soheunshim @DomenicoDeCris1. XAI work on cross-lingual alignment in speech-to-text models that disentangles phonetics and semantics. Plus: their XAI insights yield actionable improvements for low-resource language performance.
🚀 New shared task at #WMT2025 (co-located with @emnlpmeeting ): Model Compression for Machine Translation!
Can you shrink an LLM and keep translation quality high?🔧
Submit by July 3 and push the limits of efficient NLP!
👉 https://www2.statmt.org/wmt25/model-compression.html #NLP #ML #LLM #ModelCompression
More great news! 🎉
Our paper “Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution” was accepted at #Interspeech2025!
Interested in interpretability for speech models? Preprint coming soon!
✍🏼 @mgaido91, @negri_teo, M.Cettolo, @luisabentivogli