NESPOLE! System has been developed using two scenarios: the tourism scenario and the first aid medical assistance scenario. During the project life three main data collection have been carried on in order to develop the first and the second showcase. During the first year 191 dialogues have been collected. There are 62 German dialogues recorded, 61 Italian, 37 English and 31 French. Particularly an amount of 6 hours of dialogues for Italian and French, 7 hours for English, 8 hours for German has been recorded. Dialogues were about five predefined tourism scenarios. During the last year two major data collections have been carried on: the first one aimed at expanding the tourism scenario and the second one at addressing the medical domain. For the monolingual data collection five tourism scenarios were developed; 66 dialogues were recorded yielding 994.57 minutes of data: 243.52 minutes comprised in sixteen English dialogues, 246 minutes in sixteen German dialogues, 272.52 minutes in seventeen French dialogues and 232.53 minutes in seventeen Italian dialogues. The data collection on the medical domain involved Italian, English and German languages. A total of 49 dialogues were collected. The recording results in a total of 8 hours 25 minutes of audio files.
Our pick of the week by
@mgaido91
: "FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model" by Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang, Xu Li, Yufan Lin, Junwen Qiu, Jun Zhang, Lu Lu, Haizhou Li and @drwuz
#SLM #EfficientInference
Cool to see a work that adaptively chooses at inference how much to compress the input speech sequence, to control inference costs and quality based on the input, without enforcing a global trade-off to each segment: https://arxiv.org/pdf/2606.31247
@fbk_mt
Our pick of the week by
@FBKZhihangXie : "Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models" by Haoqin Sun, @Chenyang_Lyu, Shiwan Zhao, Xuanfan Ni, Xiangyu Kong, @wangly0229, Weihua Luo and Yong Qin
#SpeechLLM #LongFormSpeech #SLU
🚀 New paper: Speech-XL for long-form SpeechLLMs
📄 https://arxiv.org/abs/2602.05373
🧩 Uses Speech Summarization Tokens to compress local speech intervals into compact KV states efficiently.
✨ Improves long-form speech understanding while reducing memory and FLOPs on 10-minute audio.
Our pick of the week by
@BeatriceSavoldi
: "Accuracy: Community Perspectives on Machine Translation" by Yujun Wang,
@EhudReiter
, Shimei Pan,
@egere14
and Wei Zhao #MachineTranslation #TranslationQuality #Evaluation
📖 #PickoftheWeek @fbk_mt "Accuracy: Community Perspectives on Machine Translation"
A cool analysis of the conflicting interests of different communities around MT(AI developers, LSPs, and users)
https://arxiv.org/pdf/2606.09655
#NLP #MachineTranslation #DiverseStakeholders