Tasks that ultimately require knowledge-based multimedia techniques (content-oriented search, assessment, abstracting, etc.) are still to a major extent carried out manually. PATExpert’s overall scientific goal is to change the paradigm currently followed for patent processing from textual (viewing patents as text blocks enriched by “canned” picture material, sequences of morpho-syntactic tokens, or collections of syntactic structures) to semantic (viewing patents as multimedia knowledge objects) processing. PATExpert developed a multimedia content representation formalism based on Semantic Web technologies for selected technology areas and investigate the retrieval, classification, multilingual generation of concise patent information, assessment and visualization of patent material encoded in this formalism, taking the information needs of all user types as defined in a user typology into account. PATExpert’s technological goal was to develop a showcase that demonstrates the viability of PATExpert’s approach to content representation for real applications. The composition and the competence of the Consortium ensured the achievement of these goals.
Our pick of the week:
"Large Language Diffusion Model" by Shen Nie, Fengqi Zhu, @ZebinYou, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, @LiChongxuan
It is very cool to see how the researcher combine diffusion model and transformer blocks to train
Our pick of the week by
@mgaido91
: "FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model" by Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang, Xu Li, Yufan Lin, Junwen Qiu, Jun Zhang, Lu Lu, Haizhou Li and @drwuz
#SLM #EfficientInference
Cool to see a work that adaptively chooses at inference how much to compress the input speech sequence, to control inference costs and quality based on the input, without enforcing a global trade-off to each segment: https://arxiv.org/pdf/2606.31247
@fbk_mt
Our pick of the week by
@FBKZhihangXie : "Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models" by Haoqin Sun, @Chenyang_Lyu, Shiwan Zhao, Xuanfan Ni, Xiangyu Kong, @wangly0229, Weihua Luo and Yong Qin
#SpeechLLM #LongFormSpeech #SLU
🚀 New paper: Speech-XL for long-form SpeechLLMs
📄 https://arxiv.org/abs/2602.05373
🧩 Uses Speech Summarization Tokens to compress local speech intervals into compact KV states efficiently.
✨ Improves long-form speech understanding while reducing memory and FLOPs on 10-minute audio.