MEANING will be concerned with automatically collecting and analysing language data from the WWW on a large scale, and building more comprehensive multilingual lexical knowledge bases to support improved word sense disambiguation (WSD). Current web access applications are based on words; MEANING will open the way for access to the Multilingual Web based on concepts, providing applications with capabilities that significantly exceed those currently available. MEANING will facilitate development of concept-based open domain Internet applications (such as Question/Answering, Cross Lingual Information Retrieval, Summarisation, Text Categorisation, Event Tracking, Information Extraction, Machine Translation, etc.). Furthermore, MEANING will supply a common conceptual structure to Internet documents, thus facilitating knowledge management of web content.
Can algorithmic gender prediction ever be valid?
Check out this week's top pick by @lina_conti: "Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements" by @evandongyx & @ang3linawang.
Pick of the week by @evandongyx & @ang3linawang:
https://arxiv.org/pdf/2608.13444
Predicting gender from images or names can reveal discrimination. But the practice itself harms trans people. This paper works through when that tradeoff might be justified and how to do it responsibly.
Our pick of the week by
@dhairya_su47605
: "Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression" by @hanqi_xiao, @yilin_sung, @EliasEskin and @mohitban47
#Quantization #Interpretibility
#PickoftheWeek @fbk_mt
Super cool paper on leavaraging Interpretability for Compression!
https://arxiv.org/pdf/2504.07389
Our pick of the week:
"Large Language Diffusion Model" by Shen Nie, Fengqi Zhu, @ZebinYou, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, @LiChongxuan
It is very cool to see how the researcher combine diffusion model and transformer blocks to train
Our pick of the week by
@mgaido91
: "FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model" by Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang, Xu Li, Yufan Lin, Junwen Qiu, Jun Zhang, Lu Lu, Haizhou Li and @drwuz
#SLM #EfficientInference
Cool to see a work that adaptively chooses at inference how much to compress the input speech sequence, to control inference costs and quality based on the input, without enforcing a global trade-off to each segment: https://arxiv.org/pdf/2606.31247
@fbk_mt