REAL

Evaluating Latent Semantic Pre-training for Fine-grained Word Sense Disambiguation

Simon, Soma and Berend, Gábor (2026) Evaluating Latent Semantic Pre-training for Fine-grained Word Sense Disambiguation. In: International Conference on Natural Language and Speech Processing, 25-26 Sept 2026. (In Press)

This is the latest version of this item.

[img]
Preview
Text
icnlsp26.pdf

Download (322kB) | Preview

Abstract

Understanding word meaning in context is a central problem in Natural Language Processing, most directly addressed by word sense disambiguation (WSD). While modern contextual language models have improved WSD performance, it remains unclear how the choice of pre-training objective influences the semantic representations used in different WSD approaches. In this paper, we investigate masked latent semantic modeling (MLSM) as a semantic pre-training objective for WSD, focusing on its behavior across diverse WSD approaches. We contrast the use of MLSM pre-trained models to that of traditionally pre-trained ones for a wide range of WSD modeling scenarios, including centroid and sparse representation-based probing methods, as well as supervised disambiguation systems. Experimental results on standard WSD benchmarks show that relying on MLSM yields consistent improvements in disambiguation performance across models and datasets.

Item Type: Conference or Workshop Item (Paper)
Subjects: Q Science / természettudomány > QA Mathematics / matematika > QA75 Electronic computers. Computer science / számítástechnika, számítógéptudomány
Depositing User: Gábor Berend
Date Deposited: 28 Sep 2026 06:44
Last Modified: 28 Sep 2026 06:44
URI: https://real.mtak.hu/id/eprint/247853

Available Versions of this Item

  • Evaluating Latent Semantic Pre-training for Fine-grained Word Sense Disambiguation. (deposited 28 Sep 2026 06:44) [Currently Displayed]

Actions (login required)

View Item View Item