← All publications

ToMMeR - Efficient Entity Mention Detection from Large Language Models

Victor Morand, Nadi Tomeh, Josiane Mothe, Benjamin Piwowarski
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)2026

This work provides evidence that structured entity representations exist in early transformer layers and can be efficiently recovered with minimal parameters, and reveals that diverse architectures converge on similar mention boundaries, confirming that mention detection emerges naturally from language modeling

Abstract

Identifying which text spans refer to entities - mention detection- is both foundational for information extraction and a known performance bottleneck. We introduce ToMMeR, a lightweight model (\ensuremath75%), confirming that mention detection emerges naturally from language modeling. When extended with span classification heads, ToMMeR achieves competitive NER performance (80-87% F1 on standard benchmarks). Our work provides evidence that structured entity representations exist in early transformer layers and can be efficiently recovered with minimal parameters.

Other versions