사업성과 BK21 FOUR 산업혁신 애널리틱스 교육연구단

논문

2026 UFORank:Unified Framework of Unsupervised Keyphrase Extraction for Long Documents

페이지 정보

작성자 관리자 작성일 26-07-15 10:16

본문

Author
Doyoon Kim, Pilsung Kang
Journal
IEEE Access
Vol
14
Page
9986-10001
Year
2026

Abstract

Keyphrase extraction involves automatically identifying key phrases that represent the core content within a document. Recent advancements have improved unsupervised methods for keyphrase extraction, enabling efficient operation without requiring labeled training data. As document length increases, effective keyphrase extraction becomes increasingly crucial for information summarization and retrieval. However, existing methods often struggle with long documents due to challenges in capturing global context and semantic relationships across extended text. In this paper, we propose UFORank, a unified framework for unsupervised keyphrase extraction specifically designed for long documents. UFORank integrates three key components: 1) topic importance derived from clustering semantically similar phrases, 2) position-biased weights that consider both relative positions and frequencies of phrases within the document structure, and 3) phrase-to-topic similarity measures for enhanced relevance scoring. Additionally, UFORank employs Glow, a flow-based generative model, to improve the semantic representation quality of both phrases and documents in the embedding space. Experimental evaluations on three benchmark datasets for long document keyphrase extraction demonstrate that UFORank achieves competitive performance compared to existing state-of-the-art methods, including PromptRank and Attention-Seeker. Specifically, UFORank achieves F1-scores of 17.66%, 20.73%, and 13.17% on the respective datasets. Comprehensive ablation studies validate the individual contributions of each framework component to the overall performance gains.