Back to Search Start Over

A Case-Based Recommender System for Persian Scientific Document Indexing

Authors :
Azadeh Mohebi
Azadeh Fakhrzdaeh
Marzieh Zarinbal
Source :
Iranian Journal of Information Processing & Management, Vol 39, Iss 2, Pp 599-626 (2023)
Publication Year :
2023
Publisher :
Iranian Research Institute for Information and Technology, 2023.

Abstract

Keyword extraction is a key step in document indexing. Keywords are semantic and content-based descriptors of a document, which can be used in document retrieval and representation. In databases containing scientific documents, such as Ganj in Irannian Research Institue for Information Science and Technology (IranDoc), it is even more critical to assign meaningful keywords for documents, since the documents are from different academic disciplines and contain technical terms.As the number of scientific documents grows exponentially, having an automatic and intelligent keyword extraction technique is getting more critical. There are various keyword extraction techniques that are either based on statistical features of the text or machine learning approaches, and sometimes a combination of both. In this research, we propose a new keyword extraction method for Persian scientific documents based on recommender systems and case-based reasoning. The proposed method is designed based on case-based reasoning in which the main assumption is that similar documents share similar keywords. There are two main steps in the proposed approach: first, similar documents to a given new document are retrieved based on TFIDF and word2vec model, second, the candidate keywords are extracted from retrieved documents and ranked based on a new scoring scheme, and a set of keyword are selected from the candidate keywords based on their score. The proposed method is tested and avaluated on a set of documents of Ganj database in three different subject areas (Art, Humanities and Engineering), based on precision, recall and expert panel

Details

Language :
Persian
ISSN :
22518223 and 22518231
Volume :
39
Issue :
2
Database :
Directory of Open Access Journals
Journal :
Iranian Journal of Information Processing & Management
Publication Type :
Academic Journal
Accession number :
edsdoj.bf887682e8704ff3a2bc6e311f8b13da
Document Type :
article
Full Text :
https://doi.org/10.22034/jipm.2023.704737