Back to Search
Start Over
Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records
- Source :
- BMC Medical Informatics and Decision Making, BMC Medical Informatics and Decision Making, Vol 17, Iss 1, Pp 1-11 (2017)
- Publication Year :
- 2017
- Publisher :
- Springer Science and Business Media LLC, 2017.
-
Abstract
- Cardiovascular disease (CVD) has become the leading cause of death in China, and most of the cases can be prevented by controlling risk factors. The goal of this study was to build a corpus of CVD risk factor annotations based on Chinese electronic medical records (CEMRs). This corpus is intended to be used to develop a risk factor information extraction system that, in turn, can be applied as a foundation for the further study of the progress of risk factors and CVD. We designed a light annotation task to capture CVD risk factors with indicators, temporal attributes and assertions that were explicitly or implicitly displayed in the records. The task included: 1) preparing data; 2) creating guidelines for capturing annotations (these were created with the help of clinicians); 3) proposing an annotation method including building the guidelines draft, training the annotators and updating the guidelines, and corpus construction. Then, a risk factor annotated corpus based on de-identified discharge summaries and progress notes from 600 patients was developed. Built with the help of clinicians, this corpus has an inter-annotator agreement (IAA) F1-measure of 0.968, indicating a high reliability. To the best of our knowledge, this is the first annotated corpus concerning CVD risk factors in CEMRs and the guidelines for capturing CVD risk factor annotations from CEMRs were proposed. The obtained document-level annotations can be applied in future studies to monitor risk factors and CVD over the long term.<br />Comment: 32 pages, 3 figures, 3 tables
- Subjects :
- FOS: Computer and information sciences
0301 basic medicine
China
Information extraction
Annotation
Chinese electronic medical records
Information Storage and Retrieval
Health Informatics
Disease
Cardiovascular disease risk factors
lcsh:Computer applications to medicine. Medical informatics
computer.software_genre
Health informatics
Task (project management)
03 medical and health sciences
0302 clinical medicine
Risk Factors
Disease risk factor
Electronic Health Records
Humans
Medicine
030212 general & internal medicine
Natural Language Processing
Corpus construction
Computer Science - Computation and Language
business.industry
I.2.7
Health Policy
Medical record
Risk factor (computing)
Data science
Computer Science Applications
030104 developmental biology
Cardiovascular Diseases
lcsh:R858-859.7
Artificial intelligence
business
Computation and Language (cs.CL)
computer
Natural language processing
Research Article
Subjects
Details
- ISSN :
- 14726947
- Volume :
- 17
- Database :
- OpenAIRE
- Journal :
- BMC Medical Informatics and Decision Making
- Accession number :
- edsair.doi.dedup.....f566724e7a9dc363a8865431ba8f3256
- Full Text :
- https://doi.org/10.1186/s12911-017-0512-7