Ranked #1 on KorQuAD 2.0: A Generative Model for Long-Context Question Answering
Key point
LG AI Research achieved the #1 ranking on the Korean machine reading comprehension benchmark KorQuAD 2.0 using EXAONE LM 1.0.
Details
LG AI Research's EXAONE Lab achieved the #1 ranking on the leaderboard of KorQuAD 2.0, a Korean machine reading comprehension (MRC) evaluation dataset. This follows their #1 ranking on KorQuAD 1.0 last June.
KorQuAD 2.0 takes Wikipedia documents as input that are longer and more complex than those in the original 1.0, and include HTML tags. The model faces the challenging task of accurately generating a wide range of answer formats, from short answers to long-passage answers in Table and List forms.
To address this, LG AI Research developed the EXAONE LM 1.0 model. This model adopts an Encoder-Decoder architecture to simultaneously secure document understanding and generation capabilities, and is trained on expert data, delivering high performance even with limited computing resources.
To further enhance performance, Post-training techniques were applied. To prevent the loss of context from HTML structure, additional training was conducted using HTML documents and QA datasets with a Span Corruption Objective. In particular, the Long Span Denoising method proposed in HTML-T5 was adopted to maximize the model's ability to handle long answers.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.