All Issue

2026 Vol.14, Issue 3

Research Article

30 September 2026. pp. 1-15
Abstract
본 논문은 뉴스 기사의 구조적 특성을 반영하여 생성 요약 성능을 향상하는 방법을 제안한다. CNN/DailyMail 데이터셋에는 문장 단위 5W1H 구조 라벨이 포함되어 있지 않으므로, 본 연구에서는 GPT-5 mini 기반 5W1H 프롬프트를 활용하여 뉴스 문장 라벨을 자동 생성하였다. 생성된 라벨은 사람이 직접 구축한 정답 라벨이 아니라 약지도 기반 의사 라벨로 활용하였으며, 이를 바탕으로 BERT- BiGRU-CRF 시퀀스 라벨링 모델을 학습하였다. 학습된 모델은 뉴스 문장을 LEAD, EVENT, REACTION, OUTLOOK, CONSEQUENCE, QUOTE, NOISE의 7개 라벨로 분류하고, 핵심 라벨 문장을 선별하여 BART, BERTSUMABS, T5 생성 요약 모델의 입력으로 사용하였다. 실험 결과, 시퀀스 라벨링 모델은 GPT-5 mini 기반 의사 라벨에 대해 정확도 73.6%, F1 스코어 72.2%를 기록하였다. 생성 요약에서는 LEQ와 LEQC 방식이 전반적으로 기본 모델보다 우수한 성능을 보였으며, 특히 LEQC+BART는 ROUGE-1 47.0%, ROUGE-2 25.6%, ROUGE-L 36.8%로 가장 높은 성능을 나타냈다. 이는 뉴스의 핵심 구조를 반영한 문장 선별이 생성 요약 성능 향상에 효과적임을 보여준다.
This paper proposes a method to improve abstractive summarization performance by incorporating the structural characteristics of news articles. Since the CNN/DailyMail dataset does not provide sentence-level 5W1H structural labels, this study automatically generated news sentence labels by applying 5W1H-based prompting to GPT-5 mini. The generated labels were not treated as manually annotated gold labels, but were used as weakly supervised pseudo-labels to train a BERT-BiGRU-CRF sequence labeling model. The trained model classifies news sentences into seven categories: LEAD, EVENT, REACTION, OUTLOOK, CONSEQUENCE, QUOTE, and NOISE. Sentences-assigned core labels were then selected and used as inputs for abstractive summarization models, including BART, BERTSUMABS, and T5. Experimental results show that the sequence labeling model achieved an accuracy of 73.6% and an F1-score of 72.2% on the GPT-5 mini based pseudo-labels. In the summarization experiments, the LEQ and LEQC methods generally outperformed the baseline models. In particular, LEQC+BART achieved the highest performance, with ROUGE-1, ROUGE-2, and ROUGE-L scores of 47.0%, 25.6%, and 36.8%, respectively. These results indicate that selecting structurally important sentences is effective in improving abstractive summarization performance.
References
  1. A. Nenkova, and K. McKeown, “A Survey of Text Summarization Techniques,” Mining Text Data, pp. 43-76, 2012.

    10.1007/978-1-4614-3223-4_3
  2. M. Lewis, et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7871-7880, 2020.

    10.18653/v1/2020.acl-main.703
  3. C. Raffel, et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” Journal of Machine Learning Research, Vol. 21, No. 140, pp. 1-67, 2020.

  4. K. M. Hermann, et al., “Teaching Machines to Read and Comprehend,” Advances in Neural Information Processing Systems (NIPS), 2015.

  5. W. X. Zhao, et al., “A Survey of Large Language Models,” arXiv preprint arXiv:2303.18223, 2023.

  6. A. Ratner, et al., “Snorkel: Rapid Training Data Creation with Weak Supervision,” Proceedings of the VLDB Endowment, Vol. 11, No. 3, pp. 269-282, 2017.

    10.14778/3157794.3157797 29770249 PMC5951191
  7. Z. Tan, et al., “Large Language Models for Data Annotation and Synthesis: A Survey,” arXiv preprint arXiv:2402.13446, 2024.

  8. F. Gilardi, M. Alizadeh, and M. Kubli, “ChatGPT outperforms crowd-workers for text-annotation tasks,” Proceedings of the National Academy of Sciences, Vol. 120, No. 30, e2305016120, 2023.

    10.1073/pnas.2305016120 37463210 PMC10372638
  9. Y. Liu, and M. Lapata, “Text Summarization with Pretrained Encoders,” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pp. 3730-3740, 2019.

    10.18653/v1/D19-1387
  10. Y. Liu, et al., “Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models’ Alignment,” arXiv:2308.05374, 2023.

  11. X. Chen, and Z. Qiu, “Research on Core Function of Adjacency Pairs Prediction Based on BERT-BiGRU-CRF,” Proceedings of the 2021 2nd International Conference on Control, Robotics and Intelligent System (CCRIS '21), pp. 117-121, 2021.

    10.1145/3483845.3483866
  12. J. Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” NAACL, pp. 4171-4186, 2019.

    10.18653/v1/N19-1423
  13. J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling,” arXiv preprint arXiv:1412.3555, 2014.

  14. J. Lafferty, A. McCallum, and F.C.N. Pereira, “Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data,” Proceedings of the 18th International Conference on Machine Learning (ICML), pp. 282-289, 2001.

  15. C.Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries,” Text Summarization Branches Out, 2004.

Information
  • Publisher :The Society of Convergence Knowledge
  • Publisher(Ko) :융복합지식학회
  • Journal Title :The Society of Convergence Knowledge Transactions
  • Journal Title(Ko) :융복합지식학회논문지
  • Volume : 14
  • No :3
  • Pages :1-15