실감미디어의 범위 · 연구 포털

LLM 배심원 검증 — 방법 근거 참고문헌

모든 문헌은 2026-09-22 OpenAlex에서 제목·저자·DOI를 대조하였다.

LLM 배심원(다수 모형 판정)과 자기 모형 선호

  1. Verga, P., Hofstätter, S., Althammer, S., et al. (2024). Replacing judges with juries: Evaluating LLM generations with a panel of diverse models. arXiv, 2404.18796. https://doi.org/10.48550/arXiv.2404.18796
    • 서로 다른 계열의 여러 모형으로 구성한 배심원단(PoLL)이 단일 대형 판정 모형보다 사람 판정과 더 잘 일치하고 모형 내 편향이 작다는 근거. 본 연구가 다섯 개발사의 모형으로 배심원을 구성한 이유.
  2. Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems 36 (Datasets and Benchmarks Track). https://doi.org/10.52202/075280-2020
    • LLM 판정자와 사람 판정의 일치도, 그리고 위치·장황함·자기 선호 편향을 체계적으로 보고한 연구.
  3. Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems 37. https://doi.org/10.52202/079017-2197
    • 판정 모형이 자기 계열의 산출물을 선호한다는 근거. 선별 모형(OpenAI gpt-5-mini)과 같은 개발사의 모형을 배심원에서 뺀 이유.

LLM의 주석·문헌 선별 성능

  1. Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), e2305016120. https://doi.org/10.1073/pnas.2305016120
  2. Khraisha, Q., Van Put, S., Kappenberg, J., Warraitch, A., & Hadfield, K. (2024). Can large language models replace humans in systematic reviews? Evaluating GPT-4’s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages. Research Synthesis Methods, 15(4), 616–626. https://doi.org/10.1002/jrsm.1715
  3. Guo, E., Gupta, M., Deng, J., Park, Y.-J., Paget, M., & Naugler, C. (2024). Automated paper screening for clinical reviews using large language models: Data analysis study. Journal of Medical Internet Research, 26, e48996. https://doi.org/10.2196/48996

일치도 통계

  1. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104
  2. Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378–382. https://doi.org/10.1037/h0031619
  3. Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310

보고 지침

  1. Tricco, A. C., Lillie, E., Zarin, W., et al. (2018). PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and explanation. Annals of Internal Medicine, 169(7), 467–473. https://doi.org/10.7326/M18-0850

한계로 적을 점