실감미디어의 범위 · 연구 포털
LLM 배심원 검증 — 방법 근거 참고문헌
모든 문헌은 2026-09-22 OpenAlex에서 제목·저자·DOI를 대조하였다.
LLM 배심원(다수 모형 판정)과 자기 모형 선호
- Verga, P., Hofstätter, S., Althammer, S., et al. (2024). Replacing
judges with juries: Evaluating LLM generations with a panel of diverse
models. arXiv, 2404.18796.
https://doi.org/10.48550/arXiv.2404.18796
- 서로 다른 계열의 여러 모형으로 구성한 배심원단(PoLL)이 단일 대형 판정 모형보다 사람 판정과 더 잘 일치하고 모형 내 편향이 작다는 근거. 본 연구가 다섯 개발사의 모형으로 배심원을 구성한 이유.
- Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). Judging
LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural
Information Processing Systems 36 (Datasets and Benchmarks Track).
https://doi.org/10.52202/075280-2020
- LLM 판정자와 사람 판정의 일치도, 그리고 위치·장황함·자기 선호 편향을 체계적으로 보고한 연구.
- Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM
evaluators recognize and favor their own generations. Advances in
Neural Information Processing Systems 37.
https://doi.org/10.52202/079017-2197
- 판정 모형이 자기 계열의 산출물을 선호한다는 근거. 선별 모형(OpenAI gpt-5-mini)과 같은 개발사의 모형을 배심원에서 뺀 이유.
LLM의 주석·문헌 선별 성능
- Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), e2305016120. https://doi.org/10.1073/pnas.2305016120
- Khraisha, Q., Van Put, S., Kappenberg, J., Warraitch, A., & Hadfield, K. (2024). Can large language models replace humans in systematic reviews? Evaluating GPT-4’s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages. Research Synthesis Methods, 15(4), 616–626. https://doi.org/10.1002/jrsm.1715
- Guo, E., Gupta, M., Deng, J., Park, Y.-J., Paget, M., & Naugler, C. (2024). Automated paper screening for clinical reviews using large language models: Data analysis study. Journal of Medical Internet Research, 26, e48996. https://doi.org/10.2196/48996
일치도 통계
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104
- Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378–382. https://doi.org/10.1037/h0031619
- Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310
보고 지침
- Tricco, A. C., Lillie, E., Zarin, W., et al. (2018). PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and explanation. Annals of Internal Medicine, 169(7), 467–473. https://doi.org/10.7326/M18-0850
한계로 적을 점
- 사람 전문가 판정을 기준으로 삼지 않았다. 배심원 다수결은 ’참값’이 아니라 독립적인 모형 합의이며, 모든 모형이 같은 방향으로 틀리는 경우(공통 편향)는 잡아내지 못한다.
- 이를 줄이기 위해 개발사·학습 계열이 서로 다른 모형 다섯 종을 쓰고, 선별 모형과 같은 개발사의 모형은 제외하였다. 또한 배심원 간 일치도(Fleiss κ)와 만장일치 비율을 함께 보고한다.