
TOTAL VIEWS: 545
This study examines whether different large language models revise the same L2 learner errors in similar ways. The dataset included 100 error units from IELTS Academic Writing Task 1 responses written by Chinese EFL learners. Each error unit was revised by 3 mainstream Large language models (LLMs): GPT-4o-mini, DeepSeek-V3, and Gemini Flash to test how similar their outputs are. Two measuring methods were introduced to do the analysis: BERTScore F1 was used to measure similarity in meaning, and Jaccard Similarity was used to measure overlap in wording. The study also examined whether learner-language features were retained, reformulated, or removed after revision. The results show clear similarity across the three models, especially at the meaning level. Grammar-related and collocation errors produced the most similar revisions. The findings suggest that LLM-based revision can improve accuracy and fluency, but it can also overlook interlanguage features for the purpose of standardized writing. LLM feedback is therefore better used for comparison and reflection than as a final answer.
Large language models; L2 writing; written corrective feedback; homogenization; interlanguage; human-AI interaction
Agarwal, D., Naaman, M., & Vashistha, A. (2025). AI suggestions homogenize writing toward Western styles and diminish cultural nuances. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery.
https://doi.org/10.1145/3706598.3713564
Anderson, B. R., Shah, J. H., & Kreminski, M. (2024). Homogenization effects of large language models on human creative ideation. In Proceedings of the 16th Conference on Creativity & Cognition (pp. 413-425). Association for Computing Machinery.
https://doi.org/10.1145/3635636.3656204
Bitchener, J., & Knoch, U. (2010). The contribution of written corrective feedback to language development: A ten-month investigation. Applied Linguistics, 31(2), 193-214.
https://doi.org/10.1093/applin/amp016
Corder, S. P. (1967). The significance of learner’s errors. International Review of Applied Linguistics in Language Teaching, 5(1-4), 161-170.
https://doi.org/10.1515/iral.1967.5.1-4.161
Ellis, R. (2008). The study of second language acquisition (2nd ed.). Oxford University Press.
Ellis, R. (2009). A typology of written corrective feedback types. ELT Journal, 63(2), 97-107.
https://doi.org/10.1093/elt/ccn023
Ferris, D. R. (2011). Treatment of error in second language student writing (2nd ed.). University of Michigan Press.
Hyland, K., & Hyland, F. (2006). Feedback on second language students’ writing. Language Teaching, 39(2), 83-101.
https://doi.org/10.1017/S0261444806003399
IELTS. (2023, May 3). IELTS writing band descriptors and key assessment criteria.
https://ielts.org/news-and-insights/ielts-writing-band-descriptors-and-key-assessment-criteria
Jaccard, P. (1901). Distribution de la flore alpine dans le Bassin des Dranses et dans quelques régions voisines. Bulletin de la Société Vaudoise des Sciences Naturelles, 37, 241-272.
Moon, K., Green, A. E., & Kushlev, K. (2025). Homogenizing effect of large language models (LLMs) on creative diver-sity: An empirical comparison of human and ChatGPT writing. Computers in Human Behavior: Artificial Humans, 6, 100207.
https://doi.org/10.1016/j.chbah.2025.100207
Narreddy, C., Joordens, S., & Prompiengchai, S. (2025). Harnessing large language models for scalable and effective formative assessment in higher education: A review. Trends in Higher Education, 4(4), Article 65.
https://doi.org/10.3390/higheredu4040065
Niwattanakul, S., Singthongchai, J., Naenudorn, E., & Wanapu, S. (2013). Using of Jaccard coefficient for keywords similarity. Proceedings of the International MultiConference of Engineers and Computer Scientists, 1, 380-384.
Schmidt, R. (1990). The role of consciousness in second language learning. Applied Linguistics, 11(2), 129-158.
https://doi.org/10.1093/applin/11.2.129
Seddiki, M., & Korichi, S. (2026). The AI paradox in L2 writing: Why helpful feedback creates unhelpful dependency in higher education. Research Square.
https://doi.org/10.21203/rs.3.rs-8731897/v2
Selinker, L. (1972). Interlanguage. International Review of Applied Linguistics in Language Teaching, 10(1-4), 209-231.
https://doi.org/10.1515/iral.1972.10.1-4.209
Sourati, Z., Ziabari, A. S., & Dehghani, M. (2026). The homogenizing effect of large language models on human expression and thought. Trends in Cognitive Sciences. Advance online publication.
https://doi.org/10.1016/j.tics.2026.01.003
Sung, H., Csuros, K., & Sung, M.-C. (2025). Comparing human and LLM proofreading in L2 writing: Impact on lexical and syntactic features. In Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025) (pp. 11-23). Association for Computational Linguistics.
https://doi.org/10.18653/v1/2025.bea-1.2
Wang, J. (2025). EDCEW-LLM: Error detection and correction in English writing: A large language model-based ap-proach. Alexandria Engineering Journal, 129, 1153-1164.
https://doi.org/10.1016/j.aej.2025.08.005
Weidlich, J., Gotsch, F., Schudel, K., Marusic-Würscher, C., Mazzarella, J., Bolten, H., Bütler, D., Luger, S., Wohlfender, B., & Maag Merki, K. (2025). Teacher, peer, or AI? Comparing effects of feedback sources in higher education. Computers and Education Open, 9, Article 100300.
https://doi.org/10.1016/j.caeo.2025.100300
Wenger, E., & Kenett, Y. N. (2026). Large language models are homogeneously creative. PNAS Nexus, 5(3), pgag042.
https://doi.org/10.1093/pnasnexus/pgag042
Yan, D., & Zhang, S. (2024). L2 writer engagement with automated written corrective feedback provided by ChatGPT: A mixed-method multiple case study. Humanities and Social Sciences Communications, 11, Article 1143.
https://doi.org/10.1057/s41599-024-03543-y
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In the International Conference on Learning Representations.
https://openreview.net/forum?id=SkeHuCVFDr
Human-AI Interaction in L2 Writing Support: Homogenization in LLM Revisions of Learner Errors
How to cite this paper: Wei Ren. (2026) Human-AI Interaction in L2 Writing Support: Homogenization in LLM Revisions of Learner Errors. Journal of Humanities, Arts and Social Science, 10(7), 747-756.
DOI: http://dx.doi.org/10.26855/jhass.2026.07.001