Methodology of three step generation of verified UI-tests based on LLM
Аuthors
Moscow Aviation Institute (National Research University), 4, Volokolamskoe shosse, Moscow, А-80, GSP-3, 125993, Russia
e-mail: loksader@yandex.ru
Abstract
This work addresses the problem of automating the testing of user interfaces in modern web applications, where a gap persists between natural language requirements and automated checks, and "false-positive" tests create an illusion of quality. The issue is particularly acute in the aerospace industry when testing interfaces for telemetry monitoring and pre-flight planning, where visualization errors are critical and automation specialists are scarce. A three-stage methodology is proposed for the automatic generation of verified UI tests based on natural language specifications using the Llama 3.3 70B large language model and the Playwright framework. In the first stage, a structured UI specification accessible to non-technical experts is formed. In the second, Page Object classes encapsulating interaction logic are generated. In the third, test scenarios are created with multi-level verification, including checks for syntax, compilability, execution reliability, and coverage growth. A key element is mutation testing through DOM event interception, allowing defects to be non-invasively introduced into the interface and verifying the test's ability to detect them. As a result, multi-level filtering criteria have been developed to weed out unstable and false-positive tests, along with a mutation testing method that does not require modification of the application's source code.
Keywords:
test automation, UI tests; large language models; test generation; mutation testing; Page Object Model; PlaywrightReferences
- Maurizio Leotta, Diego Clerissi, Filippo Ricca, Cristiano Spadaro Improving Test Suites Maintainability with the Page Object Pattern: An Industrial Case Study// 2013 IEEE Sixth International Conference on Software Testing, Verification and Validation Workshops, 2013, pp. 108-113, DOI:10.1109/ICSTW.2013.19
- Junjie Wang, Yuchao Huang, Chunyang Chen, Zhe Liu, Song Wang, Qing Wang Software Testing with Large Language Models: Survey, Landscape, and Vision, 2023, DOI: 10.48550/arXiv.2307.07221
- Andrea Lops, Fedelucio Narducci, Azzurra Ragone, Michelantonio Trizio, Claudio Bartolini Lopes LLMs for Automated Unit Test Generation and Assessment in Java: The AgoneTest Framework, 2025, DOI: 10.48550/arXiv.2511.20403
- Saranya Alagarsamy, Chakkrit Tantithamthavorn, Aldeida Aleti A3Test: Assertion-Augmented Automated Test Case Generation, 2023, DOI: 10.48550/arXiv.2302.10352
- Lin Yang, Chen Yang, Shutao Gao, Weijing Wang, Bo Wang, Qihao Zhu, Xiao Chu, Jianyi Zhou, Guangtai Liang, Qianxiang Wang, Junjie Chen On the Evaluation of Large Language Models in Unit Test Generation, 2024, DOI: 10.48550/arXiv.2406.18181
- Max Schäfer, Sarah Nadi, Aryaz Eghbali, Frank Tip Adaptive Test Generation Using a Large Language Model, 2023, DOI: 10.48550/arXiv.2302.06527
- Yinghao Chen, Zehao Hu, Chen Zhi, Junxiao Han, Shuiguang Deng, Jianwei Yin ChatUniTest: A Framework for LLM-Based Test Generation, 2024, DOI: 10.48550/arXiv.2305.04764
- Eleanor Foley, Ashley Jacob, Ronit Kapoor Generating Test Cases Through Large Language Models, 2024, URL: https://digital.wpi.edu/downloads/pv63g461d
- Sutharsan Saarathy, Suresh Bathrachalam, Rajendran Bharath Self-Healing Test Automation Framework using AI and ML // International Journal of Strategic Management, 2024, vol 3, pp. 45-77, DOI: 10.47604/ijsm.2843
- Zemin Su, Cuiqin Bai, Shaowen Wei, Kangyong Liu, Yang Liu, Zhenxuan Huan Self-Healing UI Test Automation via Multi-Modal Fusion, 2025, URL: https://ieeexplore.ieee.org/abstract/document/11047603
- Zejun Wang, Kaibo Liu, Ge Li, Zhi Jin HITS: High-coverage LLM-based Unit Test Generation via Method Slicing, 2024, DOI: 10.48550/arXiv.2408.11324
- Mike Papadakis, Marinos Kintis, Jie Zhang, Yue Jia, Yves Le Traon, Mark Harman Mutation Testing Advances: An Analysis and Survey // Advances in Computers, vol 112, pp. 275-378, DOI: 10.1016/bs.adcom.2018.03.015
- René Just, Darioush Jalali, Laura Inozemtseva, Michael D. Ernst, Reid Holmes, Gordon Fraser Are mutants a valid substitute for real faults in software testing? // FSE 2014: Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering, 2014, pp. 654-665, URL: https://www.researchgate.net/publication/277664637_Are_Mutants_a_Valid_Substitute_for_Real_Faults_in...
- Ana B. Sánchez, Pedro Delgado-Pérez, Inmaculada, Medina-Bulo, Sergio Segura1 Mutation testing in the wild: findings from GitHub, 2022, URL: https://www.researchgate.net/publication/362078955_Mutation_testing_in_the_wild_findings_from_GitHub
Download

