Similar Items: Machine learning, deep learning, or large language models: An empirical study on multi-label requirements classification
- Empirical benchmarking of large language models for data science coding: a multidimensional evaluation
- Reducing labeling effort in architecture technical debt detection through active learning and explainable AI
- Large language models in model-driven engineering: a systematic mapping study
- Smelly-shot is all you need: an empirical study to compare in-context learning verses fine tuning for code smell detection
- A multi-language perspective on the robustness of LLM code generation
- An empirical evaluation of white-box and black-box test case prioritization techniques in CPSs modeled in Simulink