Similar Items: Multi-scenario benchmark for autonomous driving systems: Exposing diverse behavioral anomalies
- Does road diversity really matter in testing automated driving systems?
- CDBench: Benchmarking the mutation testing capabilities of LLMs with code defenders
- Empirical benchmarking of large language models for data science coding: a multidimensional evaluation
- Diversity’s role in collaboration and conflict: a case-study of React.js
- A multi-language perspective on the robustness of LLM code generation
- Machine learning, deep learning, or large language models: An empirical study on multi-label requirements classification