Latent Dirichlet Allocation (LDA) has been used to support many software engineering tasks. Previous studies showed that default settings lead to sub-optimal topic modeling with a dramatic impact on the performance of such approaches in terms of precision and recall. For this reason, researchers used search algorithms (e.g., genetic algorithms) to automatically configure topic models in an unsupervised fashion. While previous work showed the ability of individual search algorithms in finding near-optimal configurations, it is not clear to what extent the choice of the meta-heuristic matters for SE tasks. In this paper, we present a systematic comparison of ve different meta-heuristics to configure LDA in the context of duplicate bug reports identification. The results show that (1) no master algorithm outperforms the others for all software projects, (2) random search and PSO are the least effective meta-heuristics. Finally, the running time (s) strongly depends on the computational complexity of LDA while the internal complexity of the search algorithms plays a negligible role.
Original languageEnglish
Title of host publication11th Symposium on Search-Based Software Engineering
Number of pages16
Publication statusPublished - Aug 2019

    Research areas

  • Topic modeling, Latent Dirichlet Allocation, Search-based Software Engineering, Evolutionary algorithms, Duplicate Bug Reports

ID: 54192548