Szabó, Márk and Kovács, Ádám and Kusper, Gábor (2026) Data cleaning and spam filtering methods in the TAWOS database. ANNALES MATHEMATICAE ET INFORMATICAE, 63. pp. 132-148. ISSN 1787-6117
|
Text
132_148.pdf - Published Version Download (700kB) | Preview |
Abstract
The TAWOS dataset is a large relational database of issue-tracking data mined from open-source projects. While it is valuable for empirical software engineering, the issue texts also contain spam, placeholders, and off-topic noise that can distort downstream analytics and fine-tuning tasks. We present a four-stage filtering pipeline for cleaning TAWOS and focus on a validity scoring step that assigns each issue a 0–100 score from its title and description. We compare a deterministic rule-based classifier, OwnMetrics, against four local small language models (Llama-3.1-8B, Mistral-7B, Phi-3.5-mini, and Gemma-3-4B) prompted for JSON scores. Evaluation first uses a near-balanced labeled benchmark of 947 GitHub issues (481 spam/noise and 466 legitimate issues), collected from moderator-locked spam issues and legitimate issues from popular repositories. On this GitHub-based threshold-selection benchmark, OwnMetrics obtains the highest accuracy, 91.0%, at threshold 75, while the best LLM configurations reach 89.2% (Gemma-3-4B) and 89.1% (Mistral-7B). A separate manually labeled in-domain validation on 300 Jira/TAWOS issues (39 spam/noise and 261 non-spam) provides an in-domain sanity check consistent with the benchmark results, with OwnMetrics giving the strongest spam-class F1 among the selected configurations. On a 10,000-item sample, OwnMetrics yields zero parsing failures and an average execution time of 2.28 ms per item, whereas the LLMs require seconds per item and Meta-Llama-3.1-8B exhibits failure rates above 54%. Applying the full filtering pipeline to the processed TAWOS issue corpus flags 16.2% of records for filtering. The results show that a lightweight domain-tailored scorer can be both accurate and robust for large-scale issue cleaning.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | software repositories, issue tracking, data cleaning, spam filtering, large language models |
| Subjects: | Q Science / természettudomány > QA Mathematics / matematika > QA76 Computer software / programozás |
| Depositing User: | Tibor Gál |
| Date Deposited: | 22 Jul 2026 07:28 |
| Last Modified: | 22 Jul 2026 07:28 |
| URI: | https://real.mtak.hu/id/eprint/242826 |
Actions (login required)
![]() |
Edit Item |




