A Machine Learning-Based Approach to Enhancing Label Aggregation in Crowdsourcing
Abstract Crowdsourcing is the most effective means of obtaining labelled data for supervised machine learning. However, the varying expertise of crowd workers often results in noisy annotations. While traditional label aggregation methods attempt to handle label noise, they typically overlook the relationships between different data instances. Moreover, crowdsourced datasets often experience class imbalance, wherein predominant classes eclipse minority classes, hence exacerbating label accuracy issues. This paper proposed a Reliability-Weighted Bayesian Label Aggregation (RWBLA) to overcome the above challenges. First, the K-nearest neighbours (KNN) method is used to improve the label set for each instance by augmenting labels from its closest neighbours, resulting in multiple noisy label sets. Next, it improves the aggregation by assigning weights to the neighbouring labels based on worker reliability, adaptive distance, and label similarity for each instance. In addition, worker reliability is determined dynamically based on the neighbourhood information, and to handle the imbalance issue, label similarity is modified. In the end, a weighted Bayesian inference method is used to infer the correct label for each instance. The performance of the proposed approach is evaluated on 20 synthetics and three real-world crowdsourcing datasets. It shows that RWBLA consistently surpasses eight baseline label aggregations, improving aggregation accuracy by 3 % to 13 %. Moreover, in analyses utilizing imbalanced real-world crowdsourcing data, RWBLA surpassed state-of-art aggregation algorithms by 2 % to 6 %, underscoring its efficacy in situations with minor class imbalances.
- Referencias
- Cómo citar
- Del mismo autor
- Métricas
Bi, W., Wang, L., Kwok, J. T., & Tu, Z. (2014). Learning to Predict from Crowdsourced Data. UAI, 14, 82–91.
Chen, Z., Jiang, L., & Li, C. (2022). Label augmented and weighted majority voting for crowdsourcing. Information Sciences, 606, 397–409.
Daniel, F., Kucherbaev, P., Cappiello, C., Benatallah, B., & Allahbakhsh, M. (2018). Quality Control in Crowdsourcing. ACM Computing Surveys, 51(1), 1–40. https://doi.org/10.1145/3148148
Dawid, A. P., & Skene, A. M. (1979). Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1), 20–28.
Demartini, G., Difallah, D. E., & Cudré-Mauroux, P. (2012). ZenCrowd:leveraging probabilistic reasoning and crowdsourcing techniques for large-scale entity linking. Proceedings of the 21st International Conference on World Wide Web, 469–478. https://doi.org/10.1145/2187836.2187900
Jiang, L., Zhang, H., Tao, F., & Li, C. (2022). Learning From Crowds With Multiple Noisy Label Distribution Propagation. IEEE Transactions on Neural Networks and Learning Systems, 33(11), 6558–6568. https://doi.org/10.1109/TNNLS.2021.3082496
Karger, D., Oh, S., & Shah, D. (2011). Iterative Learning for Reliable Crowdsourcing Systems. In J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, & K. Q. Weinberger (Eds.), Advances in Neural Information Processing Systems (Vol. 24). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2011/file/c667d53acd899a97a85de0c201ba99be-Paper.pdf
Kim, H. C., & Ghahramani, Z. (2012). Bayesian classifier combination. Artificial Intelligence and Statistics, 619–627.
Kurve, A., Miller, D. J., & Kesidis, G. (2014). Multicategory crowdsourcing accounting for variable task difficulty, worker skill, and worker intention. IEEE Transactions on Knowledge and Data Engineering, 27(3), 794–809.
Li, J., Jiang, L., & Zhang, W. (2024). Label Consistency-Based Ground Truth Inference for Crowdsourcing. IEEE Transactions on Neural Networks and Learning Systems.
Li, Y., Rubinstein, B., & Cohn, T. (2019). Exploiting worker correlation for label aggregation in crowdsourcing. International Conference on Machine Learning, 3886–3895.
Raykar, V. C., Yu, S., Zhao, L. H., Valadez, G. H., Florin, C., Bogoni, L., & Moy, L. (2010). Learning from crowds. Journal of Machine Learning Research, 11(4).
Read, J., Reutemann, P., Pfahringer, B., & Holmes, G. (2016). Meka: A multi-label/multi-target extension to Weka. Journal of Machine Learning Research, 17, 1–5. https://www.jmlr.org/papers/volume17/12-164/12-164.pdf
Ren, L., Jiang, L., & Li, C. (2023). Label confidence-based noise correction for crowdsourcing. Engineering Applications of Artificial Intelligence, 117, 105624.
Ren, L., Jiang, L., Zhang, W., & Li, C. (2024). Label distribution similarity-based noise correction for crowdsourcing. Frontiers of Computer Science, 18(5), 185323.
Sheng, V. S., Zhang, J., Gu, B., & Wu, X. (2017). Majority voting and pairing with multiple noisy labeling. IEEE Transactions on Knowledge and Data Engineering, 31(7), 1355–1368.
Snow, R., O’Connor, B., Jurafsky, D., & Ng, A. (2008). Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, 254–263. https://aclanthology.org/D08-1027
Suyal, H., & Singh, A. (2024). Multilabel classification using crowdsourcing under budget constraints. Knowledge and Information Systems, 66(2), 841–877.
Suyal, H., & Singh, A. (2025). A Systematic Literature Survey of Crowdsourcing: Current Status and Future Perspectives. WIREs Data Mining and Knowledge Discovery, 15(3). https://doi.org/10.1002/widm.70037
Tao, F., Jiang, L., & Li, C. (2020). Label similarity-based weighted soft majority voting and pairing for crowdsourcing. Knowledge and Information Systems, 62, 2521–2538.
Tian, T., & Zhu, J. (2015). Max-margin majority voting for learning from crowds. Advances in Neural Information Processing Systems, 28.
Venanzi, M., Guiver, J., Kazai, G., Kohli, P., & Shokouhi, M. (2014). Community-based bayesian aggregation models for crowdsourcing. Proceedings of the 23rd International Conference on World Wide Web, 155–164.
Whitehill, J., Wu, T., Bergsma, J., Movellan, J., & Ruvolo, P. (2009). Whose vote should count more: Optimal integration of labels from labelers of unknown expertise. Advances in Neural Information Processing Systems, 22.
Wilson, D. R., & Martinez, T. R. (1997). Improved Heterogeneous Distance Functions. Journal of Artificial Intelligence Research, 6, 1–34. https://doi.org/10.1613/jair.346
Yang, Y., Zhao, Z. Q., Wu, G., Zhuo, X., Liu, Q., Bai, Q., & Li, W. (2024). A Lightweight, Effective, and Efficient Model for Label Aggregation in Crowdsourcing. ACM Transactions on Knowledge Discovery from Data, 18(4), 1–27.
Zhang, J., Jiang, X., Tian, N., & Wu, M. (2024). Label noise correction for crowdsourcing using dynamic resampling. Engineering Applications of Artificial Intelligence, 133, 108439.
Zhang, J., Sheng, V. S., Li, Q., Wu, J., & Wu, X. (2017). Consensus algorithms for biased labeling in crowdsourcing. Information Sciences, 382, 254–273.
Zhang, J., Sheng, V. S., Li, T., & Wu, X. (2018). Improving Crowdsourced Label Quality Using Noise Correction. IEEE Transactions on Neural Networks and Learning Systems, 29(5), 1675–1688. https://doi.org/10.1109/TNNLS.2017.2677468
Zhang, J., Sheng, V. S., Wu, J., & Wu, X. (2015). Multi-class ground truth inference in crowdsourcing with clustering. IEEE Transactions on Knowledge and Data Engineering, 28(4), 1080–1085.
Zhang, J., Wu, X., & Sheng, V. S. (2014). Imbalanced multiple noisy labeling. IEEE Transactions on Knowledge and Data Engineering, 27(2), 489–503.
Zhong, J., Yang, P., & Tang, K. (2017). A quality-sensitive method for learning from crowds. IEEE Transactions on Knowledge and Data Engineering, 29(12), 2643–2654.
Chen, Z., Jiang, L., & Li, C. (2022). Label augmented and weighted majority voting for crowdsourcing. Information Sciences, 606, 397–409.
Daniel, F., Kucherbaev, P., Cappiello, C., Benatallah, B., & Allahbakhsh, M. (2018). Quality Control in Crowdsourcing. ACM Computing Surveys, 51(1), 1–40. https://doi.org/10.1145/3148148
Dawid, A. P., & Skene, A. M. (1979). Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1), 20–28.
Demartini, G., Difallah, D. E., & Cudré-Mauroux, P. (2012). ZenCrowd:leveraging probabilistic reasoning and crowdsourcing techniques for large-scale entity linking. Proceedings of the 21st International Conference on World Wide Web, 469–478. https://doi.org/10.1145/2187836.2187900
Jiang, L., Zhang, H., Tao, F., & Li, C. (2022). Learning From Crowds With Multiple Noisy Label Distribution Propagation. IEEE Transactions on Neural Networks and Learning Systems, 33(11), 6558–6568. https://doi.org/10.1109/TNNLS.2021.3082496
Karger, D., Oh, S., & Shah, D. (2011). Iterative Learning for Reliable Crowdsourcing Systems. In J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, & K. Q. Weinberger (Eds.), Advances in Neural Information Processing Systems (Vol. 24). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2011/file/c667d53acd899a97a85de0c201ba99be-Paper.pdf
Kim, H. C., & Ghahramani, Z. (2012). Bayesian classifier combination. Artificial Intelligence and Statistics, 619–627.
Kurve, A., Miller, D. J., & Kesidis, G. (2014). Multicategory crowdsourcing accounting for variable task difficulty, worker skill, and worker intention. IEEE Transactions on Knowledge and Data Engineering, 27(3), 794–809.
Li, J., Jiang, L., & Zhang, W. (2024). Label Consistency-Based Ground Truth Inference for Crowdsourcing. IEEE Transactions on Neural Networks and Learning Systems.
Li, Y., Rubinstein, B., & Cohn, T. (2019). Exploiting worker correlation for label aggregation in crowdsourcing. International Conference on Machine Learning, 3886–3895.
Raykar, V. C., Yu, S., Zhao, L. H., Valadez, G. H., Florin, C., Bogoni, L., & Moy, L. (2010). Learning from crowds. Journal of Machine Learning Research, 11(4).
Read, J., Reutemann, P., Pfahringer, B., & Holmes, G. (2016). Meka: A multi-label/multi-target extension to Weka. Journal of Machine Learning Research, 17, 1–5. https://www.jmlr.org/papers/volume17/12-164/12-164.pdf
Ren, L., Jiang, L., & Li, C. (2023). Label confidence-based noise correction for crowdsourcing. Engineering Applications of Artificial Intelligence, 117, 105624.
Ren, L., Jiang, L., Zhang, W., & Li, C. (2024). Label distribution similarity-based noise correction for crowdsourcing. Frontiers of Computer Science, 18(5), 185323.
Sheng, V. S., Zhang, J., Gu, B., & Wu, X. (2017). Majority voting and pairing with multiple noisy labeling. IEEE Transactions on Knowledge and Data Engineering, 31(7), 1355–1368.
Snow, R., O’Connor, B., Jurafsky, D., & Ng, A. (2008). Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, 254–263. https://aclanthology.org/D08-1027
Suyal, H., & Singh, A. (2024). Multilabel classification using crowdsourcing under budget constraints. Knowledge and Information Systems, 66(2), 841–877.
Suyal, H., & Singh, A. (2025). A Systematic Literature Survey of Crowdsourcing: Current Status and Future Perspectives. WIREs Data Mining and Knowledge Discovery, 15(3). https://doi.org/10.1002/widm.70037
Tao, F., Jiang, L., & Li, C. (2020). Label similarity-based weighted soft majority voting and pairing for crowdsourcing. Knowledge and Information Systems, 62, 2521–2538.
Tian, T., & Zhu, J. (2015). Max-margin majority voting for learning from crowds. Advances in Neural Information Processing Systems, 28.
Venanzi, M., Guiver, J., Kazai, G., Kohli, P., & Shokouhi, M. (2014). Community-based bayesian aggregation models for crowdsourcing. Proceedings of the 23rd International Conference on World Wide Web, 155–164.
Whitehill, J., Wu, T., Bergsma, J., Movellan, J., & Ruvolo, P. (2009). Whose vote should count more: Optimal integration of labels from labelers of unknown expertise. Advances in Neural Information Processing Systems, 22.
Wilson, D. R., & Martinez, T. R. (1997). Improved Heterogeneous Distance Functions. Journal of Artificial Intelligence Research, 6, 1–34. https://doi.org/10.1613/jair.346
Yang, Y., Zhao, Z. Q., Wu, G., Zhuo, X., Liu, Q., Bai, Q., & Li, W. (2024). A Lightweight, Effective, and Efficient Model for Label Aggregation in Crowdsourcing. ACM Transactions on Knowledge Discovery from Data, 18(4), 1–27.
Zhang, J., Jiang, X., Tian, N., & Wu, M. (2024). Label noise correction for crowdsourcing using dynamic resampling. Engineering Applications of Artificial Intelligence, 133, 108439.
Zhang, J., Sheng, V. S., Li, Q., Wu, J., & Wu, X. (2017). Consensus algorithms for biased labeling in crowdsourcing. Information Sciences, 382, 254–273.
Zhang, J., Sheng, V. S., Li, T., & Wu, X. (2018). Improving Crowdsourced Label Quality Using Noise Correction. IEEE Transactions on Neural Networks and Learning Systems, 29(5), 1675–1688. https://doi.org/10.1109/TNNLS.2017.2677468
Zhang, J., Sheng, V. S., Wu, J., & Wu, X. (2015). Multi-class ground truth inference in crowdsourcing with clustering. IEEE Transactions on Knowledge and Data Engineering, 28(4), 1080–1085.
Zhang, J., Wu, X., & Sheng, V. S. (2014). Imbalanced multiple noisy labeling. IEEE Transactions on Knowledge and Data Engineering, 27(2), 489–503.
Zhong, J., Yang, P., & Tang, K. (2017). A quality-sensitive method for learning from crowds. IEEE Transactions on Knowledge and Data Engineering, 29(12), 2643–2654.
Suyal, H., & Singh, A. (2025). A Machine Learning-Based Approach to Enhancing Label Aggregation in Crowdsourcing. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 14, e32757. https://doi.org/10.14201/adcaij.32757
Downloads
Download data is not yet available.
+
−