Multilingual Fake News Detection Using Cross-Lingual Transformer Models
Keywords:
Multilingual Fake News Detection, Cross-Lingual Transformer, Claim-Evidence Alignment, Metadata Fusion, Zero-Shot LearningAbstract
The rapid spread of misinformation across multilingual digital environments has created an urgent need for robust fact-verification systems, particularly for low-resource languages with limited labeled data. This study proposes the Language-Aware Evidence and Alignment Framework (LEAF), which integrates multilingual Transformer representations, claim-evidence cosine-alignment features, publisher and claimant credibility metadata, and linguistic indicators. To improve stability in low-resource and zero-shot settings, LEAF employs a multi-objective loss function that combines classification loss with a Kullback-Leibler-divergence-based linguistic-consistency regularizer and Laplace-smoothed metadata calibration. The framework was evaluated on the multilingual X-FACT dataset. LEAF outperformed the strongest Transformer baseline across macro-F1, micro-F1, and accuracy. The largest gains were observed in low-resource languages, including improvements of 11.8 and 13.3 percentage points in macro-F1 for Urdu and Punjabi, respectively. In a cold-start evaluation involving previously unseen publishers, LEAF improved macro-F1 by 10.3 percentage points. Sensitivity analysis indicated that using three evidence documents provided the most efficient balance between predictive performance and inference latency, requiring approximately 25 ms per instance. Ablation results further showed that claim-evidence alignment was the most influential auxiliary component, particularly for resource-constrained languages. These findings indicate that combining multilingual semantic representations with evidence alignment and calibrated credibility metadata can reduce dependence on language-specific labeled data and improve cross-lingual generalization in multilingual fake news detection.
Downloads
References
Dementieva, D., Kuimov, M., & Panchenko, A. (2022). Multiverse: Multilingual evidence for fake news detection. arXiv. https://doi.org/10.48550/arXiv.2211.14279
Gupta, A., & Srikumar, V. (2021). X-Fact: A new benchmark dataset for multilingual fact checking. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers),
Mittal, S., Sundriyal, M., & Nakov, P. (2023). Lost in translation, found in spans: Identifying claims in multilingual social media. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,
Panchendrarajan, R., & Zubiaga, A. (2024). Claim detection for automated fact-checking: A survey on monolingual, multilingual and cross-lingual research. Natural Language Processing Journal, 7, 100066. https://doi.org/10.1016/j.nlp.2024.100066
Tian, L., Zhang, X., & Lau, J. H. (2021). Rumour detection via zero-shot cross-lingual transfer learning. In Machine Learning and Knowledge Discovery in Databases (pp. 603-618). Springer. https://doi.org/10.1007/978-3-030-86486-6_37

