Cross-lingual Robustness of Behavioral Coordination Signals for Detecting State-Linked Information Operations on Twitter: A Comparative Empirical Study
DOI:
https://doi.org/10.69987/AIMLR.2025.60304Keywords:
coordinated inauthentic behavior, cross-lingual transfer, information operations, user similarity networksAbstract
Coordinated inauthentic behavior on social media has become a primary vector through which state-linked information operations seek to shape public political discourse. Computational detection work to date has been built and validated mainly on English-language operations, leaving open whether the behavioral coordination signals the field relies on—co-retweet, co-hashtag, co-URL, temporal synchronization, and textual similarity—generalize to non-English political environments. This study presents a comparative empirical assessment of the cross-lingual robustness of these five canonical signals using two publicly released Twitter datasets: 3,841 IRA accounts (English-dominant, 2018 release) and 23,750 PRC-attributed accounts (Chinese-dominant, June 2020 takedown). User similarity networks are constructed following the framework standardized in prior literature, the textual-similarity backbone is swapped among mBERT, XLM-R, and LaBSE, and unsupervised detection quality is measured against time- and topic-matched organic control samples. Co-retweet and temporal synchronization transfer most reliably across languages, with F1 gaps under 0.05; textual similarity drops by 0.12 on Chinese data even with multilingual encoders. The findings offer empirical guidance for cross-border information operation monitoring and reveal the unequal language coverage of current detection toolchains. The contribution is empirical rather than methodological, the scope is restricted to two operations on a single platform, and the results are intended as a calibration baseline for English-trained pipelines redeployed in Chinese-language political environments.

