
PM
Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze
· 1 min read
ResearcharXiv cs.CL
Identifying Scientists on X
arXiv:2609.31264v1 Announce Type: new
Abstract: With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based on their user biographies and tweets. We show that we are able to classify accounts as scientists and non- scientists on two different datasets, reaching an F1 score of up to 0.88 using Random Forests with linguistic features and up to 0.96 using a contrastively fine-tuned DeBERTa model in an ensemble setup. Furthermore, we provide two datasets with X users labeled as scientists or non scientists and their respective tweets and user biographies.
Original source
This story was published by arXiv cs.CL and written by Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


