@asmelash

Detecting spammers on twitter

, , , and . Annual Collaboration, Electronic messaging, Anti-Abuse and Spam Conference (CEAS), (2010)

Abstract

With millions of users tweeting around the world, real time search systems and different types of mining tools are emerging to allow people tracking the repercussion of events and news on Twitter. However, although appealing as mech- anisms to ease the spread of news and allow users to discuss events and post their status, these services open opportu- nities for new forms of spam. Trending topics, the most talked about items on Twitter at a given point in time, have been seen as an opportunity to generate traffic and revenue. Spammers post tweets containing typical words of a trend- ing topic and URLs, usually obfuscated by URL shorteners, that lead users to completely unrelated websites. This kind of spam can contribute to de-value real time search services unless mechanisms to fight and stop spammers can be found. In this paper we consider the problem of detecting spam- mers on Twitter. We first collected a large dataset of Twit- ter that includes more than 54 million users, 1.9 billion links, and almost 1.8 billion tweets. Using tweets related to three famous trending topics from 2009, we construct a large la- beled collection of users, manually classified into spammers and non-spammers. We then identify a number of charac- teristics related to tweet content and user social behavior, which could potentially be used to detect spammers. We used these characteristics as attributes of machine learn- ing process for classifying users as either spammers or non- spammers. Our strategy succeeds at detecting much of the spammers while only a small percentage of non-spammers are misclassified. Approximately 70% of spammers and 96% of non-spammers were correctly classified. Our results also highlight the most important attributes for spam detection on Twitter.

Links and resources

Tags

community

  • @becker
  • @dimitargn
  • @asmelash
  • @nosebrain
  • @beate
  • @silencetwitter
  • @khilgenberg
@asmelash's tags highlighted