Nav: Home

New application can detect Twitter bots in any language

June 13, 2019

Thanks to fruitful collaboration between language scholars and machine learning specialists, a new application developed by researchers at the University of Eastern Finland and Linnaeus University in Sweden can detect Twitter bots independent of the language used.

In recent years, big data from various social media applications have turned the web into a user-generated repository of information in ever-increasing number of areas. Because of the relatively easy access to tweets and their metadata, Twitter has become a popular source of data for investigations of a number of phenomena. These include, for instance, various political campaigns, social and political upheavals, Twitter as a tool for emergency communication, and using social media data to predict stock market prices.

However, research using data from social media data is often skewed by the presence of bots. Bots are non-personal and automated accounts that post content to online social networks. The popularity of Twitter as an instrument in public debate has led to a situation in which it has become an ideal target of spammers and automated scripts. It has been estimated that around 5-10% of all users are bots, and that these accounts generate about 20-25% of all tweets posted.

Researchers of the digital humanities at the University of Eastern Finland and Linnaeus University in Sweden have developed a new application that relies on machine learning to detect Twitter bots. The application is able to detect autogenerated tweets independent of the language used. The researchers captured for analysis a total of 15,000 tweets in Finnish, Swedish and English. Finnish and Swedish were mainly used for training, whereas tweets in English were used to evaluate the language independence of the application. The application is light, making it possible to classify vast amounts of data quickly and relatively efficiently.

"This enhances the quality of data - and paints a more accurate picture of the reality," Professor of English Mikko Laitinen from the University of Eastern Finland notes.

According to Professor Laitinen, bots are relatively harmless, whereas trolls do harm as they spread fake news and come up with made-up stories. This is why there's a need for increasingly advanced tools for social media monitoring.

"This is a complex issue and requires interdisciplinary approaches. For instance, we linguists are working together with machine learning specialists. This type of work also calls for determination and investments in research infrastructures that serve as a platform for researchers from different fields to collaborate on."

According to Professor Laitinen, it is essential for researchers to have access to social media data.

"Currently, data are the property of American technology conglomerates, and a source of their income. In order for researchers to gain access to this data, cooperation at the national and international levels, and especially the involvement of the EU are needed."
-end-
For further information, please contact:

Professor Mikko Laitinen, mikko.laitinen@uef.fi, tel. +358 50 441 2389

Publication:

Jonas Lundberg, Jonas Nordqvist, Mikko Laitinen. Towards a language independent Twitter bot detector. Proceedings of the Digital Humanities in the Nordic Countries 4th Conference, 308-318. http://ceur-ws.org/Vol-2364/28_paper.pdf, published online on 17 May 2019.

University of Eastern Finland

Related Language Articles:

Chinese to rise as a global language
With the continuing rise of China as a global economic and trading power, there is no barrier to prevent Chinese from becoming a global language like English, according to Flinders University academic Dr Jeffrey Gil.
'She' goes missing from presidential language
MIT researchers have found that although a significant percentage of the American public believed the winner of the November 2016 presidential election would be a woman, people rarely used the pronoun 'she' when referring to the next president before the election.
How does language emerge?
How did the almost 6000 languages of the world come into being?
New research quantifies how much speakers' first language affects learning a new language
Linguistic research suggests that accents are strongly shaped by the speaker's first language they learned growing up.
Why the language-ready brain is so complex
In a review article published in Science, Peter Hagoort, professor of Cognitive Neuroscience at Radboud University and director of the Max Planck Institute for Psycholinguistics, argues for a new model of language, involving the interaction of multiple brain networks.
Do as i say: Translating language into movement
Researchers at Carnegie Mellon University have developed a computer model that can translate text describing physical movements directly into simple computer-generated animations, a first step toward someday generating movies directly from scripts.
Learning language
When it comes to learning a language, the left side of the brain has traditionally been considered the hub of language processing.
Learning a second alphabet for a first language
A part of the brain that maps letters to sounds can acquire a second, visually distinct alphabet for the same language, according to a study of English speakers published in eNeuro.
Sign language reveals the hidden logical structure, and limitations, of spoken language
Sign languages can help reveal hidden aspects of the logical structure of spoken language, but they also highlight its limitations because speech lacks the rich iconic resources that sign language uses on top of its sophisticated grammar.
Lying in a foreign language is easier
It is not easy to tell when someone is lying.
More Language News and Language Current Events

Trending Science News

Current Coronavirus (COVID-19) News

Top Science Podcasts

We have hand picked the top science podcasts of 2020.
Now Playing: TED Radio Hour

Making Amends
What makes a true apology? What does it mean to make amends for past mistakes? This hour, TED speakers explore how repairing the wrongs of the past is the first step toward healing for the future. Guests include historian and preservationist Brent Leggs, law professor Martha Minow, librarian Dawn Wacek, and playwright V (formerly Eve Ensler).
Now Playing: Science for the People

#565 The Great Wide Indoors
We're all spending a bit more time indoors this summer than we probably figured. But did you ever stop to think about why the places we live and work as designed the way they are? And how they could be designed better? We're talking with Emily Anthes about her new book "The Great Indoors: The Surprising Science of how Buildings Shape our Behavior, Health and Happiness".
Now Playing: Radiolab

The Third. A TED Talk.
Jad gives a TED talk about his life as a journalist and how Radiolab has evolved over the years. Here's how TED described it:How do you end a story? Host of Radiolab Jad Abumrad tells how his search for an answer led him home to the mountains of Tennessee, where he met an unexpected teacher: Dolly Parton.Jad Nicholas Abumrad is a Lebanese-American radio host, composer and producer. He is the founder of the syndicated public radio program Radiolab, which is broadcast on over 600 radio stations nationwide and is downloaded more than 120 million times a year as a podcast. He also created More Perfect, a podcast that tells the stories behind the Supreme Court's most famous decisions. And most recently, Dolly Parton's America, a nine-episode podcast exploring the life and times of the iconic country music star. Abumrad has received three Peabody Awards and was named a MacArthur Fellow in 2011.