Written evidence submitted by Factmata Limited, UK
Author: Dhruv Ghulati
Contributors: Zeerak Waseem, David Corney, Robert Stojnic, Rishabh Shukla
Introduction
Factmata is a UK based technology startup founded by 3 artificial intelligence researchers from University College London and the University of Sheffield. Funded by an initial grant from the Google Digital News Initiative, the team has now grown to 10 natural language processing experts, engineers, and journalists, including one of the early developers of Wikipedia. The company has raised $1m in funding from global investors to build its technology. Factmata believes that there are separate strands for the reason why fake news has proliferated today within the media environment. First of all, one must define what is fake news.
What is fake news?
Fake news has been used as a catch-all term and has captured the imagination of media, culture, government, and more since 2016, the last US presidential election being a key driver. Fake news is thought of as either politicised or inaccurate content, or deliberate disinformation, though many definitions have been proposed.
From Factmata’s perspective, we have been trying to understand the concept of fake news by speaking to industry participants including financial companies, publishers, advertising exchanges and governments. By surveying multiple leaders within companies, we have learnt that organisational views of what fake news is differ considerably. Broadly, we use our own categorisation based in the incentives:
Advertising profit driven | The sole objective of this content is to make money for the creator. Fake news and hoaxes can become viral and thus create lucrative revenue wells by attracting advertisers. |
Disruption driven | Content such as hoaxes and satire are primarily created for amusement, clicks and laughs, and to confuse people into thinking something is real. |
Deliberate malinformation driven | Content that is intended to incite hatred, bullying or abuse, and intentionally litter the internet with false information and junk. |
Propaganda driven
| Content typically funded by state actors using political or social bot networks, to spread propaganda messages through the internet. |
The success on all types of fake news listed above depends on how quickly and widely the content is shared.
Factmata believes the existing ranking and relevance models of social networks such as Twitter and Facebook enable this virality of problematic content to occur because their underlying ranking algorithms are mostly based on the quantity of comments, shares and likes as well as overall clicks on a particular story.
Only very preliminary work has been undertaken to build a semantic understanding of the facts expressed in a piece of content to estimate if it is low-quality or problematic.
Focusing on advertising-driven fake news
Factmata believes the Internet’s current advertising model funds fake news. Quality simply doesn’t pay in today’s media environment. An advertiser pays the same to bid for a user’s cookies if target characteristics are matched, purely based on pre-set prices from the publisher’s entire inventory, regardless of whether they are viewing an article that costs 50¢ or $1m to produce, or whether it was well researched and balanced or extremely biased.
Content such as hate speech and propaganda is very susceptible to online sharing, likes and views. Content with high engagement naturally attracts users and is therefore more likely to be financially encouraged by attracting advertising bids on pages.
Factmata works with major advertising exchanges and platforms to help advertisers detect and categorise this type of content, protecting brands’ reputation by hindering promotion and alignment with it. By allowing ad platforms to accurately recognise this type of toxic content, we hope to choke off the funding and incentive structure, eventually eliminating its creation and the wrongful acquisition of revenues. This is our first step in the battle to reduce online misinformation for the public.
We also believe that by building systems to algorithmically score content for quality, “good” publishing and journalism will be encouraged and thus a major incentive for fake news creation will disappear. Our broader goal is to develop a quality score for content, built as a pricing mechanism within programmatic advertising. This will facilitate a pricing mechanism based on content quality, minimising revenues for spam and clickbait.
Our next step is to form industry-wide accepted definitions, schema and taxonomy for hate speech, fake news and other forms of problematic content. Our goal is to provide this technology to publishers, advertisers, ad platforms, social networks and blogging platforms to detect, filter, or moderate such content in a more fair, accurate, and less biased manner than before by adopting open evaluation methods and explainable algorithms with no black boxes.
Factmata’s technology
Factmata has built machine learning algorithms to detect hate speech, extreme propaganda content, and fake news. Our definitions for these categories, and the definitions of each, are as follows:
1. Hate speech and abusive content - this includes content with the existence of sexist, racist or ethnic statements that use slur, attack or negatively stereotype a minority, or defend xenophobia or sexism.
2. Propaganda and extremely politically biased content - hyperpartisan news can be understood as extremely one-sided, extremely biased news articles. These articles provide an unbalanced and provocative point of view in describing events, and often contain strong sentiment in describing political parties or politicians, with positive/negative association; insulting/aggravating or slanderous statements towards people or parties; direct calls to action to support a particular faction or aggressive campaigning, and finally content which tends to cherry pick evidence to support its own biases and arguments.
3. Spoof websites and content spread by known fake news networks - here, we define fake news as those websites or articles propagating untrue information, knowing they are untrue, to deceive others. This definition is heavily focused on intent, as other types of websites containing untrue information (e.g. satire, fiction) are not included in our definition of fake news. This system picks up fake news content propagated by known fake news sites (and networks of sites), content that links to such sites, and sites that try to spoof or plagiarise the branding of reputable news publishers.
A data analysis on the size of the “fake news” problem
Factmata is working towards participating in discussions with social media platforms such as Twitter, Google and Facebook, to help in their analysis and research on the amount of cross platform fake news and problematic content that exists on their networks. However, such platforms have been publicly criticised for being unwilling to share their data on flagged news within the platform, or own internal analyses of the size of the problem.
Alongside this, Factmata is working with major advertising networks and supply side platforms to obtain representative samples of inventory types. Factmata can then estimate the number of fake or toxic internet sites that carry advertising. These are sites that content advertisers are supporting and that ordinary members of the public are being exposed to. These datasets provide a first indication of how much fake news there is on social networks. In principle, an early warning alerting system could be built for social networks to tell them when content on their networks is fake, based on an inventory of internet-wide links being registered and articles being written in real time.
Existing studies for the scale of fake news across social networks have either focused on Twitter, using simplified methods such as counting the number of Tweets containing the #fakenews hashtag and identifying the unique links; or have been focused on analysing only content from Europe or the US; or simply looking at certain lists that were already verified by fact checking networks as being fake news sites. One Oxford University study looked at 28 million feeds shared in political debates and elections in the US, UK, France, and Germany. It found a seven-to-one real news versus fake news ratio in France. The UK and Germany had a ratio of four-to-one.
Factmata’s system uses more advanced methods of looking at the semantic nature of the content itself, with access to unique annotated evaluation datasets of fake news content, hate speech sentences, and politically biased articles.
Our findings
Factmata has built pipelines to analyse 2.8bn ad impressions so far, growing to a potential 50bn impressions a day to be processed through our platform. We performed an analysis on a sample of 100k unique articles that were from domains that a single supply side platform currently advertises on.
Fake news sites: Relatively few sites contain purely ‘fake news’, but they can create individual stories that get shared widely. These sites often mix completely false stories with hyperpartisan distortions. Factmata finds these through a combination of blacklists, spoof URL detectors, shared ownership analysis and other methods.
Hateful sites: these contain language that attacks people based on their race, religion, gender, sexuality etc. Examples include:
“keep your mouth closed, no hoes, no bitches, no nothing”,
“why are you on a dating app if you hate women? literally, you’ve never met me and you’re texting me like i’m a stupid bitch … texting me and being mad rude",
“certainly, it is not hard to understand why men coming from these and similar backgrounds, propelled by islamic and traditional attitudes, unable to cope with the sight of a lightly dressed woman, teenager or little girl, and taught that women are their property, zealous of their personal or family honour, and never taught self-control, assuming male rights over all women, and contemptuous of non-muslim women on the grounds that all non-muslims are the inferiors of all muslims, choose to groom, traffic, and rape vulnerable females living in their own towns”
Note that the last does not contain any profanities or slurs but is still correctly recognised as Islamophobic.
Hyperpartisan sites: These are reporting genuine stories, but distorted to the political extremes with no attempt at balance or debate. There is often a blurred line between hyper-partisanship and fake news.
Here are our main findings:
Takeaways
The first takeaway from our work is that academic and industrial researchers need better insight into the presence and spread of misinformation on social media platforms. Factmata is devising its own methods to obtain data from social media feeds, and building its own crawling architecture, However, it would be far more effective to work directly and collaboratively with platforms by means of an industry initiative.
Another major reason for our analysis is to be the first to inform the public with a clear view on how much fake news exists online. Factmata has gathered together data from across advertising networks and supply chains, and is building its own representative dataset of what exists. From this, Factmata is hoping to build the world’s first global real-time monitor of hyper-partisan, fake and hate speech content, ensuring platforms are held accountable.
Asks and Concluding Remarks
There are a number of concluding steps that would allow fake news to be more accurately researched and measured. Once evaluation is simple and standardised, the world can move forward to build the best systems to detect and remove problematic content.
Firstly, Facebook, Twitter and Google should allow fake news specialists and researchers controlled access to their data, as part of an open and collaborative research project, totally anonymised and privacy compliant with GDPR. This will ensure open innovation and oversight of methods to tackle this problem across platforms on the internet, rather than islands of methods.
Secondly, Facebook, Twitter and Google should provide access to their annotated data of what they have found to be violations of hate speech, abuse and fake news policy to specialists and researchers to help drive faster innovation in better methods to tackle the problem through automation.
Finally, Facebook, Twitter and Google should form an official expert stakeholder group, in the same way that the EU government has done, and have others hold them accountable for their systems and product development roadmaps.
January 2018