Mining Twitter
So here recently Facebook has been in the news for breaching trust with the people they serve, pretty much you and I. Essentially the story is around sharing data with “Cambridge Analytica” and how this company used the data of about 50 Million Facebook users. Facebook claims they weren’t aware of the situation and they say they prohibit this kind of activity.
I wonder if the general public can access social media data and if so, how can we receive it in a format that can be loaded for better analysis? This short article will briefly walk you through just the basics of getting some of this interesting data from Twitter, one of the social media offerings.
Setup
To begin you will need to create a developer account with Twitter, don’t worry its free. You can sign-up for it at http://developer.twitter.com , you will find that there is a ton of information but you need to browse around the site and get a feel for what tools are available and in the process understand the Twitter platform.
You will eventually find that you will need to setup your own Twitter application, so that you can use it as an intermediary to the datasets you are looking to find.
Figure 1.0 - An example of a Twitter App
Figure 1.0 - An example of a Twitter App
Once you have the application created you will be presented with Consumer Key’s, Consumer Secrets, Access Tokens and Access Token Secrets. These keys are important as they are used to gain access to this application. This application is nothing more than an interface for you to connect to Twitter. Of course, developers can create applications that pop advertising and other fun stuff but what we are doing is very basic stuff. Also, remember when you create your own application you can name it what you want and leave pretty much everything else at their defaults.
Code to Mine
Once you have created the application you are ready to mine Twitter for something interesting. Below is the code to pull interesting data from Twitter in a structured way so that you can load it into your favorite analytics engine. There are many applications out there that can consume your data for further analysis, but the free software that comes to mind is R Studio. This package will take a bit of commitment as you will need to stand up the server and learn another programming language. You will find that this can get pretty deep, pretty fast.
get-twitter.py
— Cut here —
from tweepy.streaming import StreamListener
from tweepy import OAuthHandler
from tweepy import Stream
access_token = “YOUR OWN ACCESS TOKEN"
access_token_secret = “YOUR OWN ACCESS TOKEN SECRET"
consumer_key = “YOUR OWN CONSUMER KEY"
consumer_secret = “YOUR OWN CONSUMER SECRET"
class StdOutListener(StreamListener):
def on_data(self, data):
print data
return True
def on_error(self, status):
print status
if __name__ == '__main__':
#This handles Twitter authetification and the connection to Twitter Streaming API
l = StdOutListener()
auth = OAuthHandler(consumer_key, consumer_secret)
auth.set_access_token(access_token, access_token_secret)
stream = Stream(auth, l)
#This line filter Twitter Streams to capture data by the keywords: '#ripstephenhawking'
stream.filter(track=['#ripstephenhawking'])
—— Cut here ——
In the code example above, you can see at the last line we performed a search on #ripstephenhawking which was the Twitter hash tag that was trending at the time of writing. Once you run this script the output you receive is in a JSON format as shown below in figure 3.0. Something interesting is the size of the data it spits out for just a single Tweet. As you can see briefly the tweet below originated from Brazil from a user called “VilsonCFGoda” and he/she is from the city of “Rio Negrinho”. There is much more metadata in this single tweet but you get the gist of it.
Figure 3.0 – JSON formatted output for #ripstephenhawkings Twitter hashtag
Summary
So as you can see it really didn't take too much effort to get some large datasets with more information than would be present from a standard Twitter client. Once you get your datasets in such a raw format you can go to town and do some R-fu on it.
So as you can see it really didn't take too much effort to get some large datasets with more information than would be present from a standard Twitter client. Once you get your datasets in such a raw format you can go to town and do some R-fu on it.



Comments
Post a Comment