{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1> 1. Business Problem </h1>"},{"metadata":{},"cell_type":"markdown","source":"<h2>1.1 Problem Description </h2>"},{"metadata":{},"cell_type":"markdown","source":"__ Introduction: <br> Clickthrough rate (CTR) __\nis a ratio showing how often people who see your ad end up clicking it. Clickthrough rate (CTR) can be used to gauge how well your keywords and ads are performing.\n\n- CTR is the number of clicks that your ad receives divided by the number of times your ad is shown: clicks ÷ impressions = CTR. For example, if you had 5 clicks and 100 impressions, then your CTR would be 5%.\n\n- Each of your ads and keywords have their own CTRs that you can see listed in your account.\n- A high CTR is a good indication that users find your ads helpful and relevant. CTR also contributes to your keyword's expected CTR, which is a component of Ad Rank. Note that a good CTR is relative to what you're advertising and on which networks.\n> Credits: Google (https://support.google.com/adwords/answer/2615875?hl=en) \n\n<p> Search advertising has been one of the major revenue sources of the Internet industry for years. A key technology behind search advertising is to predict the click-through rate (pCTR) of ads, as the economic model behind search advertising requires pCTR values to rank ads and to price clicks.<b> In this task, given the training instances derived from session logs of the Tencent proprietary search engine, soso.com, participants are expected to accurately predict the pCTR of ads in the testing instances. </b></p>"},{"metadata":{},"cell_type":"markdown","source":"<h2>1.2 Source/Useful Links </h2>"},{"metadata":{},"cell_type":"markdown","source":"__ Source __ : https://www.kaggle.com/c/kddcup2012-track2 <br>\n__ Dropbox Links __: https://www.dropbox.com/sh/k84z8y9n387ptjb/AAA8O8IDFsSRhOhaLfXVZcJwa?dl=0 <br>\n__ Blog __ :https://hivemall.incubator.apache.org/userguide/regression/kddcup12tr2_dataset.html"},{"metadata":{},"cell_type":"markdown","source":"<h2> 1.3 Real-world/Business Objectives and Constraints </h2>"},{"metadata":{},"cell_type":"markdown","source":"Objective: Predict the pClick (probability of click) as accurately as possible.\n\nConstraints: Low latency, Interpretability."},{"metadata":{},"cell_type":"markdown","source":"<h1>2. Machine Learning problem </h1>"},{"metadata":{},"cell_type":"markdown","source":"<h2>2.1 Data </h2>"},{"metadata":{},"cell_type":"markdown","source":"<h3> 2.1.1 Data Overview </h3>"},{"metadata":{},"cell_type":"markdown","source":"<table style=\"width:50%;text-align:center;\">\n<caption style=\"text-align:center;\">Data Files</caption>\n<tr>\n<td><b>Filename</b></td><td><b>Available Format</b></td>\n</tr>\n<tr>\n<td>training</td><td>.txt (9.9Gb)</td>\n</tr>\n<tr>\n<td>queryid_tokensid</td><td>.txt (704Mb)</td>\n</tr>\n<tr>\n<td>purchasedkeywordid_tokensid</td><td>.txt (26Mb)</td>\n</tr>\n<tr>\n<td>titleid_tokensid</td><td>.txt (172Mb)</td>\n</tr>\n<tr>\n<td>descriptionid_tokensid</td><td>.txt (268Mb)</td>\n</tr>\n<tr>\n<td>userid_profile</td><td>.txt (284Mb)</td>\n</tr>\n</table>\n\n<table style=\"width:100%\">\n  <caption style=\"text-align:center;\">training.txt</caption>\n  <tr>\n    <th>Feature</th>\n    <th>Description</th>\n  </tr>\n  <tr>\n    <td>UserID</td>\n    <td>The unique id for each user</td>\n    </tr>\n  <tr>\n    <td>AdID</td>\n    <td>The unique id for each ad</td>\n  </tr>\n  <tr>\n    <td>QueryID</td>\n    <td>The unique id for each Query (it is a primary key in Query table(queryid_tokensid.txt))</td>\n  </tr>\n  <tr>\n    <td>Depth</td>\n    <td>The number of ads impressed in a session is known as the 'depth'. </td>\n  </tr>\n  <tr>\n    <td>Position</td>\n    <td>The order of an ad in the impression list is known as the ‘position’ of that ad.</td>\n  </tr>\n  <tr>\n    <td>Impression</td>\n    <td>The number of search sessions in which the ad (AdID) was impressed by the user (UserID) who issued the query (Query).</td>\n  </tr>\n  <tr>\n    <td>Click</td>\n    <td>The number of times, among the above impressions, the user (UserID) clicked the ad (AdID).</td>\n  </tr>\n  <tr>\n    <td>TitleId</td>\n    <td>A property of ads. This is the key of 'titleid_tokensid.txt'. [An Ad, when impressed, would be displayed as a short text known as ’title’, followed by a slightly longer text known as the ’description’, and a URL (usually shortened to save screen space) known as ’display URL’.]</td>\n  </tr>\n  <tr>\n    <td>DescId</td>\n    <td>A property of ads.  This is the key of 'descriptionid_tokensid.txt'. [An Ad, when impressed, would be displayed as a short text known as ’title’, followed by a slightly longer text known as the ’description’, and a URL (usually shortened to save screen space) known as ’display URL’.]</td>\n  </tr>\n  <tr>\n    <td>AdURL</td>\n    <td>The URL is shown together with the title and description of an ad. It is usually the shortened landing page URL of the ad, but not always. In the data file,  this URL is hashed for anonymity.</td>\n  </tr>\n  <tr>\n    <td>KeyId</td>\n    <td>A property of ads. This is the key of  'purchasedkeyword_tokensid.txt'.</td>\n  </tr>\n  <tr>\n    <td>AdvId</td>\n    <td>a property of the ad. Some advertisers consistently optimize their ads, so the title and description of their ads are more attractive than those of others’ ads.</td>\n  </tr>\n</table>\n\n___\nThere are five additional data files, as mentioned in the above section: \n\n1. queryid_tokensid.txt \n\n2. purchasedkeywordid_tokensid.txt \n\n3. titleid_tokensid.txt \n\n4. descriptionid_tokensid.txt \n\n5. userid_profile.txt \n\nEach line of the first four files maps an id to a list of tokens, corresponding to the query, keyword, ad title, and ad description, respectively. In each line, a TAB character separates the id and the token set.  A token can basically be a word in a natural language. For anonymity, each token is represented by its hash value.  Tokens are delimited by the character ‘|’. \n\nEach line of ‘userid_profile.txt’ is composed of UserID, Gender, and Age, delimited by the TAB character. Note that not every UserID in the training and the testing set will be present in ‘userid_profile.txt’. Each field is described below: \n\n1. Gender:  '1'  for male, '2' for female,  and '0'  for unknown. \n\n2. Age: '1'  for (0, 12],  '2' for (12, 18], '3' for (18, 24], '4'  for  (24, 30], '5' for (30,  40], and '6' for greater than 40. "},{"metadata":{},"cell_type":"markdown","source":"<h3> 2.1.2 Example Data point </h3>"},{"metadata":{},"cell_type":"markdown","source":"__ training.txt __\n<pre>\nClick Impression\tAdURL\t     AdId\t   AdvId  Depth\tPos\t QId\t   KeyId\tTitleId\t DescId\t UId\n0\t 1\t 4298118681424644510\t7686695\t385\t    3\t  3\t 1601\t    5521\t 7709\t  576\t 490234\n0\t 1\t 4860571499428580850\t21560664\t37484\t  2\t  2\t 2255103\t317\t     48989\t  44771\t 490234\n0\t 1\t 9704320783495875564\t21748480\t36759\t  3\t  3\t 4532751\t60721\t 685038\t  29681\t 490234\n</pre>\n\n__ queryid_tokensid.txt__\n<pre>\nQId\tQuery\n0\t12731\n1\t1545|75|31\n2\t383\n3\t518|1996\n4\t4189|75|31\n</pre>\n\n__purchasedkeywordid_tokensid.txt__\n<pre>\n</pre>\n\n__titleid_tokensid.txt__\n<pre>\nTitleId\tTitle\n0\t615|1545|75|31|1|138|1270|615|131\n1\t466|582|685|1|42|45|477|314\n2\t12731|190|513|12731|677|183\n3\t2371|3970|1|2805|4340|3|2914|10640|3688|11|834|3\n4\t165|134|460|2887|50|2|17527|1|1540|592|2181|3|...\n</pre>\n\n__descriptionid_tokensid.txt__\n<pre>\nDescId\tDescription\n0\t1545|31|40|615|1|272|18889|1|220|511|20|5270|1...\n1\t172|46|467|170|5634|5112|40|155|1965|834|21|41...\n2\t2672|6|1159|109662|123|49933|160|848|248|207|1...\n3\t13280|35|1299|26|282|477|606|1|4016|1671|771|1...\n4\t13327|99|128|494|2928|21|26500|10|11733|10|318\n</pre>\n\n__userid_profile.txt__\n<pre>\nUId\tGender\tAge\n1\t1\t5\n2\t2\t3\n3\t1\t5\n4\t1\t3\n5\t2\t1\n</pre>"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# Loading libraries...\nimport pandas as pd    \nimport numpy as np\n\n%matplotlib inline\nimport matplotlib.pyplot as plt \nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1> 3. Exploratory Data Analysis </h1>"},{"metadata":{},"cell_type":"markdown","source":"<h2> 3.1 Reading and Preparing Data </h2>"},{"metadata":{"trusted":true},"cell_type":"code","source":"import zipfile\nwith zipfile.ZipFile('/kaggle/input/kddcup2012-track2/track2.zip', 'r') as zip_ref:\n    zip_ref.extractall('/kaggle/input')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Load Training Data..\n\ncolumn  = ['Click', 'Impression', 'AdURL', 'AdId', 'AdvId', 'Depth', 'Pos', 'QId', 'KeyId', 'TitleId', 'DescId', 'UId']\norignal = pd.read_csv('/kaggle/input/track2/training.txt', sep='\\t', header=None, nrows=5000000, names=column)\norignal.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Load User Data..\n\nuser_col  = ['UId', 'Gender', 'Age']\nuser      = pd.read_csv('/kaggle/input/track2/userid_profile.txt', sep='\\t', header=None, names=user_col)\nuser.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Load Query Data..\n\nquery_col = ['QId', 'Query']\nquery     = pd.read_csv('/kaggle/input/track2/queryid_tokensid.txt', sep='\\t', header=None, names=query_col)\nquery.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Load Ad Description Data..\n\ndesc_col  = ['DescId', 'Description']\ndesc      = pd.read_csv('/kaggle/input/track2/descriptionid_tokensid.txt', sep='\\t', header=None, names=desc_col)\ndesc.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Load Ad Title Data..\n\ntitle_col = ['TitleId', 'Title']\ntitle     = pd.read_csv('/kaggle/input/track2/titleid_tokensid.txt', sep='\\t', header=None, names=title_col)\ntitle.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def count(sentence):\n    '''\n        (str) -> (int)\n        Returns no. of words in a sentence.\n    '''\n    return len(str(sentence).split('|'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Count no. of words in a query issued by a user.\n\nquery['QCount'] = query['Query'].apply(count)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"query.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del query['Query']\nquery.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Count no. of words in title of an advertisement.\n\ntitle['TCount'] = title['Title'].apply(count)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"title.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Advertisement Title isn't required now, get rid of it.\n\ndel title['Title']\ntitle.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Count no. of words in description of an advertisement.\n\ndesc['DCount'] = desc['Description'].apply(count)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Advertisement Description isn't required now, get rid of it.\n\ndel desc['Description']\ndesc.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Merging orignal with user, query, title & desc on appropriate keys to get data..\n\ndata = pd.merge(orignal, user,  on='UId')\ndata = pd.merge(data,    query, on='QId')\ndata = pd.merge(data,    title, on='TitleId')\ndata = pd.merge(data,    desc,  on='DescId')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Add target variable CTR to the dataset...\n\ndata['CTR'] = data['Click'] * 1.0 / data['Impression'] * 100\ndata.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Note: We loaded 5M datapoints initially, after merger we have around 4.95M datapoints. What does this indicate ? Actually for a lot of user ids data is missing hence merge operation gets rid of such datapoints."},{"metadata":{},"cell_type":"markdown","source":"<h2> 3.2 Analyzing features </h2>"},{"metadata":{},"cell_type":"markdown","source":"<h3> 3.2.1 Getting sense out of the data</h3>"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# CTR(ad) = #Clicks(ad)/#Impressions(ad)\n\n# Calculating net CTR for our dataset...\n\ntotal_impressions = data['Impression'].sum()\ntotal_clicks      = data['Click'].sum()\nnet_CTR           = total_clicks * 1.0 / total_impressions\n\nprint( ('Net CTR: {0}'.format(round(net_CTR*100,2))), '%')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"total = data.shape[0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# total no. of unique users in the dataset...\n\n# print round(len(data.groupby('UId')) * 1.0 / total * 100, 2), '%'\nprint( 'Total no. of unique users:', len(data.groupby('UId')))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# total no. of unique queries in the dataset...\n\n# print round(len(data.groupby('QId')) * 1.0 / total * 100, 2), '%'\nprint( 'Total no. of unique queries:', len(data.groupby('QId')))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# total no. of unique advertisements in the dataset...\n\n# print round(len(data.groupby('AdId')) * 1.0 / total * 100, 2) , '%'\nprint( 'Total no. of unique ads:', len(data.groupby('AdId')))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# total no. of unique advertisers in the dataset...\n\n# print round(len(data.groupby('AdvId')) * 1.0 / total * 100, 2), '%'\nprint( 'Total no. of unique advertisers:', len(data.groupby('AdvId')))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us study the distribution of no. of words in a search query.\n\n# Preparing Data...\n\ntemp = data[['QCount']].copy()\n\nprint( 'Maximum Length of a Query: ', temp['QCount'].max())\nprint( 'Average Length of a Query: ', temp['QCount'].mean())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, (ax1, ax2) = plt.subplots(2)\nsns.kdeplot(temp['QCount'], ax=ax1)\nsns.boxplot(x=None,y='QCount',data=temp, ax=ax2)\n\n# sns.boxplot(x=None,y='QCount',data=temp)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Clearly, data contains outliers. We will remove them in order to make our analysis more robust."},{"metadata":{"trusted":true},"cell_type":"code","source":"print( 'Avg No. of words in a Search query:', round(temp['QCount'].mean(),2))\nprint( 'Median No. of words in a Search query:', temp['QCount'].quantile(0.5))\nprint('3rd Quantile No. of words in a Search query:', temp['QCount'].quantile(0.75))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Remove outliers by considering only queries with lengh < 10.0 (chosen randomly)\n\ntemp = temp[temp['QCount'] < 10.0]\nprint( 'Maximum Length of a Query: ', temp['QCount'].max())\nprint('Average Length of a Query: ', temp['QCount'].mean())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# This is the same plot as above just a little more readable since outliers were removed...\n\nplt.figure(figsize=(10,4))\n\nplt.subplot(1, 2, 1)\nplt.hist(temp['QCount'],\n         color='green',\n         bins=25,\n         normed=True)\nplt.xlabel('No. of words in a Query')\n\nplt.subplot(1, 2, 2)\nplt.boxplot(temp['QCount'],\n            labels=['No. of words in a Query'],\n            )\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: \n\n75 % of search queries has less than 4.0 words."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us study the distribution of no. of words in Ad description.\n\n# Preparing Data...\n\ntemp = data[['DCount']].copy()\n\nprint ('Maximum Length of an Ad Description: ', temp['DCount'].max())\nprint ('Average Length of an Ad Description: ', temp['DCount'].mean())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Distribution of word count in description of an ad...\n\nplt.figure(figsize=(10,4))\n\nplt.subplot(1, 2, 1)\nplt.hist(temp['DCount'],\n         bins=100,\n         color='red',\n         normed=False)\nplt.xlabel('No. of words in a Ad Description')\n\nplt.subplot(1, 2, 2)\nplt.boxplot(temp['DCount'],\n            labels=['No. of words in a Ad Description'],\n            )\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print ('Median No. of words in a Ad description:', temp['DCount'].quantile(0.5))\nprint ('3rd Quantile No. of words in a Ad description:', temp['DCount'].quantile(0.75))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion:\n\n75 % of the Ads use <= 25.0 words for Ad description."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us study the distribution of no. of words in Ad title.\n\n# Preparing Data...\n\ntemp = data[['TCount']].copy()\n\nprint( 'Maximum Length of an Ad Title: ', temp['TCount'].max())\nprint('Average Length of an Ad Title: ', temp['TCount'].mean())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Distribution of word count in a ad title...\n\nplt.figure(figsize=(10,4))\n\nplt.subplot(1, 2, 1)\nplt.hist(temp['TCount'],\n         color='red',\n         bins=100,\n         normed=False)\nplt.xlabel('No. of words in a Ad Title')\n\nplt.subplot(1, 2, 2)\nplt.boxplot(temp['TCount'],\n            labels=['No. of words in a Ad Title'],\n            )\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print( 'Median No. of words in a Ad title:', temp['TCount'].quantile(0.5))\nprint( '3rd Quantile No. of words in a Ad title:', temp['TCount'].quantile(0.75))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion:\n\n75 % of the Ads use < = 11.0 words in their Ad titles."},{"metadata":{"trusted":true},"cell_type":"code","source":"# How is no. of words in Search query affect Ad CTR...\n\n# Preparing data...\n\ntemp = data[['QCount', 'CTR']].copy()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# We are considering only those queries which have less than 10.0 words.\n\ntemp = temp[temp['QCount'] < 10.0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp.shape[0] * 1.0 / data.shape[0] # 99.5% datapoints use less than 10 words in query...","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('QCount').agg(['mean'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(5,5))\n\nplt.bar(result.index, result[('CTR', 'mean')],\n        color='red',\n        width=0.4)\nplt.xlabel('No. of words in a Query')\nplt.ylabel('Avg. CTR')\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: As no. of words in a search query increases, typically CTR of Ads displayed falls."},{"metadata":{"trusted":true},"cell_type":"code","source":"# How is no. of words in Ad description affect Ad CTR...\n\n# Preparing data...\n\ntemp = data[['DCount', 'CTR']].copy()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp[temp['DCount'] >= 40.0].shape[0] * 1.0 / data.shape[0] # only 0.02 percent datapoints use >= 40 words in ad desc.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('DCount').agg(['mean'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(5,5))\n\nplt.bar(result.index, result[('CTR', 'mean')],\n        color='red',\n        width=0.4)\nplt.xlabel('No. of words in a Ad Description')\nplt.ylabel('Avg. CTR')\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: No. of words in Ad description doesn't give a clear picture of Ad CTR."},{"metadata":{"trusted":true},"cell_type":"code","source":"# How does no. of words in Ad Title affect Ad CTR...\n\n# Preparing data...\n\ntemp = data[['TCount', 'CTR']].copy()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp[temp['TCount'] >= 25.0].shape[0] * 1.0 / data.shape[0] # only 0.0005 percent datapoints use >= 25 words in ad title.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('TCount').agg(['mean'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(5,5))\n\nplt.bar(result.index, result[('CTR', 'mean')],\n        color='red',\n        width=0.4)\nplt.xlabel('No. of words in a Ad Title')\nplt.ylabel('Avg. CTR')\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Avg. Ad CTR is more or less distributed uniformly with no. of words in Ad title."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Does Ad Impresions affect Ad Clicks... (more impressions mean more click ?)\n\n# Preparing data...\n\ntemp = data[['AdId', 'Impression', 'Click']].copy()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('AdId').agg(['mean'])\nresult.head(6)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x = result[('Impression', 'mean')]\ny = result[('Click', 'mean')]\nplt.scatter(x,\n            y,\n            c='green',\n            s=100,\n            marker='o',\n            edgecolor=None)\nplt.xlabel('No. of Impressions')\nplt.ylabel('No. of Clicks')\nplt.title('Relationship between Ad Impressions & Clicks')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: As no. of impressions of an advertisement inc. clicks are mostly ~ 0.\n\nThis indicates a very crucial aspect of human behaviour. As a user see the same ad again & again, they are less likely to click it."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us see how Gender of a user has an impact on Ad CTR\n\n# Preparing data...\n\ntemp = data[['Gender', 'CTR']].copy()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('Gender').agg(['mean'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(5,5))\n\nplt.bar(result.index, result[('CTR', 'mean')],\n        color='red',\n        width=0.3)\nplt.xlabel('Gender')\nplt.ylabel('Avg. CTR')\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# What about Age of a user...\n\n# Preparing data...\n\ntemp = data[['Age', 'CTR']].copy()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp[temp['Age'] > 4.0].shape[0] * 1.0 / data.shape[0] # 21 percent datapoints are in age group > 4.0","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('Age').agg(['mean'])\nresult.head(6)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Recall Age categories as follows:\n\n# '1' : (0, 12]\n# '2' : (12, 18]\n# '3' : (18, 24]\n# '4' : (24, 30]\n# '5' : (30,  40]\n# '6' : > 40.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(5,5))\n\nplt.bar(result.index, result[('CTR', 'mean')],\n        color='red',\n        width=0.3)\nplt.xlabel('Age')\nplt.ylabel('Avg. CTR')\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: An user in categories 5 & 6 has higher avg. CTR as compared to users in other categories."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Since users in Age categories 5 & 6 have a higher CTR, let us try to find out \n\n# how does gender of a user in category 5 & 6 affect CTR of an ad...\n\n# Preparing data...\n\ntemp = data[['Gender', 'Age', 'CTR']].copy()\ntemp = temp[(temp['Age'] == 5) | (temp['Age'] == 6)] # filter aged users.\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp[temp['Gender'] == 2.0].shape[0] * 1.0 / data.shape[0] # 9.3 percent users are old female.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp = temp[['Gender', 'CTR']].copy()\nresult = temp.groupby('Gender').agg(['mean'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(5,5))\n\nplt.bar(result.index, result[('CTR', 'mean')],\n        color='green',\n        width=0.5)\nplt.xlabel('Gender')\nplt.ylabel('Avg. CTR')\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Female users (2) in Age categories 5 & 6 are more likely to click an Ad as opposed to their male (1) counterparts."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us try to see if Ad position affects on Ad CTR...\n\n# Preparing data...\n\ntemp = data[['Pos', 'CTR']].copy()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('Pos').agg(['mean', 'count'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,4))\n\nplt.subplot(1, 2, 1)\nplt.bar(result.index, result[('CTR', 'mean')],\n        color='red')\nplt.xlabel('Position')\nplt.ylabel('Avg. CTR')\n\n\nplt.subplot(1, 2, 2)\nplt.bar(result.index, result[('CTR', 'count')],\n        color='green')\nplt.xlabel('Position')\nplt.ylabel('Frequency of Ads')\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Clearly, the CTR for an advertisement which has a low position (more visible to user) is higher as compared to CTR of an advertisement with higher position(not directly visible).\n\nTypically advertisement have lower position. [1,2]"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us try to see if depth of a search session has an affect on CTR\n\n# Preparing data...\n\ntemp = data[['Depth', 'CTR']].copy()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('Depth').agg(['mean', 'count'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,4))\n\nplt.subplot(1, 2, 1)\nplt.bar(result.index, result[('CTR', 'mean')],\n        color='red')\nplt.xlabel('Depth')\nplt.ylabel('Avg. CTR')\n\nplt.subplot(1, 2, 2)\nplt.bar(result.index, result[('CTR', 'count')],\n        color='green')\nplt.xlabel('Depth')\nplt.ylabel('Frequency of Ads')\n\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: \n\n1. Mostly depth of a Search Session is 2.\n2. If depth if high (3) avg. CTR falls. This means if there as no. of ads in a Search Session inc. avg. CTR dec."},{"metadata":{},"cell_type":"markdown","source":"<h3> 3.2.2 Studying the role of Advertiser </h3>"},{"metadata":{},"cell_type":"markdown","source":"<h4>3.2.2.1 : Studying the role of Advertiser based on Ad CTR </h4>"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# We divide the data into two categories, one corr. to Advertisers who have a high CTR on their Ads & other who don't.\n\n# Once we know who are the Advertisers with high CTR Ads, we can study how Ads by a high CTR Adv. differs from a Adv. \n\n# with low CTR Ads.\n\n# e.g. We can know if an advertiser with high CTR ads use more words to describe ad, more words in the ad title etc...\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Preparing data...\n\ntemp = data[['AdvId', 'CTR', 'DCount', 'TCount']].copy()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('AdvId').agg(['mean'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp = pd.DataFrame()\n\ntemp['AdvId']    = result.index\ntemp['CTR']      = result[('CTR', 'mean')].get_values()\ntemp['DCount']   = result[('DCount', 'mean')].get_values()\ntemp['TCount']   = result[('TCount', 'mean')].get_values()\n\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print( 'No. of unique advertisers: ',temp.shape[0] )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# The burning question...\n\n# How to decide if an Advertiser is a high CTR Advertiser ? \n\n# Let us study the distribution of avg. Ad CTRs corr. to Advertisers...","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, (ax1, ax2) = plt.subplots(2)\nsns.kdeplot(temp['CTR'], ax=ax1)\nsns.boxplot(x=None,y='CTR',data=temp, ax=ax2)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"mean_advertiser_ctr = temp['CTR'].mean()\nprint ('Average CTR of Ads given by an advertiser: ', round(mean_advertiser_ctr, 2))\n\nmedian_advertiser_ctr = temp['CTR'].median()\nprint( 'Median CTR of Ads given by an advertiser: ', round(median_advertiser_ctr, 2))\n\nthird_quantile_advertiser_ctr = temp['CTR'].quantile(0.75)\nprint( '3rd Quantile CTR of Ads given by an advertiser: ', round(third_quantile_advertiser_ctr, 2))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us define 'High CTR Advertiser' as - an advertiser whose ad CTR > 3rd quantile Advertiser CTR","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['HighCTR'] = temp['CTR'] > third_quantile_advertiser_ctr\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['HighCTR'].value_counts() # Clearly, an imbalanced dataset...","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Out of 13921 Advertisers, only 3464 Advertisers have Ads. with CTR > 5.13 %.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Does an Advertiser with high CTR Ads uses more words to describe their Ads...\n\nsns.boxplot(x='HighCTR', y='DCount', data=temp)\n#sns.violinplot(x='HighCTR', y='DCount', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['HighCTR', 'DCount']].copy()).groupby('HighCTR').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Median no. of words in the description of an Ad for high CTR advertiser (21.47) is slightly more as compared to a low CTR advertiser(21.02)."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Does an Advertiser with high CTR Ads uses more words in Ad title...\n\nsns.boxplot(x='HighCTR', y='TCount', data=temp)\n#sns.violinplot(x='HighCTR', y='DCount', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['HighCTR', 'TCount']].copy()).groupby('HighCTR').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Median no. of words in the title of an Ad for high CTR advertiser is slightly high (8.7) than a low CTR advertiser (8.5)."},{"metadata":{},"cell_type":"markdown","source":"##### Takeaway:\n\nAdvertisers who have high CTRs use alomost same median no. of words in the title & description of their ads as an advertiser with low CTR which is intuitive. Why ? Because typically limited display space is given to every Ad irrespective of the Advertiser. Then why some Advertisers have High CTR ? \n\nThere are various reasons we can think off:\n\n1. High Quality content in Ads\n2. Product sold by an Advertiser can have high demand when data was collected."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Do high CTR Advertisers have more impressions of their advertisements (i.e are they frequent advertisers)...\n\n# Intuition says they should be, lets find out...\n\ninterim = data[['AdvId','Impression']].copy()\ninterim.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = interim.groupby('AdvId').agg(['sum', 'count'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# How to intepret above figure ? Advertisements by Adv. Id 82 were displayed 5230 times across 3651 user queries.\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['Impression'] = result[('Impression', 'sum')].get_values()\ntemp['Count']      = result[('Impression', 'count')].get_values()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# d = temp[temp['Net_Impr'] < 2]\nsns.boxplot(x='HighCTR', y='Impression', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Clearly there are outliers in data. Let us take 3rd quantile value to be robust in our estimate.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['HighCTR', 'Impression']].copy()).groupby('HighCTR').agg(['mean', 'median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: A High CTR Advertiser has higher avg. impressions(590.0) as opposed to a low CTR Advertiser (381.0). However median no. of impressions for a High CTR Advertiser is lower (49) as opposed to a low CTR Advertiser(53)."},{"metadata":{},"cell_type":"markdown","source":"<h4>3.2.2.2 : Studying the role of Advertiser based on Ad Frequency </h4>"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# We divide the data into two categories, each corr. to frequent & infrequent Advertisers.\n\n# Once we have divided the data,we can study how Ads by a frequent Advertiser differs from a infrequent Advertiser.\n\n# e.g. We can know if an frequent advertiser use more words to describe ad, more words in the ad title etc...\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# The burning question...\n\n# How to decide if an Advertiser is a frequent Advertiser or not? \n\n# There are two ways we can do this.. based on 1. Impression 2. Count\n\n# Advertiser Impression: total no. of impressions of all Ads by an Adv. \n\n# Advertiser Count: total no. training entries all Ads by an Adv.\n\n# We choose Advertiser Impression as a criteria for deciding if an Advertiser is frequent or not.\n\n# Let us study the distribution of Advertiser Impressions... \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, (ax1, ax2) = plt.subplots(2)\nsns.kdeplot(temp['Impression'], ax=ax1)\nsns.boxplot(x=None,y='Impression',data=temp, ax=ax2)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"mean_advertiser_impression = temp['Impression'].mean()\nprint ('Average Advertiser Impression: ', round(mean_advertiser_impression, 2))\n\nmedian_advertiser_impression = temp['Impression'].median()\nprint ('Median Advertiser Impression: ', round(median_advertiser_impression, 2))\n\nthird_quantile_advertiser_impression = temp['Impression'].quantile(0.75)\nprint( '3rd Quantile Advertiser Impression: ', round(third_quantile_advertiser_impression, 2))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us define 'Frequent Advertiser' as - advertiser with Advertiser Impression > 3rd quantile Advertiser Impression.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['IsFrequent'] = temp['Count'] > third_quantile_advertiser_impression\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['IsFrequent'].value_counts() # Clearly, there is an imbalance since we chose 3rd quantile Adv. Impr. as threshold.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Does a frequent advertiser has higher avg. CTR...\n\n# Intuitively they should, lets investigate.\n\nsns.violinplot(x='IsFrequent',y='CTR',data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['IsFrequent', 'CTR']].copy()).groupby('IsFrequent').agg(['mean', 'median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: A frequent advertiser has higher median CTR(3.58%) & avg. CTR(4.14%) as compared to an infrequent advertiser with 1.58% median CTR & 3.8% avg. CTR. Why it that ? This is very intuitive. How ? \n\nAn Advertiser being frequent tantamounts to higher CTR else he/she wouldn't be frequent in the first place.\nWhy on earth would an advertiser want to show their Ads if more impressions isn't generating revenue for the advertiser."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Does a frequent Advertiser use more words to describe their Ads...\n\nsns.boxplot(x='IsFrequent', y='DCount', data=temp)\n#sns.violinplot(x='HighCTR', y='DCount', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['IsFrequent', 'DCount']].copy()).groupby('IsFrequent').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Median no. of words in the Ad description for frequent advertiser is slightly high (21.5) than a infrequent advertiser(21.0)."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Does a frequent Advertiser use more words to describe their Ads...\n\nsns.boxplot(x='IsFrequent', y='TCount', data=temp)\n#sns.violinplot(x='HighCTR', y='DCount', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['IsFrequent', 'TCount']].copy()).groupby('IsFrequent').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Median no. of words in the Ad Title for frequent advertiser is slightly high (8.9) than a infrequent advertiser(8.4)."},{"metadata":{},"cell_type":"markdown","source":"<h3> 3.2.3 Studying Ads irrespective of Advertisers </h3>"},{"metadata":{},"cell_type":"markdown","source":"<h4> 3.2.3.1 Studying Ad properties based on Ad CTR </h4>"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n\n# We divide the data into two categories, one corr. to Ads with high CTR & other corr. to Ads with low CTR.\n\n# Once this is done, we can investigate how an Ad with high CTR differs from an Ad with low CTR.\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# What is my goal ? To study properties of Advertisements.\n\n# What questions I intend to answer ? \n\n# 1. How does no. of words in Ad Description vary with CTR of an Ad ? \n\n# 2. How does no. of words in Ad Title vary with CTR of an Ad ? \n\n# 3. How is the Ad frequency related to no. of words used in Ad description ? \n\n# 4. How is the Ad frequency related to no. of words used in Ad title ? \n\n# 5. How is the Ad frequency related to Ad Position ? \n\n# 6. How is the Ad frequency related to Ad Depth ? \n\n# 7. How is the Ad frequency related to Ad Clicks ? Does more Ad impressions mean more clicks ? \n\n# 8. Does frequency of an Ad has an effect on CTR of the Ad ? \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Preparing Data for Analysis....\n\ntemp = data[['AdId', 'CTR', 'Pos', 'Depth', 'QCount', 'DCount', 'TCount']].copy()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"result = temp.groupby('AdId').agg(['mean'])\nresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp = pd.DataFrame()\n\ntemp['AdId']   = result.index\ntemp['CTR']    = result[('CTR', 'mean')].get_values()\ntemp['Pos']    = result[('Pos', 'mean')].get_values()\ntemp['Depth']  = result[('Depth', 'mean')].get_values()\ntemp['DCount'] = result[('DCount', 'mean')].get_values()\ntemp['TCount'] = result[('TCount', 'mean')].get_values()\n\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"interim = data[['AdId', 'Impression', 'Click']].copy()\niresult = interim.groupby('AdId').agg(['sum'])\niresult.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['Impression'] = iresult[('Impression', 'sum')].get_values()\ntemp['Click']      = iresult[('Click', 'sum')].get_values()\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print ('No. of unique advertisements: ',temp.shape[0] )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# The burning question...\n\n# How to decide if an Ad qualifies as a high CTR Advertisment ? \n\n# Let us study the distribution of avg. Ad CTRs...\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, (ax1, ax2) = plt.subplots(2)\nsns.kdeplot(temp['CTR'], ax=ax1)\nsns.boxplot(x=None,y='CTR',data=temp, ax=ax2)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"mean_ad_ctr = temp['CTR'].mean()\nprint ('Average CTR of Ads : ', round(mean_ad_ctr, 2))\n\nmedian_ad_ctr = temp['CTR'].median()\nprint ('Median CTR of Ads : ', round(median_ad_ctr, 2))\n\nthird_quantile_ad_ctr = temp['CTR'].quantile(0.75)\nprint ('3rd Quantile CTR of Ads: ', round(third_quantile_ad_ctr, 2))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Since median & 3rd quantile avg. CTR is 0.0 using it as a threshold is meaningless. So, let us use \n\n# Avg. Ad CTR as a threshold for deciding if an ad is qualified as a high CTR ad or not.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['HighCTR'] = temp['CTR'] > mean_ad_ctr\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['HighCTR'].value_counts() # Clearly, an imbalanced dataset...","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# How does Ad CTR affect no. of words in Ad description...\n\nsns.boxplot(x='HighCTR', y='DCount', data=temp)\n#sns.violinplot(x='HighCTR', y='DCount', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['HighCTR', 'DCount']].copy()).groupby('HighCTR').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Description of an Ad with high CTR typically has a higher median no. of words (22.0) as compared to a low CTR Ad(21.0)."},{"metadata":{"trusted":true},"cell_type":"code","source":"# How does Ad CTR affect no. of words in Ad Title...\n\nsns.boxplot(x='HighCTR', y='TCount', data=temp)\n# sns.violinplot(x='HighCTR', y='TCount', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['HighCTR', 'TCount']].copy()).groupby('HighCTR').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Title of an Ad with high CTR typically has equal median no. of words (9.0) as a low CTR Ad."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Do high CTR Ads have more no. of impressions...\n\n# Intuition says they should be, lets find out...\n\nsns.boxplot(x='HighCTR', y='Impression', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['HighCTR', 'Impression']].copy()).groupby('HighCTR').agg(['mean', 'median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: A High CTR Ad has higher avg. impressions(50.8) & median impressions(9.0) as opposed to a low CTR Ad with avg. impressions of 23.90 & median impressions of 3.0"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# Part 3.2 Studying Ad properties based on Ad frequency...\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# The burning question...\n\n# How to decide if an Ad is frequent or not? \n\n# There are two ways we can do this.. based on 1. Impression 2. Count\n\n# Ad Impression: total no. of impressions of an Ad across all entries in the training file. \n\n# Ad Count: total no. of training entries in which Ad appeared.\n\n# We choose Ad Impression as a criteria for deciding if an Ad is frequent or not.\n\n# Let us study the distribution of Ad Impressions... \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, (ax1, ax2) = plt.subplots(2)\nsns.kdeplot(temp['Impression'], ax=ax1)\nsns.boxplot(x=None,y='Impression',data=temp, ax=ax2)\n# sns.boxplot(x=None,y='Impression',data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"mean_ad_impression = temp['Impression'].mean()\nprint ('Avg. impresssions of an Ad: ', round(mean_ad_impression))\n\nmedian_ad_impression = temp['Impression'].median()\nprint ('Median impresssions of an Ad: ', round(median_ad_impression))\n\nthird_quantile_ad_impression = temp['Impression'].quantile(0.75)\nprint ('3rd quantile impresssions of an Ad: ', round(third_quantile_ad_impression))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let us define 'Frequent Ad' as - an Ad with Ad Impression > 3rd quantile Ad Impression.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['IsFrequent'] = temp['Impression'] > third_quantile_ad_impression\ntemp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"temp['IsFrequent'].value_counts() # Dataset is balanced.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# How does Ad frequency affect no. of words in Ad description...\n\nsns.boxplot(x='IsFrequent', y='DCount', data=temp)\n#sns.violinplot(x='HighCTR', y='DCount', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['IsFrequent', 'DCount']].copy()).groupby('IsFrequent').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Description of a frequent Ad has a slightly higher median no. of words (22.0) as compared to a infrequent Ad (21.0)."},{"metadata":{"trusted":true},"cell_type":"code","source":"# How does Ad frequency affect no. of words in Ad title...\n\nsns.boxplot(x='IsFrequent', y='TCount', data=temp)\n#sns.violinplot(x='HighCTR', y='DCount', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(temp[['IsFrequent', 'TCount']].copy()).groupby('IsFrequent').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: Title of a frequent Ad has a almost equal median no. of words (8.7) as compared to a infrequent Ad (9.0).m"},{"metadata":{"trusted":true},"cell_type":"code","source":"# How does Ad frequency affect Ad position..\n\nsns.FacetGrid(temp, hue=\"IsFrequent\", size=5) \\\n   .map(sns.kdeplot, \"Pos\") \\\n   .add_legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: infrequent & frequenct ads usually occupy similar positions which is mostly positions 1 & 2."},{"metadata":{"trusted":true},"cell_type":"code","source":"# How does Ad frequency affect Ad Depth...\n\nsns.FacetGrid(temp, hue=\"IsFrequent\", size=5) \\\n   .map(sns.kdeplot, \"Depth\") \\\n   .add_legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Conclusion: infrequent advertisement occur mostly in search sessions with depths is 2. On the other hand occurence of a frequent advertisement is distributed normally (not exactly) across all depths."},{"metadata":{},"cell_type":"markdown","source":"Conclusion: If an advertisement is infrequent it has less no. of clicks as compared to an frequent advertisement. \nWhy is that ? It is very intuitive. How ? Ad is infrequent in the first place becuase it has low CTR hence low no. of clicks.\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Do frequent Ads have higher CTR...\n\n# Intuition says they should be, lets find out...\n\nsns.boxplot(x='IsFrequent', y='CTR', data=temp)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Clearly there are outliers so we will use median CTR to decide this...\n\n(temp[['IsFrequent', 'CTR']].copy()).groupby('IsFrequent').agg(['median'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Frequent Ads have higher median CTR (2.36) as opposed to infrequent Ads (0.0) which is intuitive."},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# Takeaways:\n\n# 1. Ads at lower positions have higher avg. CTR so we have included a feature mPosCTR.\n\n# 2. Frequent advertisers have higher avg. CTR as we included feature mAdvCTR.\n\n# 3. Frequents ads have higher avg. CTR so we have included feature mAdCTR. \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}