{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"},"cell_type":"markdown","source":"# Summary\n**This notebook illustrates the relation between the AUC score and the probability (p_hit) with which your model correctly identifies a download-click. It shall provide some intuition.**\n\n**In particular, to achieve *0.9552*, we need a probability of *p_hit = 2.7%*  (that our predictions of download-clicks are correct) for the very first clicks, that we predict as download-clicks (the ones we are most sure of)**. \n\nThis probability can and must fall to *p_hit = 0.24%* as we predict more and more clicks as download-clicks. Eventually we will and must predict that all clicks are download-clicks. 0.9552 currently corresponds to #2295 on the public leaderboard. To get 0.9827, which corresponds to #1 on the current public leaderboard, you would need to increase *p_hit = 2.7%* to *p_hit = 6.5%*\n\nConfusion potential: *p_hit* is different from the probability *p_click* that we assign to each click (the probability, that it is a download-click). *p_hit* is the observed hit-rate after we have assigned a probability to each click, and then fixed a threshold *p_threshold*, such that we predict that each click where *p_click >= p_threshold*, is a download-click.\n\n**Note: The numbers in this code are only illustrative. I didn't run any analysis on the actual training or test data. **\n\n\n# Factors to increase p_hit\nThe minimum rate of correct predictions should be around *p_hit = 0.24%* if we predict randomly (because in the training set around 0.24% of all clicks are download-clicks). \n\nSuppose that half of the clicks are fraudulent (=clicks which were made fraudulently and will surely not become downloads), and we could identify them all correctly. Then, we would just take the remainig clicks, and randomly pick any number and predict that they are download-clicks. This would yield a probability around *p_hit = 0.48%* ( = 2 * 0.24%). \n\nThe factors to increase *p_hit* are:\n* percentage of fraudulent clicks, which will surely not lead to a download\n* percentage of how many of the above fraudulent clicks we can identifiy (obviously, we will be able to identify some of those fraudulent clicks very reliably, and some others with less and less probability)\n* identify certain clicks among non-fraudulent clicks, which will yield a download more often than average (note: here we do not look at all at fraudulent clicks!)\n\n\n# How to plot the ROC curve\nLet R be the ROC curve. If (x,y) is an element of the set R, then\n* x = true positive rate = % (number of clicks correctly identified as leading to downloads / number of all clicks leading to downloads)\n* y = false positive rate = % (number of clicks falsely identified as leading to downloads / number of all clicks, which do NOT lead to downloads)\n\nThe following statements follow from the definition of the ROC curve:\n* The ROC curve starts at the point (x,y) = (0,0), thus at the beginning we must predict no clicks as downloads-clicks\n* The ROC curve ends at the point (x,y) = (1,1), thus at the end we have predicted that every click becomes an install, which leads us to have found every download-click (y=1), but at the same time we have predicted also every non-download-click as leading to a download (thus x=1)\n\nLet us start by fixing the assumed number of clicks:"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"trusted":true},"cell_type":"code","source":"number_non_download_clicks = 17960000  # illustrative number\nnumber_download_clicks = 0.0024 * number_non_download_clicks  # 0.0024 corresponds to 0.24% as observed in the training set\nnumber_all_clicks = number_non_download_clicks + number_download_clicks","execution_count":2,"outputs":[]},{"metadata":{"_uuid":"0f7ba7727293082dc3d0f9838a51e9ede9f5d0d8","_cell_guid":"c67a5160-5716-43a6-8b04-6ad44ff756d7"},"cell_type":"markdown","source":"We now fix the number *n(p)* of clicks, that we identify as download-clicks. The number will depend on the threshold *p = p_threshold*. Note that we must have *n(1) = 0* and *n(0) = number_all_clicks*.  At the same time, we define the targeted probability *p_hit* that whatever we predict as a download-click will indeed be one:"},{"metadata":{"_uuid":"75e2f037225985881b45db20412b54329a70bd0c","_cell_guid":"b29bcda3-e073-4136-a6b9-860fc8598604","collapsed":true,"trusted":true},"cell_type":"code","source":"def n(p): return (1-p**4)*number_all_clicks\ndef p_hit(p):\n    factor = 10  # use 10 for 0.9552 and 26 for 0.9827\n    prob_to_get_hit = number_download_clicks*(1+factor*p)/number_all_clicks  # this formula is an assumption, NOT a derivation\n    max_prob = number_download_clicks/n(p) if n(p)>0 else 0\n    prob_to_get_hit = min(prob_to_get_hit, max_prob)  # cant have too high a probability because otherwise will predict more correct ones as there are\n    return prob_to_get_hit\n\n# calculate x and y of ROC curve, for each probability threshold p = p_threshold\ndef x(p): return n(p) * (1-p_hit(p)) / number_non_download_clicks\ndef y(p): return n(p) * p_hit(p) / number_download_clicks","execution_count":18,"outputs":[]},{"metadata":{"_uuid":"a8989df22d071fda1dfe95671e8f682400272a09","_cell_guid":"206d62c5-2a72-4a59-b7b7-4814cd7bef78"},"cell_type":"markdown","source":"We proceed to plot the ROC-curve, by letting *p=p_threshold* run from 0 to 1. We also plot the graph for the number of clicks *n(p)*, that we predict as download-clicks, and the plot for the probability *p_hit(p)* that those are correct:"},{"metadata":{"_uuid":"933737929c79d3db1dd84b3a90d525b9157620a2","_cell_guid":"35d0661f-d48a-4e91-9d26-f9679f2b33e5","trusted":true},"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\n\ndelta = 0.0001\nx_values = []\ny_values = []\narea_under_the_roc_curve = 0\nx_equidistant = []\nfor p in np.arange(0.0, 1.0+delta, delta):\n    x_equidistant.append(p)\n    x_values.append(x(p))\n    y_values.append(y(p))\n    area_under_the_roc_curve += y(p) * (x(p) - x(p+delta))\nx_equidistant_rev = x_equidistant[::-1]\n\n\nplt.figure(figsize=(22,10))\n\nplt.subplot(131)\nplt.title('ROC-curve (area underneath is {:5.4f})'.format(area_under_the_roc_curve))\nplt.scatter(x_values, y_values, marker='.')\nplt.xlabel('x from 0 to 1 (as p goes from 1 to 0 !)')\n\nplt.subplot(132)\nplt.title('Probability p_hit that click is download-click')\nplt.scatter(x_equidistant_rev, [p_hit(p) for p in x_equidistant_rev], marker='.')\nplt.xlabel('p from 1 to 0')\nplt.xlim([1, 0])\n\nplt.subplot(133)\nplt.title('Number of clicks predicted as download-clicks')\nplt.scatter(x_equidistant_rev, [n(p) for p in x_equidistant_rev], marker='.')\nplt.xlabel('p from 1 to 0')\nplt.xlim([1, 0])\n\nplt.show()","execution_count":19,"outputs":[]},{"metadata":{"_uuid":"edb4a06db64eeee3d4848e78c625305ad369411d","_cell_guid":"41011636-dec2-445b-a0b7-03c6e75463e2"},"cell_type":"markdown","source":"Please let me know if you have any questions or comments. Corrections and other explanations are more than welcome!\n\nhth"},{"metadata":{"_uuid":"fff9e4c134ee522d2f31ddc4db74644e6da86260","_cell_guid":"e585d54e-e613-4bd6-9417-726549b69c4e","collapsed":true,"trusted":false},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}