{
  "id": 1776,
  "title": "Understanding the AUC",
  "url": "/competitions/kddcup2012-track2/discussion/1776",
  "author_name": "",
  "post_date": "2012-04-25T16:55:25.467Z",
  "votes": null,
  "comment_count": 7,
  "views": 8431,
  "content": "<p>I don't understand the AUC metric in this contest. The paper referenced in the evaluation page only talks about positive and negative examples, and here we have clicks and impressions. How's the AUC applied here ?</p>",
  "messages": [
    {
      "id": "10375",
      "postDate": "04/25/2012 16:55:25",
      "content": "<p>I don't understand the AUC metric in this contest. The paper referenced in the evaluation page only talks about positive and negative examples, and here we have clicks and impressions. How's the AUC applied here ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "10376",
      "postDate": "04/25/2012 18:07:26",
      "content": "<p>Each impression has a binary response - click or don't click. However, sometimes the ad is shown again to the same user at the same depth. In these cases, these test instances are aggregated together, and you don't know the # of impressions.</p>\r\n<p>When you make a prediction (say 0.4) on a test instance, this may correspond to making a prediction of 0.4 on 100 binary responses, of which 38 are positive responses (clicks) and 62 are negative.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "10417",
      "postDate": "04/26/2012 22:08:09",
      "content": "<p>[quote=Ben Hamner;10376]When you make a prediction (say 0.4) on a test instance, this may correspond to making a prediction of 0.4 on 100 binary responses, of which 38 are positive responses (clicks) and 62 are negative.</p>\r\n<p>[/quote]</p>\r\n<p>Yes I can see this part, but still how do you calculate AUC here ?</p>\r\n<p>For example, let's say we have:</p>\r\n<p>A: 30 clicks, 70 non-clicks, prediction 0.65</p>\r\n<p>B: 30 clicks, 70 non-clicks, prediction 0.75</p>\r\n<p>Exactly how do you calculate AUC ?</p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "10420",
      "postDate": "04/26/2012 22:27:57",
      "content": "<p>From there, you treat that as making 200 predictions:</p>\r\n<p>30 positive instances where you predict 0.65<br>\r\n70 negative instances where you predict 0.65<br>\r\n30 positive instances where you predict 0.75<br>\r\n70 negative instances where you predict 0.75,</p>\r\n<p>and calculate AUC in the standard way across the 200 predictions.</p>\r\n<pre>def scoreClickAUC(num_clicks, num_impressions, predicted_ctr):\r\n    &quot;&quot;&quot;\r\n    Calculates the area under the ROC curve (AUC) for click rates\r\n\r\n    Parameters\r\n    ----------\r\n    num_clicks : a list containing the number of clicks\r\n\r\n    num_impressions : a list containing the number of impressions\r\n\r\n    predicted_ctr : a list containing the predicted click-through rates\r\n\r\n    Returns\r\n    -------\r\n    auc : the area under the ROC curve (AUC) for click rates\r\n    &quot;&quot;&quot;\r\n    i_sorted = sorted(range(len(predicted_ctr)),key=lambda i: predicted_ctr[i],\r\n                      reverse=True)\r\n    auc_temp = 0.0\r\n    click_sum = 0.0\r\n    old_click_sum = 0.0\r\n    no_click = 0.0\r\n    no_click_sum = 0.0\r\n\r\n    # treat all instances with the same predicted_ctr as coming from the\r\n    # same bucket\r\n    last_ctr = predicted_ctr[i_sorted[0]] &#43; 1.0\r\n\r\n    for i in range(len(predicted_ctr)):\r\n        if last_ctr != predicted_ctr[i_sorted[i]]: \r\n            auc_temp &#43;= (click_sum&#43;old_click_sum) * no_click / 2.0        \r\n            old_click_sum = click_sum\r\n            no_click = 0.0\r\n            last_ctr = predicted_ctr[i_sorted[i]]\r\n        no_click &#43;= num_impressions[i_sorted[i]] - num_clicks[i_sorted[i]]\r\n        no_click_sum &#43;= num_impressions[i_sorted[i]] - num_clicks[i_sorted[i]]\r\n        click_sum &#43;= num_clicks[i_sorted[i]]\r\n    auc_temp &#43;= (click_sum&#43;old_click_sum) * no_click / 2.0\r\n    auc = auc_temp / (click_sum * no_click_sum)\r\n    return auc\r\n</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "10869",
      "postDate": "05/08/2012 15:19:29",
      "content": "<p>it may be helpful to someone (an R fast implementation of AUC, one minute to predict 20 million rows):</p>\r\n<p>#y = actual response, x predicted<span id=\"x__plain_text_marker\">&nbsp;</span></p>\r\n<pre>score.AUC &lt;- function(y, x) {<br><br>  x1 = x[y==1]; n1 = as.numeric(length(x1))<br>  x2 = x[y==0]; n2 = as.numeric(length(x2))<br>  r = rank(c(x1,x2))  <br>  auc = (sum(r[1:n1]) - n1*(n1&#43;1)/2) / (n1*n2) <br>  return (auc)<br>}</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "10994",
      "postDate": "05/13/2012 11:17:19",
      "content": "<p>Hello Ben,</p>\r\n<p>Instead of predicting Click &amp; Impression for each record in test data and then the CTR, are we allowed to directly use CTR from training dataset as our target variable in regression and output CTR for test records?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "11082",
      "postDate": "05/17/2012 16:13:42",
      "content": "<p>Hey. Happened to see your name several times even from long time ago. Ensemble..</p>\r\n<p>Assuming your predictions as {P_i}, your classification don't need to be I (P_i &gt; 0.5). It could be I (P_i &gt; T ), where T could be any number within [0,1]. Therefore, for each T, you have a type-I error and type-II error. Draw a graph of type I vs. type\r\n II error, then the area below the curve is defined as AUC, which measures the discriminant power.</p>\r\n<p>&nbsp;</p>\r\n<p>[quote=B Yang;10375]</p>\r\n<p>I don't understand the AUC metric in this contest. The paper referenced in the evaluation page only talks about positive and negative examples, and here we have clicks and impressions. How's the AUC applied here ?</p>\r\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "11122",
      "postDate": "05/19/2012 12:45:03",
      "content": "<p>it is giving consistent results for me.</p>\r\n<p>Anyway, it may just be an aproximation, but its running time is way lower!</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 10376,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "04/25/2012 18:07:26",
      "content": "<p>Each impression has a binary response - click or don't click. However, sometimes the ad is shown again to the same user at the same depth. In these cases, these test instances are aggregated together, and you don't know the # of impressions.</p>\r\n<p>When you make a prediction (say 0.4) on a test instance, this may correspond to making a prediction of 0.4 on 100 binary responses, of which 38 are positive responses (clicks) and 62 are negative.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 10417,
      "author_name": "byang1",
      "author_url": "",
      "post_date": "04/26/2012 22:08:09",
      "content": "<p>[quote=Ben Hamner;10376]When you make a prediction (say 0.4) on a test instance, this may correspond to making a prediction of 0.4 on 100 binary responses, of which 38 are positive responses (clicks) and 62 are negative.</p>\r\n<p>[/quote]</p>\r\n<p>Yes I can see this part, but still how do you calculate AUC here ?</p>\r\n<p>For example, let's say we have:</p>\r\n<p>A: 30 clicks, 70 non-clicks, prediction 0.65</p>\r\n<p>B: 30 clicks, 70 non-clicks, prediction 0.75</p>\r\n<p>Exactly how do you calculate AUC ?</p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 10420,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "04/26/2012 22:27:57",
      "content": "<p>From there, you treat that as making 200 predictions:</p>\r\n<p>30 positive instances where you predict 0.65<br>\r\n70 negative instances where you predict 0.65<br>\r\n30 positive instances where you predict 0.75<br>\r\n70 negative instances where you predict 0.75,</p>\r\n<p>and calculate AUC in the standard way across the 200 predictions.</p>\r\n<pre>def scoreClickAUC(num_clicks, num_impressions, predicted_ctr):\r\n    &quot;&quot;&quot;\r\n    Calculates the area under the ROC curve (AUC) for click rates\r\n\r\n    Parameters\r\n    ----------\r\n    num_clicks : a list containing the number of clicks\r\n\r\n    num_impressions : a list containing the number of impressions\r\n\r\n    predicted_ctr : a list containing the predicted click-through rates\r\n\r\n    Returns\r\n    -------\r\n    auc : the area under the ROC curve (AUC) for click rates\r\n    &quot;&quot;&quot;\r\n    i_sorted = sorted(range(len(predicted_ctr)),key=lambda i: predicted_ctr[i],\r\n                      reverse=True)\r\n    auc_temp = 0.0\r\n    click_sum = 0.0\r\n    old_click_sum = 0.0\r\n    no_click = 0.0\r\n    no_click_sum = 0.0\r\n\r\n    # treat all instances with the same predicted_ctr as coming from the\r\n    # same bucket\r\n    last_ctr = predicted_ctr[i_sorted[0]] &#43; 1.0\r\n\r\n    for i in range(len(predicted_ctr)):\r\n        if last_ctr != predicted_ctr[i_sorted[i]]: \r\n            auc_temp &#43;= (click_sum&#43;old_click_sum) * no_click / 2.0        \r\n            old_click_sum = click_sum\r\n            no_click = 0.0\r\n            last_ctr = predicted_ctr[i_sorted[i]]\r\n        no_click &#43;= num_impressions[i_sorted[i]] - num_clicks[i_sorted[i]]\r\n        no_click_sum &#43;= num_impressions[i_sorted[i]] - num_clicks[i_sorted[i]]\r\n        click_sum &#43;= num_clicks[i_sorted[i]]\r\n    auc_temp &#43;= (click_sum&#43;old_click_sum) * no_click / 2.0\r\n    auc = auc_temp / (click_sum * no_click_sum)\r\n    return auc\r\n</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 10869,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "05/08/2012 15:19:29",
      "content": "<p>it may be helpful to someone (an R fast implementation of AUC, one minute to predict 20 million rows):</p>\r\n<p>#y = actual response, x predicted<span id=\"x__plain_text_marker\">&nbsp;</span></p>\r\n<pre>score.AUC &lt;- function(y, x) {<br><br>  x1 = x[y==1]; n1 = as.numeric(length(x1))<br>  x2 = x[y==0]; n2 = as.numeric(length(x2))<br>  r = rank(c(x1,x2))  <br>  auc = (sum(r[1:n1]) - n1*(n1&#43;1)/2) / (n1*n2) <br>  return (auc)<br>}</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 10994,
      "author_name": "sashikanthdareddy",
      "author_url": "",
      "post_date": "05/13/2012 11:17:19",
      "content": "<p>Hello Ben,</p>\r\n<p>Instead of predicting Click &amp; Impression for each record in test data and then the CTR, are we allowed to directly use CTR from training dataset as our target variable in regression and output CTR for test records?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 11082,
      "author_name": "nanzhou",
      "author_url": "",
      "post_date": "05/17/2012 16:13:42",
      "content": "<p>Hey. Happened to see your name several times even from long time ago. Ensemble..</p>\r\n<p>Assuming your predictions as {P_i}, your classification don't need to be I (P_i &gt; 0.5). It could be I (P_i &gt; T ), where T could be any number within [0,1]. Therefore, for each T, you have a type-I error and type-II error. Draw a graph of type I vs. type\r\n II error, then the area below the curve is defined as AUC, which measures the discriminant power.</p>\r\n<p>&nbsp;</p>\r\n<p>[quote=B Yang;10375]</p>\r\n<p>I don't understand the AUC metric in this contest. The paper referenced in the evaluation page only talks about positive and negative examples, and here we have clicks and impressions. How's the AUC applied here ?</p>\r\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 11122,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "05/19/2012 12:45:03",
      "content": "<p>it is giving consistent results for me.</p>\r\n<p>Anyway, it may just be an aproximation, but its running time is way lower!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "10375": "",
    "10376": "",
    "10417": "",
    "10420": "",
    "10869": "",
    "10994": "",
    "11082": "",
    "11122": ""
  },
  "source": "meta"
}