{
  "id": 1721,
  "title": "rank by AUC only?",
  "url": "/competitions/kddcup2012-track2/discussion/1721",
  "author_name": "",
  "post_date": "2012-04-13T21:15:23.550Z",
  "votes": null,
  "comment_count": 2,
  "views": 4209,
  "content": "<p>solution file:</p>\r\n<pre>0.5,10<br>0.7,10<br>0.5,10<br>0.1,10</pre>\r\n<p>submission file</p>\r\n<pre>0.5<br>0.7<br>0.5<br>0.1</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>AUC = 0.727273</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>for</pre>\r\n<p>submission file</p>\r\n<pre>0.5<br>0.8<br>0.5<br>0.1</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>AUC = 0.727273</pre>\r\n<pre>&nbsp;</pre>\r\n<p>I can't get more then that on AUC result. Changes in the submission takes effect to AUC only when change the order of lines&nbsp;</p>\r\n<p>and order of lines correlated to solution file. And this is correct to python scoring code, but it don't measure performance as it was mention before on forum</p>\r\n<p>The perfect classifier doesn't have AUC=1.0 .</p>\r\n<p>How do you want to achieve fair rank?</p>\r\n&nbsp;\r\n<p>Another thing is that training data has 95% clicks=0 and test data is far from that as you can see how much score has ad_id_benchmark,&nbsp;</p>\r\n<p>which in my opinion is <span>ridiculous. As I understand it gives answers based on ad id history. How you can build classfier based on unique property...</span></p>\r\n<p>It is classical example of overfitting, and giving to it new data it will have 0% of correctness.</p>",
  "messages": [
    {
      "id": "10120",
      "postDate": "04/13/2012 21:15:23",
      "content": "<p>solution file:</p>\r\n<pre>0.5,10<br>0.7,10<br>0.5,10<br>0.1,10</pre>\r\n<p>submission file</p>\r\n<pre>0.5<br>0.7<br>0.5<br>0.1</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>AUC = 0.727273</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>for</pre>\r\n<p>submission file</p>\r\n<pre>0.5<br>0.8<br>0.5<br>0.1</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>AUC = 0.727273</pre>\r\n<pre>&nbsp;</pre>\r\n<p>I can't get more then that on AUC result. Changes in the submission takes effect to AUC only when change the order of lines&nbsp;</p>\r\n<p>and order of lines correlated to solution file. And this is correct to python scoring code, but it don't measure performance as it was mention before on forum</p>\r\n<p>The perfect classifier doesn't have AUC=1.0 .</p>\r\n<p>How do you want to achieve fair rank?</p>\r\n&nbsp;\r\n<p>Another thing is that training data has 95% clicks=0 and test data is far from that as you can see how much score has ad_id_benchmark,&nbsp;</p>\r\n<p>which in my opinion is <span>ridiculous. As I understand it gives answers based on ad id history. How you can build classfier based on unique property...</span></p>\r\n<p>It is classical example of overfitting, and giving to it new data it will have 0% of correctness.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "10121",
      "postDate": "04/13/2012 21:29:43",
      "content": "<p>[quote=kurak38;10120]</p>\r\n<p>The perfect classifier doesn't have AUC=1.0 .</p>\r\n<p>[/quote]No, but it is really close to 1.0 because of the sparsity of the clicks.[/quote]</p>\r\n<p>[quote=kurak38;10120]As I understand it gives answers based on ad id history. How you can build classfier based on unique property...</p>\r\n<p>It is classical example of overfitting, and giving to it new data it will have 0% of correctness.[/quote]Each ad may be shown many times to different users, and some ads are more likely to be clicked on than others. &nbsp;The ad id refers to the individual ad,\r\n which may correspond to many training or test samples. Thus it is not a unique property of the sample, and we'd expect an estimated CTR based on the ad's previous CTR to generalize to future samples. This is seen in the ad id benchmark, where predicting the\r\n CTR based only on the ad's previous CTR performs far above random on the validation instances.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "10122",
      "postDate": "04/13/2012 22:41:10",
      "content": "<p>I know that ad id is not unique, but ids in databases are. The meaning to introduce ids is to create unique property, is has no information about test case. At any data mining class it is pointed out that ids should not be considered. If you have new ad\r\n in system how you can predict anything based on new id. It is obvious that ids are poor, but why is has some many score.<br>\r\nIf you want to predict CTR of ad based on this particular ad history is it not data mining but statistic related field in my opinion.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 10121,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "04/13/2012 21:29:43",
      "content": "<p>[quote=kurak38;10120]</p>\r\n<p>The perfect classifier doesn't have AUC=1.0 .</p>\r\n<p>[/quote]No, but it is really close to 1.0 because of the sparsity of the clicks.[/quote]</p>\r\n<p>[quote=kurak38;10120]As I understand it gives answers based on ad id history. How you can build classfier based on unique property...</p>\r\n<p>It is classical example of overfitting, and giving to it new data it will have 0% of correctness.[/quote]Each ad may be shown many times to different users, and some ads are more likely to be clicked on than others. &nbsp;The ad id refers to the individual ad,\r\n which may correspond to many training or test samples. Thus it is not a unique property of the sample, and we'd expect an estimated CTR based on the ad's previous CTR to generalize to future samples. This is seen in the ad id benchmark, where predicting the\r\n CTR based only on the ad's previous CTR performs far above random on the validation instances.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 10122,
      "author_name": "kurak38",
      "author_url": "",
      "post_date": "04/13/2012 22:41:10",
      "content": "<p>I know that ad id is not unique, but ids in databases are. The meaning to introduce ids is to create unique property, is has no information about test case. At any data mining class it is pointed out that ids should not be considered. If you have new ad\r\n in system how you can predict anything based on new id. It is obvious that ids are poor, but why is has some many score.<br>\r\nIf you want to predict CTR of ad based on this particular ad history is it not data mining but statistic related field in my opinion.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "10120": "",
    "10121": "",
    "10122": ""
  },
  "source": "meta"
}