{
  "id": 20765,
  "title": "CV and LB scores",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/20765",
  "author_name": "",
  "post_date": "2016-05-06T12:27:10.843Z",
  "votes": 1,
  "comment_count": 40,
  "views": 5466,
  "content": "<p>My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?</p>",
  "messages": [
    {
      "id": "118966",
      "postDate": "05/06/2016 12:27:10",
      "content": "<p>My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?</p>",
      "rawMarkdown": "My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?",
      "votes": null
    },
    {
      "id": "118972",
      "postDate": "05/06/2016 13:04:37",
      "content": "<p>[quote=Abhishek;118966]</p>\n\n<p>My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?</p>\n\n<p>[/quote]</p>\n\n<p>CV = 0.75\nLB = 0.744</p>",
      "rawMarkdown": "[quote=Abhishek;118966]\r\n\r\nMy validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?\r\n\r\n[/quote]\r\n\r\nCV = 0.75\r\nLB = 0.744",
      "votes": null
    },
    {
      "id": "119075",
      "postDate": "05/07/2016 01:44:52",
      "content": "<p>[quote=Abhishek;118966]</p>\n\n<p>My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?</p>\n\n<p>[/quote]\nI get a similar behaviour with xgboost.</p>",
      "rawMarkdown": "[quote=Abhishek;118966]\r\n\r\nMy validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?\r\n\r\n[/quote]\r\nI get a similar behaviour with xgboost.",
      "votes": null
    },
    {
      "id": "119081",
      "postDate": "05/07/2016 02:53:07",
      "content": "<p>Not that much offset here:</p>\n\n<p>CV= 0.789824, LB=0.74889</p>",
      "rawMarkdown": "Not that much offset here:\r\n\r\nCV= 0.789824, LB=0.74889",
      "votes": null
    },
    {
      "id": "119103",
      "postDate": "05/07/2016 08:08:10",
      "content": "<p>Same problem here:</p>\n\n<p>CV=0.87, LB=0.77</p>",
      "rawMarkdown": "Same problem here:\r\n\r\nCV=0.87, LB=0.77",
      "votes": null
    },
    {
      "id": "119264",
      "postDate": "05/08/2016 17:01:55",
      "content": "<p>People are <a href=\"https://www.kaggle.com/c/santander-customer-satisfaction/forums/t/20725/kaggle-auc-sklearn-auc-was-a-factor-in-lb-shake-up\">saying</a> that kaggle AUC and sklearn AUC give different results. What is your CV using kaggle AUC?</p>",
      "rawMarkdown": "People are [saying][1] that kaggle AUC and sklearn AUC give different results. What is your CV using kaggle AUC?\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/santander-customer-satisfaction/forums/t/20725/kaggle-auc-sklearn-auc-was-a-factor-in-lb-shake-up",
      "votes": null
    },
    {
      "id": "119277",
      "postDate": "05/08/2016 19:06:41",
      "content": "<p>Public LB: 0.745</p>\n\n<p>5-fold stratified CV: 0.761</p>\n\n<p>Not very much over here... Are you guys using stratified CV?</p>",
      "rawMarkdown": "Public LB: 0.745\r\n\r\n5-fold stratified CV: 0.761\r\n\r\nNot very much over here... Are you guys using stratified CV?",
      "votes": null
    },
    {
      "id": "119279",
      "postDate": "05/08/2016 19:22:53",
      "content": "<p>Same to anokas, I am using 5-fold stratified split.</p>",
      "rawMarkdown": "Same to anokas, I am using 5-fold stratified split.",
      "votes": null
    },
    {
      "id": "119287",
      "postDate": "05/08/2016 21:36:45",
      "content": "<p>I confirm that the problem comes from the use of sklearn's AUC. </p>",
      "rawMarkdown": "I confirm that the problem comes from the use of sklearn's AUC.",
      "votes": null
    },
    {
      "id": "119364",
      "postDate": "05/09/2016 14:54:44",
      "content": "<p>[quote=alexandreyc;119287]</p>\n\n<p>I confirm that the problem comes from the use of sklearn's AUC. </p>\n\n<p>[/quote]\nI was following the bug report on sklearn but could not any release version with the fix. So - the ones who are not using sklearn AUC - which implementation are you guys using ?</p>",
      "rawMarkdown": "[quote=alexandreyc;119287]\r\n\r\nI confirm that the problem comes from the use of sklearn's AUC. \r\n\r\n[/quote]\r\nI was following the bug report on sklearn but could not any release version with the fix. So - the ones who are not using sklearn AUC - which implementation are you guys using ?",
      "votes": null
    },
    {
      "id": "119421",
      "postDate": "05/10/2016 06:26:09",
      "content": "<p>I am using this one</p>\n\n<p><a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/auc.py\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/auc.py</a></p>\n\n<p>however I still have a huge gap between LB-CV</p>",
      "rawMarkdown": "I am using this one\r\n\r\nhttps://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/auc.py\r\n\r\nhowever I still have a huge gap between LB-CV",
      "votes": null
    },
    {
      "id": "119425",
      "postDate": "05/10/2016 06:43:22",
      "content": "<p>The gap in AUC may be because the test set is sampled at a different time to the training set, meaning that the problem may be slightly different.</p>",
      "rawMarkdown": "The gap in AUC may be because the test set is sampled at a different time to the training set, meaning that the problem may be slightly different.",
      "votes": null
    },
    {
      "id": "119491",
      "postDate": "05/10/2016 19:34:14",
      "content": "<p>CV ~0.87 LB ~0.77  </p>\n\n<p>This is not AUC&#180;s fault, it's something with the data.</p>\n\n<p>@anokas\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that&#180;s part of the problem.</p>",
      "rawMarkdown": "CV ~0.87 LB ~0.77  \r\n\r\nThis is not AUC´s fault, it's something with the data.\r\n\r\n@anokas\r\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that´s part of the problem.",
      "votes": null
    },
    {
      "id": "119492",
      "postDate": "05/10/2016 20:05:37",
      "content": "<p>[quote=Snow Dog;119491]</p>\n\n<p>CV ~0.87 LB ~0.77  </p>\n\n<p>This is not AUC&#180;s fault, it's something with the data.</p>\n\n<p>@anokas\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that&#180;s part of the problem.</p>\n\n<p>[/quote]</p>\n\n<p>I mean that Avito didn't use a random split for creating the training and testing datasets, but rather that they are from different points in time eg. we are training on 2014 data and predicting 2015 data. This could cause a slight difference in data.</p>",
      "rawMarkdown": "[quote=Snow Dog;119491]\r\n\r\nCV ~0.87 LB ~0.77  \r\n\r\nThis is not AUC´s fault, it's something with the data.\r\n\r\n@anokas\r\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that´s part of the problem.\r\n\r\n[/quote]\r\n\r\nI mean that Avito didn't use a random split for creating the training and testing datasets, but rather that they are from different points in time eg. we are training on 2014 data and predicting 2015 data. This could cause a slight difference in data.",
      "votes": null
    },
    {
      "id": "119493",
      "postDate": "05/10/2016 20:13:40",
      "content": "<p>[quote=Snow Dog;119491]</p>\n\n<p>CV ~0.87 LB ~0.77  </p>\n\n<p>This is not AUC&#180;s fault, it's something with the data.</p>\n\n<p>@anokas\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that&#180;s part of the problem.</p>\n\n<p>[/quote]</p>\n\n<p>My CV and LB are quite close. <br>\nCV - 0.77\nLB - 0.76</p>",
      "rawMarkdown": "[quote=Snow Dog;119491]\r\n\r\nCV ~0.87 LB ~0.77  \r\n\r\nThis is not AUC´s fault, it's something with the data.\r\n\r\n@anokas\r\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that´s part of the problem.\r\n\r\n[/quote]\r\n\r\nMy CV and LB are quite close.  \r\nCV - 0.77\r\nLB - 0.76",
      "votes": null
    },
    {
      "id": "119494",
      "postDate": "05/10/2016 20:26:52",
      "content": "<p>@DataGeek</p>\n\n<p>Are you doing something special, besides stratifying the target? I need to use all the data and see what happens.</p>",
      "rawMarkdown": "DataGeek\r\n\r\nAre you doing something special, besides stratifying the target? I need to use all the data and see what happens.",
      "votes": null
    },
    {
      "id": "119495",
      "postDate": "05/10/2016 20:30:10",
      "content": "<p>[quote=Snow Dog;119494]</p>\n\n<p>@DataGeek</p>\n\n<p>Are you doing something special, besides stratifying the target? I need to use all the data and see what happens.</p>\n\n<p>[/quote]</p>\n\n<p>I think my CV and LB are close because I am using similarity b/w features of 2 ads instead of features directly. I haven't tested but most of the people having large differences in CV and LB must be using raw features + similarity features or just the raw features.</p>",
      "rawMarkdown": "[quote=Snow Dog;119494]\r\n\r\n@DataGeek\r\n\r\nAre you doing something special, besides stratifying the target? I need to use all the data and see what happens.\r\n\r\n\r\n[/quote]\r\n\r\nI think my CV and LB are close because I am using similarity b/w features of 2 ads instead of features directly. I haven't tested but most of the people having large differences in CV and LB must be using raw features + similarity features or just the raw features.",
      "votes": null
    },
    {
      "id": "119497",
      "postDate": "05/10/2016 20:44:46",
      "content": "<p>@Data Geek</p>\n\n<p>Thanks, going to check that! </p>",
      "rawMarkdown": "Data Geek\r\n\r\nThanks, going to check that!",
      "votes": null
    },
    {
      "id": "119711",
      "postDate": "05/12/2016 12:52:45",
      "content": "<p>[quote=Abhishek;118966]</p>\n\n<p>My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?</p>\n\n<p>[/quote]</p>\n\n<p>Hi there! I found out that shuffling is very important. Here are my results from logistic regression on 4 very simple features. First, results from 5-fold stratified split (from sklearn) without shuffling:</p>\n\n<ul>\n<li>AUC: 0.7661</li>\n<li>AUC: 0.7805</li>\n<li>AUC: 0.7343</li>\n<li>AUC: 0.6626</li>\n<li>AUC: 0.6555</li>\n</ul>\n\n<p>And results with shuffling:</p>\n\n<ul>\n<li>AUC: 0.7184</li>\n<li>AUC: 0.7179 </li>\n<li>AUC: 0.7185</li>\n<li>AUC: 0.7186</li>\n<li>AUC: 0.7193</li>\n</ul>\n\n<p>You can see, that with shuffling, results are much more stable and reliable. It's clear evidence that the data has a time structure, so I suppose you should split the data not randomly. First N% rows for traininig, remaining are for validation. It's also clear because of decrease of AUC with moving validation window further to the end of the data.</p>",
      "rawMarkdown": "[quote=Abhishek;118966]\r\n\r\nMy validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?\r\n\r\n[/quote]\r\n\r\nHi there! I found out that shuffling is very important. Here are my results from logistic regression on 4 very simple features. First, results from 5-fold stratified split (from sklearn) without shuffling:\r\n\r\n- AUC: 0.7661\r\n- AUC: 0.7805\r\n- AUC: 0.7343\r\n- AUC: 0.6626\r\n- AUC: 0.6555\r\n\r\nAnd results with shuffling:\r\n\r\n- AUC: 0.7184\r\n- AUC: 0.7179 \r\n- AUC: 0.7185\r\n- AUC: 0.7186\r\n- AUC: 0.7193\r\n\r\nYou can see, that with shuffling, results are much more stable and reliable. It's clear evidence that the data has a time structure, so I suppose you should split the data not randomly. First N% rows for traininig, remaining are for validation. It's also clear because of decrease of AUC with moving validation window further to the end of the data.",
      "votes": null
    },
    {
      "id": "119714",
      "postDate": "05/12/2016 13:01:15",
      "content": "<p>[quote=Ilya;119711]</p>\n\n<p>Hi there! I found out that shuffling is very important. Here are my results from logistic regression on 4 very simple features. First, results from 5-fold stratified split (from sklearn) without shuffling:</p>\n\n<ul>\n<li>AUC: 0.7661</li>\n<li>AUC: 0.7805</li>\n<li>AUC: 0.7343</li>\n<li>AUC: 0.6626</li>\n<li>AUC: 0.6555</li>\n</ul>\n\n<p>And results with shuffling:</p>\n\n<ul>\n<li>AUC: 0.7184</li>\n<li>AUC: 0.7179 </li>\n<li>AUC: 0.7185</li>\n<li>AUC: 0.7186</li>\n<li>AUC: 0.7193</li>\n</ul>\n\n<p>You can see, that with shuffling, results are much more stable and reliable. It's clear evidence that the data has a time structure, so I suppose you should split the data not randomly. First N% rows for traininig, remaining are for validation. It's also clear because of decrease of AUC with moving validation window further to the end of the data.</p>\n\n<p>[/quote]\nNice find, mate :)</p>",
      "rawMarkdown": "[quote=Ilya;119711]\r\n\r\nHi there! I found out that shuffling is very important. Here are my results from logistic regression on 4 very simple features. First, results from 5-fold stratified split (from sklearn) without shuffling:\r\n\r\n- AUC: 0.7661\r\n- AUC: 0.7805\r\n- AUC: 0.7343\r\n- AUC: 0.6626\r\n- AUC: 0.6555\r\n\r\nAnd results with shuffling:\r\n\r\n- AUC: 0.7184\r\n- AUC: 0.7179 \r\n- AUC: 0.7185\r\n- AUC: 0.7186\r\n- AUC: 0.7193\r\n\r\nYou can see, that with shuffling, results are much more stable and reliable. It's clear evidence that the data has a time structure, so I suppose you should split the data not randomly. First N% rows for traininig, remaining are for validation. It's also clear because of decrease of AUC with moving validation window further to the end of the data.\r\n\r\n[/quote]\r\nNice find, mate :)",
      "votes": null
    },
    {
      "id": "119789",
      "postDate": "05/12/2016 19:49:24",
      "content": "<p>I've found consistently in my testing that as you train more XGBoost rounds, the delta between leaderboard score and local CV increases, even if the deviation locally is very low and difference between train/cv auc is very low. </p>\n\n<p>I suspect this might be because the training and testing sets are from different dates, and so the content differs somewhat. Looking at the values in the dataset, there does seem to be quite a large difference between training/testing sets, while the data is very uniform within the training set, so this could be what is causing the difference between LB/CV scores.</p>\n\n<p>Just for reference, my current scores are</p>\n\n<p>CV: ~0.945 LB: ~0.905</p>",
      "rawMarkdown": "I've found consistently in my testing that as you train more XGBoost rounds, the delta between leaderboard score and local CV increases, even if the deviation locally is very low and difference between train/cv auc is very low. \r\n\r\nI suspect this might be because the training and testing sets are from different dates, and so the content differs somewhat. Looking at the values in the dataset, there does seem to be quite a large difference between training/testing sets, while the data is very uniform within the training set, so this could be what is causing the difference between LB/CV scores.\r\n\r\nJust for reference, my current scores are\r\n\r\nCV: ~0.945 LB: ~0.905",
      "votes": null
    },
    {
      "id": "119796",
      "postDate": "05/12/2016 20:13:47",
      "content": "<p>@anokas\nCongrats on being the first to beat the mark!</p>\n\n<p>I managed to decrease the gap, but still a large difference: CV: 0.82, LB: 0.76 (using half of the data).</p>",
      "rawMarkdown": "anokas\r\nCongrats on being the first to beat the mark!\r\n\r\nI managed to decrease the gap, but still a large difference: CV: 0.82, LB: 0.76 (using half of the data).",
      "votes": null
    },
    {
      "id": "120144",
      "postDate": "05/15/2016 21:02:47",
      "content": "<p>It seems that the CV - LB = 0.03~0.04. We can use this information to inter the difference between the distributions of duplicates in train/test datasets.  </p>",
      "rawMarkdown": "It seems that the CV - LB = 0.03~0.04. We can use this information to inter the difference between the distributions of duplicates in train/test datasets.",
      "votes": null
    },
    {
      "id": "120165",
      "postDate": "05/16/2016 00:05:04",
      "content": "<p>I made my first sub today using a very simple model (no attrsJSON/no images/no cat and location files/default xgb params) just to see how it would fare in LB. Got 0.91x in local validation and 0.81x in LB. I wonder if I'm overfitting or if when I start adding more features that difference will shrink.  </p>",
      "rawMarkdown": "I made my first sub today using a very simple model (no attrsJSON/no images/no cat and location files/default xgb params) just to see how it would fare in LB. Got 0.91x in local validation and 0.81x in LB. I wonder if I'm overfitting or if when I start adding more features that difference will shrink.",
      "votes": null
    },
    {
      "id": "120246",
      "postDate": "05/16/2016 17:43:18",
      "content": "<p>It seems like some of duplicates presented by more that one pair in training set, and some pair can get in train and some - in test set with random partitioning, so classifier learn to recognize concrete ads instead of really classify duplicates, so with increase of tree depth CV score increases, but LB don't increases (of course, these concrete ads didn't presented in LB dataset)</p>",
      "rawMarkdown": "It seems like some of duplicates presented by more that one pair in training set, and some pair can get in train and some - in test set with random partitioning, so classifier learn to recognize concrete ads instead of really classify duplicates, so with increase of tree depth CV score increases, but LB don't increases (of course, these concrete ads didn't presented in LB dataset)",
      "votes": null
    },
    {
      "id": "120500",
      "postDate": "05/18/2016 17:39:22",
      "content": "<p>[quote=ZavodRobotov;120246]</p>\n\n<p>It seems like some of duplicates presented by more that one pair in training set, and some pair can get in train and some - in test set with random partitioning, so classifier learn to recognize concrete ads instead of really classify duplicates, so with increase of tree depth CV score increases, but LB don't increases (of course, these concrete ads didn't presented in LB dataset)</p>\n\n<p>[/quote]\nI think you're right and that makes sense. As has DataGeek pointed out earlier, if you use raw features it can lead to some degree of overfitting because the model start recognizing the ads instead of the similarity between them. Using some raw features I currently get around 0.82x LB. Without using them, I get a lower score, but I think that is the way to go.</p>\n\n<p>Also, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.</p>",
      "rawMarkdown": "[quote=ZavodRobotov;120246]\r\n\r\nIt seems like some of duplicates presented by more that one pair in training set, and some pair can get in train and some - in test set with random partitioning, so classifier learn to recognize concrete ads instead of really classify duplicates, so with increase of tree depth CV score increases, but LB don't increases (of course, these concrete ads didn't presented in LB dataset)\r\n\r\n[/quote]\r\nI think you're right and that makes sense. As has DataGeek pointed out earlier, if you use raw features it can lead to some degree of overfitting because the model start recognizing the ads instead of the similarity between them. Using some raw features I currently get around 0.82x LB. Without using them, I get a lower score, but I think that is the way to go.\r\n\r\nAlso, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.",
      "votes": null
    },
    {
      "id": "120503",
      "postDate": "05/18/2016 17:50:08",
      "content": "<p>[quote=FernandoProcy;120500]</p>\n\n<p>Also, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.</p>\n\n<p>[/quote]</p>\n\n<p>I recommend hand-labeling, just to be on the safe side.</p>",
      "rawMarkdown": "[quote=FernandoProcy;120500]\r\n\r\nAlso, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.\r\n\r\n[/quote]\r\n\r\nI recommend hand-labeling, just to be on the safe side.",
      "votes": null
    },
    {
      "id": "120506",
      "postDate": "05/18/2016 18:31:47",
      "content": "<p>how many pictures are there ?</p>",
      "rawMarkdown": "how many pictures are there ?",
      "votes": null
    },
    {
      "id": "120508",
      "postDate": "05/18/2016 18:47:09",
      "content": "<p>[quote=eagle4;120506]</p>\n\n<p>how many pictures are there ?</p>\n\n<p>[/quote]</p>\n\n<p>~11,000,000</p>",
      "rawMarkdown": "[quote=eagle4;120506]\r\n\r\nhow many pictures are there ?\r\n\r\n[/quote]\r\n\r\n~11,000,000",
      "votes": null
    },
    {
      "id": "120509",
      "postDate": "05/18/2016 18:52:40",
      "content": "<p>@inversion: I was going to believe you since I always follow your advises on Kaggle forums...</p>\n\n<p>May I ask the leaders what the best score achieved without using the pictures ? </p>",
      "rawMarkdown": "inversion: I was going to believe you since I always follow your advises on Kaggle forums...\r\n\r\nMay I ask the leaders what the best score achieved without using the pictures ?",
      "votes": null
    },
    {
      "id": "120512",
      "postDate": "05/18/2016 19:12:59",
      "content": "<p>[quote=eagle4;120509]</p>\n\n<p>@inversion: I was going to believe you since I always follow your advises on Kaggle forums...</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, I was just messing around.</p>\n\n<p>For serious, though, I think (not 100% sure) that you can get in the top 10 without image features. (Or at least close). </p>",
      "rawMarkdown": "[quote=eagle4;120509]\r\n\r\n@inversion: I was going to believe you since I always follow your advises on Kaggle forums...\r\n\r\n[/quote]\r\n\r\nYeah, I was just messing around.\r\n\r\nFor serious, though, I think (not 100% sure) that you can get in the top 10 without image features. (Or at least close).",
      "votes": null
    },
    {
      "id": "120523",
      "postDate": "05/18/2016 20:58:40",
      "content": "<p>[quote=inversion;120503]</p>\n\n<p>[quote=FernandoProcy;120500]</p>\n\n<p>Also, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.</p>\n\n<p>[/quote]</p>\n\n<p>I recommend hand-labeling, just to be on the safe side.</p>\n\n<p>[/quote]\nThat seems a reasonable way to spend my 20's</p>",
      "rawMarkdown": "[quote=inversion;120503]\r\n\r\n[quote=FernandoProcy;120500]\r\n\r\nAlso, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.\r\n\r\n[/quote]\r\n\r\nI recommend hand-labeling, just to be on the safe side.\r\n\r\n[/quote]\r\nThat seems a reasonable way to spend my 20's",
      "votes": null
    },
    {
      "id": "120536",
      "postDate": "05/18/2016 22:21:39",
      "content": "<p>[quote=FernandoProcy;120523]</p>\n\n<p>That seems a reasonable way to spend my 20's</p>\n\n<p>[/quote]</p>\n\n<p>When I was in my 20s, all we <em>had</em> was hand labeling. Up hill. Both ways.</p>",
      "rawMarkdown": "[quote=FernandoProcy;120523]\r\n\r\nThat seems a reasonable way to spend my 20's\r\n\r\n[/quote]\r\n\r\nWhen I was in my 20s, all we _had_ was hand labeling. Up hill. Both ways.",
      "votes": null
    },
    {
      "id": "120538",
      "postDate": "05/18/2016 22:40:34",
      "content": "<p>[quote=inversion;120536]</p>\n\n<p>[quote=FernandoProcy;120523]</p>\n\n<p>That seems a reasonable way to spend my 20's</p>\n\n<p>[/quote]</p>\n\n<p>When I was in my 20s, all we <em>had</em> was hand labeling. Up hill. Both ways.</p>\n\n<p>[/quote]</p>\n\n<p>When I was in my 20s, I was drunk</p>",
      "rawMarkdown": "[quote=inversion;120536]\r\n\r\n[quote=FernandoProcy;120523]\r\n\r\nThat seems a reasonable way to spend my 20's\r\n\r\n[/quote]\r\n\r\nWhen I was in my 20s, all we _had_ was hand labeling. Up hill. Both ways.\r\n\r\n\r\n[/quote]\r\n\r\nWhen I was in my 20s, I was drunk",
      "votes": null
    },
    {
      "id": "120539",
      "postDate": "05/18/2016 22:41:00",
      "content": "<p>I started hand labeling when I was a couple of months old, especially after food sessions.</p>",
      "rawMarkdown": "I started hand labeling when I was a couple of months old, especially after food sessions.",
      "votes": null
    },
    {
      "id": "122504",
      "postDate": "06/04/2016 16:55:56",
      "content": "<p>I still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)</p>\n\n<ul>\n<li>Using stratified kfold with shuffling</li>\n<li>Using only similarity measures</li>\n<li>Using a &quot;bug free&quot; implementation of auc (in case sklearn's one is really buggy)</li>\n</ul>\n\n<p>And btw I am using XGBoost ... anybody figured it out ?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "I still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)\r\n\r\n- Using stratified kfold with shuffling\r\n- Using only similarity measures\r\n- Using a \"bug free\" implementation of auc (in case sklearn's one is really buggy)\r\n\r\nAnd btw I am using XGBoost ... anybody figured it out ?\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "122523",
      "postDate": "06/04/2016 20:06:52",
      "content": "<p>[quote=Amine Benhalloum;122504]</p>\n\n<p>I still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)</p>\n\n<ul>\n<li>Using stratified kfold with shuffling</li>\n<li>Using only similarity measures</li>\n<li>Using a &quot;bug free&quot; implementation of auc (in case sklearn's one is really buggy)</li>\n</ul>\n\n<p>And btw I am using XGBoost ... anybody figured it out ?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>This is because the training and testing sets are sampled at different point in time, and hence the nature of the items are slightly different. Your model may be fitting well to individual samples, or individual sellers that only appear in the training set.</p>",
      "rawMarkdown": "[quote=Amine Benhalloum;122504]\r\n\r\nI still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)\r\n\r\n- Using stratified kfold with shuffling\r\n- Using only similarity measures\r\n- Using a \"bug free\" implementation of auc (in case sklearn's one is really buggy)\r\n\r\nAnd btw I am using XGBoost ... anybody figured it out ?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nThis is because the training and testing sets are sampled at different point in time, and hence the nature of the items are slightly different. Your model may be fitting well to individual samples, or individual sellers that only appear in the training set.",
      "votes": null
    },
    {
      "id": "122527",
      "postDate": "06/04/2016 20:35:51",
      "content": "<p>[quote=anokas;122523]</p>\n\n<p>[quote=Amine Benhalloum;122504]</p>\n\n<p>I still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)</p>\n\n<ul>\n<li>Using stratified kfold with shuffling</li>\n<li>Using only similarity measures</li>\n<li>Using a &quot;bug free&quot; implementation of auc (in case sklearn's one is really buggy)</li>\n</ul>\n\n<p>And btw I am using XGBoost ... anybody figured it out ?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>This is because the training and testing sets are sampled at different point in time, and hence the nature of the items are slightly different. Your model may be fitting well to individual samples, or individual sellers that only appear in the training set.</p>\n\n<p>[/quote]\nThank you for your answer, so maybe even if I use only similarity based features (no raw features) there's still some &quot;leakage&quot; going on. I'll investigate further.</p>",
      "rawMarkdown": "[quote=anokas;122523]\r\n\r\n[quote=Amine Benhalloum;122504]\r\n\r\nI still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)\r\n\r\n- Using stratified kfold with shuffling\r\n- Using only similarity measures\r\n- Using a \"bug free\" implementation of auc (in case sklearn's one is really buggy)\r\n\r\nAnd btw I am using XGBoost ... anybody figured it out ?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nThis is because the training and testing sets are sampled at different point in time, and hence the nature of the items are slightly different. Your model may be fitting well to individual samples, or individual sellers that only appear in the training set.\r\n\r\n[/quote]\r\nThank you for your answer, so maybe even if I use only similarity based features (no raw features) there's still some \"leakage\" going on. I'll investigate further.",
      "votes": null
    },
    {
      "id": "124222",
      "postDate": "06/16/2016 10:31:26",
      "content": "<p>I can confirm. When my score was lower, I saw no difference between my CV and LB. But as I added more features (and they are not &quot;raw&quot; features), there is definitively something going on. I think it might have to do with the training and testing being collected at different times.</p>",
      "rawMarkdown": "I can confirm. When my score was lower, I saw no difference between my CV and LB. But as I added more features (and they are not \"raw\" features), there is definitively something going on. I think it might have to do with the training and testing being collected at different times.",
      "votes": null
    },
    {
      "id": "124229",
      "postDate": "06/16/2016 10:53:49",
      "content": "<p>what strategy people are using for CV split? one idea that i thought will be good is to create the splits between train and validation sets based on unique items. meaning to insure that the no ad that is present in the training split appears in the validation set. I am still not happy with the difference i am getting between CV mean score and LB score (CV: 0.945 and LB 0.908). </p>",
      "rawMarkdown": "what strategy people are using for CV split? one idea that i thought will be good is to create the splits between train and validation sets based on unique items. meaning to insure that the no ad that is present in the training split appears in the validation set. I am still not happy with the difference i am getting between CV mean score and LB score (CV: 0.945 and LB 0.908).",
      "votes": null
    },
    {
      "id": "124230",
      "postDate": "06/16/2016 10:57:17",
      "content": "<p>For me, it doesn't matter what strategy I employ. The dataset is so huge that whether I use bootstrapping or kfold, I get very small variance. The fact that the dataset is huge is already a big regularization factor.</p>",
      "rawMarkdown": "For me, it doesn't matter what strategy I employ. The dataset is so huge that whether I use bootstrapping or kfold, I get very small variance. The fact that the dataset is huge is already a big regularization factor.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118972,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "05/06/2016 13:04:37",
      "content": "<p>[quote=Abhishek;118966]</p>\n\n<p>My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?</p>\n\n<p>[/quote]</p>\n\n<p>CV = 0.75\nLB = 0.744</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119075,
      "author_name": "thequants",
      "author_url": "",
      "post_date": "05/07/2016 01:44:52",
      "content": "<p>[quote=Abhishek;118966]</p>\n\n<p>My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?</p>\n\n<p>[/quote]\nI get a similar behaviour with xgboost.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119081,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "05/07/2016 02:53:07",
      "content": "<p>Not that much offset here:</p>\n\n<p>CV= 0.789824, LB=0.74889</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119103,
      "author_name": "jzaratie",
      "author_url": "",
      "post_date": "05/07/2016 08:08:10",
      "content": "<p>Same problem here:</p>\n\n<p>CV=0.87, LB=0.77</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119264,
      "author_name": "ivicajovic",
      "author_url": "",
      "post_date": "05/08/2016 17:01:55",
      "content": "<p>People are <a href=\"https://www.kaggle.com/c/santander-customer-satisfaction/forums/t/20725/kaggle-auc-sklearn-auc-was-a-factor-in-lb-shake-up\">saying</a> that kaggle AUC and sklearn AUC give different results. What is your CV using kaggle AUC?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119277,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "05/08/2016 19:06:41",
      "content": "<p>Public LB: 0.745</p>\n\n<p>5-fold stratified CV: 0.761</p>\n\n<p>Not very much over here... Are you guys using stratified CV?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119279,
      "author_name": "dimitrislev",
      "author_url": "",
      "post_date": "05/08/2016 19:22:53",
      "content": "<p>Same to anokas, I am using 5-fold stratified split.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119287,
      "author_name": "alexandreyc",
      "author_url": "",
      "post_date": "05/08/2016 21:36:45",
      "content": "<p>I confirm that the problem comes from the use of sklearn's AUC. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119364,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "05/09/2016 14:54:44",
      "content": "<p>[quote=alexandreyc;119287]</p>\n\n<p>I confirm that the problem comes from the use of sklearn's AUC. </p>\n\n<p>[/quote]\nI was following the bug report on sklearn but could not any release version with the fix. So - the ones who are not using sklearn AUC - which implementation are you guys using ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119421,
      "author_name": "x75a40890",
      "author_url": "",
      "post_date": "05/10/2016 06:26:09",
      "content": "<p>I am using this one</p>\n\n<p><a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/auc.py\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/auc.py</a></p>\n\n<p>however I still have a huge gap between LB-CV</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119425,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "05/10/2016 06:43:22",
      "content": "<p>The gap in AUC may be because the test set is sampled at a different time to the training set, meaning that the problem may be slightly different.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119491,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "05/10/2016 19:34:14",
      "content": "<p>CV ~0.87 LB ~0.77  </p>\n\n<p>This is not AUC&#180;s fault, it's something with the data.</p>\n\n<p>@anokas\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that&#180;s part of the problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119492,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "05/10/2016 20:05:37",
      "content": "<p>[quote=Snow Dog;119491]</p>\n\n<p>CV ~0.87 LB ~0.77  </p>\n\n<p>This is not AUC&#180;s fault, it's something with the data.</p>\n\n<p>@anokas\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that&#180;s part of the problem.</p>\n\n<p>[/quote]</p>\n\n<p>I mean that Avito didn't use a random split for creating the training and testing datasets, but rather that they are from different points in time eg. we are training on 2014 data and predicting 2015 data. This could cause a slight difference in data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119493,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "05/10/2016 20:13:40",
      "content": "<p>[quote=Snow Dog;119491]</p>\n\n<p>CV ~0.87 LB ~0.77  </p>\n\n<p>This is not AUC&#180;s fault, it's something with the data.</p>\n\n<p>@anokas\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that&#180;s part of the problem.</p>\n\n<p>[/quote]</p>\n\n<p>My CV and LB are quite close. <br>\nCV - 0.77\nLB - 0.76</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119494,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "05/10/2016 20:26:52",
      "content": "<p>@DataGeek</p>\n\n<p>Are you doing something special, besides stratifying the target? I need to use all the data and see what happens.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119495,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "05/10/2016 20:30:10",
      "content": "<p>[quote=Snow Dog;119494]</p>\n\n<p>@DataGeek</p>\n\n<p>Are you doing something special, besides stratifying the target? I need to use all the data and see what happens.</p>\n\n<p>[/quote]</p>\n\n<p>I think my CV and LB are close because I am using similarity b/w features of 2 ads instead of features directly. I haven't tested but most of the people having large differences in CV and LB must be using raw features + similarity features or just the raw features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119497,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "05/10/2016 20:44:46",
      "content": "<p>@Data Geek</p>\n\n<p>Thanks, going to check that! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119711,
      "author_name": "iezepov",
      "author_url": "",
      "post_date": "05/12/2016 12:52:45",
      "content": "<p>[quote=Abhishek;118966]</p>\n\n<p>My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?</p>\n\n<p>[/quote]</p>\n\n<p>Hi there! I found out that shuffling is very important. Here are my results from logistic regression on 4 very simple features. First, results from 5-fold stratified split (from sklearn) without shuffling:</p>\n\n<ul>\n<li>AUC: 0.7661</li>\n<li>AUC: 0.7805</li>\n<li>AUC: 0.7343</li>\n<li>AUC: 0.6626</li>\n<li>AUC: 0.6555</li>\n</ul>\n\n<p>And results with shuffling:</p>\n\n<ul>\n<li>AUC: 0.7184</li>\n<li>AUC: 0.7179 </li>\n<li>AUC: 0.7185</li>\n<li>AUC: 0.7186</li>\n<li>AUC: 0.7193</li>\n</ul>\n\n<p>You can see, that with shuffling, results are much more stable and reliable. It's clear evidence that the data has a time structure, so I suppose you should split the data not randomly. First N% rows for traininig, remaining are for validation. It's also clear because of decrease of AUC with moving validation window further to the end of the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119714,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "05/12/2016 13:01:15",
      "content": "<p>[quote=Ilya;119711]</p>\n\n<p>Hi there! I found out that shuffling is very important. Here are my results from logistic regression on 4 very simple features. First, results from 5-fold stratified split (from sklearn) without shuffling:</p>\n\n<ul>\n<li>AUC: 0.7661</li>\n<li>AUC: 0.7805</li>\n<li>AUC: 0.7343</li>\n<li>AUC: 0.6626</li>\n<li>AUC: 0.6555</li>\n</ul>\n\n<p>And results with shuffling:</p>\n\n<ul>\n<li>AUC: 0.7184</li>\n<li>AUC: 0.7179 </li>\n<li>AUC: 0.7185</li>\n<li>AUC: 0.7186</li>\n<li>AUC: 0.7193</li>\n</ul>\n\n<p>You can see, that with shuffling, results are much more stable and reliable. It's clear evidence that the data has a time structure, so I suppose you should split the data not randomly. First N% rows for traininig, remaining are for validation. It's also clear because of decrease of AUC with moving validation window further to the end of the data.</p>\n\n<p>[/quote]\nNice find, mate :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119789,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "05/12/2016 19:49:24",
      "content": "<p>I've found consistently in my testing that as you train more XGBoost rounds, the delta between leaderboard score and local CV increases, even if the deviation locally is very low and difference between train/cv auc is very low. </p>\n\n<p>I suspect this might be because the training and testing sets are from different dates, and so the content differs somewhat. Looking at the values in the dataset, there does seem to be quite a large difference between training/testing sets, while the data is very uniform within the training set, so this could be what is causing the difference between LB/CV scores.</p>\n\n<p>Just for reference, my current scores are</p>\n\n<p>CV: ~0.945 LB: ~0.905</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119796,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "05/12/2016 20:13:47",
      "content": "<p>@anokas\nCongrats on being the first to beat the mark!</p>\n\n<p>I managed to decrease the gap, but still a large difference: CV: 0.82, LB: 0.76 (using half of the data).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120144,
      "author_name": "kelexu",
      "author_url": "",
      "post_date": "05/15/2016 21:02:47",
      "content": "<p>It seems that the CV - LB = 0.03~0.04. We can use this information to inter the difference between the distributions of duplicates in train/test datasets.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120165,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "05/16/2016 00:05:04",
      "content": "<p>I made my first sub today using a very simple model (no attrsJSON/no images/no cat and location files/default xgb params) just to see how it would fare in LB. Got 0.91x in local validation and 0.81x in LB. I wonder if I'm overfitting or if when I start adding more features that difference will shrink.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120246,
      "author_name": "zavodrobotov",
      "author_url": "",
      "post_date": "05/16/2016 17:43:18",
      "content": "<p>It seems like some of duplicates presented by more that one pair in training set, and some pair can get in train and some - in test set with random partitioning, so classifier learn to recognize concrete ads instead of really classify duplicates, so with increase of tree depth CV score increases, but LB don't increases (of course, these concrete ads didn't presented in LB dataset)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120500,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "05/18/2016 17:39:22",
      "content": "<p>[quote=ZavodRobotov;120246]</p>\n\n<p>It seems like some of duplicates presented by more that one pair in training set, and some pair can get in train and some - in test set with random partitioning, so classifier learn to recognize concrete ads instead of really classify duplicates, so with increase of tree depth CV score increases, but LB don't increases (of course, these concrete ads didn't presented in LB dataset)</p>\n\n<p>[/quote]\nI think you're right and that makes sense. As has DataGeek pointed out earlier, if you use raw features it can lead to some degree of overfitting because the model start recognizing the ads instead of the similarity between them. Using some raw features I currently get around 0.82x LB. Without using them, I get a lower score, but I think that is the way to go.</p>\n\n<p>Also, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120503,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/18/2016 17:50:08",
      "content": "<p>[quote=FernandoProcy;120500]</p>\n\n<p>Also, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.</p>\n\n<p>[/quote]</p>\n\n<p>I recommend hand-labeling, just to be on the safe side.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120506,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "05/18/2016 18:31:47",
      "content": "<p>how many pictures are there ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120508,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/18/2016 18:47:09",
      "content": "<p>[quote=eagle4;120506]</p>\n\n<p>how many pictures are there ?</p>\n\n<p>[/quote]</p>\n\n<p>~11,000,000</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120509,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "05/18/2016 18:52:40",
      "content": "<p>@inversion: I was going to believe you since I always follow your advises on Kaggle forums...</p>\n\n<p>May I ask the leaders what the best score achieved without using the pictures ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120512,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/18/2016 19:12:59",
      "content": "<p>[quote=eagle4;120509]</p>\n\n<p>@inversion: I was going to believe you since I always follow your advises on Kaggle forums...</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, I was just messing around.</p>\n\n<p>For serious, though, I think (not 100% sure) that you can get in the top 10 without image features. (Or at least close). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120523,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "05/18/2016 20:58:40",
      "content": "<p>[quote=inversion;120503]</p>\n\n<p>[quote=FernandoProcy;120500]</p>\n\n<p>Also, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.</p>\n\n<p>[/quote]</p>\n\n<p>I recommend hand-labeling, just to be on the safe side.</p>\n\n<p>[/quote]\nThat seems a reasonable way to spend my 20's</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120536,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/18/2016 22:21:39",
      "content": "<p>[quote=FernandoProcy;120523]</p>\n\n<p>That seems a reasonable way to spend my 20's</p>\n\n<p>[/quote]</p>\n\n<p>When I was in my 20s, all we <em>had</em> was hand labeling. Up hill. Both ways.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120538,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "05/18/2016 22:40:34",
      "content": "<p>[quote=inversion;120536]</p>\n\n<p>[quote=FernandoProcy;120523]</p>\n\n<p>That seems a reasonable way to spend my 20's</p>\n\n<p>[/quote]</p>\n\n<p>When I was in my 20s, all we <em>had</em> was hand labeling. Up hill. Both ways.</p>\n\n<p>[/quote]</p>\n\n<p>When I was in my 20s, I was drunk</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120539,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "05/18/2016 22:41:00",
      "content": "<p>I started hand labeling when I was a couple of months old, especially after food sessions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122504,
      "author_name": "bamine",
      "author_url": "",
      "post_date": "06/04/2016 16:55:56",
      "content": "<p>I still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)</p>\n\n<ul>\n<li>Using stratified kfold with shuffling</li>\n<li>Using only similarity measures</li>\n<li>Using a &quot;bug free&quot; implementation of auc (in case sklearn's one is really buggy)</li>\n</ul>\n\n<p>And btw I am using XGBoost ... anybody figured it out ?</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122523,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "06/04/2016 20:06:52",
      "content": "<p>[quote=Amine Benhalloum;122504]</p>\n\n<p>I still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)</p>\n\n<ul>\n<li>Using stratified kfold with shuffling</li>\n<li>Using only similarity measures</li>\n<li>Using a &quot;bug free&quot; implementation of auc (in case sklearn's one is really buggy)</li>\n</ul>\n\n<p>And btw I am using XGBoost ... anybody figured it out ?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>This is because the training and testing sets are sampled at different point in time, and hence the nature of the items are slightly different. Your model may be fitting well to individual samples, or individual sellers that only appear in the training set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122527,
      "author_name": "bamine",
      "author_url": "",
      "post_date": "06/04/2016 20:35:51",
      "content": "<p>[quote=anokas;122523]</p>\n\n<p>[quote=Amine Benhalloum;122504]</p>\n\n<p>I still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)</p>\n\n<ul>\n<li>Using stratified kfold with shuffling</li>\n<li>Using only similarity measures</li>\n<li>Using a &quot;bug free&quot; implementation of auc (in case sklearn's one is really buggy)</li>\n</ul>\n\n<p>And btw I am using XGBoost ... anybody figured it out ?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>This is because the training and testing sets are sampled at different point in time, and hence the nature of the items are slightly different. Your model may be fitting well to individual samples, or individual sellers that only appear in the training set.</p>\n\n<p>[/quote]\nThank you for your answer, so maybe even if I use only similarity based features (no raw features) there's still some &quot;leakage&quot; going on. I'll investigate further.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124222,
      "author_name": "rpmcruz",
      "author_url": "",
      "post_date": "06/16/2016 10:31:26",
      "content": "<p>I can confirm. When my score was lower, I saw no difference between my CV and LB. But as I added more features (and they are not &quot;raw&quot; features), there is definitively something going on. I think it might have to do with the training and testing being collected at different times.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124229,
      "author_name": "samehif",
      "author_url": "",
      "post_date": "06/16/2016 10:53:49",
      "content": "<p>what strategy people are using for CV split? one idea that i thought will be good is to create the splits between train and validation sets based on unique items. meaning to insure that the no ad that is present in the training split appears in the validation set. I am still not happy with the difference i am getting between CV mean score and LB score (CV: 0.945 and LB 0.908). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124230,
      "author_name": "rpmcruz",
      "author_url": "",
      "post_date": "06/16/2016 10:57:17",
      "content": "<p>For me, it doesn't matter what strategy I employ. The dataset is so huge that whether I use bootstrapping or kfold, I get very small variance. The fact that the dataset is huge is already a big regularization factor.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118966": "My validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?",
    "118972": "[quote=Abhishek;118966]\r\n\r\nMy validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?\r\n\r\n[/quote]\r\n\r\nCV = 0.75\r\nLB = 0.744",
    "119075": "[quote=Abhishek;118966]\r\n\r\nMy validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?\r\n\r\n[/quote]\r\nI get a similar behaviour with xgboost.",
    "119081": "Not that much offset here:\r\n\r\nCV= 0.789824, LB=0.74889",
    "119103": "Same problem here:\r\n\r\nCV=0.87, LB=0.77",
    "119264": "People are [saying][1] that kaggle AUC and sklearn AUC give different results. What is your CV using kaggle AUC?\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/santander-customer-satisfaction/forums/t/20725/kaggle-auc-sklearn-auc-was-a-factor-in-lb-shake-up",
    "119277": "Public LB: 0.745\r\n\r\n5-fold stratified CV: 0.761\r\n\r\nNot very much over here... Are you guys using stratified CV?",
    "119279": "Same to anokas, I am using 5-fold stratified split.",
    "119287": "I confirm that the problem comes from the use of sklearn's AUC.",
    "119364": "[quote=alexandreyc;119287]\r\n\r\nI confirm that the problem comes from the use of sklearn's AUC. \r\n\r\n[/quote]\r\nI was following the bug report on sklearn but could not any release version with the fix. So - the ones who are not using sklearn AUC - which implementation are you guys using ?",
    "119421": "I am using this one\r\n\r\nhttps://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/auc.py\r\n\r\nhowever I still have a huge gap between LB-CV",
    "119425": "The gap in AUC may be because the test set is sampled at a different time to the training set, meaning that the problem may be slightly different.",
    "119491": "CV ~0.87 LB ~0.77  \r\n\r\nThis is not AUC´s fault, it's something with the data.\r\n\r\n@anokas\r\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that´s part of the problem.",
    "119492": "[quote=Snow Dog;119491]\r\n\r\nCV ~0.87 LB ~0.77  \r\n\r\nThis is not AUC´s fault, it's something with the data.\r\n\r\n@anokas\r\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that´s part of the problem.\r\n\r\n[/quote]\r\n\r\nI mean that Avito didn't use a random split for creating the training and testing datasets, but rather that they are from different points in time eg. we are training on 2014 data and predicting 2015 data. This could cause a slight difference in data.",
    "119493": "[quote=Snow Dog;119491]\r\n\r\nCV ~0.87 LB ~0.77  \r\n\r\nThis is not AUC´s fault, it's something with the data.\r\n\r\n@anokas\r\nYou mean stratified by target? I didn't use 100% of trainning data, maybe that´s part of the problem.\r\n\r\n[/quote]\r\n\r\nMy CV and LB are quite close.  \r\nCV - 0.77\r\nLB - 0.76",
    "119494": "DataGeek\r\n\r\nAre you doing something special, besides stratifying the target? I need to use all the data and see what happens.",
    "119495": "[quote=Snow Dog;119494]\r\n\r\n@DataGeek\r\n\r\nAre you doing something special, besides stratifying the target? I need to use all the data and see what happens.\r\n\r\n\r\n[/quote]\r\n\r\nI think my CV and LB are close because I am using similarity b/w features of 2 ads instead of features directly. I haven't tested but most of the people having large differences in CV and LB must be using raw features + similarity features or just the raw features.",
    "119497": "Data Geek\r\n\r\nThanks, going to check that!",
    "119711": "[quote=Abhishek;118966]\r\n\r\nMy validation score is 0.90+ but my LB is my current score, 0.77. Is anyone else having the same problem?\r\n\r\n[/quote]\r\n\r\nHi there! I found out that shuffling is very important. Here are my results from logistic regression on 4 very simple features. First, results from 5-fold stratified split (from sklearn) without shuffling:\r\n\r\n- AUC: 0.7661\r\n- AUC: 0.7805\r\n- AUC: 0.7343\r\n- AUC: 0.6626\r\n- AUC: 0.6555\r\n\r\nAnd results with shuffling:\r\n\r\n- AUC: 0.7184\r\n- AUC: 0.7179 \r\n- AUC: 0.7185\r\n- AUC: 0.7186\r\n- AUC: 0.7193\r\n\r\nYou can see, that with shuffling, results are much more stable and reliable. It's clear evidence that the data has a time structure, so I suppose you should split the data not randomly. First N% rows for traininig, remaining are for validation. It's also clear because of decrease of AUC with moving validation window further to the end of the data.",
    "119714": "[quote=Ilya;119711]\r\n\r\nHi there! I found out that shuffling is very important. Here are my results from logistic regression on 4 very simple features. First, results from 5-fold stratified split (from sklearn) without shuffling:\r\n\r\n- AUC: 0.7661\r\n- AUC: 0.7805\r\n- AUC: 0.7343\r\n- AUC: 0.6626\r\n- AUC: 0.6555\r\n\r\nAnd results with shuffling:\r\n\r\n- AUC: 0.7184\r\n- AUC: 0.7179 \r\n- AUC: 0.7185\r\n- AUC: 0.7186\r\n- AUC: 0.7193\r\n\r\nYou can see, that with shuffling, results are much more stable and reliable. It's clear evidence that the data has a time structure, so I suppose you should split the data not randomly. First N% rows for traininig, remaining are for validation. It's also clear because of decrease of AUC with moving validation window further to the end of the data.\r\n\r\n[/quote]\r\nNice find, mate :)",
    "119789": "I've found consistently in my testing that as you train more XGBoost rounds, the delta between leaderboard score and local CV increases, even if the deviation locally is very low and difference between train/cv auc is very low. \r\n\r\nI suspect this might be because the training and testing sets are from different dates, and so the content differs somewhat. Looking at the values in the dataset, there does seem to be quite a large difference between training/testing sets, while the data is very uniform within the training set, so this could be what is causing the difference between LB/CV scores.\r\n\r\nJust for reference, my current scores are\r\n\r\nCV: ~0.945 LB: ~0.905",
    "119796": "anokas\r\nCongrats on being the first to beat the mark!\r\n\r\nI managed to decrease the gap, but still a large difference: CV: 0.82, LB: 0.76 (using half of the data).",
    "120144": "It seems that the CV - LB = 0.03~0.04. We can use this information to inter the difference between the distributions of duplicates in train/test datasets.",
    "120165": "I made my first sub today using a very simple model (no attrsJSON/no images/no cat and location files/default xgb params) just to see how it would fare in LB. Got 0.91x in local validation and 0.81x in LB. I wonder if I'm overfitting or if when I start adding more features that difference will shrink.",
    "120246": "It seems like some of duplicates presented by more that one pair in training set, and some pair can get in train and some - in test set with random partitioning, so classifier learn to recognize concrete ads instead of really classify duplicates, so with increase of tree depth CV score increases, but LB don't increases (of course, these concrete ads didn't presented in LB dataset)",
    "120500": "[quote=ZavodRobotov;120246]\r\n\r\nIt seems like some of duplicates presented by more that one pair in training set, and some pair can get in train and some - in test set with random partitioning, so classifier learn to recognize concrete ads instead of really classify duplicates, so with increase of tree depth CV score increases, but LB don't increases (of course, these concrete ads didn't presented in LB dataset)\r\n\r\n[/quote]\r\nI think you're right and that makes sense. As has DataGeek pointed out earlier, if you use raw features it can lead to some degree of overfitting because the model start recognizing the ads instead of the similarity between them. Using some raw features I currently get around 0.82x LB. Without using them, I get a lower score, but I think that is the way to go.\r\n\r\nAlso, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.",
    "120503": "[quote=FernandoProcy;120500]\r\n\r\nAlso, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.\r\n\r\n[/quote]\r\n\r\nI recommend hand-labeling, just to be on the safe side.",
    "120506": "how many pictures are there ?",
    "120508": "[quote=eagle4;120506]\r\n\r\nhow many pictures are there ?\r\n\r\n[/quote]\r\n\r\n~11,000,000",
    "120509": "inversion: I was going to believe you since I always follow your advises on Kaggle forums...\r\n\r\nMay I ask the leaders what the best score achieved without using the pictures ?",
    "120512": "[quote=eagle4;120509]\r\n\r\n@inversion: I was going to believe you since I always follow your advises on Kaggle forums...\r\n\r\n[/quote]\r\n\r\nYeah, I was just messing around.\r\n\r\nFor serious, though, I think (not 100% sure) that you can get in the top 10 without image features. (Or at least close).",
    "120523": "[quote=inversion;120503]\r\n\r\n[quote=FernandoProcy;120500]\r\n\r\nAlso, did anyone get huge improvements using images? I'm under the impression that going through 40GB+ of images won't net a significant better AUC. Hope I'm wrong, though.\r\n\r\n[/quote]\r\n\r\nI recommend hand-labeling, just to be on the safe side.\r\n\r\n[/quote]\r\nThat seems a reasonable way to spend my 20's",
    "120536": "[quote=FernandoProcy;120523]\r\n\r\nThat seems a reasonable way to spend my 20's\r\n\r\n[/quote]\r\n\r\nWhen I was in my 20s, all we _had_ was hand labeling. Up hill. Both ways.",
    "120538": "[quote=inversion;120536]\r\n\r\n[quote=FernandoProcy;120523]\r\n\r\nThat seems a reasonable way to spend my 20's\r\n\r\n[/quote]\r\n\r\nWhen I was in my 20s, all we _had_ was hand labeling. Up hill. Both ways.\r\n\r\n\r\n[/quote]\r\n\r\nWhen I was in my 20s, I was drunk",
    "120539": "I started hand labeling when I was a couple of months old, especially after food sessions.",
    "122504": "I still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)\r\n\r\n- Using stratified kfold with shuffling\r\n- Using only similarity measures\r\n- Using a \"bug free\" implementation of auc (in case sklearn's one is really buggy)\r\n\r\nAnd btw I am using XGBoost ... anybody figured it out ?\r\n\r\nThanks",
    "122523": "[quote=Amine Benhalloum;122504]\r\n\r\nI still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)\r\n\r\n- Using stratified kfold with shuffling\r\n- Using only similarity measures\r\n- Using a \"bug free\" implementation of auc (in case sklearn's one is really buggy)\r\n\r\nAnd btw I am using XGBoost ... anybody figured it out ?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nThis is because the training and testing sets are sampled at different point in time, and hence the nature of the items are slightly different. Your model may be fitting well to individual samples, or individual sellers that only appear in the training set.",
    "122527": "[quote=anokas;122523]\r\n\r\n[quote=Amine Benhalloum;122504]\r\n\r\nI still can't figure out why my CV score is so different from my LB score : (CV 0.88 - LB 0.806)\r\n\r\n- Using stratified kfold with shuffling\r\n- Using only similarity measures\r\n- Using a \"bug free\" implementation of auc (in case sklearn's one is really buggy)\r\n\r\nAnd btw I am using XGBoost ... anybody figured it out ?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nThis is because the training and testing sets are sampled at different point in time, and hence the nature of the items are slightly different. Your model may be fitting well to individual samples, or individual sellers that only appear in the training set.\r\n\r\n[/quote]\r\nThank you for your answer, so maybe even if I use only similarity based features (no raw features) there's still some \"leakage\" going on. I'll investigate further.",
    "124222": "I can confirm. When my score was lower, I saw no difference between my CV and LB. But as I added more features (and they are not \"raw\" features), there is definitively something going on. I think it might have to do with the training and testing being collected at different times.",
    "124229": "what strategy people are using for CV split? one idea that i thought will be good is to create the splits between train and validation sets based on unique items. meaning to insure that the no ad that is present in the training split appears in the validation set. I am still not happy with the difference i am getting between CV mean score and LB score (CV: 0.945 and LB 0.908).",
    "124230": "For me, it doesn't matter what strategy I employ. The dataset is so huge that whether I use bootstrapping or kfold, I get very small variance. The fact that the dataset is huge is already a big regularization factor."
  },
  "source": "meta"
}