{
  "id": 78229,
  "title": "The dangers of LB CLimbing",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/78229",
  "author_name": "Scirpus",
  "post_date": "2019-01-21T12:18:36.839000",
  "votes": 13,
  "comment_count": 16,
  "views": 0,
  "content": "<p>LB = \"13% of the test data\"  As there are only 2624 targets total - is anyone else scared to death about using the LB as an indication of success?  341 targets is a very very tiny sample for me to be confident.</p>",
  "messages": [
    {
      "id": 459226,
      "postDate": "2019-01-21T12:18:36.840Z",
      "content": "<p>LB = \"13% of the test data\"  As there are only 2624 targets total - is anyone else scared to death about using the LB as an indication of success?  341 targets is a very very tiny sample for me to be confident.</p>",
      "rawMarkdown": "LB = \"13% of the test data\"  As there are only 2624 targets total - is anyone else scared to death about using the LB as an indication of success?  341 targets is a very very tiny sample for me to be confident.",
      "votes": 13
    },
    {
      "id": 459443,
      "postDate": "2019-01-21T18:41:15.960Z",
      "content": "<p>yes, scared, but that's also the most challenging part in every modeling task, even out of competition context.</p>\n\n<p>I can think of some precautionary measures:</p>\n\n<p>(1) trust local CV\n(2) evaluate local CV with confidence intervals instead of a single number summary statistic alone.\n(3) evaluate models for, not just prediction accuracy,  but also their complexity and number of features used (lesser/fewer the better).\n(4) evaluate models with both local CV and public LB score, for example, average of CV and LB</p>\n\n<p>add more by leaving comments below :)</p>",
      "rawMarkdown": "yes, scared, but that's also the most challenging part in every modeling task, even out of competition context.\n\nI can think of some precautionary measures:\n\n(1) trust local CV\n(2) evaluate local CV with confidence intervals instead of a single number summary statistic alone.\n(3) evaluate models for, not just prediction accuracy,  but also their complexity and number of features used (lesser/fewer the better).\n(4) evaluate models with both local CV and public LB score, for example, average of CV and LB\n\nadd more by leaving comments below :)\n\n",
      "votes": 10
    },
    {
      "id": 460073,
      "postDate": "2019-01-22T23:43:46.760Z",
      "content": "<p>The shakeup potential in this competition is indeed huge.</p>\n\n<p>The problem appears to be quite sensitive on the test data selection. I currently use 25-fold CV in model development. This makes each test fold about half the size of the public LB set. With this mode of operation, I see a standard deviation in the 0.10 to 0.13 range among the test fold MAEs.</p>\n\n<p>I intend to rely on local CV. Using the public LB for blend weight optimization or the like seems dangerous.</p>",
      "rawMarkdown": "The shakeup potential in this competition is indeed huge.\n\nThe problem appears to be quite sensitive on the test data selection. I currently use 25-fold CV in model development. This makes each test fold about half the size of the public LB set. With this mode of operation, I see a standard deviation in the 0.10 to 0.13 range among the test fold MAEs.\n\nI intend to rely on local CV. Using the public LB for blend weight optimization or the like seems dangerous.",
      "votes": 5
    },
    {
      "id": 460083,
      "postDate": "2019-01-23T00:45:39.713Z",
      "content": "<p>@Scirpus, you are quite right.</p>\n\n<p>My model with the best CV scores would put me about 120 points below where I am on the public LB right now. I believe proper error analysis is the key here. We have 4 months to figure it out so lets keep the discussions going :-)</p>",
      "rawMarkdown": "@Scirpus, you are quite right.\n\nMy model with the best CV scores would put me about 120 points below where I am on the public LB right now. I believe proper error analysis is the key here. We have 4 months to figure it out so lets keep the discussions going :-)",
      "votes": 3
    },
    {
      "id": 459229,
      "postDate": "2019-01-21T12:20:44.690Z",
      "content": "<p>That's totally right. Someone had fallen off more than 3000 places, from the 2nd place on the final public LB to final private LB in the Santander Value Prediction Challenge not so long ago. \n<a href=\"https://www.kaggle.com/c/santander-value-prediction-challenge/discussion/63753#latest-373618\">https://www.kaggle.com/c/santander-value-prediction-challenge/discussion/63753#latest-373618</a></p>",
      "rawMarkdown": "That's totally right. Someone had fallen off more than 3000 places, from the 2nd place on the final public LB to final private LB in the Santander Value Prediction Challenge not so long ago. \nhttps://www.kaggle.com/c/santander-value-prediction-challenge/discussion/63753#latest-373618",
      "votes": 3,
      "replies": [
        {
          "id": 460085,
          "postDate": "2019-01-23T00:56:46.317Z",
          "content": "<p><a href=\"/khahuras\">@khahuras</a>, don't forget the most participated competition so far, i.e.  <a href=\"https://www.kaggle.com/c/home-credit-default-risk/leaderboard\">\"Home Credit Default Risk\"</a> where people jumped up so many places on the private LB. The one I took note of was a 1700+ jump up to within the top 100. I myself jumped up about 400 places. It was much higher before cheaters removal which also is the biggest number (i.e. cheaters) I have seen in a competition I have been a part of so far. Position wise, I went from private LB #93 at closing to final number 85 after they removed all the cheaters. I think the public/ private LB in that competition will make a good data for a case study.</p>",
          "rawMarkdown": "@khahuras, don't forget the most participated competition so far, i.e.  [\"Home Credit Default Risk\"][1] where people jumped up so many places on the private LB. The one I took note of was a 1700+ jump up to within the top 100. I myself jumped up about 400 places. It was much higher before cheaters removal which also is the biggest number (i.e. cheaters) I have seen in a competition I have been a part of so far. Position wise, I went from private LB #93 at closing to final number 85 after they removed all the cheaters. I think the public/ private LB in that competition will make a good data for a case study.\n\n\n  [1]: https://www.kaggle.com/c/home-credit-default-risk/leaderboard",
          "votes": 1
        }
      ]
    },
    {
      "id": 459544,
      "postDate": "2019-01-22T00:44:22.497Z",
      "content": "<p>The more I improve my CV score I do worst in the LB Score, I haven't found a good strategy to test, but maybe is just because of that 13%, how everyone is dealing with this ?</p>",
      "rawMarkdown": "The more I improve my CV score I do worst in the LB Score, I haven't found a good strategy to test, but maybe is just because of that 13%, how everyone is dealing with this ?",
      "votes": 4
    },
    {
      "id": 459270,
      "postDate": "2019-01-21T13:32:27.407Z",
      "content": "<p>I agree!\nAlso we have only 4194 samples in train (if we aggregate each 150000 rows), which could easily lead to overfitting.</p>\n\n<p>I tried sampling additional data - it significantly decreased score.</p>",
      "rawMarkdown": "I agree!\nAlso we have only 4194 samples in train (if we aggregate each 150000 rows), which could easily lead to overfitting.\n\nI tried sampling additional data - it significantly decreased score.",
      "votes": 2,
      "replies": [
        {
          "id": 459276,
          "postDate": "2019-01-21T13:45:46.903Z",
          "content": "<p>I agree it is even more scary as getting an outlier correct will significantly lower your score on a sample of 341 targets.</p>",
          "rawMarkdown": "I agree it is even more scary as getting an outlier correct will significantly lower your score on a sample of 341 targets.",
          "votes": 2
        },
        {
          "id": 459406,
          "postDate": "2019-01-21T16:57:13.603Z",
          "content": "<p>In a past competition (Mercedes manufacturing), one outlier was responsible of lots of places on the LB!\nI don't think that it will happen here: the time before earthquake is always below 15 (max 20) which means that there won't be one value with 100 seconds. And the scoring is absolute value, which gives a low importance to outliers.</p>",
          "rawMarkdown": "In a past competition (Mercedes manufacturing), one outlier was responsible of lots of places on the LB!\nI don't think that it will happen here: the time before earthquake is always below 15 (max 20) which means that there won't be one value with 100 seconds. And the scoring is absolute value, which gives a low importance to outliers.",
          "votes": 4
        }
      ]
    },
    {
      "id": 459408,
      "postDate": "2019-01-21T16:59:01.863Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true,
      "replies": [
        {
          "id": 459409,
          "postDate": "2019-01-21T17:04:51.917Z",
          "content": "<p>They probably deleted it because:</p>\n\n<p>You accused people of cheating\nYou used a pretty inappropriate title</p>",
          "rawMarkdown": "They probably deleted it because:\n\nYou accused people of cheating\nYou used a pretty inappropriate title\n\n",
          "votes": 3
        },
        {
          "id": 459414,
          "postDate": "2019-01-21T17:21:14.980Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 459427,
          "postDate": "2019-01-21T17:53:58.760Z",
          "content": "<p>Just stop now you are tiring me out - if you have genuine concerns just contact Kaggle directly rather than going on a rant.  With the greatest of respect it just makes you look foolish.</p>",
          "rawMarkdown": "Just stop now you are tiring me out - if you have genuine concerns just contact Kaggle directly rather than going on a rant.  With the greatest of respect it just makes you look foolish."
        },
        {
          "id": 459429,
          "postDate": "2019-01-21T17:57:31.887Z",
          "content": "<p>im waiting for you to win this competition.</p>",
          "rawMarkdown": "im waiting for you to win this competition.",
          "votes": 3
        },
        {
          "id": 459430,
          "postDate": "2019-01-21T17:59:35.230Z",
          "content": "<p>im waiting for you to win this competition. jkjk\nHiding public LB would not be a viable solution either. </p>",
          "rawMarkdown": "im waiting for you to win this competition. jkjk\nHiding public LB would not be a viable solution either. "
        },
        {
          "id": 461186,
          "postDate": "2019-01-25T13:48:06.293Z",
          "content": "<p>Looks like he has fallen over the Reichenbach Falls</p>",
          "rawMarkdown": "Looks like he has fallen over the Reichenbach Falls"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 459443,
      "author_name": "Elliot",
      "author_url": "",
      "post_date": "2019-01-21T18:41:15.960000",
      "content": "<p>yes, scared, but that's also the most challenging part in every modeling task, even out of competition context.</p>\n\n<p>I can think of some precautionary measures:</p>\n\n<p>(1) trust local CV\n(2) evaluate local CV with confidence intervals instead of a single number summary statistic alone.\n(3) evaluate models for, not just prediction accuracy,  but also their complexity and number of features used (lesser/fewer the better).\n(4) evaluate models with both local CV and public LB score, for example, average of CV and LB</p>\n\n<p>add more by leaving comments below :)</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 460073,
      "author_name": "Andre Naef",
      "author_url": "",
      "post_date": "2019-01-22T23:43:46.760000",
      "content": "<p>The shakeup potential in this competition is indeed huge.</p>\n\n<p>The problem appears to be quite sensitive on the test data selection. I currently use 25-fold CV in model development. This makes each test fold about half the size of the public LB set. With this mode of operation, I see a standard deviation in the 0.10 to 0.13 range among the test fold MAEs.</p>\n\n<p>I intend to rely on local CV. Using the public LB for blend weight optimization or the like seems dangerous.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 460083,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-01-23T00:45:39.713000",
      "content": "<p>@Scirpus, you are quite right.</p>\n\n<p>My model with the best CV scores would put me about 120 points below where I am on the public LB right now. I believe proper error analysis is the key here. We have 4 months to figure it out so lets keep the discussions going :-)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 459229,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-01-21T12:20:44.690000",
      "content": "<p>That's totally right. Someone had fallen off more than 3000 places, from the 2nd place on the final public LB to final private LB in the Santander Value Prediction Challenge not so long ago. \n<a href=\"https://www.kaggle.com/c/santander-value-prediction-challenge/discussion/63753#latest-373618\">https://www.kaggle.com/c/santander-value-prediction-challenge/discussion/63753#latest-373618</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 460085,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-01-23T00:56:46.317000",
          "content": "<p><a href=\"/khahuras\">@khahuras</a>, don't forget the most participated competition so far, i.e.  <a href=\"https://www.kaggle.com/c/home-credit-default-risk/leaderboard\">\"Home Credit Default Risk\"</a> where people jumped up so many places on the private LB. The one I took note of was a 1700+ jump up to within the top 100. I myself jumped up about 400 places. It was much higher before cheaters removal which also is the biggest number (i.e. cheaters) I have seen in a competition I have been a part of so far. Position wise, I went from private LB #93 at closing to final number 85 after they removed all the cheaters. I think the public/ private LB in that competition will make a good data for a case study.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 459544,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2019-01-22T00:44:22.497000",
      "content": "<p>The more I improve my CV score I do worst in the LB Score, I haven't found a good strategy to test, but maybe is just because of that 13%, how everyone is dealing with this ?</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 459270,
      "author_name": "Andrey Lukyanenko",
      "author_url": "",
      "post_date": "2019-01-21T13:32:27.407000",
      "content": "<p>I agree!\nAlso we have only 4194 samples in train (if we aggregate each 150000 rows), which could easily lead to overfitting.</p>\n\n<p>I tried sampling additional data - it significantly decreased score.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 459276,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-01-21T13:45:46.903000",
          "content": "<p>I agree it is even more scary as getting an outlier correct will significantly lower your score on a sample of 341 targets.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 459406,
          "author_name": "Zidmie",
          "author_url": "",
          "post_date": "2019-01-21T16:57:13.603000",
          "content": "<p>In a past competition (Mercedes manufacturing), one outlier was responsible of lots of places on the LB!\nI don't think that it will happen here: the time before earthquake is always below 15 (max 20) which means that there won't be one value with 100 seconds. And the scoring is absolute value, which gives a low importance to outliers.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 459408,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-21T16:59:01.863000",
      "content": "",
      "votes": -2,
      "replies": [
        {
          "id": 459409,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-01-21T17:04:51.917000",
          "content": "<p>They probably deleted it because:</p>\n\n<p>You accused people of cheating\nYou used a pretty inappropriate title</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 459414,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-21T17:21:14.980000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 459427,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-01-21T17:53:58.760000",
          "content": "<p>Just stop now you are tiring me out - if you have genuine concerns just contact Kaggle directly rather than going on a rant.  With the greatest of respect it just makes you look foolish.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 459429,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2019-01-21T17:57:31.887000",
          "content": "<p>im waiting for you to win this competition.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 459430,
          "author_name": "Elliot",
          "author_url": "",
          "post_date": "2019-01-21T17:59:35.230000",
          "content": "<p>im waiting for you to win this competition. jkjk\nHiding public LB would not be a viable solution either. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 461186,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-01-25T13:48:06.293000",
          "content": "<p>Looks like he has fallen over the Reichenbach Falls</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "459226": "LB = \"13% of the test data\"  As there are only 2624 targets total - is anyone else scared to death about using the LB as an indication of success?  341 targets is a very very tiny sample for me to be confident.",
    "459443": "yes, scared, but that's also the most challenging part in every modeling task, even out of competition context.\n\nI can think of some precautionary measures:\n\n(1) trust local CV\n(2) evaluate local CV with confidence intervals instead of a single number summary statistic alone.\n(3) evaluate models for, not just prediction accuracy,  but also their complexity and number of features used (lesser/fewer the better).\n(4) evaluate models with both local CV and public LB score, for example, average of CV and LB\n\nadd more by leaving comments below :)\n\n",
    "460073": "The shakeup potential in this competition is indeed huge.\n\nThe problem appears to be quite sensitive on the test data selection. I currently use 25-fold CV in model development. This makes each test fold about half the size of the public LB set. With this mode of operation, I see a standard deviation in the 0.10 to 0.13 range among the test fold MAEs.\n\nI intend to rely on local CV. Using the public LB for blend weight optimization or the like seems dangerous.",
    "460083": "@Scirpus, you are quite right.\n\nMy model with the best CV scores would put me about 120 points below where I am on the public LB right now. I believe proper error analysis is the key here. We have 4 months to figure it out so lets keep the discussions going :-)",
    "459229": "That's totally right. Someone had fallen off more than 3000 places, from the 2nd place on the final public LB to final private LB in the Santander Value Prediction Challenge not so long ago. \nhttps://www.kaggle.com/c/santander-value-prediction-challenge/discussion/63753#latest-373618",
    "459544": "The more I improve my CV score I do worst in the LB Score, I haven't found a good strategy to test, but maybe is just because of that 13%, how everyone is dealing with this ?",
    "459270": "I agree!\nAlso we have only 4194 samples in train (if we aggregate each 150000 rows), which could easily lead to overfitting.\n\nI tried sampling additional data - it significantly decreased score.",
    "459408": ""
  }
}