{
  "id": 30126,
  "title": "Public/Private leaderboard split",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/30126",
  "author_name": "Devin Anzelmo",
  "post_date": "2017-03-15T18:02:35.238000",
  "votes": 0,
  "comment_count": 12,
  "views": 0,
  "content": "<p>It seems I was able to overfit the public leaderboard by overfitting the training data?  My best public submission dropped 0.23 or so, and this submission had a decently good chance of being overfit to training.</p>\n\n<p>Post competition submission of much less overfit model had only 0.02 difference between public and private(though lower public score). With increasing epoch of this model the difference between public and private widened.</p>\n\n<p>So what was the split. It could potentially have been random that there was more correlation between public leaderboard and train set then private and train set but this seems like a little too much. Or is there another explanation? </p>",
  "messages": [
    {
      "id": 168111,
      "postDate": "2017-03-16T09:12:00.863Z",
      "content": "<p>In my case I had good correlations between cross-validation and public dataset. <br>\nHowever in private dataset my scores for roads and still water dropped from 0.8 to 0.3 and from 0.7 and 0.2.   </p>",
      "rawMarkdown": "In my case I had good correlations between cross-validation and public dataset.   \nHowever in private dataset my scores for roads and still water dropped from 0.8 to 0.3 and from 0.7 and 0.2.   ",
      "votes": 1
    },
    {
      "id": 167882,
      "postDate": "2017-03-15T18:26:02.970Z",
      "content": "<p>The percentage of each class was really different between public, private and training data.</p>\n\n<p>@threeplusone : you probably performed well on categories less represented in the public leaderboard</p>\n\n<p>@Devin : The split was inbalanced, and feels random yes.</p>\n\n<p>For example, cars were heavily unrepresented in the public leaderboard. I think it is one explanation for the evolutions at the top of the leaderboard.</p>",
      "rawMarkdown": "The percentage of each class was really different between public, private and training data.\n\n@threeplusone : you probably performed well on categories less represented in the public leaderboard\n\n@Devin : The split was inbalanced, and feels random yes.\n\nFor example, cars were heavily unrepresented in the public leaderboard. I think it is one explanation for the evolutions at the top of the leaderboard.",
      "votes": 1
    },
    {
      "id": 167887,
      "postDate": "2017-03-15T18:49:38.360Z",
      "content": "<p>Our score for the top submission per class. It is more or less stable, although for some classes there is a significant difference.</p>\n\n<pre><code>Public   Private\n\n0.07453 0.06290\n0.01905 0.02015\n0.08005 0.05605\n0.03281 0.03965\n0.05018 0.06984\n0.08251 0.08280\n0.09697 0.09131\n0.06081 0.05272\n0.02964 0.00331\n0.00186 0.00000\n</code></pre>",
      "rawMarkdown": "Our score for the top submission per class. It is more or less stable, although for some classes there is a significant difference.\n\n    Public\t Private\n\n    0.07453\t0.06290\n    0.01905\t0.02015\n    0.08005\t0.05605\n    0.03281\t0.03965\n    0.05018\t0.06984\n    0.08251\t0.08280\n    0.09697\t0.09131\n    0.06081\t0.05272\n    0.02964\t0.00331\n    0.00186\t0.00000\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 167892,
          "postDate": "2017-03-15T19:02:51.123Z",
          "content": "<p>I will have to do the same for the classes in my submission to track back what happened. Because I only really made one submission to leaderboard I was confused by drop. Will make the submissions to find out why. </p>",
          "rawMarkdown": "I will have to do the same for the classes in my submission to track back what happened. Because I only really made one submission to leaderboard I was confused by drop. Will make the submissions to find out why. "
        },
        {
          "id": 167895,
          "postDate": "2017-03-15T19:13:00.577Z",
          "content": "<p>Mine is a very different picture</p>\n\n<p>Private           Public</p>\n\n<p>0.28977         0.43451 </p>\n\n<p>0.28842        0.42758</p>\n\n<p>0.28249        0.41811</p>\n\n<p>0.34744        0.40776  </p>\n\n<p>And the submissions had the same class 8,9,10 so this huge \"nonlinearity\" is just from the first 7 classes (the cars imbalance can't explain it here). Well, i didn't thought that i should select manually, i let the system pick up best 2, i would never had guessed that the 4th public submission which is 7% behind on public leaderboard  gets 20% in front. Anyone has a bigger difference?</p>",
          "rawMarkdown": "Mine is a very different picture\n\nPrivate           Public\n\n0.28977\t        0.43451\t\n\n0.28842        0.42758\n\n0.28249        0.41811\n\n0.34744        0.40776\t\n\n\nAnd the submissions had the same class 8,9,10 so this huge \"nonlinearity\" is just from the first 7 classes (the cars imbalance can't explain it here). Well, i didn't thought that i should select manually, i let the system pick up best 2, i would never had guessed that the 4th public submission which is 7% behind on public leaderboard  gets 20% in front. Anyone has a bigger difference?"
        },
        {
          "id": 167898,
          "postDate": "2017-03-15T19:24:39.053Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 167979,
      "postDate": "2017-03-15T23:57:11.653Z",
      "content": "<p>I think Wendy has mostly answered the mystery on another thread. There were only 57 labeled images total. That means 25 in training set, and presumably about 6 in public and 26 in private test sets. Any tuning of model parameters against the public leaderboard is likely to lead to overfitting. </p>",
      "rawMarkdown": "I think Wendy has mostly answered the mystery on another thread. There were only 57 labeled images total. That means 25 in training set, and presumably about 6 in public and 26 in private test sets. Any tuning of model parameters against the public leaderboard is likely to lead to overfitting. ",
      "replies": [
        {
          "id": 167994,
          "postDate": "2017-03-16T00:57:18.293Z",
          "content": "<p>It definitely does account for the observed occurances. The reason for this thread was that i did not fit against the public leaderboard at all. I just did better on some classes overrepresented on the public leaderboard and worse on classes underrepresented on the public leaderboard. Frustrating especially with the week long wait for reveal. But overfitting the leaderboard is a common thing on kaggle and part of the fun, just surprising when not using the leaderboard for feedback. </p>",
          "rawMarkdown": "It definitely does account for the observed occurances. The reason for this thread was that i did not fit against the public leaderboard at all. I just did better on some classes overrepresented on the public leaderboard and worse on classes underrepresented on the public leaderboard. Frustrating especially with the week long wait for reveal. But overfitting the leaderboard is a common thing on kaggle and part of the fun, just surprising when not using the leaderboard for feedback. "
        }
      ]
    },
    {
      "id": 167886,
      "postDate": "2017-03-15T18:48:08.417Z",
      "content": "<p>Well its all explainable given that it clearly states public  leaderboard is calculated with approximately 1% of the test data and the final results will be based on the other 99%, so the final standings may be (very very) different. Well it would be nice that for upcoming competitions public leaderboard would represent at least 20 percent of test and also for best results for the sponsors if possible train should be larger than test, I mean its not like the ubiquitous ImageNet pretrained models are trained on 150k images and tested on 1.2M images but the other way around.</p>",
      "rawMarkdown": "Well its all explainable given that it clearly states public  leaderboard is calculated with approximately 1% of the test data and the final results will be based on the other 99%, so the final standings may be (very very) different. Well it would be nice that for upcoming competitions public leaderboard would represent at least 20 percent of test and also for best results for the sponsors if possible train should be larger than test, I mean its not like the ubiquitous ImageNet pretrained models are trained on 150k images and tested on 1.2M images but the other way around.",
      "replies": [
        {
          "id": 167888,
          "postDate": "2017-03-15T18:50:40.507Z",
          "content": "<p>So your saying the leaderboard bug was in fact telling the truth? A bug that is not a bug. </p>",
          "rawMarkdown": "So your saying the leaderboard bug was in fact telling the truth? A bug that is not a bug. "
        },
        {
          "id": 167889,
          "postDate": "2017-03-15T18:50:59.520Z",
          "content": "<p>Allegedly the listing of the 1%/99% split is a bug, and the true split was 18%/82%. It's discussed on the forum somewhere. The correct split used to be reported before the redesign. </p>",
          "rawMarkdown": "Allegedly the listing of the 1%/99% split is a bug, and the true split was 18%/82%. It's discussed on the forum somewhere. The correct split used to be reported before the redesign. "
        }
      ]
    },
    {
      "id": 167879,
      "postDate": "2017-03-15T18:15:58.617Z",
      "content": "<p>Weird, isn't it. The top model dropped by 9%. (58%-&gt;49%) Our model actually got better by 2%, and as a result we jumped 100 positions on the leaderboard to end 22nd. No idea why. </p>",
      "rawMarkdown": "Weird, isn't it. The top model dropped by 9%. (58%->49%) Our model actually got better by 2%, and as a result we jumped 100 positions on the leaderboard to end 22nd. No idea why. "
    },
    {
      "id": 167874,
      "postDate": "2017-03-15T18:02:35.240Z",
      "content": "<p>It seems I was able to overfit the public leaderboard by overfitting the training data?  My best public submission dropped 0.23 or so, and this submission had a decently good chance of being overfit to training.</p>\n\n<p>Post competition submission of much less overfit model had only 0.02 difference between public and private(though lower public score). With increasing epoch of this model the difference between public and private widened.</p>\n\n<p>So what was the split. It could potentially have been random that there was more correlation between public leaderboard and train set then private and train set but this seems like a little too much. Or is there another explanation? </p>",
      "rawMarkdown": "It seems I was able to overfit the public leaderboard by overfitting the training data?  My best public submission dropped 0.23 or so, and this submission had a decently good chance of being overfit to training.\n\nPost competition submission of much less overfit model had only 0.02 difference between public and private(though lower public score). With increasing epoch of this model the difference between public and private widened.\n\nSo what was the split. It could potentially have been random that there was more correlation between public leaderboard and train set then private and train set but this seems like a little too much. Or is there another explanation? "
    }
  ],
  "comments": [
    {
      "id": 168111,
      "author_name": "Guillermo Barbadillo",
      "author_url": "",
      "post_date": "2017-03-16T09:12:00.863000",
      "content": "<p>In my case I had good correlations between cross-validation and public dataset. <br>\nHowever in private dataset my scores for roads and still water dropped from 0.8 to 0.3 and from 0.7 and 0.2.   </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 167882,
      "author_name": "Max",
      "author_url": "",
      "post_date": "2017-03-15T18:26:02.970000",
      "content": "<p>The percentage of each class was really different between public, private and training data.</p>\n\n<p>@threeplusone : you probably performed well on categories less represented in the public leaderboard</p>\n\n<p>@Devin : The split was inbalanced, and feels random yes.</p>\n\n<p>For example, cars were heavily unrepresented in the public leaderboard. I think it is one explanation for the evolutions at the top of the leaderboard.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 167887,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2017-03-15T18:49:38.360000",
      "content": "<p>Our score for the top submission per class. It is more or less stable, although for some classes there is a significant difference.</p>\n\n<pre><code>Public   Private\n\n0.07453 0.06290\n0.01905 0.02015\n0.08005 0.05605\n0.03281 0.03965\n0.05018 0.06984\n0.08251 0.08280\n0.09697 0.09131\n0.06081 0.05272\n0.02964 0.00331\n0.00186 0.00000\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 167892,
          "author_name": "Devin Anzelmo",
          "author_url": "",
          "post_date": "2017-03-15T19:02:51.123000",
          "content": "<p>I will have to do the same for the classes in my submission to track back what happened. Because I only really made one submission to leaderboard I was confused by drop. Will make the submissions to find out why. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 167895,
          "author_name": "Bogdan Ghenea",
          "author_url": "",
          "post_date": "2017-03-15T19:13:00.577000",
          "content": "<p>Mine is a very different picture</p>\n\n<p>Private           Public</p>\n\n<p>0.28977         0.43451 </p>\n\n<p>0.28842        0.42758</p>\n\n<p>0.28249        0.41811</p>\n\n<p>0.34744        0.40776  </p>\n\n<p>And the submissions had the same class 8,9,10 so this huge \"nonlinearity\" is just from the first 7 classes (the cars imbalance can't explain it here). Well, i didn't thought that i should select manually, i let the system pick up best 2, i would never had guessed that the 4th public submission which is 7% behind on public leaderboard  gets 20% in front. Anyone has a bigger difference?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 167898,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-03-15T19:24:39.053000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 167979,
      "author_name": "threeplusone",
      "author_url": "",
      "post_date": "2017-03-15T23:57:11.653000",
      "content": "<p>I think Wendy has mostly answered the mystery on another thread. There were only 57 labeled images total. That means 25 in training set, and presumably about 6 in public and 26 in private test sets. Any tuning of model parameters against the public leaderboard is likely to lead to overfitting. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 167994,
          "author_name": "Devin Anzelmo",
          "author_url": "",
          "post_date": "2017-03-16T00:57:18.293000",
          "content": "<p>It definitely does account for the observed occurances. The reason for this thread was that i did not fit against the public leaderboard at all. I just did better on some classes overrepresented on the public leaderboard and worse on classes underrepresented on the public leaderboard. Frustrating especially with the week long wait for reveal. But overfitting the leaderboard is a common thing on kaggle and part of the fun, just surprising when not using the leaderboard for feedback. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 167886,
      "author_name": "Bogdan Ghenea",
      "author_url": "",
      "post_date": "2017-03-15T18:48:08.417000",
      "content": "<p>Well its all explainable given that it clearly states public  leaderboard is calculated with approximately 1% of the test data and the final results will be based on the other 99%, so the final standings may be (very very) different. Well it would be nice that for upcoming competitions public leaderboard would represent at least 20 percent of test and also for best results for the sponsors if possible train should be larger than test, I mean its not like the ubiquitous ImageNet pretrained models are trained on 150k images and tested on 1.2M images but the other way around.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 167888,
          "author_name": "Devin Anzelmo",
          "author_url": "",
          "post_date": "2017-03-15T18:50:40.507000",
          "content": "<p>So your saying the leaderboard bug was in fact telling the truth? A bug that is not a bug. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 167889,
          "author_name": "threeplusone",
          "author_url": "",
          "post_date": "2017-03-15T18:50:59.520000",
          "content": "<p>Allegedly the listing of the 1%/99% split is a bug, and the true split was 18%/82%. It's discussed on the forum somewhere. The correct split used to be reported before the redesign. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 167879,
      "author_name": "threeplusone",
      "author_url": "",
      "post_date": "2017-03-15T18:15:58.617000",
      "content": "<p>Weird, isn't it. The top model dropped by 9%. (58%-&gt;49%) Our model actually got better by 2%, and as a result we jumped 100 positions on the leaderboard to end 22nd. No idea why. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "168111": "In my case I had good correlations between cross-validation and public dataset.   \nHowever in private dataset my scores for roads and still water dropped from 0.8 to 0.3 and from 0.7 and 0.2.   ",
    "167882": "The percentage of each class was really different between public, private and training data.\n\n@threeplusone : you probably performed well on categories less represented in the public leaderboard\n\n@Devin : The split was inbalanced, and feels random yes.\n\nFor example, cars were heavily unrepresented in the public leaderboard. I think it is one explanation for the evolutions at the top of the leaderboard.",
    "167887": "Our score for the top submission per class. It is more or less stable, although for some classes there is a significant difference.\n\n    Public\t Private\n\n    0.07453\t0.06290\n    0.01905\t0.02015\n    0.08005\t0.05605\n    0.03281\t0.03965\n    0.05018\t0.06984\n    0.08251\t0.08280\n    0.09697\t0.09131\n    0.06081\t0.05272\n    0.02964\t0.00331\n    0.00186\t0.00000\n\n",
    "167979": "I think Wendy has mostly answered the mystery on another thread. There were only 57 labeled images total. That means 25 in training set, and presumably about 6 in public and 26 in private test sets. Any tuning of model parameters against the public leaderboard is likely to lead to overfitting. ",
    "167886": "Well its all explainable given that it clearly states public  leaderboard is calculated with approximately 1% of the test data and the final results will be based on the other 99%, so the final standings may be (very very) different. Well it would be nice that for upcoming competitions public leaderboard would represent at least 20 percent of test and also for best results for the sponsors if possible train should be larger than test, I mean its not like the ubiquitous ImageNet pretrained models are trained on 150k images and tested on 1.2M images but the other way around.",
    "167879": "Weird, isn't it. The top model dropped by 9%. (58%->49%) Our model actually got better by 2%, and as a result we jumped 100 positions on the leaderboard to end 22nd. No idea why. ",
    "167874": "It seems I was able to overfit the public leaderboard by overfitting the training data?  My best public submission dropped 0.23 or so, and this submission had a decently good chance of being overfit to training.\n\nPost competition submission of much less overfit model had only 0.02 difference between public and private(though lower public score). With increasing epoch of this model the difference between public and private widened.\n\nSo what was the split. It could potentially have been random that there was more correlation between public leaderboard and train set then private and train set but this seems like a little too much. Or is there another explanation? "
  }
}