{
  "id": 28771,
  "title": "How to validate ?",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/28771",
  "author_name": "",
  "post_date": "2017-02-13T16:19:18.523871700Z",
  "votes": 1,
  "comment_count": 12,
  "views": 0,
  "content": "<p>As this small dataset, currently I use all the data to train and validate, and I can get 0.47 on all the training data using the official evaluation metric and I can see nice segment on test images, but my lb is low..</p>\n\n<p>And if anyone want to cooperate on this, please contact me, tks!</p>",
  "messages": [
    {
      "id": "161407",
      "postDate": "02/13/2017 16:19:18",
      "content": "<p>As this small dataset, currently I use all the data to train and validate, and I can get 0.47 on all the training data using the official evaluation metric and I can see nice segment on test images, but my lb is low..</p>\n\n<p>And if anyone want to cooperate on this, please contact me, tks!</p>",
      "rawMarkdown": "As this small dataset, currently I use all the data to train and validate, and I can get 0.47 on all the training data using the official evaluation metric and I can see nice segment on test images, but my lb is low..\n\nAnd if anyone want to cooperate on this, please contact me, tks!",
      "votes": null
    },
    {
      "id": "161432",
      "postDate": "02/13/2017 18:32:38",
      "content": "<p>We've hit a similar problem. We get 0.51 against training data, but only 0.29 on public test. Either we're overtrained,  the test data is much harder than the training data, or we're doing something completely wrong.  </p>",
      "rawMarkdown": "We've hit a similar problem. We get 0.51 against training data, but only 0.29 on public test. Either we're overtrained,  the test data is much harder than the training data, or we're doing something completely wrong.",
      "votes": null
    },
    {
      "id": "161499",
      "postDate": "02/14/2017 02:34:01",
      "content": "<p>Proper validation is still an open question for me...</p>",
      "rawMarkdown": "Proper validation is still an open question for me...",
      "votes": null
    },
    {
      "id": "161504",
      "postDate": "02/14/2017 03:21:18",
      "content": "<p>The test data does not have the same distribution of classes as the train data. Certain classes clearly seem oversampled in the training data. This alone will cause a difference between validation and public leaderboard.(also make sure you are summing tp/fn/fp across all images and not computing your metric per image)</p>",
      "rawMarkdown": "The test data does not have the same distribution of classes as the train data. Certain classes clearly seem oversampled in the training data. This alone will cause a difference between validation and public leaderboard.(also make sure you are summing tp/fn/fp across all images and not computing your metric per image)",
      "votes": null
    },
    {
      "id": "161508",
      "postDate": "02/14/2017 03:41:36",
      "content": "<p>I find the miss-alignment between 3, A, M, P really matters, from the description A and M is 7.5m and 1.24m, that is about 6.04, but in the data, A band have height 134 and M band is 837, that is 6.24, I think this miss-alignment between pixels hurt the training process.</p>",
      "rawMarkdown": "I find the miss-alignment between 3, A, M, P really matters, from the description A and M is 7.5m and 1.24m, that is about 6.04, but in the data, A band have height 134 and M band is 837, that is 6.24, I think this miss-alignment between pixels hurt the training process.",
      "votes": null
    },
    {
      "id": "161509",
      "postDate": "02/14/2017 03:44:30",
      "content": "<p>I sum tp/fn/fp across all images, surely the testing distribution is different, but how can we overcome this?</p>",
      "rawMarkdown": "I sum tp/fn/fp across all images, surely the testing distribution is different, but how can we overcome this?",
      "votes": null
    },
    {
      "id": "161510",
      "postDate": "02/14/2017 03:48:29",
      "content": "<p>Then we can only use submissions ? I think submission is not useful  for me now.. I submit two result from the same model but use different epoches, the lb differs a lot(about 0.1..), this makes me think what's i'm training of.</p>",
      "rawMarkdown": "Then we can only use submissions ? I think submission is not useful  for me now.. I submit two result from the same model but use different epoches, the lb differs a lot(about 0.1..), this makes me think what's i'm training of.",
      "votes": null
    },
    {
      "id": "161511",
      "postDate": "02/14/2017 04:21:38",
      "content": "<p>Try to find a validation method such that if your validation score increases your leaderboard score increases. If there is a large difference between the scores(even 0.2) it may not be a sign there is anything wrong because the class distributions are different. The consistancy of the movement of the scores is more important then how close they are together.</p>\n\n<p>Though i imagine you are probably also seeing some fluctuation in the differences between validation and leaderboard as well.</p>",
      "rawMarkdown": "Try to find a validation method such that if your validation score increases your leaderboard score increases. If there is a large difference between the scores(even 0.2) it may not be a sign there is anything wrong because the class distributions are different. The consistancy of the movement of the scores is more important then how close they are together.\n\nThough i imagine you are probably also seeing some fluctuation in the differences between validation and leaderboard as well.",
      "votes": null
    },
    {
      "id": "161576",
      "postDate": "02/14/2017 16:04:46",
      "content": "<p>I noticed the scale discrepancy also. The A band images appear to be 7.75m resolution, not the claimed 7.5m. But that isn't where the misalignment comes from. Once you rescale all the images to the same size the original resolution becomes irrelevant.</p>\n\n<p>Most of the misalignment is because the 3 and (A,M,P) band images have different boundaries.  The misalignment between 3 and P in a 1km x 1km subregion can be large. But if you glue the 3-band and P-band 5x5 subregions into a single 5km x 5km region, then the discrepancy is only a few pixels, in size and alignment. </p>\n\n<p>Apparently the original data was the 5km x 5km regions, which were each chunked up into 25 subregions. But for no apparently good reason the 3 and (A,M,P) bands got chunked up slightly differently!?</p>",
      "rawMarkdown": "I noticed the scale discrepancy also. The A band images appear to be 7.75m resolution, not the claimed 7.5m. But that isn't where the misalignment comes from. Once you rescale all the images to the same size the original resolution becomes irrelevant.\n\nMost of the misalignment is because the 3 and (A,M,P) band images have different boundaries.  The misalignment between 3 and P in a 1km x 1km subregion can be large. But if you glue the 3-band and P-band 5x5 subregions into a single 5km x 5km region, then the discrepancy is only a few pixels, in size and alignment. \n\nApparently the original data was the 5km x 5km regions, which were each chunked up into 25 subregions. But for no apparently good reason the 3 and (A,M,P) bands got chunked up slightly differently!?",
      "votes": null
    },
    {
      "id": "161606",
      "postDate": "02/14/2017 20:17:44",
      "content": "<ol>\n<li>I am using a few images as a hold out set. Working kind of ok, bit not really.</li>\n<li>Visual inspection of the predictions may give an estimate of the model performance.</li>\n</ol>",
      "rawMarkdown": "1. I am using a few images as a hold out set. Working kind of ok, bit not really.\n 2. Visual inspection of the predictions may give an estimate of the model performance.",
      "votes": null
    },
    {
      "id": "162626",
      "postDate": "02/20/2017 11:12:22",
      "content": "<p>In this competition, only 1 % data is used for calculating public LB.<br>\nI think to see public LB score make overfit. We can only use holdout validation or cross validation.</p>",
      "rawMarkdown": "In this competition, only 1 % data is used for calculating public LB.<br>\nI think to see public LB score make overfit. We can only use holdout validation or cross validation.",
      "votes": null
    },
    {
      "id": "162629",
      "postDate": "02/20/2017 11:34:16",
      "content": "<p>Wendy wrote <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/28453\">here</a> that this is a display bug and the public testset is based on 18% of the testset.</p>",
      "rawMarkdown": "Wendy wrote [here][1] that this is a display bug and the public testset is based on 18% of the testset.\n\n\n  [1]: https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/28453",
      "votes": null
    },
    {
      "id": "162635",
      "postDate": "02/20/2017 12:01:26",
      "content": "<p>Oh! I overlooked this page.<br>\nThank you for your information.</p>",
      "rawMarkdown": "Oh! I overlooked this page.<br>\nThank you for your information.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 161432,
      "author_name": "threeplusone",
      "author_url": "",
      "post_date": "02/13/2017 18:32:38",
      "content": "<p>We've hit a similar problem. We get 0.51 against training data, but only 0.29 on public test. Either we're overtrained,  the test data is much harder than the training data, or we're doing something completely wrong.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 161508,
          "author_name": "zeliek",
          "author_url": "",
          "post_date": "02/14/2017 03:41:36",
          "content": "<p>I find the miss-alignment between 3, A, M, P really matters, from the description A and M is 7.5m and 1.24m, that is about 6.04, but in the data, A band have height 134 and M band is 837, that is 6.24, I think this miss-alignment between pixels hurt the training process.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 161576,
          "author_name": "threeplusone",
          "author_url": "",
          "post_date": "02/14/2017 16:04:46",
          "content": "<p>I noticed the scale discrepancy also. The A band images appear to be 7.75m resolution, not the claimed 7.5m. But that isn't where the misalignment comes from. Once you rescale all the images to the same size the original resolution becomes irrelevant.</p>\n\n<p>Most of the misalignment is because the 3 and (A,M,P) band images have different boundaries.  The misalignment between 3 and P in a 1km x 1km subregion can be large. But if you glue the 3-band and P-band 5x5 subregions into a single 5km x 5km region, then the discrepancy is only a few pixels, in size and alignment. </p>\n\n<p>Apparently the original data was the 5km x 5km regions, which were each chunked up into 25 subregions. But for no apparently good reason the 3 and (A,M,P) bands got chunked up slightly differently!?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 161499,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "02/14/2017 02:34:01",
      "content": "<p>Proper validation is still an open question for me...</p>",
      "votes": null,
      "replies": [
        {
          "id": 161510,
          "author_name": "zeliek",
          "author_url": "",
          "post_date": "02/14/2017 03:48:29",
          "content": "<p>Then we can only use submissions ? I think submission is not useful  for me now.. I submit two result from the same model but use different epoches, the lb differs a lot(about 0.1..), this makes me think what's i'm training of.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 161606,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "02/14/2017 20:17:44",
          "content": "<ol>\n<li>I am using a few images as a hold out set. Working kind of ok, bit not really.</li>\n<li>Visual inspection of the predictions may give an estimate of the model performance.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 161504,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/14/2017 03:21:18",
      "content": "<p>The test data does not have the same distribution of classes as the train data. Certain classes clearly seem oversampled in the training data. This alone will cause a difference between validation and public leaderboard.(also make sure you are summing tp/fn/fp across all images and not computing your metric per image)</p>",
      "votes": null,
      "replies": [
        {
          "id": 161509,
          "author_name": "zeliek",
          "author_url": "",
          "post_date": "02/14/2017 03:44:30",
          "content": "<p>I sum tp/fn/fp across all images, surely the testing distribution is different, but how can we overcome this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 161511,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "02/14/2017 04:21:38",
          "content": "<p>Try to find a validation method such that if your validation score increases your leaderboard score increases. If there is a large difference between the scores(even 0.2) it may not be a sign there is anything wrong because the class distributions are different. The consistancy of the movement of the scores is more important then how close they are together.</p>\n\n<p>Though i imagine you are probably also seeing some fluctuation in the differences between validation and leaderboard as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 162626,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "02/20/2017 11:12:22",
      "content": "<p>In this competition, only 1 % data is used for calculating public LB.<br>\nI think to see public LB score make overfit. We can only use holdout validation or cross validation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 162629,
          "author_name": "voltaire",
          "author_url": "",
          "post_date": "02/20/2017 11:34:16",
          "content": "<p>Wendy wrote <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/28453\">here</a> that this is a display bug and the public testset is based on 18% of the testset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 162635,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/20/2017 12:01:26",
          "content": "<p>Oh! I overlooked this page.<br>\nThank you for your information.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "161407": "As this small dataset, currently I use all the data to train and validate, and I can get 0.47 on all the training data using the official evaluation metric and I can see nice segment on test images, but my lb is low..\n\nAnd if anyone want to cooperate on this, please contact me, tks!",
    "161432": "We've hit a similar problem. We get 0.51 against training data, but only 0.29 on public test. Either we're overtrained,  the test data is much harder than the training data, or we're doing something completely wrong.",
    "161499": "Proper validation is still an open question for me...",
    "161504": "The test data does not have the same distribution of classes as the train data. Certain classes clearly seem oversampled in the training data. This alone will cause a difference between validation and public leaderboard.(also make sure you are summing tp/fn/fp across all images and not computing your metric per image)",
    "161508": "I find the miss-alignment between 3, A, M, P really matters, from the description A and M is 7.5m and 1.24m, that is about 6.04, but in the data, A band have height 134 and M band is 837, that is 6.24, I think this miss-alignment between pixels hurt the training process.",
    "161509": "I sum tp/fn/fp across all images, surely the testing distribution is different, but how can we overcome this?",
    "161510": "Then we can only use submissions ? I think submission is not useful  for me now.. I submit two result from the same model but use different epoches, the lb differs a lot(about 0.1..), this makes me think what's i'm training of.",
    "161511": "Try to find a validation method such that if your validation score increases your leaderboard score increases. If there is a large difference between the scores(even 0.2) it may not be a sign there is anything wrong because the class distributions are different. The consistancy of the movement of the scores is more important then how close they are together.\n\nThough i imagine you are probably also seeing some fluctuation in the differences between validation and leaderboard as well.",
    "161576": "I noticed the scale discrepancy also. The A band images appear to be 7.75m resolution, not the claimed 7.5m. But that isn't where the misalignment comes from. Once you rescale all the images to the same size the original resolution becomes irrelevant.\n\nMost of the misalignment is because the 3 and (A,M,P) band images have different boundaries.  The misalignment between 3 and P in a 1km x 1km subregion can be large. But if you glue the 3-band and P-band 5x5 subregions into a single 5km x 5km region, then the discrepancy is only a few pixels, in size and alignment. \n\nApparently the original data was the 5km x 5km regions, which were each chunked up into 25 subregions. But for no apparently good reason the 3 and (A,M,P) bands got chunked up slightly differently!?",
    "161606": "1. I am using a few images as a hold out set. Working kind of ok, bit not really.\n 2. Visual inspection of the predictions may give an estimate of the model performance.",
    "162626": "In this competition, only 1 % data is used for calculating public LB.<br>\nI think to see public LB score make overfit. We can only use holdout validation or cross validation.",
    "162629": "Wendy wrote [here][1] that this is a display bug and the public testset is based on 18% of the testset.\n\n\n  [1]: https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/28453",
    "162635": "Oh! I overlooked this page.<br>\nThank you for your information."
  },
  "source": "meta"
}