{
  "id": 70221,
  "title": "Inconsistency of test label generation",
  "url": "/competitions/airbus-ship-detection/discussion/70221",
  "author_name": "Iafoss",
  "post_date": "2018-11-01T03:39:01.075000",
  "votes": 12,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I think that many participants who reached ~0.725+ score have found that the test score (and even its relative change) is inconsistent with local validation. <a href=\"https://www.kaggle.com/iafoss/list-of-overlapping-images-for-validation-set\">Generation of a train set that does not overlap with validation one</a> does not really help with this discrepancy much. Below I put several examples:</p>\n\n<pre><code>Model S2:\n    public LB score: 0.733\n    false positives for empty masks (public LB): 0.012\n    f2 val (only on images with ships): 0.487\n    1 ship f2: 0.535\n    2-4 ships f2: 0.425\n    5-9 ships f2: 0.331\n    10+ f2: 0.250\n\nModel S2(+24 epochs):\n    public LB score: 0.731\n    false positives for empty masks (public LB): 0.011\n    f2 val (only on images with ships): 0.499\n    1 ship f2: 0.549\n    2-4 ships f2: 0.435\n    5-9 ships f2: 0.334\n    10+ f2: 0.260\n\nModel S3:\n    public LB score: 0.735\n    false positives for empty masks (public LB): 0.011\n    f2 val (only on images with ships): 0.550\n    1 ship f2: 0.609\n    2-4 ships f2: 0.472\n    5-9 ships f2: 0.365\n    10+ f2: 0.298\n\nModel S3(+20 epochs):\n    public LB score: 0.732\n    false positives for empty masks (public LB): 0.011\n    f2 val (only on images with ships): 0.576\n    1 ship f2: 0.640\n    2-4 ships f2: 0.489\n    5-9 ships f2: 0.377\n    10+ f2: 0.303\n</code></pre>\n\n<p>Even if validation score increases with continuing training, test score fluctuates or even decreases. Also more advanced model performs much better on validation set but has about the same test score...</p>\n\n<p>I have only one explanation for that: the new test dataset is generated with using a different method than the train data. I think many people already have realized that train labels are produced with machine labeling. Meanwhile, the test images may be labeled by humans. That is why the test set contains just ~4k images with ships, and organizers spent a month to prepare it... </p>\n\n<p>It would be interesting to hear the experience of other teams participating in this competition.</p>",
  "messages": [
    {
      "id": 413502,
      "postDate": "2018-11-01T03:39:01.077Z",
      "content": "<p>I think that many participants who reached ~0.725+ score have found that the test score (and even its relative change) is inconsistent with local validation. <a href=\"https://www.kaggle.com/iafoss/list-of-overlapping-images-for-validation-set\">Generation of a train set that does not overlap with validation one</a> does not really help with this discrepancy much. Below I put several examples:</p>\n\n<pre><code>Model S2:\n    public LB score: 0.733\n    false positives for empty masks (public LB): 0.012\n    f2 val (only on images with ships): 0.487\n    1 ship f2: 0.535\n    2-4 ships f2: 0.425\n    5-9 ships f2: 0.331\n    10+ f2: 0.250\n\nModel S2(+24 epochs):\n    public LB score: 0.731\n    false positives for empty masks (public LB): 0.011\n    f2 val (only on images with ships): 0.499\n    1 ship f2: 0.549\n    2-4 ships f2: 0.435\n    5-9 ships f2: 0.334\n    10+ f2: 0.260\n\nModel S3:\n    public LB score: 0.735\n    false positives for empty masks (public LB): 0.011\n    f2 val (only on images with ships): 0.550\n    1 ship f2: 0.609\n    2-4 ships f2: 0.472\n    5-9 ships f2: 0.365\n    10+ f2: 0.298\n\nModel S3(+20 epochs):\n    public LB score: 0.732\n    false positives for empty masks (public LB): 0.011\n    f2 val (only on images with ships): 0.576\n    1 ship f2: 0.640\n    2-4 ships f2: 0.489\n    5-9 ships f2: 0.377\n    10+ f2: 0.303\n</code></pre>\n\n<p>Even if validation score increases with continuing training, test score fluctuates or even decreases. Also more advanced model performs much better on validation set but has about the same test score...</p>\n\n<p>I have only one explanation for that: the new test dataset is generated with using a different method than the train data. I think many people already have realized that train labels are produced with machine labeling. Meanwhile, the test images may be labeled by humans. That is why the test set contains just ~4k images with ships, and organizers spent a month to prepare it... </p>\n\n<p>It would be interesting to hear the experience of other teams participating in this competition.</p>",
      "rawMarkdown": "I think that many participants who reached ~0.725+ score have found that the test score (and even its relative change) is inconsistent with local validation. [Generation of a train set that does not overlap with validation one][1] does not really help with this discrepancy much. Below I put several examples:\n\n    Model S2:\n        public LB score: 0.733\n        false positives for empty masks (public LB): 0.012\n        f2 val (only on images with ships): 0.487\n        1 ship f2: 0.535\n        2-4 ships f2: 0.425\n        5-9 ships f2: 0.331\n        10+ f2: 0.250\n    \n    Model S2(+24 epochs):\n        public LB score: 0.731\n        false positives for empty masks (public LB): 0.011\n        f2 val (only on images with ships): 0.499\n        1 ship f2: 0.549\n        2-4 ships f2: 0.435\n        5-9 ships f2: 0.334\n        10+ f2: 0.260\n    \n    Model S3:\n        public LB score: 0.735\n        false positives for empty masks (public LB): 0.011\n        f2 val (only on images with ships): 0.550\n        1 ship f2: 0.609\n        2-4 ships f2: 0.472\n        5-9 ships f2: 0.365\n        10+ f2: 0.298\n    \n    Model S3(+20 epochs):\n        public LB score: 0.732\n        false positives for empty masks (public LB): 0.011\n        f2 val (only on images with ships): 0.576\n        1 ship f2: 0.640\n        2-4 ships f2: 0.489\n        5-9 ships f2: 0.377\n        10+ f2: 0.303\n\nEven if validation score increases with continuing training, test score fluctuates or even decreases. Also more advanced model performs much better on validation set but has about the same test score...\n\nI have only one explanation for that: the new test dataset is generated with using a different method than the train data. I think many people already have realized that train labels are produced with machine labeling. Meanwhile, the test images may be labeled by humans. That is why the test set contains just ~4k images with ships, and organizers spent a month to prepare it... \n\nIt would be interesting to hear the experience of other teams participating in this competition.\n\n\n  [1]: https://www.kaggle.com/iafoss/list-of-overlapping-images-for-validation-set",
      "votes": 12
    },
    {
      "id": 418417,
      "postDate": "2018-11-09T21:56:29.600Z",
      "content": "<p>Hi, Lafoss, I am stuck now as I found my local validation still overestimate the f2score even with considering the big image problem. I am wondering how do you generate good validation dataset at least still works when LB reach 0.725 level?</p>\n\n<p>For example, I have an model with LB 0.712, but my local validation with only ship images is already 0.55 level</p>",
      "rawMarkdown": "Hi, Lafoss, I am stuck now as I found my local validation still overestimate the f2score even with considering the big image problem. I am wondering how do you generate good validation dataset at least still works when LB reach 0.725 level?\n\nFor example, I have an model with LB 0.712, but my local validation with only ship images is already 0.55 level",
      "votes": 1,
      "replies": [
        {
          "id": 418495,
          "postDate": "2018-11-10T01:36:42.747Z",
          "content": "<p>I don't think that your public LB score will increase much if you continue training... and I expect that results of your model trained for less number of epochs are better. The main challenge of this competition is the inconsistency of the second test set. Form my experience, after reaching ~0.47 F2 you start having problems... and out team is working on solving them. One thing that you should make sure that there is no leakage between train and val sets.</p>",
          "rawMarkdown": "I don't think that your public LB score will increase much if you continue training... and I expect that results of your model trained for less number of epochs are better. The main challenge of this competition is the inconsistency of the second test set. Form my experience, after reaching ~0.47 F2 you start having problems... and out team is working on solving them. One thing that you should make sure that there is no leakage between train and val sets."
        },
        {
          "id": 418514,
          "postDate": "2018-11-10T02:17:49.857Z",
          "content": "<p>How to avoid leak? I use the csv file from <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/69322\">https://www.kaggle.com/c/airbus-ship-detection/discussion/69322</a>, and split the big image id first, then add all the small image from big image to either train or split. Is there anything I am missing?</p>",
          "rawMarkdown": "How to avoid leak? I use the csv file from https://www.kaggle.com/c/airbus-ship-detection/discussion/69322, and split the big image id first, then add all the small image from big image to either train or split. Is there anything I am missing?",
          "votes": 1
        },
        {
          "id": 418559,
          "postDate": "2018-11-10T05:41:10.450Z",
          "content": "<p>It sounds reasonable</p>",
          "rawMarkdown": "It sounds reasonable"
        }
      ]
    },
    {
      "id": 413686,
      "postDate": "2018-11-01T11:13:53.147Z",
      "content": "<p>disposed old comment due to age</p>",
      "rawMarkdown": "disposed old comment due to age",
      "votes": 1,
      "replies": [
        {
          "id": 414073,
          "postDate": "2018-11-02T03:44:14.597Z",
          "content": "<p>Agree, there are lot of issues with the data, e.g. missing ship labels, incorrect rles, etc.  I am not sure it is best to dig too much on this matter since the final evaluation will be based on the ground truth they prepared after all. </p>",
          "rawMarkdown": "Agree, there are lot of issues with the data, e.g. missing ship labels, incorrect rles, etc.  I am not sure it is best to dig too much on this matter since the final evaluation will be based on the ground truth they prepared after all. "
        }
      ]
    },
    {
      "id": 413551,
      "postDate": "2018-11-01T06:05:21.780Z",
      "content": "<p>LB in the camp is generated base on only 12% of test set,  consider that the test set have only  4k+ image, the public LB have about 500 image with ship (4k * 0.12 * 0.48). so the score should be inconsistency as we seen.</p>\n\n<p>BTW, there are some odd image in train set, I find a image which a AIRPLANE yet its engine(I guess, only ~20pixel) mark as a ship.</p>",
      "rawMarkdown": "LB in the camp is generated base on only 12% of test set,  consider that the test set have only  4k+ image, the public LB have about 500 image with ship (4k * 0.12 * 0.48). so the score should be inconsistency as we seen.\n\nBTW, there are some odd image in train set, I find a image which a AIRPLANE yet its engine(I guess, only ~20pixel) mark as a ship.",
      "votes": 1,
      "replies": [
        {
          "id": 420376,
          "postDate": "2018-11-13T14:12:44.703Z",
          "content": "<p>Do you mean the test set have 4k+ imgae with ship? How do you konw ?</p>",
          "rawMarkdown": "Do you mean the test set have 4k+ imgae with ship? How do you konw ?"
        }
      ]
    },
    {
      "id": 413668,
      "postDate": "2018-11-01T10:39:29.777Z",
      "content": "<p>Are you sure about your local validation score for <code>(only on images with ships)</code>? It seems very high. I get only 0.453 (on non-empty) for my best submission (LB 0.73). My local validation scores are close to my public scores. So far, a better validation score resulted in a better public score.</p>\n\n<pre><code> I think many people already have realized that train labels are produced with machine labeling\n</code></pre>\n\n<p>I didn't realize that. Could you give an example?</p>",
      "rawMarkdown": "Are you sure about your local validation score for `(only on images with ships)`? It seems very high. I get only 0.453 (on non-empty) for my best submission (LB 0.73). My local validation scores are close to my public scores. So far, a better validation score resulted in a better public score.\n\n     I think many people already have realized that train labels are produced with machine labeling\n\nI didn't realize that. Could you give an example?",
      "votes": 2,
      "replies": [
        {
          "id": 413777,
          "postDate": "2018-11-01T14:16:22.073Z",
          "content": "<p>I second this. CV scores from the OP seem to be <strong>extremely</strong> high.</p>",
          "rawMarkdown": "I second this. CV scores from the OP seem to be **extremely** high."
        },
        {
          "id": 413833,
          "postDate": "2018-11-01T16:15:44.270Z",
          "content": "<p>You can check my response to Dmitriy Danevskiy on machine labeling. Regarding the way of validation score calculation, you may check <a href=\"https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb\">this kernel</a>. The only difference is that I use train set that does not overlap with validation one. Also I added separate calculation of score for images with different number of ships. The size of the validation set is 10% of train, so I have ~4k images with ships. My S1 model was giving something similar to your results, like ~0.47 that gave 0.728 public LB. At that moment I was quite optimistic about improvement the model to get higher score. I worked quite hard to improve the model architecture, and right now it can be even pushed to ~0.6... while it looks there is no point of doing it. There is no much difference of what is the architecture of the model, etc., everything will end up in 0.73-0.74 public LB score range...</p>",
          "rawMarkdown": "You can check my response to Dmitriy Danevskiy on machine labeling. Regarding the way of validation score calculation, you may check [this kernel][1]. The only difference is that I use train set that does not overlap with validation one. Also I added separate calculation of score for images with different number of ships. The size of the validation set is 10% of train, so I have ~4k images with ships. My S1 model was giving something similar to your results, like ~0.47 that gave 0.728 public LB. At that moment I was quite optimistic about improvement the model to get higher score. I worked quite hard to improve the model architecture, and right now it can be even pushed to ~0.6... while it looks there is no point of doing it. There is no much difference of what is the architecture of the model, etc., everything will end up in 0.73-0.74 public LB score range...\n\n\n  [1]: https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb"
        }
      ]
    },
    {
      "id": 422424,
      "postDate": "2018-11-16T07:47:50.490Z",
      "content": "<p>Actually, I think we have found a problem. As @Train has pointed out, the only method working for detection of overlaps is based on <a href=\"https://www.kaggle.com/manuscrits/create-a-validation-dataset-correcting-the-leak\">checking the masks</a> rather than images.  I think the problem is that the images are saved in jpeg format, and overlapping patches may be a little bit different from one image to another. Right now I'm training another model with this new set, and it seems to behave much better.</p>",
      "rawMarkdown": "Actually, I think we have found a problem. As @Train has pointed out, the only method working for detection of overlaps is based on [checking the masks][1] rather than images.  I think the problem is that the images are saved in jpeg format, and overlapping patches may be a little bit different from one image to another. Right now I'm training another model with this new set, and it seems to behave much better.\n\n\n  [1]: https://www.kaggle.com/manuscrits/create-a-validation-dataset-correcting-the-leak"
    },
    {
      "id": 413776,
      "postDate": "2018-11-01T14:14:31.063Z",
      "content": "<p>For me personally, ship masks don't look machine-generated at all. Masks rarely align with ship borders, which is a very strong evidence they were NOT generated by an algorithm. I believe any ship detection algorithm, either NN-based or not, would provide masks with good border alignment.</p>",
      "rawMarkdown": "For me personally, ship masks don't look machine-generated at all. Masks rarely align with ship borders, which is a very strong evidence they were NOT generated by an algorithm. I believe any ship detection algorithm, either NN-based or not, would provide masks with good border alignment.",
      "replies": [
        {
          "id": 413810,
          "postDate": "2018-11-01T15:30:08.670Z",
          "content": "<p>You can check <a href=\"https://arxiv.org/pdf/1711.09405.pdf\">this paper</a>, specifically detection of ships. Also, you can read <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/67286\">the post</a> and related discussion, where you can find examples of incorrect masks that could be only produced by machine labeling. I just checked 100 images and found 2 that are completely off, there must be much more in entire set. Also, there is large number of tiny ships where the box is misaligned, and the model performs even better than the \"ground truth\" one.  </p>",
          "rawMarkdown": "You can check [this paper][1], specifically detection of ships. Also, you can read [the post][2] and related discussion, where you can find examples of incorrect masks that could be only produced by machine labeling. I just checked 100 images and found 2 that are completely off, there must be much more in entire set. Also, there is large number of tiny ships where the box is misaligned, and the model performs even better than the \"ground truth\" one.  \n\n\n  [1]: https://arxiv.org/pdf/1711.09405.pdf\n  [2]: https://www.kaggle.com/c/airbus-ship-detection/discussion/67286",
          "votes": 1
        }
      ]
    },
    {
      "id": 413650,
      "postDate": "2018-11-01T09:55:22.303Z",
      "content": "<p>Hi lafoss，Did you use the cv to get 0.735+ </p>",
      "rawMarkdown": "Hi lafoss，Did you use the cv to get 0.735+ ",
      "replies": [
        {
          "id": 413836,
          "postDate": "2018-11-01T16:18:36.980Z",
          "content": "<p>As I wrote in the post, cv gives irrelevant results after reaching some public LB score. At least based on my submissions after reaching ~0.47+ validation score on images with ships, public LB results really stop changing.</p>",
          "rawMarkdown": "As I wrote in the post, cv gives irrelevant results after reaching some public LB score. At least based on my submissions after reaching ~0.47+ validation score on images with ships, public LB results really stop changing."
        },
        {
          "id": 421986,
          "postDate": "2018-11-15T16:39:41.463Z",
          "content": "<p>same here... beyond CV loss of .118 public Lb was dropping....i had pointed out difference between train and test in very beginning itself..</p>",
          "rawMarkdown": "same here... beyond CV loss of .118 public Lb was dropping....i had pointed out difference between train and test in very beginning itself.."
        }
      ]
    },
    {
      "id": 422564,
      "postDate": "2018-11-16T12:16:56.550Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 413708,
      "postDate": "2018-11-01T11:54:48.917Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 413712,
          "postDate": "2018-11-01T12:02:19.783Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 413815,
          "postDate": "2018-11-01T15:41:25.673Z",
          "content": "<p>It is very simple to measure false positives. You just take your prediction, replace nan by '1 2' and any non nan prediction by nan, and submit the file. In this case the score of the model will be the fraction of images predicted as ones with ships even if they are empty.</p>",
          "rawMarkdown": "It is very simple to measure false positives. You just take your prediction, replace nan by '1 2' and any non nan prediction by nan, and submit the file. In this case the score of the model will be the fraction of images predicted as ones with ships even if they are empty.",
          "votes": 3
        },
        {
          "id": 413847,
          "postDate": "2018-11-01T16:30:40.427Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 418417,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2018-11-09T21:56:29.600000",
      "content": "<p>Hi, Lafoss, I am stuck now as I found my local validation still overestimate the f2score even with considering the big image problem. I am wondering how do you generate good validation dataset at least still works when LB reach 0.725 level?</p>\n\n<p>For example, I have an model with LB 0.712, but my local validation with only ship images is already 0.55 level</p>",
      "votes": 1,
      "replies": [
        {
          "id": 418495,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-11-10T01:36:42.747000",
          "content": "<p>I don't think that your public LB score will increase much if you continue training... and I expect that results of your model trained for less number of epochs are better. The main challenge of this competition is the inconsistency of the second test set. Form my experience, after reaching ~0.47 F2 you start having problems... and out team is working on solving them. One thing that you should make sure that there is no leakage between train and val sets.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418514,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2018-11-10T02:17:49.857000",
          "content": "<p>How to avoid leak? I use the csv file from <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/69322\">https://www.kaggle.com/c/airbus-ship-detection/discussion/69322</a>, and split the big image id first, then add all the small image from big image to either train or split. Is there anything I am missing?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 418559,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-11-10T05:41:10.450000",
          "content": "<p>It sounds reasonable</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 413686,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2018-11-01T11:13:53.147000",
      "content": "<p>disposed old comment due to age</p>",
      "votes": 1,
      "replies": [
        {
          "id": 414073,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2018-11-02T03:44:14.597000",
          "content": "<p>Agree, there are lot of issues with the data, e.g. missing ship labels, incorrect rles, etc.  I am not sure it is best to dig too much on this matter since the final evaluation will be based on the ground truth they prepared after all. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 413551,
      "author_name": "sheep",
      "author_url": "",
      "post_date": "2018-11-01T06:05:21.780000",
      "content": "<p>LB in the camp is generated base on only 12% of test set,  consider that the test set have only  4k+ image, the public LB have about 500 image with ship (4k * 0.12 * 0.48). so the score should be inconsistency as we seen.</p>\n\n<p>BTW, there are some odd image in train set, I find a image which a AIRPLANE yet its engine(I guess, only ~20pixel) mark as a ship.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 420376,
          "author_name": "Kevin Zheng",
          "author_url": "",
          "post_date": "2018-11-13T14:12:44.703000",
          "content": "<p>Do you mean the test set have 4k+ imgae with ship? How do you konw ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 413668,
      "author_name": "See--",
      "author_url": "",
      "post_date": "2018-11-01T10:39:29.777000",
      "content": "<p>Are you sure about your local validation score for <code>(only on images with ships)</code>? It seems very high. I get only 0.453 (on non-empty) for my best submission (LB 0.73). My local validation scores are close to my public scores. So far, a better validation score resulted in a better public score.</p>\n\n<pre><code> I think many people already have realized that train labels are produced with machine labeling\n</code></pre>\n\n<p>I didn't realize that. Could you give an example?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 413777,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2018-11-01T14:16:22.073000",
          "content": "<p>I second this. CV scores from the OP seem to be <strong>extremely</strong> high.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413833,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-11-01T16:15:44.270000",
          "content": "<p>You can check my response to Dmitriy Danevskiy on machine labeling. Regarding the way of validation score calculation, you may check <a href=\"https://www.kaggle.com/iafoss/unet34-submission-tta-0-699-new-public-lb\">this kernel</a>. The only difference is that I use train set that does not overlap with validation one. Also I added separate calculation of score for images with different number of ships. The size of the validation set is 10% of train, so I have ~4k images with ships. My S1 model was giving something similar to your results, like ~0.47 that gave 0.728 public LB. At that moment I was quite optimistic about improvement the model to get higher score. I worked quite hard to improve the model architecture, and right now it can be even pushed to ~0.6... while it looks there is no point of doing it. There is no much difference of what is the architecture of the model, etc., everything will end up in 0.73-0.74 public LB score range...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422424,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2018-11-16T07:47:50.490000",
      "content": "<p>Actually, I think we have found a problem. As @Train has pointed out, the only method working for detection of overlaps is based on <a href=\"https://www.kaggle.com/manuscrits/create-a-validation-dataset-correcting-the-leak\">checking the masks</a> rather than images.  I think the problem is that the images are saved in jpeg format, and overlapping patches may be a little bit different from one image to another. Right now I'm training another model with this new set, and it seems to behave much better.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 413776,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2018-11-01T14:14:31.063000",
      "content": "<p>For me personally, ship masks don't look machine-generated at all. Masks rarely align with ship borders, which is a very strong evidence they were NOT generated by an algorithm. I believe any ship detection algorithm, either NN-based or not, would provide masks with good border alignment.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 413810,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-11-01T15:30:08.670000",
          "content": "<p>You can check <a href=\"https://arxiv.org/pdf/1711.09405.pdf\">this paper</a>, specifically detection of ships. Also, you can read <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/67286\">the post</a> and related discussion, where you can find examples of incorrect masks that could be only produced by machine labeling. I just checked 100 images and found 2 that are completely off, there must be much more in entire set. Also, there is large number of tiny ships where the box is misaligned, and the model performs even better than the \"ground truth\" one.  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 413650,
      "author_name": "TPloveYXT520",
      "author_url": "",
      "post_date": "2018-11-01T09:55:22.303000",
      "content": "<p>Hi lafoss，Did you use the cv to get 0.735+ </p>",
      "votes": 0,
      "replies": [
        {
          "id": 413836,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-11-01T16:18:36.980000",
          "content": "<p>As I wrote in the post, cv gives irrelevant results after reaching some public LB score. At least based on my submissions after reaching ~0.47+ validation score on images with ships, public LB results really stop changing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421986,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2018-11-15T16:39:41.463000",
          "content": "<p>same here... beyond CV loss of .118 public Lb was dropping....i had pointed out difference between train and test in very beginning itself..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422564,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-16T12:16:56.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 413708,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-01T11:54:48.917000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 413712,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-01T12:02:19.783000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 413815,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-11-01T15:41:25.673000",
          "content": "<p>It is very simple to measure false positives. You just take your prediction, replace nan by '1 2' and any non nan prediction by nan, and submit the file. In this case the score of the model will be the fraction of images predicted as ones with ships even if they are empty.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 413847,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-01T16:30:40.427000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "413502": "I think that many participants who reached ~0.725+ score have found that the test score (and even its relative change) is inconsistent with local validation. [Generation of a train set that does not overlap with validation one][1] does not really help with this discrepancy much. Below I put several examples:\n\n    Model S2:\n        public LB score: 0.733\n        false positives for empty masks (public LB): 0.012\n        f2 val (only on images with ships): 0.487\n        1 ship f2: 0.535\n        2-4 ships f2: 0.425\n        5-9 ships f2: 0.331\n        10+ f2: 0.250\n    \n    Model S2(+24 epochs):\n        public LB score: 0.731\n        false positives for empty masks (public LB): 0.011\n        f2 val (only on images with ships): 0.499\n        1 ship f2: 0.549\n        2-4 ships f2: 0.435\n        5-9 ships f2: 0.334\n        10+ f2: 0.260\n    \n    Model S3:\n        public LB score: 0.735\n        false positives for empty masks (public LB): 0.011\n        f2 val (only on images with ships): 0.550\n        1 ship f2: 0.609\n        2-4 ships f2: 0.472\n        5-9 ships f2: 0.365\n        10+ f2: 0.298\n    \n    Model S3(+20 epochs):\n        public LB score: 0.732\n        false positives for empty masks (public LB): 0.011\n        f2 val (only on images with ships): 0.576\n        1 ship f2: 0.640\n        2-4 ships f2: 0.489\n        5-9 ships f2: 0.377\n        10+ f2: 0.303\n\nEven if validation score increases with continuing training, test score fluctuates or even decreases. Also more advanced model performs much better on validation set but has about the same test score...\n\nI have only one explanation for that: the new test dataset is generated with using a different method than the train data. I think many people already have realized that train labels are produced with machine labeling. Meanwhile, the test images may be labeled by humans. That is why the test set contains just ~4k images with ships, and organizers spent a month to prepare it... \n\nIt would be interesting to hear the experience of other teams participating in this competition.\n\n\n  [1]: https://www.kaggle.com/iafoss/list-of-overlapping-images-for-validation-set",
    "418417": "Hi, Lafoss, I am stuck now as I found my local validation still overestimate the f2score even with considering the big image problem. I am wondering how do you generate good validation dataset at least still works when LB reach 0.725 level?\n\nFor example, I have an model with LB 0.712, but my local validation with only ship images is already 0.55 level",
    "413686": "disposed old comment due to age",
    "413551": "LB in the camp is generated base on only 12% of test set,  consider that the test set have only  4k+ image, the public LB have about 500 image with ship (4k * 0.12 * 0.48). so the score should be inconsistency as we seen.\n\nBTW, there are some odd image in train set, I find a image which a AIRPLANE yet its engine(I guess, only ~20pixel) mark as a ship.",
    "413668": "Are you sure about your local validation score for `(only on images with ships)`? It seems very high. I get only 0.453 (on non-empty) for my best submission (LB 0.73). My local validation scores are close to my public scores. So far, a better validation score resulted in a better public score.\n\n     I think many people already have realized that train labels are produced with machine labeling\n\nI didn't realize that. Could you give an example?",
    "422424": "Actually, I think we have found a problem. As @Train has pointed out, the only method working for detection of overlaps is based on [checking the masks][1] rather than images.  I think the problem is that the images are saved in jpeg format, and overlapping patches may be a little bit different from one image to another. Right now I'm training another model with this new set, and it seems to behave much better.\n\n\n  [1]: https://www.kaggle.com/manuscrits/create-a-validation-dataset-correcting-the-leak",
    "413776": "For me personally, ship masks don't look machine-generated at all. Masks rarely align with ship borders, which is a very strong evidence they were NOT generated by an algorithm. I believe any ship detection algorithm, either NN-based or not, would provide masks with good border alignment.",
    "413650": "Hi lafoss，Did you use the cv to get 0.735+ ",
    "422564": "",
    "413708": ""
  }
}