{
  "id": 68082,
  "title": "New training data",
  "url": "/competitions/airbus-ship-detection/discussion/68082",
  "author_name": "Rüdiger Jungbeck",
  "post_date": "2018-10-09T06:43:45.394000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>From what I see, training_v2 is a combination of the previous training and test images.\nThere was a serious leak between the previous training and test images (that is why many reached a score of 1.000) caused by using identical patches for generating both.\nSo when we now use this combination (trainings_v2) and split it somehow in training and validation images (as most of us will do), trainings and validation images will start having the same problem. So validation will become useless (and the models become overfitting)\nSo it might be better to stick with the original training images (maybe cleaned up up for the corrupted images). There is not much to be learned from the old test images, that can't be achieved with augmentation (it is based in the same patches).</p>",
  "messages": [
    {
      "id": 401018,
      "postDate": "2018-10-09T09:29:43.353Z",
      "content": "<p>I agree, they should re-generate the training set from scratch without the overlapping, shouldn't take more than a few days I guess, since all the labels are already there, and I'm willing to wait. Forcing the competitors to clean up the mess to create their own set of clean, non overlapped images is so unnecessary.</p>",
      "rawMarkdown": "I agree, they should re-generate the training set from scratch without the overlapping, shouldn't take more than a few days I guess, since all the labels are already there, and I'm willing to wait. Forcing the competitors to clean up the mess to create their own set of clean, non overlapped images is so unnecessary.",
      "votes": 3,
      "replies": [
        {
          "id": 402239,
          "postDate": "2018-10-11T11:58:51.197Z",
          "content": "<p>indeed, but the organisators mentionned this information has been lost</p>",
          "rawMarkdown": "indeed, but the organisators mentionned this information has been lost"
        },
        {
          "id": 402247,
          "postDate": "2018-10-11T12:14:31.920Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 400977,
      "postDate": "2018-10-09T08:19:45.323Z",
      "content": "<p>Well, the old training data was also mutually overlapping. You have to work to get a clean set of local validation data.</p>",
      "rawMarkdown": "Well, the old training data was also mutually overlapping. You have to work to get a clean set of local validation data.",
      "votes": 3
    },
    {
      "id": 400932,
      "postDate": "2018-10-09T06:43:45.393Z",
      "content": "<p>From what I see, training_v2 is a combination of the previous training and test images.\nThere was a serious leak between the previous training and test images (that is why many reached a score of 1.000) caused by using identical patches for generating both.\nSo when we now use this combination (trainings_v2) and split it somehow in training and validation images (as most of us will do), trainings and validation images will start having the same problem. So validation will become useless (and the models become overfitting)\nSo it might be better to stick with the original training images (maybe cleaned up up for the corrupted images). There is not much to be learned from the old test images, that can't be achieved with augmentation (it is based in the same patches).</p>",
      "rawMarkdown": "From what I see, training_v2 is a combination of the previous training and test images.\nThere was a serious leak between the previous training and test images (that is why many reached a score of 1.000) caused by using identical patches for generating both.\nSo when we now use this combination (trainings_v2) and split it somehow in training and validation images (as most of us will do), trainings and validation images will start having the same problem. So validation will become useless (and the models become overfitting)\nSo it might be better to stick with the original training images (maybe cleaned up up for the corrupted images). There is not much to be learned from the old test images, that can't be achieved with augmentation (it is based in the same patches).",
      "votes": 2
    },
    {
      "id": 402147,
      "postDate": "2018-10-11T08:38:00.637Z",
      "content": "<p>The new training set seems indeed to be a very simple merge of old training and test set. Just checking the first few lines in the new training csv file, the hashed filenames also seems to have stayed the same.</p>\n\n<p>So using the old training set could be achieved by simple switching to the old training csv file, not much other changes required in most cases.</p>",
      "rawMarkdown": "The new training set seems indeed to be a very simple merge of old training and test set. Just checking the first few lines in the new training csv file, the hashed filenames also seems to have stayed the same.\n\nSo using the old training set could be achieved by simple switching to the old training csv file, not much other changes required in most cases."
    }
  ],
  "comments": [
    {
      "id": 401018,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2018-10-09T09:29:43.353000",
      "content": "<p>I agree, they should re-generate the training set from scratch without the overlapping, shouldn't take more than a few days I guess, since all the labels are already there, and I'm willing to wait. Forcing the competitors to clean up the mess to create their own set of clean, non overlapped images is so unnecessary.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 402239,
          "author_name": "Max",
          "author_url": "",
          "post_date": "2018-10-11T11:58:51.197000",
          "content": "<p>indeed, but the organisators mentionned this information has been lost</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 402247,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-10-11T12:14:31.920000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 400977,
      "author_name": "PeterSorensen",
      "author_url": "",
      "post_date": "2018-10-09T08:19:45.323000",
      "content": "<p>Well, the old training data was also mutually overlapping. You have to work to get a clean set of local validation data.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 402147,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2018-10-11T08:38:00.637000",
      "content": "<p>The new training set seems indeed to be a very simple merge of old training and test set. Just checking the first few lines in the new training csv file, the hashed filenames also seems to have stayed the same.</p>\n\n<p>So using the old training set could be achieved by simple switching to the old training csv file, not much other changes required in most cases.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "401018": "I agree, they should re-generate the training set from scratch without the overlapping, shouldn't take more than a few days I guess, since all the labels are already there, and I'm willing to wait. Forcing the competitors to clean up the mess to create their own set of clean, non overlapped images is so unnecessary.",
    "400977": "Well, the old training data was also mutually overlapping. You have to work to get a clean set of local validation data.",
    "400932": "From what I see, training_v2 is a combination of the previous training and test images.\nThere was a serious leak between the previous training and test images (that is why many reached a score of 1.000) caused by using identical patches for generating both.\nSo when we now use this combination (trainings_v2) and split it somehow in training and validation images (as most of us will do), trainings and validation images will start having the same problem. So validation will become useless (and the models become overfitting)\nSo it might be better to stick with the original training images (maybe cleaned up up for the corrupted images). There is not much to be learned from the old test images, that can't be achieved with augmentation (it is based in the same patches).",
    "402147": "The new training set seems indeed to be a very simple merge of old training and test set. Just checking the first few lines in the new training csv file, the hashed filenames also seems to have stayed the same.\n\nSo using the old training set could be achieved by simple switching to the old training csv file, not much other changes required in most cases."
  }
}