{
  "id": 68042,
  "title": "Relaunch Data Available",
  "url": "/competitions/airbus-ship-detection/discussion/68042",
  "author_name": "inversion",
  "post_date": "2018-10-08T18:16:38.171000",
  "votes": 24,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi Everyone -</p>\n\n<p>The relaunch data is now available all the new files/folders have a <code>_v2</code> postfix. I have a number of administrative items to perform (updating the data pages, rescoring with the new submission file, etc.) Until that's all complete, the new sample submission won't score correctly. </p>\n\n<p>I'll provide updates here when the details are finished.</p>\n\n<p><strong>UPDATE:</strong></p>\n\n<ul>\n<li>The <code>train_v2</code> folder contains all of the original Train and Test images combined, and the new <code>train_ship_segmentations_v2.csv</code> contains all the RLE for the original Train and Test images.</li>\n<li>Since the <code>test_v2</code> image set is significantly smaller than the original Test set, a number of un-scored files were added to discourage hand labeling. (Reminder, hand labeling of the Test set is not permitted.) </li>\n</ul>",
  "messages": [
    {
      "id": 400689,
      "postDate": "2018-10-08T18:16:38.173Z",
      "content": "<p>Hi Everyone -</p>\n\n<p>The relaunch data is now available all the new files/folders have a <code>_v2</code> postfix. I have a number of administrative items to perform (updating the data pages, rescoring with the new submission file, etc.) Until that's all complete, the new sample submission won't score correctly. </p>\n\n<p>I'll provide updates here when the details are finished.</p>\n\n<p><strong>UPDATE:</strong></p>\n\n<ul>\n<li>The <code>train_v2</code> folder contains all of the original Train and Test images combined, and the new <code>train_ship_segmentations_v2.csv</code> contains all the RLE for the original Train and Test images.</li>\n<li>Since the <code>test_v2</code> image set is significantly smaller than the original Test set, a number of un-scored files were added to discourage hand labeling. (Reminder, hand labeling of the Test set is not permitted.) </li>\n</ul>",
      "rawMarkdown": "Hi Everyone -\n\nThe relaunch data is now available all the new files/folders have a `_v2` postfix. I have a number of administrative items to perform (updating the data pages, rescoring with the new submission file, etc.) Until that's all complete, the new sample submission won't score correctly. \n\nI'll provide updates here when the details are finished.\n\n**UPDATE:**\n\n - The `train_v2` folder contains all of the original Train and Test images combined, and the new `train_ship_segmentations_v2.csv` contains all the RLE for the original Train and Test images.\n - Since the `test_v2` image set is significantly smaller than the original Test set, a number of un-scored files were added to discourage hand labeling. (Reminder, hand labeling of the Test set is not permitted.) ",
      "votes": 24
    },
    {
      "id": 401159,
      "postDate": "2018-10-09T14:31:44.313Z",
      "content": "<p>Dear Kagglers!</p>\n\n<p>At Airbus, we are very happy that the competition has now been relaunched with a fully new test dataset. I would like to give you some information on how it has been prepared and what changes to expect.</p>\n\n<p>We have tagged areas of satellite images acquired recently from various areas at sea all over the world. The tagging has been validated through a two level validation scheme. Although we have done our best to make very good annotations, there are probably still some errors or approximations. This is of course annoying but we do not expect these small errors to affect the global scoring system.    </p>\n\n<p>We have then extracted tagged images of size 768 by 768 pixel from much larger satellite acquisitions. Many images without ships have been removed so that the final number of \"empty\" images is not too high. This is a change from previous test dataset that provided a very high score for an empty submission. Our idea is to offer more space for improvement and to give less importance to correctly predicting empty images over correctly detecting ships on images.</p>\n\n<p>We have also removed the overlap across the images. Although the images in the train dataset still contain overlap, the images in the test dataset do not have overlap. You should not encounter the same ship multiple times (or very, very rarely). So no need to search for overlaps in the test dataset!</p>\n\n<p>With the new end date in one month from now, I wish you all a lot of success and a lot of fun in the competition.\nHave a good day,</p>\n\n<p>Jeff.</p>",
      "rawMarkdown": "Dear Kagglers!\n\nAt Airbus, we are very happy that the competition has now been relaunched with a fully new test dataset. I would like to give you some information on how it has been prepared and what changes to expect.\n\nWe have tagged areas of satellite images acquired recently from various areas at sea all over the world. The tagging has been validated through a two level validation scheme. Although we have done our best to make very good annotations, there are probably still some errors or approximations. This is of course annoying but we do not expect these small errors to affect the global scoring system.    \n\nWe have then extracted tagged images of size 768 by 768 pixel from much larger satellite acquisitions. Many images without ships have been removed so that the final number of \"empty\" images is not too high. This is a change from previous test dataset that provided a very high score for an empty submission. Our idea is to offer more space for improvement and to give less importance to correctly predicting empty images over correctly detecting ships on images.\n\nWe have also removed the overlap across the images. Although the images in the train dataset still contain overlap, the images in the test dataset do not have overlap. You should not encounter the same ship multiple times (or very, very rarely). So no need to search for overlaps in the test dataset!\n\nWith the new end date in one month from now, I wish you all a lot of success and a lot of fun in the competition.\nHave a good day,\n\nJeff.",
      "votes": 8,
      "replies": [
        {
          "id": 401263,
          "postDate": "2018-10-09T17:54:41.680Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 415688,
          "postDate": "2018-11-05T14:02:29.947Z",
          "content": "<p>Hi，May I ask the submission number limit is 3 month * 5 ？ or 1 month * 5 ？</p>",
          "rawMarkdown": "Hi，May I ask the submission number limit is 3 month * 5 ？ or 1 month * 5 ？"
        },
        {
          "id": 415711,
          "postDate": "2018-11-05T14:35:53.230Z",
          "content": "<p><a href=\"/shentao\">@shentao</a> - The total limit is from the start of the original competition date.</p>",
          "rawMarkdown": "@shentao - The total limit is from the start of the original competition date."
        }
      ]
    },
    {
      "id": 400729,
      "postDate": "2018-10-08T19:26:36.360Z",
      "content": "<p>Thank you for uploading the new data. I saw that you also added another train data \"train_v2.zip\". Could you explain what is the difference between this data and one provided originally? Did you generate a new data or just combine previously posted train.zip and test.zip into one file?</p>",
      "rawMarkdown": "Thank you for uploading the new data. I saw that you also added another train data \"train_v2.zip\". Could you explain what is the difference between this data and one provided originally? Did you generate a new data or just combine previously posted train.zip and test.zip into one file?",
      "votes": 6,
      "replies": [
        {
          "id": 400908,
          "postDate": "2018-10-09T05:54:50.560Z",
          "content": "<p>The size of the new file looks like it was merged from train.zip and test.zip. AFAIK the old train.zip already contained about 98%  (?) of data from test.zip so if he'd generated everything from scratch without that overlapping there was no way the new training set can be that big.</p>",
          "rawMarkdown": "The size of the new file looks like it was merged from train.zip and test.zip. AFAIK the old train.zip already contained about 98%  (?) of data from test.zip so if he'd generated everything from scratch without that overlapping there was no way the new training set can be that big.",
          "votes": 1
        },
        {
          "id": 407082,
          "postDate": "2018-10-20T10:52:21.447Z",
          "content": "<p>The overlapping is still present in the train set. Only the test set should be without overlapping.</p>",
          "rawMarkdown": "The overlapping is still present in the train set. Only the test set should be without overlapping.",
          "votes": 1
        }
      ]
    },
    {
      "id": 409886,
      "postDate": "2018-10-25T02:27:18.893Z",
      "content": "<p>The reduction in test size is really appreciated by me.  I don't do any hand labeling or probing of the leader board; but I also like to use my PC for something besides watching a completion bar slowly go to 100% over a couple of hours.  </p>",
      "rawMarkdown": "The reduction in test size is really appreciated by me.  I don't do any hand labeling or probing of the leader board; but I also like to use my PC for something besides watching a completion bar slowly go to 100% over a couple of hours.  "
    },
    {
      "id": 404133,
      "postDate": "2018-10-15T09:53:02.280Z",
      "content": "<p>There is a way to get old scores?</p>",
      "rawMarkdown": "There is a way to get old scores?"
    },
    {
      "id": 400993,
      "postDate": "2018-10-09T08:52:11.913Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> can you also please provide how the data was collected and checked?</p>",
      "rawMarkdown": "@inversion can you also please provide how the data was collected and checked?"
    },
    {
      "id": 400901,
      "postDate": "2018-10-09T05:40:42.650Z",
      "content": "<p>+</p>",
      "rawMarkdown": "+"
    },
    {
      "id": 400724,
      "postDate": "2018-10-08T19:16:07.143Z",
      "content": "<p>Thanks for getting the data leakage solved and getting this competition back on track, much appreciated!</p>\n\n<p>Also I'm curious to see if my models do as well as before or I have (without realizing) introduced some severe overfitting. </p>",
      "rawMarkdown": "Thanks for getting the data leakage solved and getting this competition back on track, much appreciated!\n\nAlso I'm curious to see if my models do as well as before or I have (without realizing) introduced some severe overfitting. "
    }
  ],
  "comments": [
    {
      "id": 401159,
      "author_name": "Jeff Faudi",
      "author_url": "",
      "post_date": "2018-10-09T14:31:44.313000",
      "content": "<p>Dear Kagglers!</p>\n\n<p>At Airbus, we are very happy that the competition has now been relaunched with a fully new test dataset. I would like to give you some information on how it has been prepared and what changes to expect.</p>\n\n<p>We have tagged areas of satellite images acquired recently from various areas at sea all over the world. The tagging has been validated through a two level validation scheme. Although we have done our best to make very good annotations, there are probably still some errors or approximations. This is of course annoying but we do not expect these small errors to affect the global scoring system.    </p>\n\n<p>We have then extracted tagged images of size 768 by 768 pixel from much larger satellite acquisitions. Many images without ships have been removed so that the final number of \"empty\" images is not too high. This is a change from previous test dataset that provided a very high score for an empty submission. Our idea is to offer more space for improvement and to give less importance to correctly predicting empty images over correctly detecting ships on images.</p>\n\n<p>We have also removed the overlap across the images. Although the images in the train dataset still contain overlap, the images in the test dataset do not have overlap. You should not encounter the same ship multiple times (or very, very rarely). So no need to search for overlaps in the test dataset!</p>\n\n<p>With the new end date in one month from now, I wish you all a lot of success and a lot of fun in the competition.\nHave a good day,</p>\n\n<p>Jeff.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 401263,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-10-09T17:54:41.680000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415688,
          "author_name": "SeuTao",
          "author_url": "",
          "post_date": "2018-11-05T14:02:29.947000",
          "content": "<p>Hi，May I ask the submission number limit is 3 month * 5 ？ or 1 month * 5 ？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415711,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-05T14:35:53.230000",
          "content": "<p><a href=\"/shentao\">@shentao</a> - The total limit is from the start of the original competition date.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 400729,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2018-10-08T19:26:36.360000",
      "content": "<p>Thank you for uploading the new data. I saw that you also added another train data \"train_v2.zip\". Could you explain what is the difference between this data and one provided originally? Did you generate a new data or just combine previously posted train.zip and test.zip into one file?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 400908,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2018-10-09T05:54:50.560000",
          "content": "<p>The size of the new file looks like it was merged from train.zip and test.zip. AFAIK the old train.zip already contained about 98%  (?) of data from test.zip so if he'd generated everything from scratch without that overlapping there was no way the new training set can be that big.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 407082,
          "author_name": "Maxime Riché",
          "author_url": "",
          "post_date": "2018-10-20T10:52:21.447000",
          "content": "<p>The overlapping is still present in the train set. Only the test set should be without overlapping.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 409886,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2018-10-25T02:27:18.893000",
      "content": "<p>The reduction in test size is really appreciated by me.  I don't do any hand labeling or probing of the leader board; but I also like to use my PC for something besides watching a completion bar slowly go to 100% over a couple of hours.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 404133,
      "author_name": "Elisa",
      "author_url": "",
      "post_date": "2018-10-15T09:53:02.280000",
      "content": "<p>There is a way to get old scores?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 400993,
      "author_name": "Kostiantyn Maksymov",
      "author_url": "",
      "post_date": "2018-10-09T08:52:11.913000",
      "content": "<p><a href=\"/inversion\">@inversion</a> can you also please provide how the data was collected and checked?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 400901,
      "author_name": "Alexander Veysov",
      "author_url": "",
      "post_date": "2018-10-09T05:40:42.650000",
      "content": "<p>+</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 400724,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2018-10-08T19:16:07.143000",
      "content": "<p>Thanks for getting the data leakage solved and getting this competition back on track, much appreciated!</p>\n\n<p>Also I'm curious to see if my models do as well as before or I have (without realizing) introduced some severe overfitting. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "400689": "Hi Everyone -\n\nThe relaunch data is now available all the new files/folders have a `_v2` postfix. I have a number of administrative items to perform (updating the data pages, rescoring with the new submission file, etc.) Until that's all complete, the new sample submission won't score correctly. \n\nI'll provide updates here when the details are finished.\n\n**UPDATE:**\n\n - The `train_v2` folder contains all of the original Train and Test images combined, and the new `train_ship_segmentations_v2.csv` contains all the RLE for the original Train and Test images.\n - Since the `test_v2` image set is significantly smaller than the original Test set, a number of un-scored files were added to discourage hand labeling. (Reminder, hand labeling of the Test set is not permitted.) ",
    "401159": "Dear Kagglers!\n\nAt Airbus, we are very happy that the competition has now been relaunched with a fully new test dataset. I would like to give you some information on how it has been prepared and what changes to expect.\n\nWe have tagged areas of satellite images acquired recently from various areas at sea all over the world. The tagging has been validated through a two level validation scheme. Although we have done our best to make very good annotations, there are probably still some errors or approximations. This is of course annoying but we do not expect these small errors to affect the global scoring system.    \n\nWe have then extracted tagged images of size 768 by 768 pixel from much larger satellite acquisitions. Many images without ships have been removed so that the final number of \"empty\" images is not too high. This is a change from previous test dataset that provided a very high score for an empty submission. Our idea is to offer more space for improvement and to give less importance to correctly predicting empty images over correctly detecting ships on images.\n\nWe have also removed the overlap across the images. Although the images in the train dataset still contain overlap, the images in the test dataset do not have overlap. You should not encounter the same ship multiple times (or very, very rarely). So no need to search for overlaps in the test dataset!\n\nWith the new end date in one month from now, I wish you all a lot of success and a lot of fun in the competition.\nHave a good day,\n\nJeff.",
    "400729": "Thank you for uploading the new data. I saw that you also added another train data \"train_v2.zip\". Could you explain what is the difference between this data and one provided originally? Did you generate a new data or just combine previously posted train.zip and test.zip into one file?",
    "409886": "The reduction in test size is really appreciated by me.  I don't do any hand labeling or probing of the leader board; but I also like to use my PC for something besides watching a completion bar slowly go to 100% over a couple of hours.  ",
    "404133": "There is a way to get old scores?",
    "400993": "@inversion can you also please provide how the data was collected and checked?",
    "400901": "+",
    "400724": "Thanks for getting the data leakage solved and getting this competition back on track, much appreciated!\n\nAlso I'm curious to see if my models do as well as before or I have (without realizing) introduced some severe overfitting. "
  }
}