{
  "id": 64675,
  "title": "Inversion, could you please provide overlap information?",
  "url": "/competitions/airbus-ship-detection/discussion/64675",
  "author_name": "",
  "post_date": "2018-08-31T11:35:37.965391700Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>@Inversion,</p>\n\n<p>First of all, thank you for acknowledging mistakes made and for working hard to correct them.</p>\n\n<p>As I understand it the new training set will be comprised of multiple images, which, for a large part, are (partially) overlapping each other.\nWithout knowing which images overlap each other, I don't think it is possible to make any sort of meaningful cv or hold-out. \nCould you please provide the information which images (partly) overlap each other?</p>\n\n<p>This would be greatly appreciated by me and I guess others as well.</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "379458",
      "postDate": "08/31/2018 11:35:37",
      "content": "<p>@Inversion,</p>\n\n<p>First of all, thank you for acknowledging mistakes made and for working hard to correct them.</p>\n\n<p>As I understand it the new training set will be comprised of multiple images, which, for a large part, are (partially) overlapping each other.\nWithout knowing which images overlap each other, I don't think it is possible to make any sort of meaningful cv or hold-out. \nCould you please provide the information which images (partly) overlap each other?</p>\n\n<p>This would be greatly appreciated by me and I guess others as well.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Inversion,\n\nFirst of all, thank you for acknowledging mistakes made and for working hard to correct them.\n\nAs I understand it the new training set will be comprised of multiple images, which, for a large part, are (partially) overlapping each other.\nWithout knowing which images overlap each other, I don't think it is possible to make any sort of meaningful cv or hold-out. \nCould you please provide the information which images (partly) overlap each other?\n\nThis would be greatly appreciated by me and I guess others as well.\n\nThanks",
      "votes": null
    },
    {
      "id": "379910",
      "postDate": "09/01/2018 07:39:30",
      "content": "<p>I guess overlap information is wanted by more people: <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/64588#\">https://www.kaggle.com/c/airbus-ship-detection/discussion/64588#</a></p>",
      "rawMarkdown": "I guess overlap information is wanted by more people: [https://www.kaggle.com/c/airbus-ship-detection/discussion/64588#][1]\n\n\n  [1]: https://www.kaggle.com/c/airbus-ship-detection/discussion/64588#",
      "votes": null
    },
    {
      "id": "380429",
      "postDate": "09/02/2018 16:14:03",
      "content": "<p>I wonder why Kaggle or Airbus does this on purpose. </p>\n\n<p>If the images are sliced from bigger images (which they are), why not share it with the competitors? </p>\n\n<p>I do not see how it benefits Kaggle or Airbus to obfuscate the training (and test) sets. Determined competitors who re-assemble the jigsaw would be better suited to do more augmentations, etc. but the fact that one needs to spend time in an artificially created problem just sucks, to put it in blunt words.</p>\n\n<p>I would think that if the goal for Airbus is to get the best algorithm/model/etc. they would give participants as much information as possible and help them as much as possible; after all it's in their own interest.</p>",
      "rawMarkdown": "I wonder why Kaggle or Airbus does this on purpose. \n\nIf the images are sliced from bigger images (which they are), why not share it with the competitors? \n\nI do not see how it benefits Kaggle or Airbus to obfuscate the training (and test) sets. Determined competitors who re-assemble the jigsaw would be better suited to do more augmentations, etc. but the fact that one needs to spend time in an artificially created problem just sucks, to put it in blunt words.\n\nI would think that if the goal for Airbus is to get the best algorithm/model/etc. they would give participants as much information as possible and help them as much as possible; after all it's in their own interest.",
      "votes": null
    },
    {
      "id": "380450",
      "postDate": "09/02/2018 17:12:15",
      "content": "<p>Hi Andres,</p>\n\n<p>You are right. Our images are sliced from much bigger images typically 6,000 pixels by 60,000 pixels. Let's call the big satellite image an acquisition. In order to be able to run fast processing in the cloud, we need to be able to process chips extracted from an acquisition in parallel. We have selected of chips of 256x256 pixels which is the standard for web-mapping. </p>\n\n<p>When looking into this competition, we realised that on 256x256 chips a lot of boats will be cut at the border of the chip. So we decided to provide chips of 768x768 pixels but still with a 256 pixel step. There would be duplication in the training set but all boats could be seen in full in at least once. Of course, it was not planned that the overlapping images from the same acquisition would be both in train and validation.</p>\n\n<p>Now, that we are building a new validation dataset we will use a 768x768 pixel size and a 768 pixel step. Note that in order to include a lot of variety, we have tagging areas of various size and location across the globe. So, we might still provide some chips with overlap in order to stop exactly at the limit of the tagged area (in order to include all ships but not risk to include untagged ships). </p>\n\n<p>I hopes it helps. Please, do not hesitate to suggest potential ideas or improvements.</p>\n\n<p>Cheers,\nJeff</p>",
      "rawMarkdown": "Hi Andres,\n\nYou are right. Our images are sliced from much bigger images typically 6,000 pixels by 60,000 pixels. Let's call the big satellite image an acquisition. In order to be able to run fast processing in the cloud, we need to be able to process chips extracted from an acquisition in parallel. We have selected of chips of 256x256 pixels which is the standard for web-mapping. \n\nWhen looking into this competition, we realised that on 256x256 chips a lot of boats will be cut at the border of the chip. So we decided to provide chips of 768x768 pixels but still with a 256 pixel step. There would be duplication in the training set but all boats could be seen in full in at least once. Of course, it was not planned that the overlapping images from the same acquisition would be both in train and validation.\n\nNow, that we are building a new validation dataset we will use a 768x768 pixel size and a 768 pixel step. Note that in order to include a lot of variety, we have tagging areas of various size and location across the globe. So, we might still provide some chips with overlap in order to stop exactly at the limit of the tagged area (in order to include all ships but not risk to include untagged ships). \n\nI hopes it helps. Please, do not hesitate to suggest potential ideas or improvements.\n\nCheers,\nJeff",
      "votes": null
    },
    {
      "id": "380456",
      "postDate": "09/02/2018 17:28:11",
      "content": "<p>My two cents:</p>\n\n<ol>\n<li>As Jules suggested, share the location of imageids (chips) that compose a big acquisition. We could do more augmentations that will help training, including but not limited to training 768px chips. </li>\n<li>Regarding the processing speed concern, you are already addressing it with a nice price. My only comment is that the wording (at least for me) is open for interpretation (e.g. top teams?, fastest inference with fastest accuracy? the way I read it is the fastest kernel will win unless you decide to create a bucket with the top 10% of top N fastest and among them the most accurate). </li>\n</ol>",
      "rawMarkdown": "My two cents:\n\n 1. As Jules suggested, share the location of imageids (chips) that compose a big acquisition. We could do more augmentations that will help training, including but not limited to training 768px chips. \n 2. Regarding the processing speed concern, you are already addressing it with a nice price. My only comment is that the wording (at least for me) is open for interpretation (e.g. top teams?, fastest inference with fastest accuracy? the way I read it is the fastest kernel will win unless you decide to create a bucket with the top 10% of top N fastest and among them the most accurate).",
      "votes": null
    },
    {
      "id": "380747",
      "postDate": "09/03/2018 11:23:07",
      "content": "<p>Hi Jeff,</p>\n\n<p>Thank you for explaining the reason of the use of overlapping images (chips). \nI can see how Airbus could use the predictions of overlapping 768*768 chips to get better results.</p>\n\n<p>As I mentioned before; the information if one chip (image) overlaps with another is advantageous to know for the purpose of creating meaningful cv's or hold-out sets. But the information if chips (images) are <strong>adjacent</strong> to each other or overlap in the test and the train set could also be used to make better predictions!  </p>\n\n<p>The above means that not providing this information to the competitors creates an artificially problem, as Andres rightly mentioned. The ambitious competitor will now be forced to make models to predict overlap and adjacency of images which are of course not useful for Airbus.</p>",
      "rawMarkdown": "Hi Jeff,\n\nThank you for explaining the reason of the use of overlapping images (chips). \nI can see how Airbus could use the predictions of overlapping 768*768 chips to get better results.\n\nAs I mentioned before; the information if one chip (image) overlaps with another is advantageous to know for the purpose of creating meaningful cv's or hold-out sets. But the information if chips (images) are **adjacent** to each other or overlap in the test and the train set could also be used to make better predictions!  \n\nThe above means that not providing this information to the competitors creates an artificially problem, as Andres rightly mentioned. The ambitious competitor will now be forced to make models to predict overlap and adjacency of images which are of course not useful for Airbus.",
      "votes": null
    },
    {
      "id": "381480",
      "postDate": "09/04/2018 17:28:20",
      "content": "<p>Thank you Andres and Jules for your comments. </p>\n\n<p>Unfortunately for the training dataset, I am afraid that we have lost the information on the overlaps as each extracted image has been renamed with a generated id. You also have to know that the areas to tag are very different in sizes/shapes and that we have removed a lot of images with no ships as well as those containing too much ships or too much ground. This was done last year and this is now to late to correct. I will keep your comments in mind when we create new datasets.</p>\n\n<p>Kind regards,\nJeff.</p>",
      "rawMarkdown": "Thank you Andres and Jules for your comments. \n\nUnfortunately for the training dataset, I am afraid that we have lost the information on the overlaps as each extracted image has been renamed with a generated id. You also have to know that the areas to tag are very different in sizes/shapes and that we have removed a lot of images with no ships as well as those containing too much ships or too much ground. This was done last year and this is now to late to correct. I will keep your comments in mind when we create new datasets.\n\nKind regards,\nJeff.",
      "votes": null
    },
    {
      "id": "382707",
      "postDate": "09/07/2018 00:26:10",
      "content": "<p>As result we have awful dataset:\n 1. Too many corrupted images\n 2. Can't normal train model, because for validation we must separate the overlapped images\n 3. Set increase size - save redundancy information</p>\n\n<p>Too much ground on images - not problem. </p>",
      "rawMarkdown": "As result we have awful dataset:\n 1. Too many corrupted images\n 2. Can't normal train model, because for validation we must separate the overlapped images\n 3. Set increase size - save redundancy information\n\nToo much ground on images - not problem.",
      "votes": null
    },
    {
      "id": "385411",
      "postDate": "09/11/2018 00:52:22",
      "content": "<p>Hi Jeff,\nIs it possible to provide us the original acquisitions? That would make finding overlaps easier.</p>",
      "rawMarkdown": "Hi Jeff,\nIs it possible to provide us the original acquisitions? That would make finding overlaps easier.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 379910,
      "author_name": "julesvanligtenberg",
      "author_url": "",
      "post_date": "09/01/2018 07:39:30",
      "content": "<p>I guess overlap information is wanted by more people: <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/64588#\">https://www.kaggle.com/c/airbus-ship-detection/discussion/64588#</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 380429,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "09/02/2018 16:14:03",
      "content": "<p>I wonder why Kaggle or Airbus does this on purpose. </p>\n\n<p>If the images are sliced from bigger images (which they are), why not share it with the competitors? </p>\n\n<p>I do not see how it benefits Kaggle or Airbus to obfuscate the training (and test) sets. Determined competitors who re-assemble the jigsaw would be better suited to do more augmentations, etc. but the fact that one needs to spend time in an artificially created problem just sucks, to put it in blunt words.</p>\n\n<p>I would think that if the goal for Airbus is to get the best algorithm/model/etc. they would give participants as much information as possible and help them as much as possible; after all it's in their own interest.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 380450,
      "author_name": "jeffaudi",
      "author_url": "",
      "post_date": "09/02/2018 17:12:15",
      "content": "<p>Hi Andres,</p>\n\n<p>You are right. Our images are sliced from much bigger images typically 6,000 pixels by 60,000 pixels. Let's call the big satellite image an acquisition. In order to be able to run fast processing in the cloud, we need to be able to process chips extracted from an acquisition in parallel. We have selected of chips of 256x256 pixels which is the standard for web-mapping. </p>\n\n<p>When looking into this competition, we realised that on 256x256 chips a lot of boats will be cut at the border of the chip. So we decided to provide chips of 768x768 pixels but still with a 256 pixel step. There would be duplication in the training set but all boats could be seen in full in at least once. Of course, it was not planned that the overlapping images from the same acquisition would be both in train and validation.</p>\n\n<p>Now, that we are building a new validation dataset we will use a 768x768 pixel size and a 768 pixel step. Note that in order to include a lot of variety, we have tagging areas of various size and location across the globe. So, we might still provide some chips with overlap in order to stop exactly at the limit of the tagged area (in order to include all ships but not risk to include untagged ships). </p>\n\n<p>I hopes it helps. Please, do not hesitate to suggest potential ideas or improvements.</p>\n\n<p>Cheers,\nJeff</p>",
      "votes": null,
      "replies": [
        {
          "id": 380456,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "09/02/2018 17:28:11",
          "content": "<p>My two cents:</p>\n\n<ol>\n<li>As Jules suggested, share the location of imageids (chips) that compose a big acquisition. We could do more augmentations that will help training, including but not limited to training 768px chips. </li>\n<li>Regarding the processing speed concern, you are already addressing it with a nice price. My only comment is that the wording (at least for me) is open for interpretation (e.g. top teams?, fastest inference with fastest accuracy? the way I read it is the fastest kernel will win unless you decide to create a bucket with the top 10% of top N fastest and among them the most accurate). </li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 380747,
          "author_name": "julesvanligtenberg",
          "author_url": "",
          "post_date": "09/03/2018 11:23:07",
          "content": "<p>Hi Jeff,</p>\n\n<p>Thank you for explaining the reason of the use of overlapping images (chips). \nI can see how Airbus could use the predictions of overlapping 768*768 chips to get better results.</p>\n\n<p>As I mentioned before; the information if one chip (image) overlaps with another is advantageous to know for the purpose of creating meaningful cv's or hold-out sets. But the information if chips (images) are <strong>adjacent</strong> to each other or overlap in the test and the train set could also be used to make better predictions!  </p>\n\n<p>The above means that not providing this information to the competitors creates an artificially problem, as Andres rightly mentioned. The ambitious competitor will now be forced to make models to predict overlap and adjacency of images which are of course not useful for Airbus.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 381480,
          "author_name": "jeffaudi",
          "author_url": "",
          "post_date": "09/04/2018 17:28:20",
          "content": "<p>Thank you Andres and Jules for your comments. </p>\n\n<p>Unfortunately for the training dataset, I am afraid that we have lost the information on the overlaps as each extracted image has been renamed with a generated id. You also have to know that the areas to tag are very different in sizes/shapes and that we have removed a lot of images with no ships as well as those containing too much ships or too much ground. This was done last year and this is now to late to correct. I will keep your comments in mind when we create new datasets.</p>\n\n<p>Kind regards,\nJeff.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 382707,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "09/07/2018 00:26:10",
          "content": "<p>As result we have awful dataset:\n 1. Too many corrupted images\n 2. Can't normal train model, because for validation we must separate the overlapped images\n 3. Set increase size - save redundancy information</p>\n\n<p>Too much ground on images - not problem. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385411,
          "author_name": "waskita",
          "author_url": "",
          "post_date": "09/11/2018 00:52:22",
          "content": "<p>Hi Jeff,\nIs it possible to provide us the original acquisitions? That would make finding overlaps easier.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "379458": "Inversion,\n\nFirst of all, thank you for acknowledging mistakes made and for working hard to correct them.\n\nAs I understand it the new training set will be comprised of multiple images, which, for a large part, are (partially) overlapping each other.\nWithout knowing which images overlap each other, I don't think it is possible to make any sort of meaningful cv or hold-out. \nCould you please provide the information which images (partly) overlap each other?\n\nThis would be greatly appreciated by me and I guess others as well.\n\nThanks",
    "379910": "I guess overlap information is wanted by more people: [https://www.kaggle.com/c/airbus-ship-detection/discussion/64588#][1]\n\n\n  [1]: https://www.kaggle.com/c/airbus-ship-detection/discussion/64588#",
    "380429": "I wonder why Kaggle or Airbus does this on purpose. \n\nIf the images are sliced from bigger images (which they are), why not share it with the competitors? \n\nI do not see how it benefits Kaggle or Airbus to obfuscate the training (and test) sets. Determined competitors who re-assemble the jigsaw would be better suited to do more augmentations, etc. but the fact that one needs to spend time in an artificially created problem just sucks, to put it in blunt words.\n\nI would think that if the goal for Airbus is to get the best algorithm/model/etc. they would give participants as much information as possible and help them as much as possible; after all it's in their own interest.",
    "380450": "Hi Andres,\n\nYou are right. Our images are sliced from much bigger images typically 6,000 pixels by 60,000 pixels. Let's call the big satellite image an acquisition. In order to be able to run fast processing in the cloud, we need to be able to process chips extracted from an acquisition in parallel. We have selected of chips of 256x256 pixels which is the standard for web-mapping. \n\nWhen looking into this competition, we realised that on 256x256 chips a lot of boats will be cut at the border of the chip. So we decided to provide chips of 768x768 pixels but still with a 256 pixel step. There would be duplication in the training set but all boats could be seen in full in at least once. Of course, it was not planned that the overlapping images from the same acquisition would be both in train and validation.\n\nNow, that we are building a new validation dataset we will use a 768x768 pixel size and a 768 pixel step. Note that in order to include a lot of variety, we have tagging areas of various size and location across the globe. So, we might still provide some chips with overlap in order to stop exactly at the limit of the tagged area (in order to include all ships but not risk to include untagged ships). \n\nI hopes it helps. Please, do not hesitate to suggest potential ideas or improvements.\n\nCheers,\nJeff",
    "380456": "My two cents:\n\n 1. As Jules suggested, share the location of imageids (chips) that compose a big acquisition. We could do more augmentations that will help training, including but not limited to training 768px chips. \n 2. Regarding the processing speed concern, you are already addressing it with a nice price. My only comment is that the wording (at least for me) is open for interpretation (e.g. top teams?, fastest inference with fastest accuracy? the way I read it is the fastest kernel will win unless you decide to create a bucket with the top 10% of top N fastest and among them the most accurate).",
    "380747": "Hi Jeff,\n\nThank you for explaining the reason of the use of overlapping images (chips). \nI can see how Airbus could use the predictions of overlapping 768*768 chips to get better results.\n\nAs I mentioned before; the information if one chip (image) overlaps with another is advantageous to know for the purpose of creating meaningful cv's or hold-out sets. But the information if chips (images) are **adjacent** to each other or overlap in the test and the train set could also be used to make better predictions!  \n\nThe above means that not providing this information to the competitors creates an artificially problem, as Andres rightly mentioned. The ambitious competitor will now be forced to make models to predict overlap and adjacency of images which are of course not useful for Airbus.",
    "381480": "Thank you Andres and Jules for your comments. \n\nUnfortunately for the training dataset, I am afraid that we have lost the information on the overlaps as each extracted image has been renamed with a generated id. You also have to know that the areas to tag are very different in sizes/shapes and that we have removed a lot of images with no ships as well as those containing too much ships or too much ground. This was done last year and this is now to late to correct. I will keep your comments in mind when we create new datasets.\n\nKind regards,\nJeff.",
    "382707": "As result we have awful dataset:\n 1. Too many corrupted images\n 2. Can't normal train model, because for validation we must separate the overlapped images\n 3. Set increase size - save redundancy information\n\nToo much ground on images - not problem.",
    "385411": "Hi Jeff,\nIs it possible to provide us the original acquisitions? That would make finding overlaps easier."
  },
  "source": "meta"
}