{
  "id": 81406,
  "title": "[0.93 IoU] Fluke detection using fastai, more train data & bbox extraction",
  "url": "/competitions/humpback-whale-identification/discussion/81406",
  "author_name": "",
  "post_date": "2019-02-21T08:48:55.842753200Z",
  "votes": 34,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I added more hand annotated <a href=\"https://github.com/radekosmulski/whale/blob/master/data/annotations.json\">train data</a> to my <a href=\"https://github.com/radekosmulski/whale/blob/master/fluke_detection_redux.ipynb\">fluke detector</a>, which now gives a significantly higher IoU.</p>\n\n<p>I also added code that runs the detector on entire train and test sets and extract predicted bounding boxes to images of a specified size. You can find the notebook with the code <a href=\"https://github.com/radekosmulski/whale/blob/master/extract_bboxes.ipynb\">here</a>.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/475828/11366/flukes1.png\" alt=\"enter image description here\"></p>\n\n<p>In the notebook I also go into a bit of discussion on how Pytorch Datasets &amp; Dataloaders work so might be of interest if you are planning on doing something a bit more custom.</p>\n\n<p>Only a few more days left in the competition, but it seems training on extracted bounding boxes vs full image can make quite a difference! (quickly ran a small classification experiment and the improvement seems to be quite significant.</p>\n\n<p>Best of luck! :)</p>\n\n<p>EDIT: adding csv of predicted bounding boxes mapped to image size as per request</p>",
  "messages": [
    {
      "id": "475828",
      "postDate": "02/21/2019 08:48:55",
      "content": "<p>I added more hand annotated <a href=\"https://github.com/radekosmulski/whale/blob/master/data/annotations.json\">train data</a> to my <a href=\"https://github.com/radekosmulski/whale/blob/master/fluke_detection_redux.ipynb\">fluke detector</a>, which now gives a significantly higher IoU.</p>\n\n<p>I also added code that runs the detector on entire train and test sets and extract predicted bounding boxes to images of a specified size. You can find the notebook with the code <a href=\"https://github.com/radekosmulski/whale/blob/master/extract_bboxes.ipynb\">here</a>.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/475828/11366/flukes1.png\" alt=\"enter image description here\"></p>\n\n<p>In the notebook I also go into a bit of discussion on how Pytorch Datasets &amp; Dataloaders work so might be of interest if you are planning on doing something a bit more custom.</p>\n\n<p>Only a few more days left in the competition, but it seems training on extracted bounding boxes vs full image can make quite a difference! (quickly ran a small classification experiment and the improvement seems to be quite significant.</p>\n\n<p>Best of luck! :)</p>\n\n<p>EDIT: adding csv of predicted bounding boxes mapped to image size as per request</p>",
      "rawMarkdown": "I added more hand annotated [train data](https://github.com/radekosmulski/whale/blob/master/data/annotations.json) to my [fluke detector](https://github.com/radekosmulski/whale/blob/master/fluke_detection_redux.ipynb), which now gives a significantly higher IoU.\n\nI also added code that runs the detector on entire train and test sets and extract predicted bounding boxes to images of a specified size. You can find the notebook with the code [here](https://github.com/radekosmulski/whale/blob/master/extract_bboxes.ipynb).\n\n![enter image description here][1]\n\nIn the notebook I also go into a bit of discussion on how Pytorch Datasets &amp; Dataloaders work so might be of interest if you are planning on doing something a bit more custom.\n\nOnly a few more days left in the competition, but it seems training on extracted bounding boxes vs full image can make quite a difference! (quickly ran a small classification experiment and the improvement seems to be quite significant.\n\nBest of luck! :)\n\nEDIT: adding csv of predicted bounding boxes mapped to image size as per request\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/475828/11366/flukes1.png",
      "votes": null
    },
    {
      "id": "475936",
      "postDate": "02/21/2019 11:52:40",
      "content": "<p>At first, thanks for your great work! Could you possibly upload the file with bbox coordinates for each image?</p>",
      "rawMarkdown": "At first, thanks for your great work! Could you possibly upload the file with bbox coordinates for each image?",
      "votes": null
    },
    {
      "id": "476160",
      "postDate": "02/21/2019 17:09:17",
      "content": "<p>I’ve been waiting for radek to appear on the lb with a 1.001</p>",
      "rawMarkdown": "I’ve been waiting for radek to appear on the lb with a 1.001",
      "votes": null
    },
    {
      "id": "476397",
      "postDate": "02/22/2019 04:25:37",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!",
      "votes": null
    },
    {
      "id": "476455",
      "postDate": "02/22/2019 06:59:04",
      "content": "<p>Wow, Radek! Thank you! I am learning a lot from your notebooks. </p>",
      "rawMarkdown": "Wow, Radek! Thank you! I am learning a lot from your notebooks.",
      "votes": null
    },
    {
      "id": "476549",
      "postDate": "02/22/2019 10:18:14",
      "content": "<p>File added to the original post! There are a couple of images (you can find their filenames in the notebook) where the predicted bounding boxes don't make sense (negative width or height). There are 4 such images in train and 1 in test.</p>\n\n<p>The coordinates follow this format: y_up_left, x_up_left, y_low_right, x_low_right. Meaning the first coordinate is the pixel value of the upper left hand corner of the image, x_up_left is the x value of that corner and the remaining two coordinates are for the lower right hand corner. Makes for easy extraction from an image represented as an array.</p>\n\n<p>Best of luck in the competition!</p>",
      "rawMarkdown": "File added to the original post! There are a couple of images (you can find their filenames in the notebook) where the predicted bounding boxes don't make sense (negative width or height). There are 4 such images in train and 1 in test.\n\nThe coordinates follow this format: y_up_left, x_up_left, y_low_right, x_low_right. Meaning the first coordinate is the pixel value of the upper left hand corner of the image, x_up_left is the x value of that corner and the remaining two coordinates are for the lower right hand corner. Makes for easy extraction from an image represented as an array.\n\nBest of luck in the competition!",
      "votes": null
    },
    {
      "id": "476973",
      "postDate": "02/23/2019 15:18:10",
      "content": "<p>very nice work and good example notebook :) \nThank you. I appreciate it. I will use it as a reference.</p>",
      "rawMarkdown": "very nice work and good example notebook :) \nThank you. I appreciate it. I will use it as a reference.",
      "votes": null
    },
    {
      "id": "477214",
      "postDate": "02/24/2019 06:11:10",
      "content": "<p>Nice work!! Thanks.</p>",
      "rawMarkdown": "Nice work!! Thanks.",
      "votes": null
    },
    {
      "id": "477615",
      "postDate": "02/25/2019 01:11:08",
      "content": "<p>Thanks Radek, great work!! Have anyone compared this with the one from this kernel : <a href=\"https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\">https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes</a> ? </p>",
      "rawMarkdown": "Thanks Radek, great work!! Have anyone compared this with the one from this kernel : https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes ?",
      "votes": null
    },
    {
      "id": "477653",
      "postDate": "02/25/2019 03:25:38",
      "content": "<p>Hi <a href=\"/radek1\">@radek1</a>, thank you for sharing.\nI found 5 images have problem (image size will be zero if simply followed), here's update from me.</p>\n\n<pre><code>                x0   y0   x1   y1\nfilename                         \nb4cb30afd.jpg  370  540  600  600\n85a95e7a8.jpg  420  430  510  470\nb370e1339.jpg  270  560  570  660\nd4cb9d6e4.jpg  490  350  550  380\n</code></pre>\n\n<p><strike>6a72d84ca.jpg  490  470  550  500</strike> was removed due to test sample, attached is updated.</p>",
      "rawMarkdown": "Hi @radek1, thank you for sharing.\nI found 5 images have problem (image size will be zero if simply followed), here's update from me.\n\n                    x0   y0   x1   y1\n    filename                         \n    b4cb30afd.jpg  370  540  600  600\n    85a95e7a8.jpg  420  430  510  470\n    b370e1339.jpg  270  560  570  660\n    d4cb9d6e4.jpg  490  350  550  380\n\n<strike>6a72d84ca.jpg  490  470  550  500</strike> was removed due to test sample, attached is updated.",
      "votes": null
    },
    {
      "id": "479019",
      "postDate": "02/26/2019 23:23:32",
      "content": "<p>How did you get the coordinates?</p>\n\n<p>'6a72d84ca.jpg' is a test image. Manual annotations can not be used for test images.</p>",
      "rawMarkdown": "How did you get the coordinates?\n\n'6a72d84ca.jpg' is a test image. Manual annotations can not be used for test images.",
      "votes": null
    },
    {
      "id": "479215",
      "postDate": "02/27/2019 02:52:26",
      "content": "<p>i read a new paper today. maybe can use for future kaggle challenge.</p>\n\n<p>It can can be used for predicting bounding box without annotated data. </p>\n\n<p><a href=\"https://arxiv.org/pdf/1902.09968.pdf\">https://arxiv.org/pdf/1902.09968.pdf</a></p>\n\n<p>\"Mining Objects: Fully Unsupervised Object Discovery and Localization From a Single Image\"</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/479215/11436/unpervsied.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "i read a new paper today. maybe can use for future kaggle challenge.\n\nIt can can be used for predicting bounding box without annotated data. \n\nhttps://arxiv.org/pdf/1902.09968.pdf\n\n\"Mining Objects: Fully Unsupervised Object Discovery and Localization From a Single Image\"\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/479215/11436/unpervsied.png",
      "votes": null
    },
    {
      "id": "479982",
      "postDate": "02/27/2019 16:39:34",
      "content": "<p>By hand. And I see your point, then we can use coods.csv with bounding_boxes.csv for invalid data which we can generate: <a href=\"https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\">https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes</a></p>\n\n<p>Or my local code is using this now:</p>\n\n<pre><code>def get_coods_fixed(filename):\n    x0, y0, x1, y1 = bbdf.loc[filename].values\n    if x1 &lt;= x0 or y1 &lt;= y0:\n        print(f'Warninig {x0} {y0} {x1} {y1} was fixed for {filename}.')\n    x0, x1 = (x0, x1) if x0 &lt; x1 else (x1, x0)\n    y0, y1 = (y0, y1) if y0 &lt; y1 else (y1, y0)\n    return x0, y0, x1, y1\n</code></pre>",
      "rawMarkdown": "By hand. And I see your point, then we can use coods.csv with bounding_boxes.csv for invalid data which we can generate: https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\n\nOr my local code is using this now:\n\n    def get_coods_fixed(filename):\n        x0, y0, x1, y1 = bbdf.loc[filename].values\n        if x1 &lt;= x0 or y1 &lt;= y0:\n            print(f'Warninig {x0} {y0} {x1} {y1} was fixed for {filename}.')\n        x0, x1 = (x0, x1) if x0 &lt; x1 else (x1, x0)\n        y0, y1 = (y0, y1) if y0 &lt; y1 else (y1, y0)\n        return x0, y0, x1, y1",
      "votes": null
    },
    {
      "id": "480697",
      "postDate": "02/28/2019 14:37:12",
      "content": "<p>Thanks for the code. It seems good to use bounding_boxes.csv for invalid data.</p>",
      "rawMarkdown": "Thanks for the code. It seems good to use bounding_boxes.csv for invalid data.",
      "votes": null
    },
    {
      "id": "495092",
      "postDate": "03/20/2019 16:17:24",
      "content": "<p>Hi Radek, the title <strong>IOU 0.93</strong>, is that the training iou or validation iou?  I found that there is a huge gap between training and validation IoU in my experiment, which is ~0.93 and ~0.70..., (what is your valid IOU :)</p>",
      "rawMarkdown": "Hi Radek, the title **IOU 0.93**, is that the training iou or validation iou?  I found that there is a huge gap between training and validation IoU in my experiment, which is ~0.93 and ~0.70..., (what is your valid IOU :)",
      "votes": null
    },
    {
      "id": "495156",
      "postDate": "03/20/2019 18:24:09",
      "content": "<p>That is the validation IoU</p>",
      "rawMarkdown": "That is the validation IoU",
      "votes": null
    },
    {
      "id": "496284",
      "postDate": "03/22/2019 02:54:53",
      "content": "<p>uh.. i made some mistakes... and now Validation IoU is ~0.92, which is satisfactory!\nduring my trial and error, i found something interesting that may be useful for new comers:\n1, If image size is large, like 448, IoU drops a little (0.92-&gt;0.89), i guess bounding box is more difficult to regress if the target is <strong>finer</strong>, because bbox/224 is coarser than bbox/448.\n2, When using CNN with linear output for regression, the targets should not have large deviation.  If we use the bbox coordinate (range from 0~224) directly, it works worse. the reason is that the weight of linear can be very large to fit this range and results in large shake during learning, it is obviously more difficult to learn than squeezing the range to [-1, 1]. I guess this is the reason why Faster RCNN use template based method for decreasing the search space.\nAnyway, thanks Radek!</p>",
      "rawMarkdown": "uh.. i made some mistakes... and now Validation IoU is ~0.92, which is satisfactory!\nduring my trial and error, i found something interesting that may be useful for new comers:\n1, If image size is large, like 448, IoU drops a little (0.92-&gt;0.89), i guess bounding box is more difficult to regress if the target is **finer**, because bbox/224 is coarser than bbox/448.\n2, When using CNN with linear output for regression, the targets should not have large deviation.  If we use the bbox coordinate (range from 0~224) directly, it works worse. the reason is that the weight of linear can be very large to fit this range and results in large shake during learning, it is obviously more difficult to learn than squeezing the range to [-1, 1]. I guess this is the reason why Faster RCNN use template based method for decreasing the search space.\nAnyway, thanks Radek!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 475936,
      "author_name": "vlad0922",
      "author_url": "",
      "post_date": "02/21/2019 11:52:40",
      "content": "<p>At first, thanks for your great work! Could you possibly upload the file with bbox coordinates for each image?</p>",
      "votes": null,
      "replies": [
        {
          "id": 476549,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "02/22/2019 10:18:14",
          "content": "<p>File added to the original post! There are a couple of images (you can find their filenames in the notebook) where the predicted bounding boxes don't make sense (negative width or height). There are 4 such images in train and 1 in test.</p>\n\n<p>The coordinates follow this format: y_up_left, x_up_left, y_low_right, x_low_right. Meaning the first coordinate is the pixel value of the upper left hand corner of the image, x_up_left is the x value of that corner and the remaining two coordinates are for the lower right hand corner. Makes for easy extraction from an image represented as an array.</p>\n\n<p>Best of luck in the competition!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 476160,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "02/21/2019 17:09:17",
      "content": "<p>I’ve been waiting for radek to appear on the lb with a 1.001</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 476397,
      "author_name": "ducvm123",
      "author_url": "",
      "post_date": "02/22/2019 04:25:37",
      "content": "<p>thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 476455,
      "author_name": "amqdnguyen",
      "author_url": "",
      "post_date": "02/22/2019 06:59:04",
      "content": "<p>Wow, Radek! Thank you! I am learning a lot from your notebooks. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 476973,
      "author_name": "kerukun",
      "author_url": "",
      "post_date": "02/23/2019 15:18:10",
      "content": "<p>very nice work and good example notebook :) \nThank you. I appreciate it. I will use it as a reference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 477214,
      "author_name": "rajshree07",
      "author_url": "",
      "post_date": "02/24/2019 06:11:10",
      "content": "<p>Nice work!! Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 477615,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "02/25/2019 01:11:08",
      "content": "<p>Thanks Radek, great work!! Have anyone compared this with the one from this kernel : <a href=\"https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\">https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes</a> ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 477653,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "02/25/2019 03:25:38",
      "content": "<p>Hi <a href=\"/radek1\">@radek1</a>, thank you for sharing.\nI found 5 images have problem (image size will be zero if simply followed), here's update from me.</p>\n\n<pre><code>                x0   y0   x1   y1\nfilename                         \nb4cb30afd.jpg  370  540  600  600\n85a95e7a8.jpg  420  430  510  470\nb370e1339.jpg  270  560  570  660\nd4cb9d6e4.jpg  490  350  550  380\n</code></pre>\n\n<p><strike>6a72d84ca.jpg  490  470  550  500</strike> was removed due to test sample, attached is updated.</p>",
      "votes": null,
      "replies": [
        {
          "id": 479019,
          "author_name": "momi64",
          "author_url": "",
          "post_date": "02/26/2019 23:23:32",
          "content": "<p>How did you get the coordinates?</p>\n\n<p>'6a72d84ca.jpg' is a test image. Manual annotations can not be used for test images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 479982,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "02/27/2019 16:39:34",
          "content": "<p>By hand. And I see your point, then we can use coods.csv with bounding_boxes.csv for invalid data which we can generate: <a href=\"https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\">https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes</a></p>\n\n<p>Or my local code is using this now:</p>\n\n<pre><code>def get_coods_fixed(filename):\n    x0, y0, x1, y1 = bbdf.loc[filename].values\n    if x1 &lt;= x0 or y1 &lt;= y0:\n        print(f'Warninig {x0} {y0} {x1} {y1} was fixed for {filename}.')\n    x0, x1 = (x0, x1) if x0 &lt; x1 else (x1, x0)\n    y0, y1 = (y0, y1) if y0 &lt; y1 else (y1, y0)\n    return x0, y0, x1, y1\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 480697,
          "author_name": "momi64",
          "author_url": "",
          "post_date": "02/28/2019 14:37:12",
          "content": "<p>Thanks for the code. It seems good to use bounding_boxes.csv for invalid data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 479215,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/27/2019 02:52:26",
      "content": "<p>i read a new paper today. maybe can use for future kaggle challenge.</p>\n\n<p>It can can be used for predicting bounding box without annotated data. </p>\n\n<p><a href=\"https://arxiv.org/pdf/1902.09968.pdf\">https://arxiv.org/pdf/1902.09968.pdf</a></p>\n\n<p>\"Mining Objects: Fully Unsupervised Object Discovery and Localization From a Single Image\"</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/479215/11436/unpervsied.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 495092,
      "author_name": "gengshi",
      "author_url": "",
      "post_date": "03/20/2019 16:17:24",
      "content": "<p>Hi Radek, the title <strong>IOU 0.93</strong>, is that the training iou or validation iou?  I found that there is a huge gap between training and validation IoU in my experiment, which is ~0.93 and ~0.70..., (what is your valid IOU :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 495156,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "03/20/2019 18:24:09",
          "content": "<p>That is the validation IoU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496284,
          "author_name": "gengshi",
          "author_url": "",
          "post_date": "03/22/2019 02:54:53",
          "content": "<p>uh.. i made some mistakes... and now Validation IoU is ~0.92, which is satisfactory!\nduring my trial and error, i found something interesting that may be useful for new comers:\n1, If image size is large, like 448, IoU drops a little (0.92-&gt;0.89), i guess bounding box is more difficult to regress if the target is <strong>finer</strong>, because bbox/224 is coarser than bbox/448.\n2, When using CNN with linear output for regression, the targets should not have large deviation.  If we use the bbox coordinate (range from 0~224) directly, it works worse. the reason is that the weight of linear can be very large to fit this range and results in large shake during learning, it is obviously more difficult to learn than squeezing the range to [-1, 1]. I guess this is the reason why Faster RCNN use template based method for decreasing the search space.\nAnyway, thanks Radek!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "475828": "I added more hand annotated [train data](https://github.com/radekosmulski/whale/blob/master/data/annotations.json) to my [fluke detector](https://github.com/radekosmulski/whale/blob/master/fluke_detection_redux.ipynb), which now gives a significantly higher IoU.\n\nI also added code that runs the detector on entire train and test sets and extract predicted bounding boxes to images of a specified size. You can find the notebook with the code [here](https://github.com/radekosmulski/whale/blob/master/extract_bboxes.ipynb).\n\n![enter image description here][1]\n\nIn the notebook I also go into a bit of discussion on how Pytorch Datasets &amp; Dataloaders work so might be of interest if you are planning on doing something a bit more custom.\n\nOnly a few more days left in the competition, but it seems training on extracted bounding boxes vs full image can make quite a difference! (quickly ran a small classification experiment and the improvement seems to be quite significant.\n\nBest of luck! :)\n\nEDIT: adding csv of predicted bounding boxes mapped to image size as per request\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/475828/11366/flukes1.png",
    "475936": "At first, thanks for your great work! Could you possibly upload the file with bbox coordinates for each image?",
    "476160": "I’ve been waiting for radek to appear on the lb with a 1.001",
    "476397": "thanks!",
    "476455": "Wow, Radek! Thank you! I am learning a lot from your notebooks.",
    "476549": "File added to the original post! There are a couple of images (you can find their filenames in the notebook) where the predicted bounding boxes don't make sense (negative width or height). There are 4 such images in train and 1 in test.\n\nThe coordinates follow this format: y_up_left, x_up_left, y_low_right, x_low_right. Meaning the first coordinate is the pixel value of the upper left hand corner of the image, x_up_left is the x value of that corner and the remaining two coordinates are for the lower right hand corner. Makes for easy extraction from an image represented as an array.\n\nBest of luck in the competition!",
    "476973": "very nice work and good example notebook :) \nThank you. I appreciate it. I will use it as a reference.",
    "477214": "Nice work!! Thanks.",
    "477615": "Thanks Radek, great work!! Have anyone compared this with the one from this kernel : https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes ?",
    "477653": "Hi @radek1, thank you for sharing.\nI found 5 images have problem (image size will be zero if simply followed), here's update from me.\n\n                    x0   y0   x1   y1\n    filename                         \n    b4cb30afd.jpg  370  540  600  600\n    85a95e7a8.jpg  420  430  510  470\n    b370e1339.jpg  270  560  570  660\n    d4cb9d6e4.jpg  490  350  550  380\n\n<strike>6a72d84ca.jpg  490  470  550  500</strike> was removed due to test sample, attached is updated.",
    "479019": "How did you get the coordinates?\n\n'6a72d84ca.jpg' is a test image. Manual annotations can not be used for test images.",
    "479215": "i read a new paper today. maybe can use for future kaggle challenge.\n\nIt can can be used for predicting bounding box without annotated data. \n\nhttps://arxiv.org/pdf/1902.09968.pdf\n\n\"Mining Objects: Fully Unsupervised Object Discovery and Localization From a Single Image\"\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/479215/11436/unpervsied.png",
    "479982": "By hand. And I see your point, then we can use coods.csv with bounding_boxes.csv for invalid data which we can generate: https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\n\nOr my local code is using this now:\n\n    def get_coods_fixed(filename):\n        x0, y0, x1, y1 = bbdf.loc[filename].values\n        if x1 &lt;= x0 or y1 &lt;= y0:\n            print(f'Warninig {x0} {y0} {x1} {y1} was fixed for {filename}.')\n        x0, x1 = (x0, x1) if x0 &lt; x1 else (x1, x0)\n        y0, y1 = (y0, y1) if y0 &lt; y1 else (y1, y0)\n        return x0, y0, x1, y1",
    "480697": "Thanks for the code. It seems good to use bounding_boxes.csv for invalid data.",
    "495092": "Hi Radek, the title **IOU 0.93**, is that the training iou or validation iou?  I found that there is a huge gap between training and validation IoU in my experiment, which is ~0.93 and ~0.70..., (what is your valid IOU :)",
    "495156": "That is the validation IoU",
    "496284": "uh.. i made some mistakes... and now Validation IoU is ~0.92, which is satisfactory!\nduring my trial and error, i found something interesting that may be useful for new comers:\n1, If image size is large, like 448, IoU drops a little (0.92-&gt;0.89), i guess bounding box is more difficult to regress if the target is **finer**, because bbox/224 is coarser than bbox/448.\n2, When using CNN with linear output for regression, the targets should not have large deviation.  If we use the bbox coordinate (range from 0~224) directly, it works worse. the reason is that the weight of linear can be very large to fit this range and results in large shake during learning, it is obviously more difficult to learn than squeezing the range to [-1, 1]. I guess this is the reason why Faster RCNN use template based method for decreasing the search space.\nAnyway, thanks Radek!"
  },
  "source": "meta"
}