{
  "id": 305503,
  "title": "cropped&resized(512x512) dataset using detic",
  "url": "/competitions/happy-whale-and-dolphin/discussion/305503",
  "author_name": "phalanx",
  "post_date": "2022-02-05T15:40:16.976000",
  "votes": 212,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I create cropped dataset using <a href=\"https://github.com/facebookresearch/Detic\" target=\"_blank\">Detic</a>.<br>\nDetic train classification head(CLIP embeddings) on the imagenet-2iK in weakly-supervied manner. <br>\nOn detection dataset like coco, we train model in general detection manner. On the other hand, on imagenet-21K, we choose the largest rpn-proposal and calcualte classification loss.<br>\nIn this way, detection model can cover wider semantic domain. So, I thought it would be best to use this model for this object detection.<br>\n<img src=\"https://i.imgur.com/CXJkQCF.png\" alt=\"\"></p>\n<p><strong>How to create this dataset:</strong></p>\n<ol>\n<li>Detect objects in the class of dolphins, whales, and marine life.</li>\n<li>Select the largest bbox from the detection result and enlarge bbox by 1.2 times. then, crop&amp;resize it.</li>\n<li>If not found, just resize the image without cropping.</li>\n</ol>\n<p><strong>Problem</strong><br>\nSince the model may predict smaller proposals with higher confidence, some detection results like the one in the image below are included.<br>\nIf you are interested in false-positive result, please confirm csv file.<br>\n<img src=\"https://i.imgur.com/3cAFvvl.png\" alt=\"\"><br>\ndataset link: <a href=\"https://www.kaggle.com/phalanx/whale2-cropped-dataset\" target=\"_blank\">https://www.kaggle.com/phalanx/whale2-cropped-dataset</a><br>\n<img src=\"https://i.imgur.com/uQjAQ6i.png\" alt=\"https://i.imgur.com/uQjAQ6i.png\"></p>",
  "messages": [
    {
      "id": 1677239,
      "postDate": "2022-02-05T15:40:16.977Z",
      "content": "<p>I create cropped dataset using <a href=\"https://github.com/facebookresearch/Detic\" target=\"_blank\">Detic</a>.<br>\nDetic train classification head(CLIP embeddings) on the imagenet-2iK in weakly-supervied manner. <br>\nOn detection dataset like coco, we train model in general detection manner. On the other hand, on imagenet-21K, we choose the largest rpn-proposal and calcualte classification loss.<br>\nIn this way, detection model can cover wider semantic domain. So, I thought it would be best to use this model for this object detection.<br>\n<img src=\"https://i.imgur.com/CXJkQCF.png\" alt=\"\"></p>\n<p><strong>How to create this dataset:</strong></p>\n<ol>\n<li>Detect objects in the class of dolphins, whales, and marine life.</li>\n<li>Select the largest bbox from the detection result and enlarge bbox by 1.2 times. then, crop&amp;resize it.</li>\n<li>If not found, just resize the image without cropping.</li>\n</ol>\n<p><strong>Problem</strong><br>\nSince the model may predict smaller proposals with higher confidence, some detection results like the one in the image below are included.<br>\nIf you are interested in false-positive result, please confirm csv file.<br>\n<img src=\"https://i.imgur.com/3cAFvvl.png\" alt=\"\"><br>\ndataset link: <a href=\"https://www.kaggle.com/phalanx/whale2-cropped-dataset\" target=\"_blank\">https://www.kaggle.com/phalanx/whale2-cropped-dataset</a><br>\n<img src=\"https://i.imgur.com/uQjAQ6i.png\" alt=\"https://i.imgur.com/uQjAQ6i.png\"></p>",
      "rawMarkdown": "I create cropped dataset using [Detic](https://github.com/facebookresearch/Detic).\nDetic train classification head(CLIP embeddings) on the imagenet-2iK in weakly-supervied manner. \nOn detection dataset like coco, we train model in general detection manner. On the other hand, on imagenet-21K, we choose the largest rpn-proposal and calcualte classification loss.\nIn this way, detection model can cover wider semantic domain. So, I thought it would be best to use this model for this object detection.\n![](https://i.imgur.com/CXJkQCF.png)\n\n**How to create this dataset:**\n1. Detect objects in the class of dolphins, whales, and marine life.\n2. Select the largest bbox from the detection result and enlarge bbox by 1.2 times. then, crop&resize it.\n3. If not found, just resize the image without cropping.\n\n**Problem**\nSince the model may predict smaller proposals with higher confidence, some detection results like the one in the image below are included.\nIf you are interested in false-positive result, please confirm csv file.\n![](https://i.imgur.com/3cAFvvl.png)\ndataset link: [https://www.kaggle.com/phalanx/whale2-cropped-dataset](https://www.kaggle.com/phalanx/whale2-cropped-dataset)\n![https://i.imgur.com/uQjAQ6i.png](https://i.imgur.com/uQjAQ6i.png)",
      "votes": 211
    },
    {
      "id": 1677265,
      "postDate": "2022-02-05T15:57:30.053Z",
      "content": "<p>Did you see any boost in your scores when using cropped version data vs the normal version ? Also I see that many entries of bbox are missing the training.csv.  <br>\nAlso beluga whales do not have top dorsal fins like dolphins, so that species maybe an outlier in itself and we need to think how to tackle that. </p>",
      "rawMarkdown": "Did you see any boost in your scores when using cropped version data vs the normal version ? Also I see that many entries of bbox are missing the training.csv.  \nAlso beluga whales do not have top dorsal fins like dolphins, so that species maybe an outlier in itself and we need to think how to tackle that. ",
      "votes": 7
    },
    {
      "id": 1698268,
      "postDate": "2022-02-20T08:51:58.290Z",
      "content": "<p>I tried using them in one of the high scoring kernels. They improves CV + LB a lot! Nice 🤘 <a href=\"https://www.kaggle.com/lextoumbourou/happywhale-effnet-b6-fork-with-detic-crop\" target=\"_blank\">https://www.kaggle.com/lextoumbourou/happywhale-effnet-b6-fork-with-detic-crop</a></p>",
      "rawMarkdown": "I tried using them in one of the high scoring kernels. They improves CV + LB a lot! Nice 🤘 https://www.kaggle.com/lextoumbourou/happywhale-effnet-b6-fork-with-detic-crop",
      "votes": 5
    },
    {
      "id": 1677437,
      "postDate": "2022-02-05T18:06:39.280Z",
      "content": "<p><a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> wouldn't that be <strong>false-negative</strong>? I can that there are some images with no bbox but can't see any image with the wrong bbox.</p>",
      "rawMarkdown": "@phalanx wouldn't that be **false-negative**? I can that there are some images with no bbox but can't see any image with the wrong bbox.",
      "votes": 4,
      "replies": [
        {
          "id": 1677874,
          "postDate": "2022-02-06T04:48:20.137Z",
          "content": "<p>Look at <code>01615264fa5c3e.jpg</code>(bottom-right image).<br>\nIf there is a correct bbox, the IoU between the predicted result and the correct bbox is very low. I called it false-positive.<br>\n<img src=\"https://i.imgur.com/ANypwMt.png\" alt=\"\"></p>",
          "rawMarkdown": "Look at `01615264fa5c3e.jpg`(bottom-right image).\nIf there is a correct bbox, the IoU between the predicted result and the correct bbox is very low. I called it false-positive.\n![](https://i.imgur.com/ANypwMt.png)",
          "votes": 4
        }
      ]
    },
    {
      "id": 1752149,
      "postDate": "2022-04-11T13:19:03.480Z",
      "content": "<p>Thanks for the dataset. Can you share the code that create the dataset <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> ?</p>",
      "rawMarkdown": "Thanks for the dataset. Can you share the code that create the dataset @phalanx ?\n",
      "votes": 1
    },
    {
      "id": 1753066,
      "postDate": "2022-04-12T12:48:24.117Z",
      "content": "<p>Great work. I am intersested in replicate your work for other datasets in the future. I have a couple of questions:</p>\n<ol>\n<li><p>How did you manage to get all the predictions (bonding boxes) for the whole dataset? Did you use their <a href=\"https://github.com/facebookresearch/Detic/tree/2bb899f93c5181ea10a7397eb6be88719ae43373#demo\" target=\"_blank\">demo</a>?</p></li>\n<li><p>I can see DETIC is based on the Detectron2 framework and they use the Detectron2's <a href=\"https://detectron2.readthedocs.io/en/latest/modules/engine.html#detectron2.engine.defaults.DefaultPredictor\" target=\"_blank\">DefaultPredictor</a> in the demo to get the boxes, but this <em>runs for a single input image.</em> Did you use this method to run the inference?</p></li>\n</ol>",
      "rawMarkdown": "Great work. I am intersested in replicate your work for other datasets in the future. I have a couple of questions:\n\n1. How did you manage to get all the predictions (bonding boxes) for the whole dataset? Did you use their [demo](https://github.com/facebookresearch/Detic/tree/2bb899f93c5181ea10a7397eb6be88719ae43373#demo)?\n\n2. I can see DETIC is based on the Detectron2 framework and they use the Detectron2's [DefaultPredictor](https://detectron2.readthedocs.io/en/latest/modules/engine.html#detectron2.engine.defaults.DefaultPredictor) in the demo to get the boxes, but this *runs for a single input image.* Did you use this method to run the inference?"
    },
    {
      "id": 1750348,
      "postDate": "2022-04-09T15:03:02.300Z",
      "content": "<p>nice work!</p>",
      "rawMarkdown": "nice work!"
    },
    {
      "id": 1703420,
      "postDate": "2022-02-24T13:26:05.070Z",
      "content": "<p>Thanks for your effort, phalanx.<br>\nI'm following your twiteer.</p>",
      "rawMarkdown": "Thanks for your effort, phalanx.\nI'm following your twiteer."
    },
    {
      "id": 1700086,
      "postDate": "2022-02-21T16:37:31.087Z",
      "content": "<p>This is a very good approach to get rid of the noise. Although, I anticipate a lot of poorly detected objects.</p>",
      "rawMarkdown": "This is a very good approach to get rid of the noise. Although, I anticipate a lot of poorly detected objects."
    },
    {
      "id": 1699247,
      "postDate": "2022-02-21T03:49:08.967Z",
      "content": "<p>Looks really useful, thanks for sharing!</p>",
      "rawMarkdown": "Looks really useful, thanks for sharing!"
    },
    {
      "id": 1699197,
      "postDate": "2022-02-21T02:21:48.733Z",
      "content": "<p>thanks for sharing.  just one question,  Is there time to create private dataset for competition?</p>",
      "rawMarkdown": "thanks for sharing.  just one question,  Is there time to create private dataset for competition?"
    },
    {
      "id": 1682821,
      "postDate": "2022-02-09T11:58:33.037Z",
      "content": "<p>Great. It was really helpful to get started.</p>",
      "rawMarkdown": "Great. It was really helpful to get started."
    },
    {
      "id": 1681006,
      "postDate": "2022-02-08T07:02:09.247Z",
      "content": "<p>thanks for your idea<br>\nI will try to do it</p>",
      "rawMarkdown": "thanks for your idea\nI will try to do it"
    },
    {
      "id": 1678951,
      "postDate": "2022-02-06T21:38:47.580Z",
      "content": "<p><a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> do you have any idea how many false-positives there are?</p>",
      "rawMarkdown": "@phalanx do you have any idea how many false-positives there are?"
    },
    {
      "id": 1677916,
      "postDate": "2022-02-06T06:16:26.730Z",
      "content": "<p>thanks for your idea<br>\nI will try to do it </p>",
      "rawMarkdown": "thanks for your idea\nI will try to do it "
    },
    {
      "id": 1677266,
      "postDate": "2022-02-05T15:58:41.130Z",
      "content": "<p>Wow, this is very nice for the beginnig! But I think later on we will definitely need good object detection to achieve good score</p>",
      "rawMarkdown": "Wow, this is very nice for the beginnig! But I think later on we will definitely need good object detection to achieve good score",
      "replies": [
        {
          "id": 1677344,
          "postDate": "2022-02-05T16:43:54.943Z",
          "content": "<p>more than detection I think segmentation will be better for the task, the background is creating some kind of noise.<br>\nIn my initial model I have seen model is using background to cluster images.</p>",
          "rawMarkdown": "more than detection I think segmentation will be better for the task, the background is creating some kind of noise.\nIn my initial model I have seen model is using background to cluster images.",
          "votes": 1
        },
        {
          "id": 1677378,
          "postDate": "2022-02-05T17:03:46.747Z",
          "content": "<p>Hmm, I think this is a problem only for initial models. I'm sure with longer training/better loss functions the background issues would be largely resolved, we have quite a lot of data this time.</p>",
          "rawMarkdown": "Hmm, I think this is a problem only for initial models. I'm sure with longer training/better loss functions the background issues would be largely resolved, we have quite a lot of data this time.",
          "votes": 2
        },
        {
          "id": 1680302,
          "postDate": "2022-02-07T18:11:54.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> interesting way to look at it. I agree that you should ideally be able to identify with just the fin and nothing else. Or to put it in another way, more importantly, the model should learn to ignore the background.</p>\n<p>However, the background could even be a data leak where the watercolor could hint to a location an individual is always observed or such. </p>",
          "rawMarkdown": "@mrinath interesting way to look at it. I agree that you should ideally be able to identify with just the fin and nothing else. Or to put it in another way, more importantly, the model should learn to ignore the background.\n\nHowever, the background could even be a data leak where the watercolor could hint to a location an individual is always observed or such. ",
          "votes": 4
        },
        {
          "id": 1681704,
          "postDate": "2022-02-08T16:38:46.853Z",
          "content": "<p>The background leak will be both bad and good I think, there may be a possibility that two different whales/dolphins were from the same background. the model if it uses the background to classify them, will give us a wrong prediction.</p>",
          "rawMarkdown": "The background leak will be both bad and good I think, there may be a possibility that two different whales/dolphins were from the same background. the model if it uses the background to classify them, will give us a wrong prediction.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1677274,
      "postDate": "2022-02-05T16:04:44.273Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true
    },
    {
      "id": 1680871,
      "postDate": "2022-02-08T04:12:41.970Z",
      "content": "<p>Great! Thank you for sharing sir.</p>",
      "rawMarkdown": "Great! Thank you for sharing sir.",
      "votes": 1
    },
    {
      "id": 1680747,
      "postDate": "2022-02-08T02:27:16.593Z",
      "content": "<p>Thanks for idea <br>\nits really helpful</p>",
      "rawMarkdown": "Thanks for idea \nits really helpful",
      "votes": 1
    },
    {
      "id": 1750711,
      "postDate": "2022-04-10T01:32:27.017Z",
      "content": "<p>Thank you I will try it.</p>",
      "rawMarkdown": "Thank you I will try it."
    },
    {
      "id": 1704118,
      "postDate": "2022-02-25T07:22:08.667Z",
      "content": "<p>Thanks for the share. </p>",
      "rawMarkdown": "Thanks for the share. "
    },
    {
      "id": 1681137,
      "postDate": "2022-02-08T09:01:20.323Z",
      "content": "<p>thank you for sharing… </p>",
      "rawMarkdown": "thank you for sharing... "
    }
  ],
  "comments": [
    {
      "id": 1677265,
      "author_name": "Atharva Phatak",
      "author_url": "",
      "post_date": "2022-02-05T15:57:30.053000",
      "content": "<p>Did you see any boost in your scores when using cropped version data vs the normal version ? Also I see that many entries of bbox are missing the training.csv.  <br>\nAlso beluga whales do not have top dorsal fins like dolphins, so that species maybe an outlier in itself and we need to think how to tackle that. </p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1698268,
      "author_name": "Lex Toumbourou",
      "author_url": "",
      "post_date": "2022-02-20T08:51:58.290000",
      "content": "<p>I tried using them in one of the high scoring kernels. They improves CV + LB a lot! Nice 🤘 <a href=\"https://www.kaggle.com/lextoumbourou/happywhale-effnet-b6-fork-with-detic-crop\" target=\"_blank\">https://www.kaggle.com/lextoumbourou/happywhale-effnet-b6-fork-with-detic-crop</a></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1677437,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2022-02-05T18:06:39.280000",
      "content": "<p><a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> wouldn't that be <strong>false-negative</strong>? I can that there are some images with no bbox but can't see any image with the wrong bbox.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1677874,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2022-02-06T04:48:20.137000",
          "content": "<p>Look at <code>01615264fa5c3e.jpg</code>(bottom-right image).<br>\nIf there is a correct bbox, the IoU between the predicted result and the correct bbox is very low. I called it false-positive.<br>\n<img src=\"https://i.imgur.com/ANypwMt.png\" alt=\"\"></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1752149,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2022-04-11T13:19:03.480000",
      "content": "<p>Thanks for the dataset. Can you share the code that create the dataset <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> ?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1753066,
      "author_name": "Javier Abellán Abenza",
      "author_url": "",
      "post_date": "2022-04-12T12:48:24.117000",
      "content": "<p>Great work. I am intersested in replicate your work for other datasets in the future. I have a couple of questions:</p>\n<ol>\n<li><p>How did you manage to get all the predictions (bonding boxes) for the whole dataset? Did you use their <a href=\"https://github.com/facebookresearch/Detic/tree/2bb899f93c5181ea10a7397eb6be88719ae43373#demo\" target=\"_blank\">demo</a>?</p></li>\n<li><p>I can see DETIC is based on the Detectron2 framework and they use the Detectron2's <a href=\"https://detectron2.readthedocs.io/en/latest/modules/engine.html#detectron2.engine.defaults.DefaultPredictor\" target=\"_blank\">DefaultPredictor</a> in the demo to get the boxes, but this <em>runs for a single input image.</em> Did you use this method to run the inference?</p></li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1750348,
      "author_name": "shangyu xie",
      "author_url": "",
      "post_date": "2022-04-09T15:03:02.300000",
      "content": "<p>nice work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1703420,
      "author_name": "Bamba Tomoyasu",
      "author_url": "",
      "post_date": "2022-02-24T13:26:05.070000",
      "content": "<p>Thanks for your effort, phalanx.<br>\nI'm following your twiteer.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1700086,
      "author_name": "Filemon",
      "author_url": "",
      "post_date": "2022-02-21T16:37:31.087000",
      "content": "<p>This is a very good approach to get rid of the noise. Although, I anticipate a lot of poorly detected objects.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1699247,
      "author_name": "Ivan Aerlic",
      "author_url": "",
      "post_date": "2022-02-21T03:49:08.967000",
      "content": "<p>Looks really useful, thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1699197,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-02-21T02:21:48.733000",
      "content": "<p>thanks for sharing.  just one question,  Is there time to create private dataset for competition?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1682821,
      "author_name": "EitaF",
      "author_url": "",
      "post_date": "2022-02-09T11:58:33.037000",
      "content": "<p>Great. It was really helpful to get started.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1681006,
      "author_name": "Salih BALCI",
      "author_url": "",
      "post_date": "2022-02-08T07:02:09.247000",
      "content": "<p>thanks for your idea<br>\nI will try to do it</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1678951,
      "author_name": "ztsv-av",
      "author_url": "",
      "post_date": "2022-02-06T21:38:47.580000",
      "content": "<p><a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> do you have any idea how many false-positives there are?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1677916,
      "author_name": "zhiliang yang",
      "author_url": "",
      "post_date": "2022-02-06T06:16:26.730000",
      "content": "<p>thanks for your idea<br>\nI will try to do it </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1677266,
      "author_name": "Oleg Sidorshin",
      "author_url": "",
      "post_date": "2022-02-05T15:58:41.130000",
      "content": "<p>Wow, this is very nice for the beginnig! But I think later on we will definitely need good object detection to achieve good score</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1677344,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-02-05T16:43:54.943000",
          "content": "<p>more than detection I think segmentation will be better for the task, the background is creating some kind of noise.<br>\nIn my initial model I have seen model is using background to cluster images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1677378,
          "author_name": "Oleg Sidorshin",
          "author_url": "",
          "post_date": "2022-02-05T17:03:46.747000",
          "content": "<p>Hmm, I think this is a problem only for initial models. I'm sure with longer training/better loss functions the background issues would be largely resolved, we have quite a lot of data this time.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1680302,
          "author_name": "datta",
          "author_url": "",
          "post_date": "2022-02-07T18:11:54.867000",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> interesting way to look at it. I agree that you should ideally be able to identify with just the fin and nothing else. Or to put it in another way, more importantly, the model should learn to ignore the background.</p>\n<p>However, the background could even be a data leak where the watercolor could hint to a location an individual is always observed or such. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1681704,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-02-08T16:38:46.853000",
          "content": "<p>The background leak will be both bad and good I think, there may be a possibility that two different whales/dolphins were from the same background. the model if it uses the background to classify them, will give us a wrong prediction.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1677274,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-05T16:04:44.273000",
      "content": "",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1680871,
      "author_name": "Thanh Nguyen",
      "author_url": "",
      "post_date": "2022-02-08T04:12:41.970000",
      "content": "<p>Great! Thank you for sharing sir.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1680747,
      "author_name": "Anshul Khadse",
      "author_url": "",
      "post_date": "2022-02-08T02:27:16.593000",
      "content": "<p>Thanks for idea <br>\nits really helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1750711,
      "author_name": "fujii",
      "author_url": "",
      "post_date": "2022-04-10T01:32:27.017000",
      "content": "<p>Thank you I will try it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1704118,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2022-02-25T07:22:08.667000",
      "content": "<p>Thanks for the share. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1681137,
      "author_name": "Ahsan Zaman",
      "author_url": "",
      "post_date": "2022-02-08T09:01:20.323000",
      "content": "<p>thank you for sharing… </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1677239": "I create cropped dataset using [Detic](https://github.com/facebookresearch/Detic).\nDetic train classification head(CLIP embeddings) on the imagenet-2iK in weakly-supervied manner. \nOn detection dataset like coco, we train model in general detection manner. On the other hand, on imagenet-21K, we choose the largest rpn-proposal and calcualte classification loss.\nIn this way, detection model can cover wider semantic domain. So, I thought it would be best to use this model for this object detection.\n![](https://i.imgur.com/CXJkQCF.png)\n\n**How to create this dataset:**\n1. Detect objects in the class of dolphins, whales, and marine life.\n2. Select the largest bbox from the detection result and enlarge bbox by 1.2 times. then, crop&resize it.\n3. If not found, just resize the image without cropping.\n\n**Problem**\nSince the model may predict smaller proposals with higher confidence, some detection results like the one in the image below are included.\nIf you are interested in false-positive result, please confirm csv file.\n![](https://i.imgur.com/3cAFvvl.png)\ndataset link: [https://www.kaggle.com/phalanx/whale2-cropped-dataset](https://www.kaggle.com/phalanx/whale2-cropped-dataset)\n![https://i.imgur.com/uQjAQ6i.png](https://i.imgur.com/uQjAQ6i.png)",
    "1677265": "Did you see any boost in your scores when using cropped version data vs the normal version ? Also I see that many entries of bbox are missing the training.csv.  \nAlso beluga whales do not have top dorsal fins like dolphins, so that species maybe an outlier in itself and we need to think how to tackle that. ",
    "1698268": "I tried using them in one of the high scoring kernels. They improves CV + LB a lot! Nice 🤘 https://www.kaggle.com/lextoumbourou/happywhale-effnet-b6-fork-with-detic-crop",
    "1677437": "@phalanx wouldn't that be **false-negative**? I can that there are some images with no bbox but can't see any image with the wrong bbox.",
    "1752149": "Thanks for the dataset. Can you share the code that create the dataset @phalanx ?\n",
    "1753066": "Great work. I am intersested in replicate your work for other datasets in the future. I have a couple of questions:\n\n1. How did you manage to get all the predictions (bonding boxes) for the whole dataset? Did you use their [demo](https://github.com/facebookresearch/Detic/tree/2bb899f93c5181ea10a7397eb6be88719ae43373#demo)?\n\n2. I can see DETIC is based on the Detectron2 framework and they use the Detectron2's [DefaultPredictor](https://detectron2.readthedocs.io/en/latest/modules/engine.html#detectron2.engine.defaults.DefaultPredictor) in the demo to get the boxes, but this *runs for a single input image.* Did you use this method to run the inference?",
    "1750348": "nice work!",
    "1703420": "Thanks for your effort, phalanx.\nI'm following your twiteer.",
    "1700086": "This is a very good approach to get rid of the noise. Although, I anticipate a lot of poorly detected objects.",
    "1699247": "Looks really useful, thanks for sharing!",
    "1699197": "thanks for sharing.  just one question,  Is there time to create private dataset for competition?",
    "1682821": "Great. It was really helpful to get started.",
    "1681006": "thanks for your idea\nI will try to do it",
    "1678951": "@phalanx do you have any idea how many false-positives there are?",
    "1677916": "thanks for your idea\nI will try to do it ",
    "1677266": "Wow, this is very nice for the beginnig! But I think later on we will definitely need good object detection to achieve good score",
    "1677274": "",
    "1680871": "Great! Thank you for sharing sir.",
    "1680747": "Thanks for idea \nits really helpful",
    "1750711": "Thank you I will try it.",
    "1704118": "Thanks for the share. ",
    "1681137": "thank you for sharing... "
  }
}