{
  "id": 76276,
  "title": "Some new_whale images belong to known IDs",
  "url": "/competitions/humpback-whale-identification/discussion/76276",
  "author_name": "",
  "post_date": "2018-12-31T13:14:08.695768600Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>After a quick manual review, I found some images in the train set which are labeled as <code>new_whale</code> but I'm pretty sure that they belong to some of the known IDs.</p>\n\n<p>Like many of the public solutions, I didn't use the <code>new_whale</code> images to train my model. I create a set of potential misslabels by running my best model on the <code>new_whale</code> images and flagging the image with the top1 predicted class (thresholding at 0.8).</p>\n\n<p>I share here the full list of potential misslabels (<code>maybe_not_so_new.json</code>) and the 13 images (out of only ~50 reviewed, is a manual process and not so easy to verify) that I'm pretty sure belong to the top1 predicted class (<code>pretty_sure_not_so_new.json</code>).</p>\n\n<p>Please consider sharing your <code>pretty_sure_not_so_new</code> updates :)</p>\n\n<p>I also add an example of 2 images from the train set:\n- new_whale <code>721a359d6.jpg</code>\n- w_3ef0017 <code>f93a2f4ef.jpg</code></p>",
  "messages": [
    {
      "id": "448209",
      "postDate": "12/31/2018 13:14:08",
      "content": "<p>After a quick manual review, I found some images in the train set which are labeled as <code>new_whale</code> but I'm pretty sure that they belong to some of the known IDs.</p>\n\n<p>Like many of the public solutions, I didn't use the <code>new_whale</code> images to train my model. I create a set of potential misslabels by running my best model on the <code>new_whale</code> images and flagging the image with the top1 predicted class (thresholding at 0.8).</p>\n\n<p>I share here the full list of potential misslabels (<code>maybe_not_so_new.json</code>) and the 13 images (out of only ~50 reviewed, is a manual process and not so easy to verify) that I'm pretty sure belong to the top1 predicted class (<code>pretty_sure_not_so_new.json</code>).</p>\n\n<p>Please consider sharing your <code>pretty_sure_not_so_new</code> updates :)</p>\n\n<p>I also add an example of 2 images from the train set:\n- new_whale <code>721a359d6.jpg</code>\n- w_3ef0017 <code>f93a2f4ef.jpg</code></p>",
      "rawMarkdown": "After a quick manual review, I found some images in the train set which are labeled as `new_whale` but I'm pretty sure that they belong to some of the known IDs.\n\nLike many of the public solutions, I didn't use the `new_whale` images to train my model. I create a set of potential misslabels by running my best model on the `new_whale` images and flagging the image with the top1 predicted class (thresholding at 0.8).\n\nI share here the full list of potential misslabels (`maybe_not_so_new.json`) and the 13 images (out of only ~50 reviewed, is a manual process and not so easy to verify) that I'm pretty sure belong to the top1 predicted class (`pretty_sure_not_so_new.json`).\n\nPlease consider sharing your `pretty_sure_not_so_new` updates :)\n\nI also add an example of 2 images from the train set:\n- new_whale `721a359d6.jpg`\n- w_3ef0017 `f93a2f4ef.jpg`",
      "votes": null
    },
    {
      "id": "450340",
      "postDate": "01/04/2019 18:01:01",
      "content": "<p>Nice findings ! I wonder whether test set also include mislabels or not.  One possible approach may be putting <code>new_whale</code> in each predictions initiatively. They correspond to really new whale and mislabels.</p>",
      "rawMarkdown": "Nice findings ! I wonder whether test set also include mislabels or not.  One possible approach may be putting `new_whale` in each predictions initiatively. They correspond to really new whale and mislabels.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 450340,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "01/04/2019 18:01:01",
      "content": "<p>Nice findings ! I wonder whether test set also include mislabels or not.  One possible approach may be putting <code>new_whale</code> in each predictions initiatively. They correspond to really new whale and mislabels.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "448209": "After a quick manual review, I found some images in the train set which are labeled as `new_whale` but I'm pretty sure that they belong to some of the known IDs.\n\nLike many of the public solutions, I didn't use the `new_whale` images to train my model. I create a set of potential misslabels by running my best model on the `new_whale` images and flagging the image with the top1 predicted class (thresholding at 0.8).\n\nI share here the full list of potential misslabels (`maybe_not_so_new.json`) and the 13 images (out of only ~50 reviewed, is a manual process and not so easy to verify) that I'm pretty sure belong to the top1 predicted class (`pretty_sure_not_so_new.json`).\n\nPlease consider sharing your `pretty_sure_not_so_new` updates :)\n\nI also add an example of 2 images from the train set:\n- new_whale `721a359d6.jpg`\n- w_3ef0017 `f93a2f4ef.jpg`",
    "450340": "Nice findings ! I wonder whether test set also include mislabels or not.  One possible approach may be putting `new_whale` in each predictions initiatively. They correspond to really new whale and mislabels."
  },
  "source": "meta"
}