{
  "id": 314838,
  "title": "Why do we always add a \"new individual\" in the most popular kernals",
  "url": "/competitions/happy-whale-and-dolphin/discussion/314838",
  "author_name": "",
  "post_date": "2022-03-24T17:11:56.827255100Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Is there any reason that we always add a \"new individual\" in our prediction either in the first place or the second place? </p>\n<p><code>def get_predictions(test_df,threshold=0.2):\n    predictions = {}\n    for i,row in tqdm(test_df.iterrows()):\n        if row.image in predictions:\n            if len(predictions[row.image])==5:\n                continue\n            predictions[row.image].append(row.target)\n        elif row.confidence&gt;threshold:\n            predictions[row.image] = [row.target,'new_individual']\n        else:\n            predictions[row.image] = ['new_individual',row.target]</code></p>",
  "messages": [
    {
      "id": "1733867",
      "postDate": "03/24/2022 17:11:56",
      "content": "<p>Is there any reason that we always add a \"new individual\" in our prediction either in the first place or the second place? </p>\n<p><code>def get_predictions(test_df,threshold=0.2):\n    predictions = {}\n    for i,row in tqdm(test_df.iterrows()):\n        if row.image in predictions:\n            if len(predictions[row.image])==5:\n                continue\n            predictions[row.image].append(row.target)\n        elif row.confidence&gt;threshold:\n            predictions[row.image] = [row.target,'new_individual']\n        else:\n            predictions[row.image] = ['new_individual',row.target]</code></p>",
      "rawMarkdown": "Is there any reason that we always add a \"new individual\" in our prediction either in the first place or the second place? \n\n`def get_predictions(test_df,threshold=0.2):\n    predictions = {}\n    for i,row in tqdm(test_df.iterrows()):\n        if row.image in predictions:\n            if len(predictions[row.image])==5:\n                continue\n            predictions[row.image].append(row.target)\n        elif row.confidence>threshold:\n            predictions[row.image] = [row.target,'new_individual']\n        else:\n            predictions[row.image] = ['new_individual',row.target]`",
      "votes": null
    },
    {
      "id": "1735239",
      "postDate": "03/26/2022 03:19:13",
      "content": "<p>As shown by <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/305428\" target=\"_blank\">LB probing</a>, 11% of the images in the public test set are of whales or dolphins never seen in the training set. Therefore, when your model gives you only one strong match, it is a good strategy to say as a 2nd guess that it is a new individual. Same thing when you do not have any strong match at all, you can just assume the animal has never been seen before. Of course, there is always a risk that the distribution of new individuals in the private test set is different so this threshold could have a dramatic impact on your final score if set manually. Hope this makes sense!</p>",
      "rawMarkdown": "As shown by [LB probing](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/305428), 11% of the images in the public test set are of whales or dolphins never seen in the training set. Therefore, when your model gives you only one strong match, it is a good strategy to say as a 2nd guess that it is a new individual. Same thing when you do not have any strong match at all, you can just assume the animal has never been seen before. Of course, there is always a risk that the distribution of new individuals in the private test set is different so this threshold could have a dramatic impact on your final score if set manually. Hope this makes sense!",
      "votes": null
    },
    {
      "id": "1735541",
      "postDate": "03/26/2022 11:33:57",
      "content": "<p>Thank you for your reply! I believe the logic they use is only comparing the first target confidence with the threshold. In other words, the 'new individual' will always appear in either the first or second term no matter how confident our second target guess is, which seems strange to me. </p>",
      "rawMarkdown": "Thank you for your reply! I believe the logic they use is only comparing the first target confidence with the threshold. In other words, the 'new individual' will always appear in either the first or second term no matter how confident our second target guess is, which seems strange to me.",
      "votes": null
    },
    {
      "id": "1736942",
      "postDate": "03/28/2022 01:16:09",
      "content": "<p>weight biased</p>",
      "rawMarkdown": "weight biased",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1735239,
      "author_name": "frlemarchand",
      "author_url": "",
      "post_date": "03/26/2022 03:19:13",
      "content": "<p>As shown by <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/305428\" target=\"_blank\">LB probing</a>, 11% of the images in the public test set are of whales or dolphins never seen in the training set. Therefore, when your model gives you only one strong match, it is a good strategy to say as a 2nd guess that it is a new individual. Same thing when you do not have any strong match at all, you can just assume the animal has never been seen before. Of course, there is always a risk that the distribution of new individuals in the private test set is different so this threshold could have a dramatic impact on your final score if set manually. Hope this makes sense!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1735541,
          "author_name": "runjiali",
          "author_url": "",
          "post_date": "03/26/2022 11:33:57",
          "content": "<p>Thank you for your reply! I believe the logic they use is only comparing the first target confidence with the threshold. In other words, the 'new individual' will always appear in either the first or second term no matter how confident our second target guess is, which seems strange to me. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1736942,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "03/28/2022 01:16:09",
          "content": "<p>weight biased</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1733867": "Is there any reason that we always add a \"new individual\" in our prediction either in the first place or the second place? \n\n`def get_predictions(test_df,threshold=0.2):\n    predictions = {}\n    for i,row in tqdm(test_df.iterrows()):\n        if row.image in predictions:\n            if len(predictions[row.image])==5:\n                continue\n            predictions[row.image].append(row.target)\n        elif row.confidence>threshold:\n            predictions[row.image] = [row.target,'new_individual']\n        else:\n            predictions[row.image] = ['new_individual',row.target]`",
    "1735239": "As shown by [LB probing](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/305428), 11% of the images in the public test set are of whales or dolphins never seen in the training set. Therefore, when your model gives you only one strong match, it is a good strategy to say as a 2nd guess that it is a new individual. Same thing when you do not have any strong match at all, you can just assume the animal has never been seen before. Of course, there is always a risk that the distribution of new individuals in the private test set is different so this threshold could have a dramatic impact on your final score if set manually. Hope this makes sense!",
    "1735541": "Thank you for your reply! I believe the logic they use is only comparing the first target confidence with the threshold. In other words, the 'new individual' will always appear in either the first or second term no matter how confident our second target guess is, which seems strange to me.",
    "1736942": "weight biased"
  },
  "source": "meta"
}