{
  "id": 304950,
  "title": "Doubt about Species to Individual Id Mapping ?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/304950",
  "author_name": "",
  "post_date": "2022-02-03T04:54:24.033242300Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Based on my understanding of the dataset this looks like a Multilabel Classification Problem to me where for each image we have to accurately predict the species of that whales and dolphins .<br>\nBut in the train dataset a species can have multiple individual ids present .<br>\nEg Species \"beluga\" has following individual ids 0079a7854ab4 , 00c2feba7e03 , 00f58de6056c and so on . <br>\nSo how do we have to predict appropriate Individual_Id for a given species since a species can have multiple Individual Ids .<br>\nAlso please do correct my understanding if its wrong.<br>\nThanks </p>",
  "messages": [
    {
      "id": "1673910",
      "postDate": "02/03/2022 04:54:24",
      "content": "<p>Based on my understanding of the dataset this looks like a Multilabel Classification Problem to me where for each image we have to accurately predict the species of that whales and dolphins .<br>\nBut in the train dataset a species can have multiple individual ids present .<br>\nEg Species \"beluga\" has following individual ids 0079a7854ab4 , 00c2feba7e03 , 00f58de6056c and so on . <br>\nSo how do we have to predict appropriate Individual_Id for a given species since a species can have multiple Individual Ids .<br>\nAlso please do correct my understanding if its wrong.<br>\nThanks </p>",
      "rawMarkdown": "Based on my understanding of the dataset this looks like a Multilabel Classification Problem to me where for each image we have to accurately predict the species of that whales and dolphins .\nBut in the train dataset a species can have multiple individual ids present .\nEg Species \"beluga\" has following individual ids 0079a7854ab4 , 00c2feba7e03 , 00f58de6056c and so on . \nSo how do we have to predict appropriate Individual_Id for a given species since a species can have multiple Individual Ids .\nAlso please do correct my understanding if its wrong.\nThanks",
      "votes": null
    },
    {
      "id": "1673928",
      "postDate": "02/03/2022 05:15:35",
      "content": "<p>You only have to predict the individual ids and not the species. For eg, id <code>0079a7854ab4</code> refers to one individual beluga </p>",
      "rawMarkdown": "You only have to predict the individual ids and not the species. For eg, id `0079a7854ab4` refers to one individual beluga",
      "votes": null
    },
    {
      "id": "1673939",
      "postDate": "02/03/2022 05:29:51",
      "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a>  so this means that for any image we have to predict top 5 Individual Ids amongst 15587 Individual Ids , so the number of classes here are 15587  ?</p>",
      "rawMarkdown": "debarshichanda  so this means that for any image we have to predict top 5 Individual Ids amongst 15587 Individual Ids , so the number of classes here are 15587  ?",
      "votes": null
    },
    {
      "id": "1673947",
      "postDate": "02/03/2022 05:39:31",
      "content": "<p><a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a> yes is right 🙂</p>",
      "rawMarkdown": "sayedathar11 yes is right 🙂",
      "votes": null
    },
    {
      "id": "1673962",
      "postDate": "02/03/2022 05:52:11",
      "content": "<p>and you can predict new_individual, which is to say a prediction that the individual is not represented in the train dataset. This mimics reality where an animal may (and often is) new to the research group</p>",
      "rawMarkdown": "and you can predict new_individual, which is to say a prediction that the individual is not represented in the train dataset. This mimics reality where an animal may (and often is) new to the research group",
      "votes": null
    },
    {
      "id": "1674457",
      "postDate": "02/03/2022 14:23:30",
      "content": "<p>There are many approaches to solve this task. What about multilabelling, I think we should try to do Multi-Labeling, which will be helpful to reduce the \"region\" of searching for similar whales. For example, we have 15000+ different whales, but there are only 100+ whales of some species, so we will search the similar whales there and the error will be decreased. Just my thoughts, if you have others, reply to me.</p>",
      "rawMarkdown": "There are many approaches to solve this task. What about multilabelling, I think we should try to do Multi-Labeling, which will be helpful to reduce the \"region\" of searching for similar whales. For example, we have 15000+ different whales, but there are only 100+ whales of some species, so we will search the similar whales there and the error will be decreased. Just my thoughts, if you have others, reply to me.",
      "votes": null
    },
    {
      "id": "1674525",
      "postDate": "02/03/2022 15:11:56",
      "content": "<p>one potential problem here is that some individuals may be easier to detect directly rather than labeling to species first. For example, fin, sei and bryde's whales all have very similar dorsal fins, where we can expect more variation between individuals than between species.</p>",
      "rawMarkdown": "one potential problem here is that some individuals may be easier to detect directly rather than labeling to species first. For example, fin, sei and bryde's whales all have very similar dorsal fins, where we can expect more variation between individuals than between species.",
      "votes": null
    },
    {
      "id": "1674574",
      "postDate": "02/03/2022 15:50:51",
      "content": "<p>Due to the nature of the problem, you may have more success (or just an easier time) looking at this like a clustering/similarity problem. (Kind of like the previous Shopee competition).</p>\n<p>In the previous competition, the HappyWhale team implemented a solution that leveraged <a href=\"https://arxiv.org/abs/1801.07698\" target=\"_blank\"><strong>ArcFace</strong></a>. This encourages us to think about the problem in a way that we might think about facial recognition or some other few-shot classification task. </p>\n<h2><br></h2>\n<p>My basic approach would be:</p>\n<hr>\n<ol>\n<li>Train some sort of image model to generate embeddings of the images<ul>\n<li>You could probably start with a model pre-trained on ImageNet as a baseline</li></ul></li>\n<li>Create a pipeline that allows for an image to be passed into the model so an embedding is generated, this embedding will then be 'compared' with all the previously generated embeddings (using a similarity metric, ArcFace, clustering, etc.) and you then rank/average the top most likely other prediciton to identify which individuals in the training corpus the new individual is closest to.</li>\n<li>If the new individual/image is close to an existing cluster/individual, then you could assume it is the <strong>same individual</strong>.</li>\n<li>If the new individual/image is NOT close to an existing cluster/individual, you may be able to assume it is a <strong>new individual</strong></li>\n</ol>\n<hr>\n<p>Take all this with a grain of salt as I have barely begun to look into this competition, but I will be building a baseline along these lines over the next few days. Hope this helps!</p>",
      "rawMarkdown": "Due to the nature of the problem, you may have more success (or just an easier time) looking at this like a clustering/similarity problem. (Kind of like the previous Shopee competition).\n\nIn the previous competition, the HappyWhale team implemented a solution that leveraged [**ArcFace**](https://arxiv.org/abs/1801.07698). This encourages us to think about the problem in a way that we might think about facial recognition or some other few-shot classification task. \n\n<br>\n---\n\nMy basic approach would be:\n\n---\n\n1. Train some sort of image model to generate embeddings of the images\n  * You could probably start with a model pre-trained on ImageNet as a baseline\n2. Create a pipeline that allows for an image to be passed into the model so an embedding is generated, this embedding will then be 'compared' with all the previously generated embeddings (using a similarity metric, ArcFace, clustering, etc.) and you then rank/average the top most likely other prediciton to identify which individuals in the training corpus the new individual is closest to.\n3. If the new individual/image is close to an existing cluster/individual, then you could assume it is the **same individual**.\n4. If the new individual/image is NOT close to an existing cluster/individual, you may be able to assume it is a **new individual**\n\n---\n\nTake all this with a grain of salt as I have barely begun to look into this competition, but I will be building a baseline along these lines over the next few days. Hope this helps!",
      "votes": null
    },
    {
      "id": "1674603",
      "postDate": "02/03/2022 16:20:30",
      "content": "<p>Yes, there are many problems that will be disclosed, but we should try it!</p>",
      "rawMarkdown": "Yes, there are many problems that will be disclosed, but we should try it!",
      "votes": null
    },
    {
      "id": "1675431",
      "postDate": "02/04/2022 09:09:01",
      "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>  , Thanks for your comment , I actually didn't participated in Shoppe Competition so didn't thought of approach , also haven't look into details of previous Happywhale competition , but your reasoning does sounds logical to me . I will definitely learn this and  try this approach out  and see how it works .</p>",
      "rawMarkdown": "dschettler8845  , Thanks for your comment , I actually didn't participated in Shoppe Competition so didn't thought of approach , also haven't look into details of previous Happywhale competition , but your reasoning does sounds logical to me . I will definitely learn this and  try this approach out  and see how it works .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1673928,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "02/03/2022 05:15:35",
      "content": "<p>You only have to predict the individual ids and not the species. For eg, id <code>0079a7854ab4</code> refers to one individual beluga </p>",
      "votes": null,
      "replies": [
        {
          "id": 1673939,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "02/03/2022 05:29:51",
          "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a>  so this means that for any image we have to predict top 5 Individual Ids amongst 15587 Individual Ids , so the number of classes here are 15587  ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1673947,
          "author_name": "rdizzl3",
          "author_url": "",
          "post_date": "02/03/2022 05:39:31",
          "content": "<p><a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a> yes is right 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1673962,
          "author_name": "tedcheese",
          "author_url": "",
          "post_date": "02/03/2022 05:52:11",
          "content": "<p>and you can predict new_individual, which is to say a prediction that the individual is not represented in the train dataset. This mimics reality where an animal may (and often is) new to the research group</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1674457,
      "author_name": "vad13irt",
      "author_url": "",
      "post_date": "02/03/2022 14:23:30",
      "content": "<p>There are many approaches to solve this task. What about multilabelling, I think we should try to do Multi-Labeling, which will be helpful to reduce the \"region\" of searching for similar whales. For example, we have 15000+ different whales, but there are only 100+ whales of some species, so we will search the similar whales there and the error will be decreased. Just my thoughts, if you have others, reply to me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1674525,
          "author_name": "tedcheese",
          "author_url": "",
          "post_date": "02/03/2022 15:11:56",
          "content": "<p>one potential problem here is that some individuals may be easier to detect directly rather than labeling to species first. For example, fin, sei and bryde's whales all have very similar dorsal fins, where we can expect more variation between individuals than between species.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1674603,
          "author_name": "vad13irt",
          "author_url": "",
          "post_date": "02/03/2022 16:20:30",
          "content": "<p>Yes, there are many problems that will be disclosed, but we should try it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1674574,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "02/03/2022 15:50:51",
      "content": "<p>Due to the nature of the problem, you may have more success (or just an easier time) looking at this like a clustering/similarity problem. (Kind of like the previous Shopee competition).</p>\n<p>In the previous competition, the HappyWhale team implemented a solution that leveraged <a href=\"https://arxiv.org/abs/1801.07698\" target=\"_blank\"><strong>ArcFace</strong></a>. This encourages us to think about the problem in a way that we might think about facial recognition or some other few-shot classification task. </p>\n<h2><br></h2>\n<p>My basic approach would be:</p>\n<hr>\n<ol>\n<li>Train some sort of image model to generate embeddings of the images<ul>\n<li>You could probably start with a model pre-trained on ImageNet as a baseline</li></ul></li>\n<li>Create a pipeline that allows for an image to be passed into the model so an embedding is generated, this embedding will then be 'compared' with all the previously generated embeddings (using a similarity metric, ArcFace, clustering, etc.) and you then rank/average the top most likely other prediciton to identify which individuals in the training corpus the new individual is closest to.</li>\n<li>If the new individual/image is close to an existing cluster/individual, then you could assume it is the <strong>same individual</strong>.</li>\n<li>If the new individual/image is NOT close to an existing cluster/individual, you may be able to assume it is a <strong>new individual</strong></li>\n</ol>\n<hr>\n<p>Take all this with a grain of salt as I have barely begun to look into this competition, but I will be building a baseline along these lines over the next few days. Hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1675431,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "02/04/2022 09:09:01",
          "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>  , Thanks for your comment , I actually didn't participated in Shoppe Competition so didn't thought of approach , also haven't look into details of previous Happywhale competition , but your reasoning does sounds logical to me . I will definitely learn this and  try this approach out  and see how it works .</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1673910": "Based on my understanding of the dataset this looks like a Multilabel Classification Problem to me where for each image we have to accurately predict the species of that whales and dolphins .\nBut in the train dataset a species can have multiple individual ids present .\nEg Species \"beluga\" has following individual ids 0079a7854ab4 , 00c2feba7e03 , 00f58de6056c and so on . \nSo how do we have to predict appropriate Individual_Id for a given species since a species can have multiple Individual Ids .\nAlso please do correct my understanding if its wrong.\nThanks",
    "1673928": "You only have to predict the individual ids and not the species. For eg, id `0079a7854ab4` refers to one individual beluga",
    "1673939": "debarshichanda  so this means that for any image we have to predict top 5 Individual Ids amongst 15587 Individual Ids , so the number of classes here are 15587  ?",
    "1673947": "sayedathar11 yes is right 🙂",
    "1673962": "and you can predict new_individual, which is to say a prediction that the individual is not represented in the train dataset. This mimics reality where an animal may (and often is) new to the research group",
    "1674457": "There are many approaches to solve this task. What about multilabelling, I think we should try to do Multi-Labeling, which will be helpful to reduce the \"region\" of searching for similar whales. For example, we have 15000+ different whales, but there are only 100+ whales of some species, so we will search the similar whales there and the error will be decreased. Just my thoughts, if you have others, reply to me.",
    "1674525": "one potential problem here is that some individuals may be easier to detect directly rather than labeling to species first. For example, fin, sei and bryde's whales all have very similar dorsal fins, where we can expect more variation between individuals than between species.",
    "1674574": "Due to the nature of the problem, you may have more success (or just an easier time) looking at this like a clustering/similarity problem. (Kind of like the previous Shopee competition).\n\nIn the previous competition, the HappyWhale team implemented a solution that leveraged [**ArcFace**](https://arxiv.org/abs/1801.07698). This encourages us to think about the problem in a way that we might think about facial recognition or some other few-shot classification task. \n\n<br>\n---\n\nMy basic approach would be:\n\n---\n\n1. Train some sort of image model to generate embeddings of the images\n  * You could probably start with a model pre-trained on ImageNet as a baseline\n2. Create a pipeline that allows for an image to be passed into the model so an embedding is generated, this embedding will then be 'compared' with all the previously generated embeddings (using a similarity metric, ArcFace, clustering, etc.) and you then rank/average the top most likely other prediciton to identify which individuals in the training corpus the new individual is closest to.\n3. If the new individual/image is close to an existing cluster/individual, then you could assume it is the **same individual**.\n4. If the new individual/image is NOT close to an existing cluster/individual, you may be able to assume it is a **new individual**\n\n---\n\nTake all this with a grain of salt as I have barely begun to look into this competition, but I will be building a baseline along these lines over the next few days. Hope this helps!",
    "1674603": "Yes, there are many problems that will be disclosed, but we should try it!",
    "1675431": "dschettler8845  , Thanks for your comment , I actually didn't participated in Shoppe Competition so didn't thought of approach , also haven't look into details of previous Happywhale competition , but your reasoning does sounds logical to me . I will definitely learn this and  try this approach out  and see how it works ."
  },
  "source": "meta"
}