{
  "id": 306442,
  "title": "Low number of samples per individual",
  "url": "/competitions/happy-whale-and-dolphin/discussion/306442",
  "author_name": "",
  "post_date": "2022-02-09T12:08:43.051439300Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>There are so many great notebook and resources out here for this competition, so I am pretty sure this kind of analysis has been done before. I think one of the main challenges for this dataset is the low number of samples for most of the individuals.</p>\n<p>My analysis showed, that:<br>\n<strong>59.40%</strong> of all individuals are represented by 1 sample in the training set.<br>\n<strong>79.23%</strong> of all individuals are represented by a maximum of 2 samples in the training set.<br>\n<strong>84.19%</strong> of all individuals are represented by a maximum of 3 samples in the training set.</p>\n<p>For all deep learning models, this is a pretty low number of samples to train with. <br>\nThe data is really great to train a species classifier, but moving on to recognizing individuals seems pretty tough. I mean pretty tough to train of course, but also very hard to evaluate, e.g. if you only have 1 sample for more than half of the <em>classes</em>.</p>\n<p>While I have been doing a lot of work in general object detection and classification, I am pretty new to the field of individual recognition. Can somebody give me a hint of what is the current SOTA for this kind of problem?</p>",
  "messages": [
    {
      "id": "1682833",
      "postDate": "02/09/2022 12:08:43",
      "content": "<p>There are so many great notebook and resources out here for this competition, so I am pretty sure this kind of analysis has been done before. I think one of the main challenges for this dataset is the low number of samples for most of the individuals.</p>\n<p>My analysis showed, that:<br>\n<strong>59.40%</strong> of all individuals are represented by 1 sample in the training set.<br>\n<strong>79.23%</strong> of all individuals are represented by a maximum of 2 samples in the training set.<br>\n<strong>84.19%</strong> of all individuals are represented by a maximum of 3 samples in the training set.</p>\n<p>For all deep learning models, this is a pretty low number of samples to train with. <br>\nThe data is really great to train a species classifier, but moving on to recognizing individuals seems pretty tough. I mean pretty tough to train of course, but also very hard to evaluate, e.g. if you only have 1 sample for more than half of the <em>classes</em>.</p>\n<p>While I have been doing a lot of work in general object detection and classification, I am pretty new to the field of individual recognition. Can somebody give me a hint of what is the current SOTA for this kind of problem?</p>",
      "rawMarkdown": "There are so many great notebook and resources out here for this competition, so I am pretty sure this kind of analysis has been done before. I think one of the main challenges for this dataset is the low number of samples for most of the individuals.\n\nMy analysis showed, that:\n**59.40%** of all individuals are represented by 1 sample in the training set.\n**79.23%** of all individuals are represented by a maximum of 2 samples in the training set.\n**84.19%** of all individuals are represented by a maximum of 3 samples in the training set.\n\nFor all deep learning models, this is a pretty low number of samples to train with. \nThe data is really great to train a species classifier, but moving on to recognizing individuals seems pretty tough. I mean pretty tough to train of course, but also very hard to evaluate, e.g. if you only have 1 sample for more than half of the *classes*.\n\nWhile I have been doing a lot of work in general object detection and classification, I am pretty new to the field of individual recognition. Can somebody give me a hint of what is the current SOTA for this kind of problem?",
      "votes": null
    },
    {
      "id": "1683353",
      "postDate": "02/09/2022 18:49:11",
      "content": "<p>Kaggle has a number of conceptually similar competitions where you can browse winning strategies. The recent ones I found to be most helpful were: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/humpback-whale-identification/\" target=\"_blank\">Humpback Whale Identification</a> (3000+ whales)</li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2021/\" target=\"_blank\">Google Landmark Recognition 2021</a> (81k landmarks)</li>\n<li><a href=\"https://www.kaggle.com/c/hotel-id-2021-fgvc8\" target=\"_blank\">Hotel-ID to Combat Human Trafficking 2021</a> (7700 hotels)</li>\n</ul>",
      "rawMarkdown": "Kaggle has a number of conceptually similar competitions where you can browse winning strategies. The recent ones I found to be most helpful were: \n- [Humpback Whale Identification](https://www.kaggle.com/c/humpback-whale-identification/) (3000+ whales)\n- [Google Landmark Recognition 2021](https://www.kaggle.com/c/landmark-recognition-2021/) (81k landmarks)\n- [Hotel-ID to Combat Human Trafficking 2021](https://www.kaggle.com/c/hotel-id-2021-fgvc8) (7700 hotels)",
      "votes": null
    },
    {
      "id": "1686055",
      "postDate": "02/11/2022 18:12:34",
      "content": "<p>I would suggest you to kindly check out the winning solutions from the previous whale challenge. That's really close to this one in terms of data too 🐳</p>\n<p>Siamese networks or forms of \"one-shot learning\" might apply well here. </p>\n<p>Good Luck! :)</p>",
      "rawMarkdown": "I would suggest you to kindly check out the winning solutions from the previous whale challenge. That's really close to this one in terms of data too 🐳\n\nSiamese networks or forms of \"one-shot learning\" might apply well here. \n\nGood Luck! :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1683353,
      "author_name": "rturley",
      "author_url": "",
      "post_date": "02/09/2022 18:49:11",
      "content": "<p>Kaggle has a number of conceptually similar competitions where you can browse winning strategies. The recent ones I found to be most helpful were: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/humpback-whale-identification/\" target=\"_blank\">Humpback Whale Identification</a> (3000+ whales)</li>\n<li><a href=\"https://www.kaggle.com/c/landmark-recognition-2021/\" target=\"_blank\">Google Landmark Recognition 2021</a> (81k landmarks)</li>\n<li><a href=\"https://www.kaggle.com/c/hotel-id-2021-fgvc8\" target=\"_blank\">Hotel-ID to Combat Human Trafficking 2021</a> (7700 hotels)</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1686055,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/11/2022 18:12:34",
      "content": "<p>I would suggest you to kindly check out the winning solutions from the previous whale challenge. That's really close to this one in terms of data too 🐳</p>\n<p>Siamese networks or forms of \"one-shot learning\" might apply well here. </p>\n<p>Good Luck! :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1682833": "There are so many great notebook and resources out here for this competition, so I am pretty sure this kind of analysis has been done before. I think one of the main challenges for this dataset is the low number of samples for most of the individuals.\n\nMy analysis showed, that:\n**59.40%** of all individuals are represented by 1 sample in the training set.\n**79.23%** of all individuals are represented by a maximum of 2 samples in the training set.\n**84.19%** of all individuals are represented by a maximum of 3 samples in the training set.\n\nFor all deep learning models, this is a pretty low number of samples to train with. \nThe data is really great to train a species classifier, but moving on to recognizing individuals seems pretty tough. I mean pretty tough to train of course, but also very hard to evaluate, e.g. if you only have 1 sample for more than half of the *classes*.\n\nWhile I have been doing a lot of work in general object detection and classification, I am pretty new to the field of individual recognition. Can somebody give me a hint of what is the current SOTA for this kind of problem?",
    "1683353": "Kaggle has a number of conceptually similar competitions where you can browse winning strategies. The recent ones I found to be most helpful were: \n- [Humpback Whale Identification](https://www.kaggle.com/c/humpback-whale-identification/) (3000+ whales)\n- [Google Landmark Recognition 2021](https://www.kaggle.com/c/landmark-recognition-2021/) (81k landmarks)\n- [Hotel-ID to Combat Human Trafficking 2021](https://www.kaggle.com/c/hotel-id-2021-fgvc8) (7700 hotels)",
    "1686055": "I would suggest you to kindly check out the winning solutions from the previous whale challenge. That's really close to this one in terms of data too 🐳\n\nSiamese networks or forms of \"one-shot learning\" might apply well here. \n\nGood Luck! :)"
  },
  "source": "meta"
}