{
  "id": 305179,
  "title": "Doubt about rare Individual Ids ",
  "url": "/competitions/happy-whale-and-dolphin/discussion/305179",
  "author_name": "",
  "post_date": "2022-02-04T05:43:25.113628500Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Greetings everyone. I want to ask how are you all using individual ids for training which are present just once or twice in the whole dataset ? Should we remove those rare ids ?</p>",
  "messages": [
    {
      "id": "1675227",
      "postDate": "02/04/2022 05:43:25",
      "content": "<p>Greetings everyone. I want to ask how are you all using individual ids for training which are present just once or twice in the whole dataset ? Should we remove those rare ids ?</p>",
      "rawMarkdown": "Greetings everyone. I want to ask how are you all using individual ids for training which are present just once or twice in the whole dataset ? Should we remove those rare ids ?",
      "votes": null
    },
    {
      "id": "1675452",
      "postDate": "02/04/2022 09:22:52",
      "content": "<p>I would not recommend removing rare individuals. You can make a profit from even 1 sample per individual. For example, you could try to use something like contrastive loss (<a href=\"https://medium.com/@maksym.bekuzarov/losses-explained-contrastive-loss-f8f57fe32246\" target=\"_blank\">check out this article</a>). <br>\nHere is a good starter notebook by <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> - <a href=\"https://www.kaggle.com/debarshichanda/pytorch-happywhale-siamese-starter\" target=\"_blank\">[Pytorch] HappyWhale Siamese Starter</a> which implements siamese network.</p>",
      "rawMarkdown": "I would not recommend removing rare individuals. You can make a profit from even 1 sample per individual. For example, you could try to use something like contrastive loss ([check out this article](https://medium.com/@maksym.bekuzarov/losses-explained-contrastive-loss-f8f57fe32246)). \nHere is a good starter notebook by @debarshichanda - [[Pytorch] HappyWhale Siamese Starter](https://www.kaggle.com/debarshichanda/pytorch-happywhale-siamese-starter) which implements siamese network.",
      "votes": null
    },
    {
      "id": "1678646",
      "postDate": "02/06/2022 17:25:00",
      "content": "<p>The previous Kaggle competition had a large/similar class imbalance. You should also checkout the solutions from there. </p>",
      "rawMarkdown": "The previous Kaggle competition had a large/similar class imbalance. You should also checkout the solutions from there.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1675452,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "02/04/2022 09:22:52",
      "content": "<p>I would not recommend removing rare individuals. You can make a profit from even 1 sample per individual. For example, you could try to use something like contrastive loss (<a href=\"https://medium.com/@maksym.bekuzarov/losses-explained-contrastive-loss-f8f57fe32246\" target=\"_blank\">check out this article</a>). <br>\nHere is a good starter notebook by <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> - <a href=\"https://www.kaggle.com/debarshichanda/pytorch-happywhale-siamese-starter\" target=\"_blank\">[Pytorch] HappyWhale Siamese Starter</a> which implements siamese network.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1678646,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/06/2022 17:25:00",
      "content": "<p>The previous Kaggle competition had a large/similar class imbalance. You should also checkout the solutions from there. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1675227": "Greetings everyone. I want to ask how are you all using individual ids for training which are present just once or twice in the whole dataset ? Should we remove those rare ids ?",
    "1675452": "I would not recommend removing rare individuals. You can make a profit from even 1 sample per individual. For example, you could try to use something like contrastive loss ([check out this article](https://medium.com/@maksym.bekuzarov/losses-explained-contrastive-loss-f8f57fe32246)). \nHere is a good starter notebook by @debarshichanda - [[Pytorch] HappyWhale Siamese Starter](https://www.kaggle.com/debarshichanda/pytorch-happywhale-siamese-starter) which implements siamese network.",
    "1678646": "The previous Kaggle competition had a large/similar class imbalance. You should also checkout the solutions from there."
  },
  "source": "meta"
}