{
  "id": 321836,
  "title": "Why train on the whole data?",
  "url": "/competitions/birdclef-2022/discussion/321836",
  "author_name": "",
  "post_date": "2022-04-29T00:10:02.675960200Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I notice many top-voted notebooks are training on the whole dataset (152 classes), e.g., <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">BirdCLEF2022 : use 2nd label f0\n</a> and <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66\" target=\"_blank\">BirdCLEF '21 - 2nd place model - Submit [0.66]</a>, but as discussed <a href=\"https://www.kaggle.com/c/birdclef-2022/discussion/307938\" target=\"_blank\">here</a>, only 21 classes will be scored. Why not just train on these 21 classes?</p>",
  "messages": [
    {
      "id": "1771140",
      "postDate": "04/29/2022 00:10:02",
      "content": "<p>I notice many top-voted notebooks are training on the whole dataset (152 classes), e.g., <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">BirdCLEF2022 : use 2nd label f0\n</a> and <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66\" target=\"_blank\">BirdCLEF '21 - 2nd place model - Submit [0.66]</a>, but as discussed <a href=\"https://www.kaggle.com/c/birdclef-2022/discussion/307938\" target=\"_blank\">here</a>, only 21 classes will be scored. Why not just train on these 21 classes?</p>",
      "rawMarkdown": "I notice many top-voted notebooks are training on the whole dataset (152 classes), e.g., [BirdCLEF2022 : use 2nd label f0\n](https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0) and [BirdCLEF '21 - 2nd place model - Submit [0.66]](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66), but as discussed [here](https://www.kaggle.com/c/birdclef-2022/discussion/307938), only 21 classes will be scored. Why not just train on these 21 classes?",
      "votes": null
    },
    {
      "id": "1771170",
      "postDate": "04/29/2022 01:04:38",
      "content": "<p>First, make no mistake, the soundscape for evaluation may contain bird sounds other than the 21 species. Our classifier needs to distinguish these recordings from scored species.</p>\n<p>Second, we need to consider that the number of data samples used for training is very small for some species. As you will see if you actually train with only 21 species (e.g., sample 21 species from last year's data and split the training data into a train/val/test = 1:4.5:4.5 ratio), training with such a small data sample will easily cause over-fitting.</p>\n<p>The third perspective is that of representation learning. In general, knowing how to distinguish species other than 21 species may help distinguish 21 species; a task that distinguishes more than 21 species is more difficult than a task that distinguishes 21 species. A model trained on a more difficult task can be expected to perform better after training because the model need to focus on more subtle features of the calls.</p>\n<p>Finally, the above are general considerations and it is not obvious that they would apply to this competition's task. The best solution is to actually experiment.</p>",
      "rawMarkdown": "First, make no mistake, the soundscape for evaluation may contain bird sounds other than the 21 species. Our classifier needs to distinguish these recordings from scored species.\n\nSecond, we need to consider that the number of data samples used for training is very small for some species. As you will see if you actually train with only 21 species (e.g., sample 21 species from last year's data and split the training data into a train/val/test = 1:4.5:4.5 ratio), training with such a small data sample will easily cause over-fitting.\n\nThe third perspective is that of representation learning. In general, knowing how to distinguish species other than 21 species may help distinguish 21 species; a task that distinguishes more than 21 species is more difficult than a task that distinguishes 21 species. A model trained on a more difficult task can be expected to perform better after training because the model need to focus on more subtle features of the calls.\n\nFinally, the above are general considerations and it is not obvious that they would apply to this competition's task. The best solution is to actually experiment.",
      "votes": null
    },
    {
      "id": "1771177",
      "postDate": "04/29/2022 01:10:38",
      "content": "<p>Of course, there are disadvantages to increasing the number of species to be identified. An example is the case where the model conflates the calls of two relatively similar bird species: if one species is evaluated and one is not, the evaluation score will decrease as the probability of predicting the species to be evaluated decreases.</p>\n<p>An idea worth trying in this case is to pre-train the model with all species, and fine-tune it with only 21 species.</p>",
      "rawMarkdown": "Of course, there are disadvantages to increasing the number of species to be identified. An example is the case where the model conflates the calls of two relatively similar bird species: if one species is evaluated and one is not, the evaluation score will decrease as the probability of predicting the species to be evaluated decreases.\n\nAn idea worth trying in this case is to pre-train the model with all species, and fine-tune it with only 21 species.",
      "votes": null
    },
    {
      "id": "1771289",
      "postDate": "04/29/2022 04:47:47",
      "content": "<p>Thanks for the reply, very helpful!</p>",
      "rawMarkdown": "Thanks for the reply, very helpful!",
      "votes": null
    },
    {
      "id": "1772222",
      "postDate": "04/30/2022 02:34:56",
      "content": "<p>I am planning on training a model on 22 classes: 21 scored classes + 1 class, which contains all other non-scored classes. I will then compare this model with the same model trained on 152 classes and see which one is better. I will let you know here how it will go!</p>",
      "rawMarkdown": "I am planning on training a model on 22 classes: 21 scored classes + 1 class, which contains all other non-scored classes. I will then compare this model with the same model trained on 152 classes and see which one is better. I will let you know here how it will go!",
      "votes": null
    },
    {
      "id": "1772275",
      "postDate": "04/30/2022 04:48:51",
      "content": "<p>Sounds great! Thanks!</p>",
      "rawMarkdown": "Sounds great! Thanks!",
      "votes": null
    },
    {
      "id": "1772861",
      "postDate": "04/30/2022 16:06:29",
      "content": "<p>Thank you for question.</p>\n<p>Actually, I did 22 classes (21 scored and else) training, and got 0.62 public score. (whole class training gave me 0.71)</p>\n<p>This is because ( I think ) what <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> said.</p>",
      "rawMarkdown": "Thank you for question.\n\nActually, I did 22 classes (21 scored and else) training, and got 0.62 public score. (whole class training gave me 0.71)\n\nThis is because ( I think ) what @tatamikenn said.",
      "votes": null
    },
    {
      "id": "1773334",
      "postDate": "05/01/2022 04:09:34",
      "content": "<p>Got it, thanks!</p>",
      "rawMarkdown": "Got it, thanks!",
      "votes": null
    },
    {
      "id": "1776565",
      "postDate": "05/04/2022 03:40:34",
      "content": "<p>I thought sound data whose primary label is in 21 classes is too small to get good weight.<br>\nThe number of whole sound data is 14852 and the number of sound data whose primary label is in 21 classes is only 1266!</p>",
      "rawMarkdown": "I thought sound data whose primary label is in 21 classes is too small to get good weight.\nThe number of whole sound data is 14852 and the number of sound data whose primary label is in 21 classes is only 1266!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1771170,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "04/29/2022 01:04:38",
      "content": "<p>First, make no mistake, the soundscape for evaluation may contain bird sounds other than the 21 species. Our classifier needs to distinguish these recordings from scored species.</p>\n<p>Second, we need to consider that the number of data samples used for training is very small for some species. As you will see if you actually train with only 21 species (e.g., sample 21 species from last year's data and split the training data into a train/val/test = 1:4.5:4.5 ratio), training with such a small data sample will easily cause over-fitting.</p>\n<p>The third perspective is that of representation learning. In general, knowing how to distinguish species other than 21 species may help distinguish 21 species; a task that distinguishes more than 21 species is more difficult than a task that distinguishes 21 species. A model trained on a more difficult task can be expected to perform better after training because the model need to focus on more subtle features of the calls.</p>\n<p>Finally, the above are general considerations and it is not obvious that they would apply to this competition's task. The best solution is to actually experiment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1771177,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "04/29/2022 01:10:38",
          "content": "<p>Of course, there are disadvantages to increasing the number of species to be identified. An example is the case where the model conflates the calls of two relatively similar bird species: if one species is evaluated and one is not, the evaluation score will decrease as the probability of predicting the species to be evaluated decreases.</p>\n<p>An idea worth trying in this case is to pre-train the model with all species, and fine-tune it with only 21 species.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771289,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "04/29/2022 04:47:47",
          "content": "<p>Thanks for the reply, very helpful!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1772222,
      "author_name": "antonzv",
      "author_url": "",
      "post_date": "04/30/2022 02:34:56",
      "content": "<p>I am planning on training a model on 22 classes: 21 scored classes + 1 class, which contains all other non-scored classes. I will then compare this model with the same model trained on 152 classes and see which one is better. I will let you know here how it will go!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1772275,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "04/30/2022 04:48:51",
          "content": "<p>Sounds great! Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1772861,
      "author_name": "kaerunantoka",
      "author_url": "",
      "post_date": "04/30/2022 16:06:29",
      "content": "<p>Thank you for question.</p>\n<p>Actually, I did 22 classes (21 scored and else) training, and got 0.62 public score. (whole class training gave me 0.71)</p>\n<p>This is because ( I think ) what <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> said.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1773334,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "05/01/2022 04:09:34",
          "content": "<p>Got it, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1776565,
      "author_name": "kotanoda",
      "author_url": "",
      "post_date": "05/04/2022 03:40:34",
      "content": "<p>I thought sound data whose primary label is in 21 classes is too small to get good weight.<br>\nThe number of whole sound data is 14852 and the number of sound data whose primary label is in 21 classes is only 1266!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1771140": "I notice many top-voted notebooks are training on the whole dataset (152 classes), e.g., [BirdCLEF2022 : use 2nd label f0\n](https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0) and [BirdCLEF '21 - 2nd place model - Submit [0.66]](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66), but as discussed [here](https://www.kaggle.com/c/birdclef-2022/discussion/307938), only 21 classes will be scored. Why not just train on these 21 classes?",
    "1771170": "First, make no mistake, the soundscape for evaluation may contain bird sounds other than the 21 species. Our classifier needs to distinguish these recordings from scored species.\n\nSecond, we need to consider that the number of data samples used for training is very small for some species. As you will see if you actually train with only 21 species (e.g., sample 21 species from last year's data and split the training data into a train/val/test = 1:4.5:4.5 ratio), training with such a small data sample will easily cause over-fitting.\n\nThe third perspective is that of representation learning. In general, knowing how to distinguish species other than 21 species may help distinguish 21 species; a task that distinguishes more than 21 species is more difficult than a task that distinguishes 21 species. A model trained on a more difficult task can be expected to perform better after training because the model need to focus on more subtle features of the calls.\n\nFinally, the above are general considerations and it is not obvious that they would apply to this competition's task. The best solution is to actually experiment.",
    "1771177": "Of course, there are disadvantages to increasing the number of species to be identified. An example is the case where the model conflates the calls of two relatively similar bird species: if one species is evaluated and one is not, the evaluation score will decrease as the probability of predicting the species to be evaluated decreases.\n\nAn idea worth trying in this case is to pre-train the model with all species, and fine-tune it with only 21 species.",
    "1771289": "Thanks for the reply, very helpful!",
    "1772222": "I am planning on training a model on 22 classes: 21 scored classes + 1 class, which contains all other non-scored classes. I will then compare this model with the same model trained on 152 classes and see which one is better. I will let you know here how it will go!",
    "1772275": "Sounds great! Thanks!",
    "1772861": "Thank you for question.\n\nActually, I did 22 classes (21 scored and else) training, and got 0.62 public score. (whole class training gave me 0.71)\n\nThis is because ( I think ) what @tatamikenn said.",
    "1773334": "Got it, thanks!",
    "1776565": "I thought sound data whose primary label is in 21 classes is too small to get good weight.\nThe number of whole sound data is 14852 and the number of sound data whose primary label is in 21 classes is only 1266!"
  },
  "source": "meta"
}