{
  "id": 409200,
  "title": "Increasing the competition metric does not reflect on LB (?)",
  "url": "/competitions/birdclef-2023/discussion/409200",
  "author_name": "",
  "post_date": "2023-05-10T01:49:49.119411600Z",
  "votes": 3,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I probably have something missing here and I could not figure out what. Even though I increased the competition metric on my local validation, I can't see the same improvements on LB. Even though I increased the competition metric on my local validation, I can't see the same improvements on LB. I got stuck on .78 for a single fold model. I got different models with .78, .81, .84 p-cmap5 scores respectively but they all have .78 on the LB. I track the f1-score as well and can see the same improvements.</p>\n<p>I am curious to hear if anyone else has the same problem and I am open to suggestions for further investigation.</p>\n<p>My best guess is this model does not see some of the classes in the training and fails to predict during the test time. But still the improvements for the other classes should reflect somehow on the LB as well IMHO.</p>",
  "messages": [
    {
      "id": "2253178",
      "postDate": "05/10/2023 01:49:49",
      "content": "<p>Hi everyone,</p>\n<p>I probably have something missing here and I could not figure out what. Even though I increased the competition metric on my local validation, I can't see the same improvements on LB. Even though I increased the competition metric on my local validation, I can't see the same improvements on LB. I got stuck on .78 for a single fold model. I got different models with .78, .81, .84 p-cmap5 scores respectively but they all have .78 on the LB. I track the f1-score as well and can see the same improvements.</p>\n<p>I am curious to hear if anyone else has the same problem and I am open to suggestions for further investigation.</p>\n<p>My best guess is this model does not see some of the classes in the training and fails to predict during the test time. But still the improvements for the other classes should reflect somehow on the LB as well IMHO.</p>",
      "rawMarkdown": "Hi everyone,\n\nI probably have something missing here and I could not figure out what. Even though I increased the competition metric on my local validation, I can't see the same improvements on LB. Even though I increased the competition metric on my local validation, I can't see the same improvements on LB. I got stuck on .78 for a single fold model. I got different models with .78, .81, .84 p-cmap5 scores respectively but they all have .78 on the LB. I track the f1-score as well and can see the same improvements.\n\nI am curious to hear if anyone else has the same problem and I am open to suggestions for further investigation.\n\nMy best guess is this model does not see some of the classes in the training and fails to predict during the test time. But still the improvements for the other classes should reflect somehow on the LB as well IMHO.",
      "votes": null
    },
    {
      "id": "2253181",
      "postDate": "05/10/2023 01:55:40",
      "content": "<p>Yes same problem here, stuck on 0.77 on LB, but not for long !<br>\nI think the key to this comp is a reflective CV, a simple 80-20 split doesn't really translate the LB, maybe only taking birds that were detected around Kenya for validation helps </p>",
      "rawMarkdown": "Yes same problem here, stuck on 0.77 on LB, but not for long !\nI think the key to this comp is a reflective CV, a simple 80-20 split doesn't really translate the LB, maybe only taking birds that were detected around Kenya for validation helps",
      "votes": null
    },
    {
      "id": "2253186",
      "postDate": "05/10/2023 02:12:13",
      "content": "<p>Yes, indeed. Apparently, I am missing that. thanks for the suggestion!</p>",
      "rawMarkdown": "Yes, indeed. Apparently, I am missing that. thanks for the suggestion!",
      "votes": null
    },
    {
      "id": "2253490",
      "postDate": "05/10/2023 08:28:15",
      "content": "<p>hi <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a>, Is stratifying using location better???</p>",
      "rawMarkdown": "hi @janmpia, Is stratifying using location better???",
      "votes": null
    },
    {
      "id": "2253605",
      "postDate": "05/10/2023 10:24:58",
      "content": "<p>I too have the same problem, single model score is 0.78 but using previous year data for pretraining doesn't improves the score.</p>",
      "rawMarkdown": "I too have the same problem, single model score is 0.78 but using previous year data for pretraining doesn't improves the score.",
      "votes": null
    },
    {
      "id": "2253662",
      "postDate": "05/10/2023 11:36:20",
      "content": "<p>hi <a href=\"https://www.kaggle.com/gowrishankarp\" target=\"_blank\">@gowrishankarp</a>, I haven't test honnestly I just suggested trying this approach, I will later tho !</p>",
      "rawMarkdown": "hi @gowrishankarp, I haven't test honnestly I just suggested trying this approach, I will later tho !",
      "votes": null
    },
    {
      "id": "2255724",
      "postDate": "05/12/2023 00:23:46",
      "content": "<p>Okay, I managed to reach .79 with a single model with some additional changes. Still not performing as strong as it can be imo, the changes I made now reflected to LB as well for the first time. </p>",
      "rawMarkdown": "Okay, I managed to reach .79 with a single model with some additional changes. Still not performing as strong as it can be imo, the changes I made now reflected to LB as well for the first time.",
      "votes": null
    },
    {
      "id": "2255749",
      "postDate": "05/12/2023 01:11:15",
      "content": "<p>I'm not getting as good of a score as you are, but I managed to find a way for my CV to be more realistic, and part of it was due to the fact that I was only training using CrossEntropy and taking only the primary label, I don't know if that can help but that way I'm more exited with improvements in CV for sure !<br>\n<a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> </p>",
      "rawMarkdown": "I'm not getting as good of a score as you are, but I managed to find a way for my CV to be more realistic, and part of it was due to the fact that I was only training using CrossEntropy and taking only the primary label, I don't know if that can help but that way I'm more exited with improvements in CV for sure !\n@snnclsr",
      "votes": null
    },
    {
      "id": "2255751",
      "postDate": "05/12/2023 01:12:45",
      "content": "<p>It didn't help at all for me <a href=\"https://www.kaggle.com/gowrishankarp\" target=\"_blank\">@gowrishankarp</a> </p>",
      "rawMarkdown": "It didn't help at all for me @gowrishankarp",
      "votes": null
    },
    {
      "id": "2260820",
      "postDate": "05/15/2023 22:52:03",
      "content": "<p>Absolutely, this problem is still eluding me after plugging away at it for 2.5 months!  I think my training routine is really strong, I have local cmap scores around 0.93, including loads of data augmentation but still getting as low as 0.76 on the LB, which is not much better than just submitting all 0's!   Using CE loss is working a bit better for me generally than BCE, but this doesn't seem to be the magic bullet.</p>\n<p>I'm running out of ideas at this point.  I suspect one of the following:</p>\n<ul>\n<li>A bug I can't spot, or at least a subtle difference in pre-processing on evaluation</li>\n<li>Overtraining, because the loss metric isn't very suitable</li>\n<li>Local variations in dialect</li>\n<li>The test set has a very different distribution of bird types to the training data</li>\n</ul>\n<p>I'll stop working on elaborate plans for now, and focus on the above if I've got time.</p>\n<p>I'm learning a lot of hard lessons in this comp, that I'm going to apply in real world conservation work. But it's still a bit brutal watching my LB placing slipping down the ranks 😆</p>",
      "rawMarkdown": "Absolutely, this problem is still eluding me after plugging away at it for 2.5 months!  I think my training routine is really strong, I have local cmap scores around 0.93, including loads of data augmentation but still getting as low as 0.76 on the LB, which is not much better than just submitting all 0's!   Using CE loss is working a bit better for me generally than BCE, but this doesn't seem to be the magic bullet.\n\nI'm running out of ideas at this point.  I suspect one of the following:\n- A bug I can't spot, or at least a subtle difference in pre-processing on evaluation\n- Overtraining, because the loss metric isn't very suitable\n- Local variations in dialect\n- The test set has a very different distribution of bird types to the training data\n\nI'll stop working on elaborate plans for now, and focus on the above if I've got time.\n\nI'm learning a lot of hard lessons in this comp, that I'm going to apply in real world conservation work. But it's still a bit brutal watching my LB placing slipping down the ranks 😆",
      "votes": null
    },
    {
      "id": "2260833",
      "postDate": "05/15/2023 23:25:55",
      "content": "<p>I mean in your case the gap is really huge, tbh 😅 I would, by default, assume there is something wrong with the evaluation scheme or in the submission notebook. </p>\n<p>For myself, I stopped checking the metric and am mainly focusing on the loss value (BCE). So far it reflects better than the competition metric. My current best with a single model is .8 and with a 5-fold to reach my current LB score. </p>",
      "rawMarkdown": "I mean in your case the gap is really huge, tbh 😅 I would, by default, assume there is something wrong with the evaluation scheme or in the submission notebook. \n\nFor myself, I stopped checking the metric and am mainly focusing on the loss value (BCE). So far it reflects better than the competition metric. My current best with a single model is .8 and with a 5-fold to reach my current LB score.",
      "votes": null
    },
    {
      "id": "2262546",
      "postDate": "05/17/2023 02:23:03",
      "content": "<p>Thanks Sinan,  interesting that you're not trusting your competition metric either.  Mine is so high it just seems unbelievable, I get to 80+ after just a couple of training epochs, both LRAP &amp; CMAP, and yet the loss continues to improve for up to 20.  I'm inclined to think the route cause of my problems lies in this difference.  On the other hand, weights from epoch 2, 7 &amp; 18 all give much the same score on the LB, so there is nothing to discern if I'm under-training or over-training.</p>",
      "rawMarkdown": "Thanks Sinan,  interesting that you're not trusting your competition metric either.  Mine is so high it just seems unbelievable, I get to 80+ after just a couple of training epochs, both LRAP & CMAP, and yet the loss continues to improve for up to 20.  I'm inclined to think the route cause of my problems lies in this difference.  On the other hand, weights from epoch 2, 7 & 18 all give much the same score on the LB, so there is nothing to discern if I'm under-training or over-training.",
      "votes": null
    },
    {
      "id": "2263339",
      "postDate": "05/17/2023 14:10:37",
      "content": "<p>In my last experiments, f1 seems more reliable than the competition metric. I'm planning to stick to that until the end of the competition together with the loss value.</p>",
      "rawMarkdown": "In my last experiments, f1 seems more reliable than the competition metric. I'm planning to stick to that until the end of the competition together with the loss value.",
      "votes": null
    },
    {
      "id": "2263873",
      "postDate": "05/18/2023 01:49:45",
      "content": "<p>Thanks, I'll give it a try!</p>",
      "rawMarkdown": "Thanks, I'll give it a try!",
      "votes": null
    },
    {
      "id": "2264583",
      "postDate": "05/18/2023 14:53:14",
      "content": "<p>I'll try f1 too, thx for sharing such information <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> </p>",
      "rawMarkdown": "I'll try f1 too, thx for sharing such information @snnclsr",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2253181,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "05/10/2023 01:55:40",
      "content": "<p>Yes same problem here, stuck on 0.77 on LB, but not for long !<br>\nI think the key to this comp is a reflective CV, a simple 80-20 split doesn't really translate the LB, maybe only taking birds that were detected around Kenya for validation helps </p>",
      "votes": null,
      "replies": [
        {
          "id": 2253186,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "05/10/2023 02:12:13",
          "content": "<p>Yes, indeed. Apparently, I am missing that. thanks for the suggestion!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2253490,
          "author_name": "gowrishankarp",
          "author_url": "",
          "post_date": "05/10/2023 08:28:15",
          "content": "<p>hi <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a>, Is stratifying using location better???</p>",
          "votes": null,
          "replies": [
            {
              "id": 2253662,
              "author_name": "janmpia",
              "author_url": "",
              "post_date": "05/10/2023 11:36:20",
              "content": "<p>hi <a href=\"https://www.kaggle.com/gowrishankarp\" target=\"_blank\">@gowrishankarp</a>, I haven't test honnestly I just suggested trying this approach, I will later tho !</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2255751,
              "author_name": "janmpia",
              "author_url": "",
              "post_date": "05/12/2023 01:12:45",
              "content": "<p>It didn't help at all for me <a href=\"https://www.kaggle.com/gowrishankarp\" target=\"_blank\">@gowrishankarp</a> </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2253605,
      "author_name": "himanshunayal",
      "author_url": "",
      "post_date": "05/10/2023 10:24:58",
      "content": "<p>I too have the same problem, single model score is 0.78 but using previous year data for pretraining doesn't improves the score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2255724,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "05/12/2023 00:23:46",
      "content": "<p>Okay, I managed to reach .79 with a single model with some additional changes. Still not performing as strong as it can be imo, the changes I made now reflected to LB as well for the first time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2255749,
          "author_name": "janmpia",
          "author_url": "",
          "post_date": "05/12/2023 01:11:15",
          "content": "<p>I'm not getting as good of a score as you are, but I managed to find a way for my CV to be more realistic, and part of it was due to the fact that I was only training using CrossEntropy and taking only the primary label, I don't know if that can help but that way I'm more exited with improvements in CV for sure !<br>\n<a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2260820,
      "author_name": "ollypowell",
      "author_url": "",
      "post_date": "05/15/2023 22:52:03",
      "content": "<p>Absolutely, this problem is still eluding me after plugging away at it for 2.5 months!  I think my training routine is really strong, I have local cmap scores around 0.93, including loads of data augmentation but still getting as low as 0.76 on the LB, which is not much better than just submitting all 0's!   Using CE loss is working a bit better for me generally than BCE, but this doesn't seem to be the magic bullet.</p>\n<p>I'm running out of ideas at this point.  I suspect one of the following:</p>\n<ul>\n<li>A bug I can't spot, or at least a subtle difference in pre-processing on evaluation</li>\n<li>Overtraining, because the loss metric isn't very suitable</li>\n<li>Local variations in dialect</li>\n<li>The test set has a very different distribution of bird types to the training data</li>\n</ul>\n<p>I'll stop working on elaborate plans for now, and focus on the above if I've got time.</p>\n<p>I'm learning a lot of hard lessons in this comp, that I'm going to apply in real world conservation work. But it's still a bit brutal watching my LB placing slipping down the ranks 😆</p>",
      "votes": null,
      "replies": [
        {
          "id": 2260833,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "05/15/2023 23:25:55",
          "content": "<p>I mean in your case the gap is really huge, tbh 😅 I would, by default, assume there is something wrong with the evaluation scheme or in the submission notebook. </p>\n<p>For myself, I stopped checking the metric and am mainly focusing on the loss value (BCE). So far it reflects better than the competition metric. My current best with a single model is .8 and with a 5-fold to reach my current LB score. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2262546,
              "author_name": "ollypowell",
              "author_url": "",
              "post_date": "05/17/2023 02:23:03",
              "content": "<p>Thanks Sinan,  interesting that you're not trusting your competition metric either.  Mine is so high it just seems unbelievable, I get to 80+ after just a couple of training epochs, both LRAP &amp; CMAP, and yet the loss continues to improve for up to 20.  I'm inclined to think the route cause of my problems lies in this difference.  On the other hand, weights from epoch 2, 7 &amp; 18 all give much the same score on the LB, so there is nothing to discern if I'm under-training or over-training.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2263339,
                  "author_name": "snnclsr",
                  "author_url": "",
                  "post_date": "05/17/2023 14:10:37",
                  "content": "<p>In my last experiments, f1 seems more reliable than the competition metric. I'm planning to stick to that until the end of the competition together with the loss value.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2263873,
                      "author_name": "ollypowell",
                      "author_url": "",
                      "post_date": "05/18/2023 01:49:45",
                      "content": "<p>Thanks, I'll give it a try!</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2264583,
                      "author_name": "janmpia",
                      "author_url": "",
                      "post_date": "05/18/2023 14:53:14",
                      "content": "<p>I'll try f1 too, thx for sharing such information <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2253178": "Hi everyone,\n\nI probably have something missing here and I could not figure out what. Even though I increased the competition metric on my local validation, I can't see the same improvements on LB. Even though I increased the competition metric on my local validation, I can't see the same improvements on LB. I got stuck on .78 for a single fold model. I got different models with .78, .81, .84 p-cmap5 scores respectively but they all have .78 on the LB. I track the f1-score as well and can see the same improvements.\n\nI am curious to hear if anyone else has the same problem and I am open to suggestions for further investigation.\n\nMy best guess is this model does not see some of the classes in the training and fails to predict during the test time. But still the improvements for the other classes should reflect somehow on the LB as well IMHO.",
    "2253181": "Yes same problem here, stuck on 0.77 on LB, but not for long !\nI think the key to this comp is a reflective CV, a simple 80-20 split doesn't really translate the LB, maybe only taking birds that were detected around Kenya for validation helps",
    "2253186": "Yes, indeed. Apparently, I am missing that. thanks for the suggestion!",
    "2253490": "hi @janmpia, Is stratifying using location better???",
    "2253605": "I too have the same problem, single model score is 0.78 but using previous year data for pretraining doesn't improves the score.",
    "2253662": "hi @gowrishankarp, I haven't test honnestly I just suggested trying this approach, I will later tho !",
    "2255724": "Okay, I managed to reach .79 with a single model with some additional changes. Still not performing as strong as it can be imo, the changes I made now reflected to LB as well for the first time.",
    "2255749": "I'm not getting as good of a score as you are, but I managed to find a way for my CV to be more realistic, and part of it was due to the fact that I was only training using CrossEntropy and taking only the primary label, I don't know if that can help but that way I'm more exited with improvements in CV for sure !\n@snnclsr",
    "2255751": "It didn't help at all for me @gowrishankarp",
    "2260820": "Absolutely, this problem is still eluding me after plugging away at it for 2.5 months!  I think my training routine is really strong, I have local cmap scores around 0.93, including loads of data augmentation but still getting as low as 0.76 on the LB, which is not much better than just submitting all 0's!   Using CE loss is working a bit better for me generally than BCE, but this doesn't seem to be the magic bullet.\n\nI'm running out of ideas at this point.  I suspect one of the following:\n- A bug I can't spot, or at least a subtle difference in pre-processing on evaluation\n- Overtraining, because the loss metric isn't very suitable\n- Local variations in dialect\n- The test set has a very different distribution of bird types to the training data\n\nI'll stop working on elaborate plans for now, and focus on the above if I've got time.\n\nI'm learning a lot of hard lessons in this comp, that I'm going to apply in real world conservation work. But it's still a bit brutal watching my LB placing slipping down the ranks 😆",
    "2260833": "I mean in your case the gap is really huge, tbh 😅 I would, by default, assume there is something wrong with the evaluation scheme or in the submission notebook. \n\nFor myself, I stopped checking the metric and am mainly focusing on the loss value (BCE). So far it reflects better than the competition metric. My current best with a single model is .8 and with a 5-fold to reach my current LB score.",
    "2262546": "Thanks Sinan,  interesting that you're not trusting your competition metric either.  Mine is so high it just seems unbelievable, I get to 80+ after just a couple of training epochs, both LRAP & CMAP, and yet the loss continues to improve for up to 20.  I'm inclined to think the route cause of my problems lies in this difference.  On the other hand, weights from epoch 2, 7 & 18 all give much the same score on the LB, so there is nothing to discern if I'm under-training or over-training.",
    "2263339": "In my last experiments, f1 seems more reliable than the competition metric. I'm planning to stick to that until the end of the competition together with the loss value.",
    "2263873": "Thanks, I'll give it a try!",
    "2264583": "I'll try f1 too, thx for sharing such information @snnclsr"
  },
  "source": "meta"
}