{
  "id": 496552,
  "title": "Bizzarely high low CV and AUC?",
  "url": "/competitions/birdclef-2024/discussion/496552",
  "author_name": "",
  "post_date": "2024-04-21T15:02:34.059217200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This is problem is also noted in: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/493317\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/493317</a></p>\n<p>Not only do I have a very high AUC of 0.96+, but I also have a very low BCEwithlogits loss, resulting in extremely low gradients.</p>\n<p>However the competition lb have much lower AUC. I wonder has anyone found any robust relations between local CV and online lb?</p>",
  "messages": [
    {
      "id": "2766230",
      "postDate": "04/21/2024 15:02:34",
      "content": "<p>This is problem is also noted in: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/493317\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/493317</a></p>\n<p>Not only do I have a very high AUC of 0.96+, but I also have a very low BCEwithlogits loss, resulting in extremely low gradients.</p>\n<p>However the competition lb have much lower AUC. I wonder has anyone found any robust relations between local CV and online lb?</p>",
      "rawMarkdown": "This is problem is also noted in: https://www.kaggle.com/competitions/birdclef-2024/discussion/493317\n\nNot only do I have a very high AUC of 0.96+, but I also have a very low BCEwithlogits loss, resulting in extremely low gradients.\n\nHowever the competition lb have much lower AUC. I wonder has anyone found any robust relations between local CV and online lb?",
      "votes": null
    },
    {
      "id": "2766274",
      "postDate": "04/21/2024 15:36:16",
      "content": "<p>I'm having a similar issue. I scored 57th last year, and all my local testing indicates my models are quite good. As one test, I downloaded ~30 non-training recordings per species from Xeno-Canto. I trained a model for 19 epochs and tested epoch 19 and epoch 3 on my local data. Epoch 19 found 14532 segments with the right species and score &gt; .7. Epoch 3 found 9465. But when I submit them, epoch 3 scores .58 and epoch 19 scores .57. It makes no sense that epoch 19 would be lower, and when I examine the test recordings, the results all seem really good, so why is my score so low??  </p>",
      "rawMarkdown": "I'm having a similar issue. I scored 57th last year, and all my local testing indicates my models are quite good. As one test, I downloaded ~30 non-training recordings per species from Xeno-Canto. I trained a model for 19 epochs and tested epoch 19 and epoch 3 on my local data. Epoch 19 found 14532 segments with the right species and score > .7. Epoch 3 found 9465. But when I submit them, epoch 3 scores .58 and epoch 19 scores .57. It makes no sense that epoch 19 would be lower, and when I examine the test recordings, the results all seem really good, so why is my score so low??",
      "votes": null
    },
    {
      "id": "2766278",
      "postDate": "04/21/2024 15:40:10",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> See my previous comment. Is it possible for someone to review my submission and see if there is a problem with the scoring code? </p>",
      "rawMarkdown": "stefankahl See my previous comment. Is it possible for someone to review my submission and see if there is a problem with the scoring code?",
      "votes": null
    },
    {
      "id": "2766831",
      "postDate": "04/21/2024 23:15:11",
      "content": "<p>well this year a different evaluation function is used, which might affect the score. also this year the soundscape is taken in a different environment, so last years model might not 100% fit to this year </p>",
      "rawMarkdown": "well this year a different evaluation function is used, which might affect the score. also this year the soundscape is taken in a different environment, so last years model might not 100% fit to this year",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2766274,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "04/21/2024 15:36:16",
      "content": "<p>I'm having a similar issue. I scored 57th last year, and all my local testing indicates my models are quite good. As one test, I downloaded ~30 non-training recordings per species from Xeno-Canto. I trained a model for 19 epochs and tested epoch 19 and epoch 3 on my local data. Epoch 19 found 14532 segments with the right species and score &gt; .7. Epoch 3 found 9465. But when I submit them, epoch 3 scores .58 and epoch 19 scores .57. It makes no sense that epoch 19 would be lower, and when I examine the test recordings, the results all seem really good, so why is my score so low??  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2766278,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "04/21/2024 15:40:10",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> See my previous comment. Is it possible for someone to review my submission and see if there is a problem with the scoring code? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2766831,
          "author_name": "llleeeoooh",
          "author_url": "",
          "post_date": "04/21/2024 23:15:11",
          "content": "<p>well this year a different evaluation function is used, which might affect the score. also this year the soundscape is taken in a different environment, so last years model might not 100% fit to this year </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2766230": "This is problem is also noted in: https://www.kaggle.com/competitions/birdclef-2024/discussion/493317\n\nNot only do I have a very high AUC of 0.96+, but I also have a very low BCEwithlogits loss, resulting in extremely low gradients.\n\nHowever the competition lb have much lower AUC. I wonder has anyone found any robust relations between local CV and online lb?",
    "2766274": "I'm having a similar issue. I scored 57th last year, and all my local testing indicates my models are quite good. As one test, I downloaded ~30 non-training recordings per species from Xeno-Canto. I trained a model for 19 epochs and tested epoch 19 and epoch 3 on my local data. Epoch 19 found 14532 segments with the right species and score > .7. Epoch 3 found 9465. But when I submit them, epoch 3 scores .58 and epoch 19 scores .57. It makes no sense that epoch 19 would be lower, and when I examine the test recordings, the results all seem really good, so why is my score so low??",
    "2766278": "stefankahl See my previous comment. Is it possible for someone to review my submission and see if there is a problem with the scoring code?",
    "2766831": "well this year a different evaluation function is used, which might affect the score. also this year the soundscape is taken in a different environment, so last years model might not 100% fit to this year"
  },
  "source": "meta"
}