{
  "id": 178414,
  "title": "What's wrong with the LB",
  "url": "/competitions/birdsong-recognition/discussion/178414",
  "author_name": "",
  "post_date": "2020-08-29T20:21:13.812768700Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Participating in this competition I have noticed that I more and more often watch only on my CV and don't expect any informativity from the LB. </p>\n<p>Why? It's not about a potential shake-up which might cause a huge difference between public and private LB. I've already written posts about potential shake-ups in the previous competitions. You can find it <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173438\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/171963\" target=\"_blank\">here</a>. Coincidentally, those competitions had a high probability of shake-up (which by the way was later justified), so I wrote there my opinion about that. I decided not to write the same discussion here. I think it makes sense only in cases where a shake-up might exist. It's not here. An audio classification task itself in the form of kernel-competition doesn't cause shake-up. Except in the case of a big mess in a private dataset, but this is already hosts' responsibility.</p>\n<p>In this case, the uselessness of the LB consists in the impossibility of its normal interpretation. The problem is that all participants have almost identical scores. The most popular one is 0.568. The count of that is below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5007869%2Fd3d61ca21a5fcf7a64173e3cd14caefc%2FVslbwXgaguk.jpg?generation=1598731527391083&amp;alt=media\" alt=\"\"><br>\nBy the way, the \"all are nocall\" submission gives 0.544 scores on public LB.</p>\n<p>Initially, I thought that the cause of that is using the best public notebook by many participants. However, later I've changed my mind. I submitted a few solutions and get the same score which is 0,568. Meanwhile, a lot of times I rose in approximately 100 places but at the same time, my score wasn't changing at all. Usually, working on any competition I have a corresponding row in the table of train/inference results for each model which contains a score from LB. But now it doesn't make any sense just cause all scores are the same yet. In this case that public notebook is just an effect but not a cause.</p>\n<p>How to deal with that? Another metric for LB? This may solve the current problem, but I can't immediately remember a metric that would be better than the current one. More digits after the decimal point? That sounds much better.</p>",
  "messages": [
    {
      "id": "990752",
      "postDate": "08/29/2020 20:21:13",
      "content": "<p>Participating in this competition I have noticed that I more and more often watch only on my CV and don't expect any informativity from the LB. </p>\n<p>Why? It's not about a potential shake-up which might cause a huge difference between public and private LB. I've already written posts about potential shake-ups in the previous competitions. You can find it <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173438\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/171963\" target=\"_blank\">here</a>. Coincidentally, those competitions had a high probability of shake-up (which by the way was later justified), so I wrote there my opinion about that. I decided not to write the same discussion here. I think it makes sense only in cases where a shake-up might exist. It's not here. An audio classification task itself in the form of kernel-competition doesn't cause shake-up. Except in the case of a big mess in a private dataset, but this is already hosts' responsibility.</p>\n<p>In this case, the uselessness of the LB consists in the impossibility of its normal interpretation. The problem is that all participants have almost identical scores. The most popular one is 0.568. The count of that is below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5007869%2Fd3d61ca21a5fcf7a64173e3cd14caefc%2FVslbwXgaguk.jpg?generation=1598731527391083&amp;alt=media\" alt=\"\"><br>\nBy the way, the \"all are nocall\" submission gives 0.544 scores on public LB.</p>\n<p>Initially, I thought that the cause of that is using the best public notebook by many participants. However, later I've changed my mind. I submitted a few solutions and get the same score which is 0,568. Meanwhile, a lot of times I rose in approximately 100 places but at the same time, my score wasn't changing at all. Usually, working on any competition I have a corresponding row in the table of train/inference results for each model which contains a score from LB. But now it doesn't make any sense just cause all scores are the same yet. In this case that public notebook is just an effect but not a cause.</p>\n<p>How to deal with that? Another metric for LB? This may solve the current problem, but I can't immediately remember a metric that would be better than the current one. More digits after the decimal point? That sounds much better.</p>",
      "rawMarkdown": "Participating in this competition I have noticed that I more and more often watch only on my CV and don't expect any informativity from the LB. \n\nWhy? It's not about a potential shake-up which might cause a huge difference between public and private LB. I've already written posts about potential shake-ups in the previous competitions. You can find it [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173438) and [here](https://www.kaggle.com/c/global-wheat-detection/discussion/171963). Coincidentally, those competitions had a high probability of shake-up (which by the way was later justified), so I wrote there my opinion about that. I decided not to write the same discussion here. I think it makes sense only in cases where a shake-up might exist. It's not here. An audio classification task itself in the form of kernel-competition doesn't cause shake-up. Except in the case of a big mess in a private dataset, but this is already hosts' responsibility.\n\nIn this case, the uselessness of the LB consists in the impossibility of its normal interpretation. The problem is that all participants have almost identical scores. The most popular one is 0.568. The count of that is below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5007869%2Fd3d61ca21a5fcf7a64173e3cd14caefc%2FVslbwXgaguk.jpg?generation=1598731527391083&alt=media)\nBy the way, the \"all are nocall\" submission gives 0.544 scores on public LB.\n\nInitially, I thought that the cause of that is using the best public notebook by many participants. However, later I've changed my mind. I submitted a few solutions and get the same score which is 0,568. Meanwhile, a lot of times I rose in approximately 100 places but at the same time, my score wasn't changing at all. Usually, working on any competition I have a corresponding row in the table of train/inference results for each model which contains a score from LB. But now it doesn't make any sense just cause all scores are the same yet. In this case that public notebook is just an effect but not a cause.\n\nHow to deal with that? Another metric for LB? This may solve the current problem, but I can't immediately remember a metric that would be better than the current one. More digits after the decimal point? That sounds much better.",
      "votes": null
    },
    {
      "id": "991276",
      "postDate": "08/30/2020 09:30:37",
      "content": "<p>This competition is really strange..  There are many people have the same points.</p>",
      "rawMarkdown": "This competition is really strange..  There are many people have the same points.",
      "votes": null
    },
    {
      "id": "991395",
      "postDate": "08/30/2020 11:38:53",
      "content": "<p>I dont see anything strange here, people just submitted high score public kernel, same as many other competitions</p>",
      "rawMarkdown": "I dont see anything strange here, people just submitted high score public kernel, same as many other competitions",
      "votes": null
    },
    {
      "id": "991699",
      "postDate": "08/30/2020 15:48:39",
      "content": "<p>It could be due to fact that most highest scored public notebook has a score of 0.568.</p>",
      "rawMarkdown": "It could be due to fact that most highest scored public notebook has a score of 0.568.",
      "votes": null
    },
    {
      "id": "992014",
      "postDate": "08/30/2020 20:04:38",
      "content": "<p>Most of these scores are just forks of:</p>\n<p><a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\" target=\"_blank\">https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast</a></p>",
      "rawMarkdown": "Most of these scores are just forks of:\n\nhttps://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast",
      "votes": null
    },
    {
      "id": "992149",
      "postDate": "08/31/2020 01:46:03",
      "content": "<p>Then the shake up happens and the hard working ones always come out on top.</p>",
      "rawMarkdown": "Then the shake up happens and the hard working ones always come out on top.",
      "votes": null
    },
    {
      "id": "992157",
      "postDate": "08/31/2020 02:02:28",
      "content": "<p>Emmm. It happens a lot. And i have try few tricks, seems not work in this competition.May be because the huge difference between train/test data</p>",
      "rawMarkdown": "Emmm. It happens a lot. And i have try few tricks, seems not work in this competition.May be because the huge difference between train/test data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 991276,
      "author_name": "gzl0506",
      "author_url": "",
      "post_date": "08/30/2020 09:30:37",
      "content": "<p>This competition is really strange..  There are many people have the same points.</p>",
      "votes": null,
      "replies": [
        {
          "id": 991395,
          "author_name": "ajinomoto132",
          "author_url": "",
          "post_date": "08/30/2020 11:38:53",
          "content": "<p>I dont see anything strange here, people just submitted high score public kernel, same as many other competitions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 992149,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "08/31/2020 01:46:03",
          "content": "<p>Then the shake up happens and the hard working ones always come out on top.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 992157,
          "author_name": "gzl0506",
          "author_url": "",
          "post_date": "08/31/2020 02:02:28",
          "content": "<p>Emmm. It happens a lot. And i have try few tricks, seems not work in this competition.May be because the huge difference between train/test data</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 991699,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "08/30/2020 15:48:39",
      "content": "<p>It could be due to fact that most highest scored public notebook has a score of 0.568.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 992014,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/30/2020 20:04:38",
      "content": "<p>Most of these scores are just forks of:</p>\n<p><a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\" target=\"_blank\">https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "990752": "Participating in this competition I have noticed that I more and more often watch only on my CV and don't expect any informativity from the LB. \n\nWhy? It's not about a potential shake-up which might cause a huge difference between public and private LB. I've already written posts about potential shake-ups in the previous competitions. You can find it [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173438) and [here](https://www.kaggle.com/c/global-wheat-detection/discussion/171963). Coincidentally, those competitions had a high probability of shake-up (which by the way was later justified), so I wrote there my opinion about that. I decided not to write the same discussion here. I think it makes sense only in cases where a shake-up might exist. It's not here. An audio classification task itself in the form of kernel-competition doesn't cause shake-up. Except in the case of a big mess in a private dataset, but this is already hosts' responsibility.\n\nIn this case, the uselessness of the LB consists in the impossibility of its normal interpretation. The problem is that all participants have almost identical scores. The most popular one is 0.568. The count of that is below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5007869%2Fd3d61ca21a5fcf7a64173e3cd14caefc%2FVslbwXgaguk.jpg?generation=1598731527391083&alt=media)\nBy the way, the \"all are nocall\" submission gives 0.544 scores on public LB.\n\nInitially, I thought that the cause of that is using the best public notebook by many participants. However, later I've changed my mind. I submitted a few solutions and get the same score which is 0,568. Meanwhile, a lot of times I rose in approximately 100 places but at the same time, my score wasn't changing at all. Usually, working on any competition I have a corresponding row in the table of train/inference results for each model which contains a score from LB. But now it doesn't make any sense just cause all scores are the same yet. In this case that public notebook is just an effect but not a cause.\n\nHow to deal with that? Another metric for LB? This may solve the current problem, but I can't immediately remember a metric that would be better than the current one. More digits after the decimal point? That sounds much better.",
    "991276": "This competition is really strange..  There are many people have the same points.",
    "991395": "I dont see anything strange here, people just submitted high score public kernel, same as many other competitions",
    "991699": "It could be due to fact that most highest scored public notebook has a score of 0.568.",
    "992014": "Most of these scores are just forks of:\n\nhttps://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast",
    "992149": "Then the shake up happens and the hard working ones always come out on top.",
    "992157": "Emmm. It happens a lot. And i have try few tricks, seems not work in this competition.May be because the huge difference between train/test data"
  },
  "source": "meta"
}