{
  "id": 268426,
  "title": "LB Probing strategy, Good or Bad and Why?",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/268426",
  "author_name": "",
  "post_date": "2021-08-27T11:07:06.318100100Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am relatively new to Kaggle. Today I saw a LB 1.0 in this competition. After some quick researches I found that, this is done via a method called LB probing, and searching this on other competitions, it is turns out that it can be a strategy to gain higher ranks in competitions.  </p>\n<p>As far as I know it is about achieving information about test sets via several submissions. I am just wondering that is it a good idea or is it even worth to know about the test data of the competition?  If it is a good idea, so why Kaggle does not publish this info? If it is not why this is allowed to do?</p>\n<p>Or may be this question: Is knowing about test data, and redirecting your model to get best results on that, has any value for the community? </p>\n<p>Thanks in advance.</p>",
  "messages": [
    {
      "id": "1492662",
      "postDate": "08/27/2021 11:07:06",
      "content": "<p>I am relatively new to Kaggle. Today I saw a LB 1.0 in this competition. After some quick researches I found that, this is done via a method called LB probing, and searching this on other competitions, it is turns out that it can be a strategy to gain higher ranks in competitions.  </p>\n<p>As far as I know it is about achieving information about test sets via several submissions. I am just wondering that is it a good idea or is it even worth to know about the test data of the competition?  If it is a good idea, so why Kaggle does not publish this info? If it is not why this is allowed to do?</p>\n<p>Or may be this question: Is knowing about test data, and redirecting your model to get best results on that, has any value for the community? </p>\n<p>Thanks in advance.</p>",
      "rawMarkdown": "I am relatively new to Kaggle. Today I saw a LB 1.0 in this competition. After some quick researches I found that, this is done via a method called LB probing, and searching this on other competitions, it is turns out that it can be a strategy to gain higher ranks in competitions.  \n\nAs far as I know it is about achieving information about test sets via several submissions. I am just wondering that is it a good idea or is it even worth to know about the test data of the competition?  If it is a good idea, so why Kaggle does not publish this info? If it is not why this is allowed to do?\n\nOr may be this question: Is knowing about test data, and redirecting your model to get best results on that, has any value for the community? \n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "1493982",
      "postDate": "08/28/2021 10:09:16",
      "content": "<p>Hi , I saw one discussion about LB 1 but I couldn't understand how we can do that. you said it's a LB probing method.<br>\nplease share any resource about this method and how we can do this. thanks </p>",
      "rawMarkdown": "Hi , I saw one discussion about LB 1 but I couldn't understand how we can do that. you said it's a LB probing method.\nplease share any resource about this method and how we can do this. thanks",
      "votes": null
    },
    {
      "id": "1494021",
      "postDate": "08/28/2021 10:36:31",
      "content": "<p><a href=\"https://towardsdatascience.com/how-to-lb-probe-on-kaggle-c0aa21458bfe\" target=\"_blank\">This article</a> describes how to apply LB Probing on Kaggle. </p>",
      "rawMarkdown": "[This article](https://towardsdatascience.com/how-to-lb-probe-on-kaggle-c0aa21458bfe) describes how to apply LB Probing on Kaggle.",
      "votes": null
    },
    {
      "id": "1494032",
      "postDate": "08/28/2021 10:51:58",
      "content": "<p>Do you know any notebook about this on kaggle?</p>",
      "rawMarkdown": "Do you know any notebook about this on kaggle?",
      "votes": null
    },
    {
      "id": "1494169",
      "postDate": "08/28/2021 13:20:25",
      "content": "<p>Yes. Search on google \"Lb probing on Kaggle\", and you will find a lot of notebooks about this. </p>",
      "rawMarkdown": "Yes. Search on google \"Lb probing on Kaggle\", and you will find a lot of notebooks about this.",
      "votes": null
    },
    {
      "id": "1494245",
      "postDate": "08/28/2021 14:20:37",
      "content": "<p>Ok.Thanks.</p>",
      "rawMarkdown": "Ok.Thanks.",
      "votes": null
    },
    {
      "id": "1494439",
      "postDate": "08/28/2021 17:21:49",
      "content": "<ol>\n<li>Yes, you <strong>can</strong> use leaderboard probing and hand labelling in the submission csv. However, It is against the rules of the competition so even if you score high  you won't be able to receive any prize</li>\n<li>Only 22% of the data is available so even if you use LB probing you won't be able to predict 88% of the data so there really is no point.</li>\n<li>It's almost impossible to acquire an LB of 1 with AUC-ROC.</li>\n</ol>\n<p>I guess since the test set is so small <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> used an LB probing method to get a more accurate prediction of how his model is performing compared to the leaderboard. The only problem is that he has confused a lot of participants😅</p>",
      "rawMarkdown": "1. Yes, you **can** use leaderboard probing and hand labelling in the submission csv. However, It is against the rules of the competition so even if you score high  you won't be able to receive any prize\n2. Only 22% of the data is available so even if you use LB probing you won't be able to predict 88% of the data so there really is no point.\n3. It's almost impossible to acquire an LB of 1 with AUC-ROC.\n\nI guess since the test set is so small @onodera used an LB probing method to get a more accurate prediction of how his model is performing compared to the leaderboard. The only problem is that he has confused a lot of participants😅",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1493982,
      "author_name": "mohammadhosein1998",
      "author_url": "",
      "post_date": "08/28/2021 10:09:16",
      "content": "<p>Hi , I saw one discussion about LB 1 but I couldn't understand how we can do that. you said it's a LB probing method.<br>\nplease share any resource about this method and how we can do this. thanks </p>",
      "votes": null,
      "replies": [
        {
          "id": 1494021,
          "author_name": "kavehshahhosseini",
          "author_url": "",
          "post_date": "08/28/2021 10:36:31",
          "content": "<p><a href=\"https://towardsdatascience.com/how-to-lb-probe-on-kaggle-c0aa21458bfe\" target=\"_blank\">This article</a> describes how to apply LB Probing on Kaggle. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494032,
          "author_name": "mohammadhosein1998",
          "author_url": "",
          "post_date": "08/28/2021 10:51:58",
          "content": "<p>Do you know any notebook about this on kaggle?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494169,
          "author_name": "kavehshahhosseini",
          "author_url": "",
          "post_date": "08/28/2021 13:20:25",
          "content": "<p>Yes. Search on google \"Lb probing on Kaggle\", and you will find a lot of notebooks about this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494245,
          "author_name": "mohammadhosein1998",
          "author_url": "",
          "post_date": "08/28/2021 14:20:37",
          "content": "<p>Ok.Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1494439,
      "author_name": "aristotle609",
      "author_url": "",
      "post_date": "08/28/2021 17:21:49",
      "content": "<ol>\n<li>Yes, you <strong>can</strong> use leaderboard probing and hand labelling in the submission csv. However, It is against the rules of the competition so even if you score high  you won't be able to receive any prize</li>\n<li>Only 22% of the data is available so even if you use LB probing you won't be able to predict 88% of the data so there really is no point.</li>\n<li>It's almost impossible to acquire an LB of 1 with AUC-ROC.</li>\n</ol>\n<p>I guess since the test set is so small <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> used an LB probing method to get a more accurate prediction of how his model is performing compared to the leaderboard. The only problem is that he has confused a lot of participants😅</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1492662": "I am relatively new to Kaggle. Today I saw a LB 1.0 in this competition. After some quick researches I found that, this is done via a method called LB probing, and searching this on other competitions, it is turns out that it can be a strategy to gain higher ranks in competitions.  \n\nAs far as I know it is about achieving information about test sets via several submissions. I am just wondering that is it a good idea or is it even worth to know about the test data of the competition?  If it is a good idea, so why Kaggle does not publish this info? If it is not why this is allowed to do?\n\nOr may be this question: Is knowing about test data, and redirecting your model to get best results on that, has any value for the community? \n\nThanks in advance.",
    "1493982": "Hi , I saw one discussion about LB 1 but I couldn't understand how we can do that. you said it's a LB probing method.\nplease share any resource about this method and how we can do this. thanks",
    "1494021": "[This article](https://towardsdatascience.com/how-to-lb-probe-on-kaggle-c0aa21458bfe) describes how to apply LB Probing on Kaggle.",
    "1494032": "Do you know any notebook about this on kaggle?",
    "1494169": "Yes. Search on google \"Lb probing on Kaggle\", and you will find a lot of notebooks about this.",
    "1494245": "Ok.Thanks.",
    "1494439": "1. Yes, you **can** use leaderboard probing and hand labelling in the submission csv. However, It is against the rules of the competition so even if you score high  you won't be able to receive any prize\n2. Only 22% of the data is available so even if you use LB probing you won't be able to predict 88% of the data so there really is no point.\n3. It's almost impossible to acquire an LB of 1 with AUC-ROC.\n\nI guess since the test set is so small @onodera used an LB probing method to get a more accurate prediction of how his model is performing compared to the leaderboard. The only problem is that he has confused a lot of participants😅"
  },
  "source": "meta"
}