{
  "id": 69753,
  "title": "LB scores",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/69753",
  "author_name": "Yee Ng",
  "post_date": "2018-10-26T19:17:24.464000",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>If there are 3000 images on stage 2 test set, and the LB is based on only 1% of the test set, doesn't that mean that its only based on 30 images? The sample size seems to be too small to make any useful evaluation...</p>",
  "messages": [
    {
      "id": 410861,
      "postDate": "2018-10-26T19:17:24.463Z",
      "content": "<p>If there are 3000 images on stage 2 test set, and the LB is based on only 1% of the test set, doesn't that mean that its only based on 30 images? The sample size seems to be too small to make any useful evaluation...</p>",
      "rawMarkdown": "If there are 3000 images on stage 2 test set, and the LB is based on only 1% of the test set, doesn't that mean that its only based on 30 images? The sample size seems to be too small to make any useful evaluation...",
      "votes": 3
    },
    {
      "id": 411265,
      "postDate": "2018-10-27T18:21:34.647Z",
      "content": "<p>Current public LB is too far from our validation score. Anybody who tries to overfit this LB, may end up terribly during final scoring.</p>",
      "rawMarkdown": "Current public LB is too far from our validation score. Anybody who tries to overfit this LB, may end up terribly during final scoring.",
      "votes": 1
    },
    {
      "id": 411116,
      "postDate": "2018-10-27T12:59:37.927Z",
      "content": "<p>Julia addresses this question here: <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/69745#410882\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/69745#410882</a></p>\n\n<p>The point of the public LB is to basically make sure your submission works. Otherwise, it's not intended to give you any sense of how well your model will perform on  the stage 2 test data. </p>",
      "rawMarkdown": "Julia addresses this question here: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/69745#410882\n\nThe point of the public LB is to basically make sure your submission works. Otherwise, it's not intended to give you any sense of how well your model will perform on  the stage 2 test data. ",
      "votes": 2
    },
    {
      "id": 411023,
      "postDate": "2018-10-27T07:56:13.603Z",
      "content": "<p>I did an experiment. I took stage 1 test data and divided it in groups by 30 images. Having GT labels, we can score these groups. I got scores ranging from 0.05 to 0.45, so 30 images really don't give too much.</p>",
      "rawMarkdown": "I did an experiment. I took stage 1 test data and divided it in groups by 30 images. Having GT labels, we can score these groups. I got scores ranging from 0.05 to 0.45, so 30 images really don't give too much.",
      "votes": 2,
      "replies": [
        {
          "id": 411033,
          "postDate": "2018-10-27T08:36:09.727Z",
          "content": "<p>No. Trust your validation scores</p>",
          "rawMarkdown": "No. Trust your validation scores"
        },
        {
          "id": 411052,
          "postDate": "2018-10-27T09:48:24.883Z",
          "content": "<p>The group size of 150 images would be much better. Then scores are \"only\" from 0.16 to 0.24.</p>",
          "rawMarkdown": "The group size of 150 images would be much better. Then scores are \"only\" from 0.16 to 0.24."
        }
      ]
    },
    {
      "id": 411333,
      "postDate": "2018-10-27T21:43:06.773Z",
      "content": "<p>@Yee Seng Ng, you are right the current LB is useful only to check that our submissions are accepted.</p>\n\n<p>I used it to check a few results that finished training after the close of Stage-1. I submitted those to see where on the stage-1 LB they stood before it was wiped out. So I submitted them to the current LB but obviously, I would not choose any of them as a final submission.</p>",
      "rawMarkdown": "@Yee Seng Ng, you are right the current LB is useful only to check that our submissions are accepted.\n\nI used it to check a few results that finished training after the close of Stage-1. I submitted those to see where on the stage-1 LB they stood before it was wiped out. So I submitted them to the current LB but obviously, I would not choose any of them as a final submission."
    },
    {
      "id": 411264,
      "postDate": "2018-10-27T18:19:19.247Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 411117,
      "postDate": "2018-10-27T12:59:51.440Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 411265,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2018-10-27T18:21:34.647000",
      "content": "<p>Current public LB is too far from our validation score. Anybody who tries to overfit this LB, may end up terribly during final scoring.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 411116,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2018-10-27T12:59:37.927000",
      "content": "<p>Julia addresses this question here: <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/69745#410882\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/69745#410882</a></p>\n\n<p>The point of the public LB is to basically make sure your submission works. Otherwise, it's not intended to give you any sense of how well your model will perform on  the stage 2 test data. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 411023,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-10-27T07:56:13.603000",
      "content": "<p>I did an experiment. I took stage 1 test data and divided it in groups by 30 images. Having GT labels, we can score these groups. I got scores ranging from 0.05 to 0.45, so 30 images really don't give too much.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 411033,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2018-10-27T08:36:09.727000",
          "content": "<p>No. Trust your validation scores</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 411052,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-10-27T09:48:24.883000",
          "content": "<p>The group size of 150 images would be much better. Then scores are \"only\" from 0.16 to 0.24.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 411333,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-10-27T21:43:06.773000",
      "content": "<p>@Yee Seng Ng, you are right the current LB is useful only to check that our submissions are accepted.</p>\n\n<p>I used it to check a few results that finished training after the close of Stage-1. I submitted those to see where on the stage-1 LB they stood before it was wiped out. So I submitted them to the current LB but obviously, I would not choose any of them as a final submission.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 411264,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-27T18:19:19.247000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 411117,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-27T12:59:51.440000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "410861": "If there are 3000 images on stage 2 test set, and the LB is based on only 1% of the test set, doesn't that mean that its only based on 30 images? The sample size seems to be too small to make any useful evaluation...",
    "411265": "Current public LB is too far from our validation score. Anybody who tries to overfit this LB, may end up terribly during final scoring.",
    "411116": "Julia addresses this question here: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/69745#410882\n\nThe point of the public LB is to basically make sure your submission works. Otherwise, it's not intended to give you any sense of how well your model will perform on  the stage 2 test data. ",
    "411023": "I did an experiment. I took stage 1 test data and divided it in groups by 30 images. Having GT labels, we can score these groups. I got scores ranging from 0.05 to 0.45, so 30 images really don't give too much.",
    "411333": "@Yee Seng Ng, you are right the current LB is useful only to check that our submissions are accepted.\n\nI used it to check a few results that finished training after the close of Stage-1. I submitted those to see where on the stage-1 LB they stood before it was wiped out. So I submitted them to the current LB but obviously, I would not choose any of them as a final submission.",
    "411264": "",
    "411117": ""
  }
}