{
  "id": 233290,
  "title": "Discussion about datasets (train/test/public/private)",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/233290",
  "author_name": "",
  "post_date": "2021-04-18T12:16:46.197771900Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>So what I know are</p>\n<ol>\n<li>There is a training set of 15 TIFFs</li>\n<li>There is a folder entitled 'test' on the submission notebooks and in the downloaded data containing 5 TIFFs</li>\n<li>There is \"public test set\" used for the public leaderboard</li>\n<li>There is a \"private test set\" used for the private leaderboard</li>\n</ol>\n<p>My understanding is our code gets run on (3) and (4) whenever we submit, and we are at present shown our scores for (3), and then (4) will be immediately revealed when the competition ends.</p>\n<p>What I DON'T know is are (2) and (3) the same dataset? IE is the public leader board the performance on the 5 TIFFs in the 'test' folder.</p>\n<p>The reason this is very important is it means you cannot use your public leader board score as a realistic measure of performance if you are either handlabelling or pseudolabelling them, because they're no longer a test set.</p>\n<p>Is my understanding above correct? Do you think the public leaderboard performance is the 5 test cases we already have access to?</p>",
  "messages": [
    {
      "id": "1277115",
      "postDate": "04/18/2021 12:16:46",
      "content": "<p>So what I know are</p>\n<ol>\n<li>There is a training set of 15 TIFFs</li>\n<li>There is a folder entitled 'test' on the submission notebooks and in the downloaded data containing 5 TIFFs</li>\n<li>There is \"public test set\" used for the public leaderboard</li>\n<li>There is a \"private test set\" used for the private leaderboard</li>\n</ol>\n<p>My understanding is our code gets run on (3) and (4) whenever we submit, and we are at present shown our scores for (3), and then (4) will be immediately revealed when the competition ends.</p>\n<p>What I DON'T know is are (2) and (3) the same dataset? IE is the public leader board the performance on the 5 TIFFs in the 'test' folder.</p>\n<p>The reason this is very important is it means you cannot use your public leader board score as a realistic measure of performance if you are either handlabelling or pseudolabelling them, because they're no longer a test set.</p>\n<p>Is my understanding above correct? Do you think the public leaderboard performance is the 5 test cases we already have access to?</p>",
      "rawMarkdown": "So what I know are\n\n1. There is a training set of 15 TIFFs\n2. There is a folder entitled 'test' on the submission notebooks and in the downloaded data containing 5 TIFFs\n3.  There is \"public test set\" used for the public leaderboard\n4.  There is a \"private test set\" used for the private leaderboard\n\nMy understanding is our code gets run on (3) and (4) whenever we submit, and we are at present shown our scores for (3), and then (4) will be immediately revealed when the competition ends.\n\nWhat I DON'T know is are (2) and (3) the same dataset? IE is the public leader board the performance on the 5 TIFFs in the 'test' folder.\n\nThe reason this is very important is it means you cannot use your public leader board score as a realistic measure of performance if you are either handlabelling or pseudolabelling them, because they're no longer a test set.\n\nIs my understanding above correct? Do you think the public leaderboard performance is the 5 test cases we already have access to?",
      "votes": null
    },
    {
      "id": "1277312",
      "postDate": "04/18/2021 16:10:46",
      "content": "<p>(2) and (3) are the same; they are the test images. The public leaderboard is performance on those images and many people here are including those in their training.</p>\n<p>I haven't included the public test sets in my training but eventually I have to because it means I have less training data than others.</p>",
      "rawMarkdown": "(2) and (3) are the same; they are the test images. The public leaderboard is performance on those images and many people here are including those in their training.\n\nI haven't included the public test sets in my training but eventually I have to because it means I have less training data than others.",
      "votes": null
    },
    {
      "id": "1277351",
      "postDate": "04/18/2021 16:35:55",
      "content": "<p>Thanks, I suspected as much. It means the public/private leaderboard mix up is likely to be particularly spicy.</p>",
      "rawMarkdown": "Thanks, I suspected as much. It means the public/private leaderboard mix up is likely to be particularly spicy.",
      "votes": null
    },
    {
      "id": "1279284",
      "postDate": "04/20/2021 19:04:56",
      "content": "<p>I think (2) and (3) are the same, there are a number of discussions here saying that you basically can label all the (3) manually and score very high on LB, but it won't perform well on (4). Basically you will have to focus on your own CV scores despite the public LB to have an idea about the quality of your model</p>",
      "rawMarkdown": "I think (2) and (3) are the same, there are a number of discussions here saying that you basically can label all the (3) manually and score very high on LB, but it won't perform well on (4). Basically you will have to focus on your own CV scores despite the public LB to have an idea about the quality of your model",
      "votes": null
    },
    {
      "id": "1279289",
      "postDate": "04/20/2021 19:07:59",
      "content": "<p>There is also some external data (2000 images) that you can try to implement, it didn't made that much of a difference on my scores but i'm following the same logic as yours.</p>\n<p>Here's the link<br>\n<a href=\"https://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024\" target=\"_blank\">https://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024</a></p>",
      "rawMarkdown": "There is also some external data (2000 images) that you can try to implement, it didn't made that much of a difference on my scores but i'm following the same logic as yours.\n\nHere's the link\nhttps://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024",
      "votes": null
    },
    {
      "id": "1279970",
      "postDate": "04/21/2021 12:35:15",
      "content": "<p><a href=\"https://www.kaggle.com/victorasso\" target=\"_blank\">@victorasso</a>  is his 1024 ds based on reduce 1  and tiles sz =1024 ?</p>",
      "rawMarkdown": "victorasso  is his 1024 ds based on reduce 1  and tiles sz =1024 ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1277312,
      "author_name": "erikdali",
      "author_url": "",
      "post_date": "04/18/2021 16:10:46",
      "content": "<p>(2) and (3) are the same; they are the test images. The public leaderboard is performance on those images and many people here are including those in their training.</p>\n<p>I haven't included the public test sets in my training but eventually I have to because it means I have less training data than others.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1277351,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "04/18/2021 16:35:55",
          "content": "<p>Thanks, I suspected as much. It means the public/private leaderboard mix up is likely to be particularly spicy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279289,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "04/20/2021 19:07:59",
          "content": "<p>There is also some external data (2000 images) that you can try to implement, it didn't made that much of a difference on my scores but i'm following the same logic as yours.</p>\n<p>Here's the link<br>\n<a href=\"https://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024\" target=\"_blank\">https://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279970,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "04/21/2021 12:35:15",
          "content": "<p><a href=\"https://www.kaggle.com/victorasso\" target=\"_blank\">@victorasso</a>  is his 1024 ds based on reduce 1  and tiles sz =1024 ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1279284,
      "author_name": "victorasso",
      "author_url": "",
      "post_date": "04/20/2021 19:04:56",
      "content": "<p>I think (2) and (3) are the same, there are a number of discussions here saying that you basically can label all the (3) manually and score very high on LB, but it won't perform well on (4). Basically you will have to focus on your own CV scores despite the public LB to have an idea about the quality of your model</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1277115": "So what I know are\n\n1. There is a training set of 15 TIFFs\n2. There is a folder entitled 'test' on the submission notebooks and in the downloaded data containing 5 TIFFs\n3.  There is \"public test set\" used for the public leaderboard\n4.  There is a \"private test set\" used for the private leaderboard\n\nMy understanding is our code gets run on (3) and (4) whenever we submit, and we are at present shown our scores for (3), and then (4) will be immediately revealed when the competition ends.\n\nWhat I DON'T know is are (2) and (3) the same dataset? IE is the public leader board the performance on the 5 TIFFs in the 'test' folder.\n\nThe reason this is very important is it means you cannot use your public leader board score as a realistic measure of performance if you are either handlabelling or pseudolabelling them, because they're no longer a test set.\n\nIs my understanding above correct? Do you think the public leaderboard performance is the 5 test cases we already have access to?",
    "1277312": "(2) and (3) are the same; they are the test images. The public leaderboard is performance on those images and many people here are including those in their training.\n\nI haven't included the public test sets in my training but eventually I have to because it means I have less training data than others.",
    "1277351": "Thanks, I suspected as much. It means the public/private leaderboard mix up is likely to be particularly spicy.",
    "1279284": "I think (2) and (3) are the same, there are a number of discussions here saying that you basically can label all the (3) manually and score very high on LB, but it won't perform well on (4). Basically you will have to focus on your own CV scores despite the public LB to have an idea about the quality of your model",
    "1279289": "There is also some external data (2000 images) that you can try to implement, it didn't made that much of a difference on my scores but i'm following the same logic as yours.\n\nHere's the link\nhttps://www.kaggle.com/baesiann/glomeruli-hubmap-external-1024x1024",
    "1279970": "victorasso  is his 1024 ds based on reduce 1  and tiles sz =1024 ?"
  },
  "source": "meta"
}