{
  "id": 413038,
  "title": "Question about WSI, Dataset and Split",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/413038",
  "author_name": "",
  "post_date": "2023-05-26T13:17:31.197068100Z",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'm confused about how the data is split.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2F33a68aee38a679bdec1b22c1f9f1fe1b%2F2023-05-26%2021.08.09.png?generation=1685106505018368&amp;alt=media\" alt=\"\"></p>\n<p>In the data description, it says that two of the WSIs make up the training set. But in tile_meta.csv, there are all WSIs except for WSI 5.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2F441643b46aad922bd40ea4c0a92bfdce%2F2023-05-26%2021.10.04.png?generation=1685106711424214&amp;alt=media\" alt=\"\"></p>\n<p>So I guess WSI 5 is the private test set and WSI 6-14 are the un-annotated Dataset 3. Then there are four annotated WSIs (1-4), not two.</p>\n<p>Can someone clarify this? Thanks. </p>",
  "messages": [
    {
      "id": "2275067",
      "postDate": "05/26/2023 13:17:31",
      "content": "<p>I'm confused about how the data is split.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2F33a68aee38a679bdec1b22c1f9f1fe1b%2F2023-05-26%2021.08.09.png?generation=1685106505018368&amp;alt=media\" alt=\"\"></p>\n<p>In the data description, it says that two of the WSIs make up the training set. But in tile_meta.csv, there are all WSIs except for WSI 5.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2F441643b46aad922bd40ea4c0a92bfdce%2F2023-05-26%2021.10.04.png?generation=1685106711424214&amp;alt=media\" alt=\"\"></p>\n<p>So I guess WSI 5 is the private test set and WSI 6-14 are the un-annotated Dataset 3. Then there are four annotated WSIs (1-4), not two.</p>\n<p>Can someone clarify this? Thanks. </p>",
      "rawMarkdown": "I'm confused about how the data is split.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2F33a68aee38a679bdec1b22c1f9f1fe1b%2F2023-05-26%2021.08.09.png?generation=1685106505018368&alt=media)\n\nIn the data description, it says that two of the WSIs make up the training set. But in tile_meta.csv, there are all WSIs except for WSI 5.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2F441643b46aad922bd40ea4c0a92bfdce%2F2023-05-26%2021.10.04.png?generation=1685106711424214&alt=media)\n\nSo I guess WSI 5 is the private test set and WSI 6-14 are the un-annotated Dataset 3. Then there are four annotated WSIs (1-4), not two.\n\nCan someone clarify this? Thanks.",
      "votes": null
    },
    {
      "id": "2276159",
      "postDate": "05/26/2023 14:49:22",
      "content": "<p>Hello, as stated in the data description, two annotated WSIs are part of the training set and two are part of the public test set.</p>",
      "rawMarkdown": "Hello, as stated in the data description, two annotated WSIs are part of the training set and two are part of the public test set.",
      "votes": null
    },
    {
      "id": "2278488",
      "postDate": "05/28/2023 18:14:46",
      "content": "<p>Hi PT0X0E!<br>\nAs stated on <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">the competition data page</a>:</p>\n<blockquote>\n  <p>- All of the test set tiles are from Dataset 1.<br>\n      - Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.<br>\n      - The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.</p>\n</blockquote>\n<p>This means that some tiles of the WSIs that are assigned to the public test set (WSIs 3 and 4) are actually available in the training set; these are the tiles that belong to dataset 2, i.e., the tiles whose annotations were not reviewed by experts. </p>\n<p>The following pivot table illustrates how the tiles in the downloadable data are distributed: </p>\n<ol>\n<li>Tiles of the WSI training set whose annotations have been reviewed by experts --&gt; WSIs 1 and 2 | Dataset 1</li>\n<li>Tiles of the WSI training set whose annotations have <strong>not</strong> been reviewed by experts --&gt; WSIs 1 and 2 | Dataset 2</li>\n<li>Tiles of the WSI training set which have <strong>not</strong> been annotated at all --&gt; WSIs 6 through 14 | Dataset 3</li>\n<li>Tiles of the WSI public test set whose annotations have <strong>not</strong> been reviewed by experts --&gt; WSIs 3 and 4 | Dataset 2</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14817755%2Fac0986fd4ee861c25e022384b57338dd%2F2023-05-28_21-12.png?generation=1685297651554772&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi PT0X0E!\nAs stated on [the competition data page](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data):\n>  \n    - All of the test set tiles are from Dataset 1.\n    - Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.\n    - The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.\n\nThis means that some tiles of the WSIs that are assigned to the public test set (WSIs 3 and 4) are actually available in the training set; these are the tiles that belong to dataset 2, i.e., the tiles whose annotations were not reviewed by experts. \n\nThe following pivot table illustrates how the tiles in the downloadable data are distributed: \n1. Tiles of the WSI training set whose annotations have been reviewed by experts --> WSIs 1 and 2 | Dataset 1\n2. Tiles of the WSI training set whose annotations have **not** been reviewed by experts --> WSIs 1 and 2 | Dataset 2\n3. Tiles of the WSI training set which have **not** been annotated at all --> WSIs 6 through 14 | Dataset 3\n4. Tiles of the WSI public test set whose annotations have **not** been reviewed by experts --> WSIs 3 and 4 | Dataset 2\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14817755%2Fac0986fd4ee861c25e022384b57338dd%2F2023-05-28_21-12.png?generation=1685297651554772&alt=media)",
      "votes": null
    },
    {
      "id": "2279515",
      "postDate": "05/29/2023 12:33:17",
      "content": "<p>Hi Mohamed, <br>\nThanks so much for the explanation. It was just \"Two of the WSIs make up the training set, two WSIs make up the public test set\" that confused me. <br>\nI made a picture of it. Do you think this is correct?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2Fd0fad19fa693732c2cae00c41453adf4%2F2023-05-29%2020.32.10.png?generation=1685363555970493&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi Mohamed, \nThanks so much for the explanation. It was just \"Two of the WSIs make up the training set, two WSIs make up the public test set\" that confused me. \nI made a picture of it. Do you think this is correct?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2Fd0fad19fa693732c2cae00c41453adf4%2F2023-05-29%2020.32.10.png?generation=1685363555970493&alt=media)",
      "votes": null
    },
    {
      "id": "2279777",
      "postDate": "05/29/2023 16:06:51",
      "content": "<p>Yes. I think this is accurate.</p>\n<p>Except WSI5 is also presumably divided into dataset 1 and 2. The tiles that belong to dataset 1 are the ones that will be counted in the private test set. The tiles that belong to dataset 2 will be thrown in the trash I suppose :) (they can't test our models against non-verified annotations, can they?)</p>\n<p>I think the competition hosts decided not give us the dataset 2 of the private test set to really test our models's performance against a complete WSI that it has never seen, not just tiles that it has never seen. I also think that this will be greatly contribute to the LB shake at the end of the competition.</p>",
      "rawMarkdown": "Yes. I think this is accurate.\n\nExcept WSI5 is also presumably divided into dataset 1 and 2. The tiles that belong to dataset 1 are the ones that will be counted in the private test set. The tiles that belong to dataset 2 will be thrown in the trash I suppose :) (they can't test our models against non-verified annotations, can they?)\n\nI think the competition hosts decided not give us the dataset 2 of the private test set to really test our models's performance against a complete WSI that it has never seen, not just tiles that it has never seen. I also think that this will be greatly contribute to the LB shake at the end of the competition.",
      "votes": null
    },
    {
      "id": "2283077",
      "postDate": "06/01/2023 03:48:28",
      "content": "<p>那个Data的解释，第一次看真的会懵逼…</p>",
      "rawMarkdown": "那个Data的解释，第一次看真的会懵逼...",
      "votes": null
    },
    {
      "id": "2304190",
      "postDate": "06/15/2023 18:57:15",
      "content": "<p>This could be correct but data definition might have a problem.</p>\n<ul>\n<li>All of the test set tiles are from Dataset 1.</li>\n<li>The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.</li>\n</ul>\n<p>How both of those points can work at the same time? I don't get it. <a href=\"https://www.kaggle.com/yashvrdnjain\" target=\"_blank\">@yashvrdnjain</a> </p>",
      "rawMarkdown": "This could be correct but data definition might have a problem.\n* All of the test set tiles are from Dataset 1.\n* The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.\n\nHow both of those points can work at the same time? I don't get it. @yashvrdnjain",
      "votes": null
    },
    {
      "id": "2304294",
      "postDate": "06/15/2023 22:29:34",
      "content": "<p>The way I understand it, although they won't use dataset 2 of the private test WSI for the evaluation of our models, they still won't provide it for the training data to really put our models to the test against a WSI that they weren't exposed to a single tile of. This should ensure the winning models will have good generalization capabilities, and make it more likely that they will generalize to future WSIs yet to be collected.</p>",
      "rawMarkdown": "The way I understand it, although they won't use dataset 2 of the private test WSI for the evaluation of our models, they still won't provide it for the training data to really put our models to the test against a WSI that they weren't exposed to a single tile of. This should ensure the winning models will have good generalization capabilities, and make it more likely that they will generalize to future WSIs yet to be collected.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2276159,
      "author_name": "yashvrdnjain",
      "author_url": "",
      "post_date": "05/26/2023 14:49:22",
      "content": "<p>Hello, as stated in the data description, two annotated WSIs are part of the training set and two are part of the public test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2278488,
      "author_name": "momakmd",
      "author_url": "",
      "post_date": "05/28/2023 18:14:46",
      "content": "<p>Hi PT0X0E!<br>\nAs stated on <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">the competition data page</a>:</p>\n<blockquote>\n  <p>- All of the test set tiles are from Dataset 1.<br>\n      - Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.<br>\n      - The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.</p>\n</blockquote>\n<p>This means that some tiles of the WSIs that are assigned to the public test set (WSIs 3 and 4) are actually available in the training set; these are the tiles that belong to dataset 2, i.e., the tiles whose annotations were not reviewed by experts. </p>\n<p>The following pivot table illustrates how the tiles in the downloadable data are distributed: </p>\n<ol>\n<li>Tiles of the WSI training set whose annotations have been reviewed by experts --&gt; WSIs 1 and 2 | Dataset 1</li>\n<li>Tiles of the WSI training set whose annotations have <strong>not</strong> been reviewed by experts --&gt; WSIs 1 and 2 | Dataset 2</li>\n<li>Tiles of the WSI training set which have <strong>not</strong> been annotated at all --&gt; WSIs 6 through 14 | Dataset 3</li>\n<li>Tiles of the WSI public test set whose annotations have <strong>not</strong> been reviewed by experts --&gt; WSIs 3 and 4 | Dataset 2</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14817755%2Fac0986fd4ee861c25e022384b57338dd%2F2023-05-28_21-12.png?generation=1685297651554772&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2279515,
          "author_name": "pt0x0e",
          "author_url": "",
          "post_date": "05/29/2023 12:33:17",
          "content": "<p>Hi Mohamed, <br>\nThanks so much for the explanation. It was just \"Two of the WSIs make up the training set, two WSIs make up the public test set\" that confused me. <br>\nI made a picture of it. Do you think this is correct?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2Fd0fad19fa693732c2cae00c41453adf4%2F2023-05-29%2020.32.10.png?generation=1685363555970493&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2279777,
              "author_name": "momakmd",
              "author_url": "",
              "post_date": "05/29/2023 16:06:51",
              "content": "<p>Yes. I think this is accurate.</p>\n<p>Except WSI5 is also presumably divided into dataset 1 and 2. The tiles that belong to dataset 1 are the ones that will be counted in the private test set. The tiles that belong to dataset 2 will be thrown in the trash I suppose :) (they can't test our models against non-verified annotations, can they?)</p>\n<p>I think the competition hosts decided not give us the dataset 2 of the private test set to really test our models's performance against a complete WSI that it has never seen, not just tiles that it has never seen. I also think that this will be greatly contribute to the LB shake at the end of the competition.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2304190,
              "author_name": "gunesevitan",
              "author_url": "",
              "post_date": "06/15/2023 18:57:15",
              "content": "<p>This could be correct but data definition might have a problem.</p>\n<ul>\n<li>All of the test set tiles are from Dataset 1.</li>\n<li>The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.</li>\n</ul>\n<p>How both of those points can work at the same time? I don't get it. <a href=\"https://www.kaggle.com/yashvrdnjain\" target=\"_blank\">@yashvrdnjain</a> </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2304294,
                  "author_name": "momakmd",
                  "author_url": "",
                  "post_date": "06/15/2023 22:29:34",
                  "content": "<p>The way I understand it, although they won't use dataset 2 of the private test WSI for the evaluation of our models, they still won't provide it for the training data to really put our models to the test against a WSI that they weren't exposed to a single tile of. This should ensure the winning models will have good generalization capabilities, and make it more likely that they will generalize to future WSIs yet to be collected.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2283077,
      "author_name": "zhuomao",
      "author_url": "",
      "post_date": "06/01/2023 03:48:28",
      "content": "<p>那个Data的解释，第一次看真的会懵逼…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2275067": "I'm confused about how the data is split.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2F33a68aee38a679bdec1b22c1f9f1fe1b%2F2023-05-26%2021.08.09.png?generation=1685106505018368&alt=media)\n\nIn the data description, it says that two of the WSIs make up the training set. But in tile_meta.csv, there are all WSIs except for WSI 5.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2F441643b46aad922bd40ea4c0a92bfdce%2F2023-05-26%2021.10.04.png?generation=1685106711424214&alt=media)\n\nSo I guess WSI 5 is the private test set and WSI 6-14 are the un-annotated Dataset 3. Then there are four annotated WSIs (1-4), not two.\n\nCan someone clarify this? Thanks.",
    "2276159": "Hello, as stated in the data description, two annotated WSIs are part of the training set and two are part of the public test set.",
    "2278488": "Hi PT0X0E!\nAs stated on [the competition data page](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data):\n>  \n    - All of the test set tiles are from Dataset 1.\n    - Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.\n    - The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.\n\nThis means that some tiles of the WSIs that are assigned to the public test set (WSIs 3 and 4) are actually available in the training set; these are the tiles that belong to dataset 2, i.e., the tiles whose annotations were not reviewed by experts. \n\nThe following pivot table illustrates how the tiles in the downloadable data are distributed: \n1. Tiles of the WSI training set whose annotations have been reviewed by experts --> WSIs 1 and 2 | Dataset 1\n2. Tiles of the WSI training set whose annotations have **not** been reviewed by experts --> WSIs 1 and 2 | Dataset 2\n3. Tiles of the WSI training set which have **not** been annotated at all --> WSIs 6 through 14 | Dataset 3\n4. Tiles of the WSI public test set whose annotations have **not** been reviewed by experts --> WSIs 3 and 4 | Dataset 2\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14817755%2Fac0986fd4ee861c25e022384b57338dd%2F2023-05-28_21-12.png?generation=1685297651554772&alt=media)",
    "2279515": "Hi Mohamed, \nThanks so much for the explanation. It was just \"Two of the WSIs make up the training set, two WSIs make up the public test set\" that confused me. \nI made a picture of it. Do you think this is correct?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12313384%2Fd0fad19fa693732c2cae00c41453adf4%2F2023-05-29%2020.32.10.png?generation=1685363555970493&alt=media)",
    "2279777": "Yes. I think this is accurate.\n\nExcept WSI5 is also presumably divided into dataset 1 and 2. The tiles that belong to dataset 1 are the ones that will be counted in the private test set. The tiles that belong to dataset 2 will be thrown in the trash I suppose :) (they can't test our models against non-verified annotations, can they?)\n\nI think the competition hosts decided not give us the dataset 2 of the private test set to really test our models's performance against a complete WSI that it has never seen, not just tiles that it has never seen. I also think that this will be greatly contribute to the LB shake at the end of the competition.",
    "2283077": "那个Data的解释，第一次看真的会懵逼...",
    "2304190": "This could be correct but data definition might have a problem.\n* All of the test set tiles are from Dataset 1.\n* The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.\n\nHow both of those points can work at the same time? I don't get it. @yashvrdnjain",
    "2304294": "The way I understand it, although they won't use dataset 2 of the private test WSI for the evaluation of our models, they still won't provide it for the training data to really put our models to the test against a WSI that they weren't exposed to a single tile of. This should ensure the winning models will have good generalization capabilities, and make it more likely that they will generalize to future WSIs yet to be collected."
  },
  "source": "meta"
}