{
  "id": 419344,
  "title": "Question: how to setup CV?",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/419344",
  "author_name": "",
  "post_date": "2023-06-25T12:16:45.748944100Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I was curious to hear your thoughts about the CV setup. Here is my current understanding:</p>\n<p><strong>WSIs 1-4:</strong> these two WSIs comprise the public set and the training dataset.<br>\n<strong>WSI 5:</strong> this is the original private set?<br>\n<strong>WSIs 6-14:</strong> this WSIs are not annotated so we don't include them.</p>\n<hr>\n<p>So 4 fold CV with each fold tested on different WSI like:</p>\n<ul>\n<li>Fold 0: train on WSI 1, 2, 3 and test on WSI 4</li>\n<li>Fold 1: train on WSI 1, 2, 4 and test on WSI 3</li>\n<li>Fold 2: train on WSI 1, 3, 4 and test on WSI 2</li>\n<li>Fold 3: train on WSI 2, 3, 4 and test on WSI 1</li>\n</ul>\n<p>This way we will be able to replicate the test setup by only including a complete set of tiles of a single WSI. Does that make sense? </p>",
  "messages": [
    {
      "id": "2317056",
      "postDate": "06/25/2023 12:16:45",
      "content": "<p>Hi everyone,</p>\n<p>I was curious to hear your thoughts about the CV setup. Here is my current understanding:</p>\n<p><strong>WSIs 1-4:</strong> these two WSIs comprise the public set and the training dataset.<br>\n<strong>WSI 5:</strong> this is the original private set?<br>\n<strong>WSIs 6-14:</strong> this WSIs are not annotated so we don't include them.</p>\n<hr>\n<p>So 4 fold CV with each fold tested on different WSI like:</p>\n<ul>\n<li>Fold 0: train on WSI 1, 2, 3 and test on WSI 4</li>\n<li>Fold 1: train on WSI 1, 2, 4 and test on WSI 3</li>\n<li>Fold 2: train on WSI 1, 3, 4 and test on WSI 2</li>\n<li>Fold 3: train on WSI 2, 3, 4 and test on WSI 1</li>\n</ul>\n<p>This way we will be able to replicate the test setup by only including a complete set of tiles of a single WSI. Does that make sense? </p>",
      "rawMarkdown": "Hi everyone,\n\nI was curious to hear your thoughts about the CV setup. Here is my current understanding:\n\n**WSIs 1-4:** these two WSIs comprise the public set and the training dataset.\n**WSI 5:** this is the original private set?\n**WSIs 6-14:** this WSIs are not annotated so we don't include them.\n\n---\n\nSo 4 fold CV with each fold tested on different WSI like:\n- Fold 0: train on WSI 1, 2, 3 and test on WSI 4\n- Fold 1: train on WSI 1, 2, 4 and test on WSI 3\n- Fold 2: train on WSI 1, 3, 4 and test on WSI 2\n- Fold 3: train on WSI 2, 3, 4 and test on WSI 1\n\nThis way we will be able to replicate the test setup by only including a complete set of tiles of a single WSI. Does that make sense?",
      "votes": null
    },
    {
      "id": "2317143",
      "postDate": "06/25/2023 13:25:52",
      "content": "<p>issue is really the hidden test wsi5, which i haven't have a solution yet.</p>\n<p>public test is dataset1-wsi3,4.<br>\n(you have training dataset2-wsi3,4)</p>\n<hr>\n<p>if you can think of a way to have good results for:</p>\n<ul>\n<li>train dataset2-wsi1,2 (poor annotation of blood vessel) </li>\n<li>valid dataset1-wsi1,2 (good annotation of blood vessel + unsure),</li>\n</ul>\n<p>you probably can get good public test. and if you have good results for public test, you can \"improve labels of dataset1-wsi3,4, i.e. make them more dataset2 like)</p>\n<hr>\n<p>this is not proven, but i think:</p>\n<ul>\n<li>dataset1 is first created. (there is no unsure)</li>\n<li>dataset1 annotations  is split into \"sure\" and \"unsure\" and this is dataset2 </li>\n</ul>\n<hr>\n<p>since you have acess to tile_meta.csv at test, you can probe what is the number of tiles (e.g. what dataset, what wsi)</p>\n<hr>\n<p>to use wsi 6-14, you need to check if wsi 6-14 is representative of wsi 3,4,5.</p>",
      "rawMarkdown": "issue is really the hidden test wsi5, which i haven't have a solution yet.\n\npublic test is dataset1-wsi3,4.\n(you have training dataset2-wsi3,4)\n\n---\n\nif you can think of a way to have good results for:\n - train dataset2-wsi1,2 (poor annotation of blood vessel) \n - valid dataset1-wsi1,2 (good annotation of blood vessel + unsure),\n\nyou probably can get good public test. and if you have good results for public test, you can \"improve labels of dataset1-wsi3,4, i.e. make them more dataset2 like)\n\n\n---\n\nthis is not proven, but i think:\n- dataset1 is first created. (there is no unsure)\n- dataset1 annotations  is split into \"sure\" and \"unsure\" and this is dataset2 \n\n---\n\nsince you have acess to tile\\_meta.csv at test, you can probe what is the number of tiles (e.g. what dataset, what wsi)\n\n---\nto use wsi 6-14, you need to check if wsi 6-14 is representative of wsi 3,4,5.",
      "votes": null
    },
    {
      "id": "2317369",
      "postDate": "06/25/2023 16:42:30",
      "content": "<p>thank you for your answer. I got some additional questions if you don't mind:</p>\n<blockquote>\n  <p>public test is dataset1-wsi3,4.</p>\n</blockquote>\n<p>here, why the WSI is (3 and 4) and not (1 and 2). </p>\n<hr>\n<blockquote>\n  <p>train dataset2-wsi1,2 (poor annotation of blood vessel)<br>\n  valid dataset1-wsi1,2 (good annotation of blood vessel + unsure),</p>\n</blockquote>\n<p>In this split, what happened to 3 and 4? Should we not use them until having a model that can improve the labels?</p>\n<p>I think our perception of WSIs 1, 2, and 3, 4 is different from each other. I am confused about which one is correct now :D</p>\n<p>According to this split, I think 1 and 2 belong to the public set and 3 and 4 are the training set.<br>\n<img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/hubmap_ds_split.png\" alt=\"\"></p>",
      "rawMarkdown": "thank you for your answer. I got some additional questions if you don't mind:\n\n> public test is dataset1-wsi3,4.\n\nhere, why the WSI is (3 and 4) and not (1 and 2). \n\n---\n\n> train dataset2-wsi1,2 (poor annotation of blood vessel)\nvalid dataset1-wsi1,2 (good annotation of blood vessel + unsure),\n\nIn this split, what happened to 3 and 4? Should we not use them until having a model that can improve the labels?\n\nI think our perception of WSIs 1, 2, and 3, 4 is different from each other. I am confused about which one is correct now :D\n\nAccording to this split, I think 1 and 2 belong to the public set and 3 and 4 are the training set.\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/hubmap_ds_split.png)",
      "votes": null
    },
    {
      "id": "2317642",
      "postDate": "06/25/2023 20:42:57",
      "content": "<pre><code>The competition data comprises tiles extracted   Whole Slide Images (WSI)    datasets. Tiles  Dataset  have annotations that have been expert reviewed. Dataset  comprises  remaining tiles  these same WSIs  contain sparse annotations that have  been expert reviewed.\n\n- All   test  tiles are  Dataset \n- Two   WSIs make up  training ,  WSIs make up  public test ,   WSI makes up   test .\n- The training data includes Dataset  tiles   public test WSI, but     test WSI.\n</code></pre>\n<p>see also: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038</a></p>\n<hr>\n<p>distinguish between training data and training wsi/tile/set.</p>\n<p>\"The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets\"</p>\n<p>i read this as:<br>\n\"The competition data comprises tiles extracted from five Whole Slide Images (WSI), each split into two datasets\"</p>\n<p>so there is wsi3-dataset.1, wsi3-dataset.2.</p>\n<p>you don't download wsi3-dataset.1, so it must be the test set.</p>\n<p><strong>rather than guessing, one can probe the tile_meta.csv at test to verify</strong></p>",
      "rawMarkdown": "```\nThe competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\n\n- All of the test set tiles are from Dataset 1.\n- Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.\n- The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.\n\n```\n\nsee also: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038\n\n\n---\n\ndistinguish between training data and training wsi/tile/set.\n\n\"The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets\"\n\ni read this as:\n\"The competition data comprises tiles extracted from five Whole Slide Images (WSI), each split into two datasets\"\n\nso there is wsi3-dataset.1, wsi3-dataset.2.\n\nyou don't download wsi3-dataset.1, so it must be the test set.\n\n**rather than guessing, one can probe the tile_meta.csv at test to verify**",
      "votes": null
    },
    {
      "id": "2317920",
      "postDate": "06/26/2023 05:04:04",
      "content": "<p>I use 4 folds leave one WSI out validation like you described.</p>",
      "rawMarkdown": "I use 4 folds leave one WSI out validation like you described.",
      "votes": null
    },
    {
      "id": "2319866",
      "postDate": "06/27/2023 10:49:58",
      "content": "<p>Thank you! After having a look further I plan to proceed with this approach as well. Just now a bit curious about improving labels with this approach without abusing it :D </p>",
      "rawMarkdown": "Thank you! After having a look further I plan to proceed with this approach as well. Just now a bit curious about improving labels with this approach without abusing it :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2317143,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/25/2023 13:25:52",
      "content": "<p>issue is really the hidden test wsi5, which i haven't have a solution yet.</p>\n<p>public test is dataset1-wsi3,4.<br>\n(you have training dataset2-wsi3,4)</p>\n<hr>\n<p>if you can think of a way to have good results for:</p>\n<ul>\n<li>train dataset2-wsi1,2 (poor annotation of blood vessel) </li>\n<li>valid dataset1-wsi1,2 (good annotation of blood vessel + unsure),</li>\n</ul>\n<p>you probably can get good public test. and if you have good results for public test, you can \"improve labels of dataset1-wsi3,4, i.e. make them more dataset2 like)</p>\n<hr>\n<p>this is not proven, but i think:</p>\n<ul>\n<li>dataset1 is first created. (there is no unsure)</li>\n<li>dataset1 annotations  is split into \"sure\" and \"unsure\" and this is dataset2 </li>\n</ul>\n<hr>\n<p>since you have acess to tile_meta.csv at test, you can probe what is the number of tiles (e.g. what dataset, what wsi)</p>\n<hr>\n<p>to use wsi 6-14, you need to check if wsi 6-14 is representative of wsi 3,4,5.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2317369,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "06/25/2023 16:42:30",
          "content": "<p>thank you for your answer. I got some additional questions if you don't mind:</p>\n<blockquote>\n  <p>public test is dataset1-wsi3,4.</p>\n</blockquote>\n<p>here, why the WSI is (3 and 4) and not (1 and 2). </p>\n<hr>\n<blockquote>\n  <p>train dataset2-wsi1,2 (poor annotation of blood vessel)<br>\n  valid dataset1-wsi1,2 (good annotation of blood vessel + unsure),</p>\n</blockquote>\n<p>In this split, what happened to 3 and 4? Should we not use them until having a model that can improve the labels?</p>\n<p>I think our perception of WSIs 1, 2, and 3, 4 is different from each other. I am confused about which one is correct now :D</p>\n<p>According to this split, I think 1 and 2 belong to the public set and 3 and 4 are the training set.<br>\n<img src=\"https://raw.githubusercontent.com/snnclsr/kaggle_images/main/hubmap_ds_split.png\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2317642,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/25/2023 20:42:57",
              "content": "<pre><code>The competition data comprises tiles extracted   Whole Slide Images (WSI)    datasets. Tiles  Dataset  have annotations that have been expert reviewed. Dataset  comprises  remaining tiles  these same WSIs  contain sparse annotations that have  been expert reviewed.\n\n- All   test  tiles are  Dataset \n- Two   WSIs make up  training ,  WSIs make up  public test ,   WSI makes up   test .\n- The training data includes Dataset  tiles   public test WSI, but     test WSI.\n</code></pre>\n<p>see also: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038</a></p>\n<hr>\n<p>distinguish between training data and training wsi/tile/set.</p>\n<p>\"The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets\"</p>\n<p>i read this as:<br>\n\"The competition data comprises tiles extracted from five Whole Slide Images (WSI), each split into two datasets\"</p>\n<p>so there is wsi3-dataset.1, wsi3-dataset.2.</p>\n<p>you don't download wsi3-dataset.1, so it must be the test set.</p>\n<p><strong>rather than guessing, one can probe the tile_meta.csv at test to verify</strong></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2317920,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "06/26/2023 05:04:04",
      "content": "<p>I use 4 folds leave one WSI out validation like you described.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2319866,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "06/27/2023 10:49:58",
          "content": "<p>Thank you! After having a look further I plan to proceed with this approach as well. Just now a bit curious about improving labels with this approach without abusing it :D </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2317056": "Hi everyone,\n\nI was curious to hear your thoughts about the CV setup. Here is my current understanding:\n\n**WSIs 1-4:** these two WSIs comprise the public set and the training dataset.\n**WSI 5:** this is the original private set?\n**WSIs 6-14:** this WSIs are not annotated so we don't include them.\n\n---\n\nSo 4 fold CV with each fold tested on different WSI like:\n- Fold 0: train on WSI 1, 2, 3 and test on WSI 4\n- Fold 1: train on WSI 1, 2, 4 and test on WSI 3\n- Fold 2: train on WSI 1, 3, 4 and test on WSI 2\n- Fold 3: train on WSI 2, 3, 4 and test on WSI 1\n\nThis way we will be able to replicate the test setup by only including a complete set of tiles of a single WSI. Does that make sense?",
    "2317143": "issue is really the hidden test wsi5, which i haven't have a solution yet.\n\npublic test is dataset1-wsi3,4.\n(you have training dataset2-wsi3,4)\n\n---\n\nif you can think of a way to have good results for:\n - train dataset2-wsi1,2 (poor annotation of blood vessel) \n - valid dataset1-wsi1,2 (good annotation of blood vessel + unsure),\n\nyou probably can get good public test. and if you have good results for public test, you can \"improve labels of dataset1-wsi3,4, i.e. make them more dataset2 like)\n\n\n---\n\nthis is not proven, but i think:\n- dataset1 is first created. (there is no unsure)\n- dataset1 annotations  is split into \"sure\" and \"unsure\" and this is dataset2 \n\n---\n\nsince you have acess to tile\\_meta.csv at test, you can probe what is the number of tiles (e.g. what dataset, what wsi)\n\n---\nto use wsi 6-14, you need to check if wsi 6-14 is representative of wsi 3,4,5.",
    "2317369": "thank you for your answer. I got some additional questions if you don't mind:\n\n> public test is dataset1-wsi3,4.\n\nhere, why the WSI is (3 and 4) and not (1 and 2). \n\n---\n\n> train dataset2-wsi1,2 (poor annotation of blood vessel)\nvalid dataset1-wsi1,2 (good annotation of blood vessel + unsure),\n\nIn this split, what happened to 3 and 4? Should we not use them until having a model that can improve the labels?\n\nI think our perception of WSIs 1, 2, and 3, 4 is different from each other. I am confused about which one is correct now :D\n\nAccording to this split, I think 1 and 2 belong to the public set and 3 and 4 are the training set.\n![](https://raw.githubusercontent.com/snnclsr/kaggle_images/main/hubmap_ds_split.png)",
    "2317642": "```\nThe competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\n\n- All of the test set tiles are from Dataset 1.\n- Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.\n- The training data includes Dataset 2 tiles from the public test WSI, but not from the private test WSI.\n\n```\n\nsee also: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038\n\n\n---\n\ndistinguish between training data and training wsi/tile/set.\n\n\"The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets\"\n\ni read this as:\n\"The competition data comprises tiles extracted from five Whole Slide Images (WSI), each split into two datasets\"\n\nso there is wsi3-dataset.1, wsi3-dataset.2.\n\nyou don't download wsi3-dataset.1, so it must be the test set.\n\n**rather than guessing, one can probe the tile_meta.csv at test to verify**",
    "2317920": "I use 4 folds leave one WSI out validation like you described.",
    "2319866": "Thank you! After having a look further I plan to proceed with this approach as well. Just now a bit curious about improving labels with this approach without abusing it :D"
  },
  "source": "meta"
}