{
  "id": 417267,
  "title": "how to avoid overfit to WSI 1&2?",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/417267",
  "author_name": "",
  "post_date": "2023-06-15T01:37:01.797290900Z",
  "votes": 16,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Since the private test WSI is unknown, I would like to separate the WSI used for learning and the WSI not used for local validation, but only two WSIs annotated by experts are provided.</p>\n<p>If you have any good ideas to withstand shake, please let me know.</p>",
  "messages": [
    {
      "id": "2302945",
      "postDate": "06/15/2023 01:37:01",
      "content": "<p>Since the private test WSI is unknown, I would like to separate the WSI used for learning and the WSI not used for local validation, but only two WSIs annotated by experts are provided.</p>\n<p>If you have any good ideas to withstand shake, please let me know.</p>",
      "rawMarkdown": "Since the private test WSI is unknown, I would like to separate the WSI used for learning and the WSI not used for local validation, but only two WSIs annotated by experts are provided.\n\n\nIf you have any good ideas to withstand shake, please let me know.",
      "votes": null
    },
    {
      "id": "2304699",
      "postDate": "06/16/2023 07:11:51",
      "content": "<p>I haven't started modelling yet but I think leaving one WSI as validation on each fold makes sense. </p>",
      "rawMarkdown": "I haven't started modelling yet but I think leaving one WSI as validation on each fold makes sense.",
      "votes": null
    },
    {
      "id": "2305922",
      "postDate": "06/17/2023 02:15:01",
      "content": "<p>only datset1,only vessel(not unsure)</p>\n<p>gt_wsi2_val_wsi2:map0.5:0.95=0.363<br>\ngt_wsi2_val_wsi1:map0.5:0.95=0.159<br>\ngt_wsi1_val_wsi1:map0.5:0.95=0.475<br>\ngt_wsi1_val_wsi2:map0.5:0.95=0.12</p>\n<p>This may be not surprising since there is very little data, but I have experimented with my current model with an unknown WSI and got these results</p>",
      "rawMarkdown": "only datset1,only vessel(not unsure)\n\ngt_wsi2_val_wsi2:map0.5:0.95=0.363\ngt_wsi2_val_wsi1:map0.5:0.95=0.159\ngt_wsi1_val_wsi1:map0.5:0.95=0.475\ngt_wsi1_val_wsi2:map0.5:0.95=0.12\n\nThis may be not surprising since there is very little data, but I have experimented with my current model with an unknown WSI and got these results",
      "votes": null
    },
    {
      "id": "2306394",
      "postDate": "06/17/2023 09:00:29",
      "content": "<p>Have you try combine dataset 1 and dataset 2? </p>",
      "rawMarkdown": "Have you try combine dataset 1 and dataset 2?",
      "votes": null
    },
    {
      "id": "2315012",
      "postDate": "06/23/2023 18:48:33",
      "content": "<p>you first need to find data smiliarty … (i.e. domain shift)</p>\n<p>you have two choice:</p>\n<ol>\n<li>think of a way to normalise the data such that source_wsi 1,2,3,4, 6,7, …. are the same</li>\n<li>measure the smiliarity of 5 (hidden private test) with respect to 1,2,…..</li>\n</ol>\n<p>you can for example train a classifier to classify if a given tile is of class  1,2,3,4, 6,7, ….<br>\nthen probe the hidden private test to see how smiliar is it.</p>\n<p>or you can embed them to some common feature space using tsne, umap etc …<br>\n(e.g. we find 1,2,3,4, 6,7, … are within xx distance from data center. we probe that hidden 5 is also within xx distance)</p>\n<hr>\n<p>if you know the domain shift, then just need need to train a model that is \"robust within the shift\" (e.g. augmentation, adaption layer, etc)</p>\n<p>or you can use online learning + zeroshot.<br>\n(i.e. your submission code train using hidden test in self/unspervised way like mask auto encoder) ….</p>",
      "rawMarkdown": "you first need to find data smiliarty ... (i.e. domain shift)\n\nyou have two choice:\n1. think of a way to normalise the data such that source\\_wsi 1,2,3,4, 6,7, .... are the same\n2. measure the smiliarity of 5 (hidden private test) with respect to 1,2,.....\n\nyou can for example train a classifier to classify if a given tile is of class  1,2,3,4, 6,7, ....\nthen probe the hidden private test to see how smiliar is it.\n\nor you can embed them to some common feature space using tsne, umap etc ...\n(e.g. we find 1,2,3,4, 6,7, ... are within xx distance from data center. we probe that hidden 5 is also within xx distance)\n\n---\n\nif you know the domain shift, then just need need to train a model that is \"robust within the shift\" (e.g. augmentation, adaption layer, etc)\n\nor you can use online learning + zeroshot.\n(i.e. your submission code train using hidden test in self/unspervised way like mask auto encoder) ....",
      "votes": null
    },
    {
      "id": "2316473",
      "postDate": "06/25/2023 01:35:37",
      "content": "<p>gt_wsi2_val_wsi2:map0.5:0.95=0.363<br>\ngt_wsi1_val_wsi1:map0.5:0.95=0.475</p>\n<p>my results using yolov7 at 512:<br>\ngt_wsi2_val_wsi2:map0.5:0.95=0.303<br>\ngt_wsi1_val_wsi1:map0.5:0.95=0.480</p>\n<p>my results using yolov7 at 1024:<br>\ngt_wsi2_val_wsi2:map0.5:0.95=0.321<br>\ngt_wsi1_val_wsi1:map0.5:0.95=0.488</p>\n<p>the results are consistent for different fold.<br>\nhow did you end up \"gt_wsi2_val_wsi2:map0.5:0.95=0.363\" ?</p>",
      "rawMarkdown": "gt_wsi2_val_wsi2:map0.5:0.95=0.363\ngt_wsi1_val_wsi1:map0.5:0.95=0.475\n\nmy results using yolov7 at 512:\ngt_wsi2_val_wsi2:map0.5:0.95=0.303\ngt_wsi1_val_wsi1:map0.5:0.95=0.480\n\nmy results using yolov7 at 1024:\ngt_wsi2_val_wsi2:map0.5:0.95=0.321\ngt_wsi1_val_wsi1:map0.5:0.95=0.488\n\nthe results are consistent for different fold.\nhow did you end up \"gt_wsi2_val_wsi2:map0.5:0.95=0.363\" ?",
      "votes": null
    },
    {
      "id": "2316564",
      "postDate": "06/25/2023 05:04:13",
      "content": "<p><code>yolo segment train model=/home/abe/humap/yolodata/yolov8x-seg.pt epochs=300 imgsz=512 device=1 name=WSI1 flipud=0.5 degrees=90 mixup=0.5</code></p>",
      "rawMarkdown": "`yolo segment train model=/home/abe/humap/yolodata/yolov8x-seg.pt epochs=300 imgsz=512 device=1 name=WSI1 flipud=0.5 degrees=90 mixup=0.5`",
      "votes": null
    },
    {
      "id": "2316607",
      "postDate": "06/25/2023 05:51:52",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3627759%2F0fdefa2939aad959ad50f10fd48080e1%2Fhubmap_tsne_p30_seed_20230_wsi_type.png?generation=1687672085136194&amp;alt=media\" alt=\"\"></p>\n<p>embed all train images by <a href=\"https://github.com/mahmoodlab/HIPT\" target=\"_blank\">https://github.com/mahmoodlab/HIPT</a> →TSNE</p>\n<p>The color is divided to some extent, and I think that there is likely to be a domain shift.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3627759%2F0fdefa2939aad959ad50f10fd48080e1%2Fhubmap_tsne_p30_seed_20230_wsi_type.png?generation=1687672085136194&alt=media)\n\nembed all train images by https://github.com/mahmoodlab/HIPT →TSNE\n\nThe color is divided to some extent, and I think that there is likely to be a domain shift.",
      "votes": null
    },
    {
      "id": "2316705",
      "postDate": "06/25/2023 07:16:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a>, which features correspond to axis x and y?</p>",
      "rawMarkdown": "Hi @abebe9849, which features correspond to axis x and y?",
      "votes": null
    },
    {
      "id": "2316859",
      "postDate": "06/25/2023 09:28:14",
      "content": "<p>512*512*3-&gt; 256*256*3(resize)-&gt;384(embed by ViT-S)-&gt;2(tsne)</p>\n<p>2dim after prereduction by TSNE corresponds to xy</p>",
      "rawMarkdown": "512\\*512\\*3-> 256\\*256\\*3(resize)->384(embed by ViT-S)->2(tsne)\n\n2dim after prereduction by TSNE corresponds to xy",
      "votes": null
    },
    {
      "id": "2316875",
      "postDate": "06/25/2023 09:40:23",
      "content": "<p>512*512*3-&gt; 256*256*3(resize)-&gt;384(embed by ViT-S)-&gt;2(tsne)</p>\n<pre><code>  slash    show\n</code></pre>",
      "rawMarkdown": "512\\*512\\*3-> 256\\*256\\*3(resize)->384(embed by ViT-S)->2(tsne)\n\n```\nput a slash before \"\\*\" to show\n\n```",
      "votes": null
    },
    {
      "id": "2317168",
      "postDate": "06/25/2023 13:39:33",
      "content": "<p>in my experiments, yolov7-seg is better then yolov8x-seg.<br>\nIt is the 300 epochs that improve my previous results (100 epoch)</p>\n<p>now:</p>\n<p>my results using yolov7-seg at 512:<br>\ngt_wsi2_val_wsi2:map0.5:0.95=0.331 to 358 (difference fold)</p>\n<p>my results using yolov8x-seg at 512:<br>\ngt_wsi2_val_wsi2:map0.5:0.95=0.331 to 353 (difference fold)</p>",
      "rawMarkdown": "in my experiments, yolov7-seg is better then yolov8x-seg.\nIt is the 300 epochs that improve my previous results (100 epoch)\n\nnow:\n\nmy results using yolov7-seg at 512:\ngt_wsi2_val_wsi2:map0.5:0.95=0.331 to 358 (difference fold)\n\n\nmy results using yolov8x-seg at 512:\ngt_wsi2_val_wsi2:map0.5:0.95=0.331 to 353 (difference fold)",
      "votes": null
    },
    {
      "id": "2317979",
      "postDate": "06/26/2023 05:45:37",
      "content": "<p>How did you split K-fold, I think it's better to split by WSI, since private test is completely a difference WSI</p>",
      "rawMarkdown": "How did you split K-fold, I think it's better to split by WSI, since private test is completely a difference WSI",
      "votes": null
    },
    {
      "id": "2321612",
      "postDate": "06/28/2023 17:37:19",
      "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> I would like to reproduce your results and check if normalization procedures help to gather dots of different colors in a single group. I want to ask you a few questions.<br>\nHow do You compute embeddings? Do you use only ViT-256 (vit256_small_dino weights) model or maybe ViT-4k too?<br>\nWhat are your TSNE parameters?<br>\nThanks!</p>\n<p>Edit:<br>\nTSNE is very sensitive to initialization and other parameters.<br>\n<code>TSNE(n_components=2, learning_rate='auto', init='random', perplexity=3)</code><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5506850%2F6b23cd24c1b3098c45be4944c7609464%2Fr2.png?generation=1688026623129144&amp;alt=media\" alt=\"\"></p>\n<p><code>TSNE(n_components=2)</code><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5506850%2F63484c4aa1cac0b50acd3a3e4e49a5da%2Fr1.png?generation=1688026648949098&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "abebe9849 I would like to reproduce your results and check if normalization procedures help to gather dots of different colors in a single group. I want to ask you a few questions.\nHow do You compute embeddings? Do you use only ViT-256 (vit256_small_dino weights) model or maybe ViT-4k too?\nWhat are your TSNE parameters?\nThanks!\n\nEdit:\nTSNE is very sensitive to initialization and other parameters.\n`TSNE(n_components=2, learning_rate='auto', init='random', perplexity=3)`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5506850%2F6b23cd24c1b3098c45be4944c7609464%2Fr2.png?generation=1688026623129144&alt=media)\n\n`TSNE(n_components=2)`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5506850%2F63484c4aa1cac0b50acd3a3e4e49a5da%2Fr1.png?generation=1688026648949098&alt=media)",
      "votes": null
    },
    {
      "id": "2324693",
      "postDate": "06/30/2023 19:31:13",
      "content": "<p>only ViT-256<br>\nperplexity 30</p>",
      "rawMarkdown": "only ViT-256\nperplexity 30",
      "votes": null
    },
    {
      "id": "2325497",
      "postDate": "07/01/2023 12:11:23",
      "content": "<p>check my post.<br>\ni have \"found\" the whole slide for wsi.3.</p>\n<p>see if you can use self-supervised learning (without label) to improve results for private wsi.3.</p>\n<p>i think all test (public and private) are from the 30 slides from previous hubmap competition.<br>\nif so the unlabelled data of dataset.3 should help for private test wsi.5 </p>",
      "rawMarkdown": "check my post.\ni have \"found\" the whole slide for wsi.3.\n\nsee if you can use self-supervised learning (without label) to improve results for private wsi.3.\n\ni think all test (public and private) are from the 30 slides from previous hubmap competition.\nif so the unlabelled data of dataset.3 should help for private test wsi.5",
      "votes": null
    },
    {
      "id": "2334945",
      "postDate": "07/08/2023 06:30:38",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb86588abe96ad3f82e1b6945f141a627%2FSelection_999(2575).png?generation=1688797793176032&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> </p>\n<p>you probably can be better embedding score by usingopen ousrced histotology pretrained model likw thsi</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb86588abe96ad3f82e1b6945f141a627%2FSelection_999(2575).png?generation=1688797793176032&alt=media)\n\n@abebe9849 \n\nyou probably can be better embedding score by usingopen ousrced histotology pretrained model likw thsi",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2304699,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "06/16/2023 07:11:51",
      "content": "<p>I haven't started modelling yet but I think leaving one WSI as validation on each fold makes sense. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2305922,
          "author_name": "abebe9849",
          "author_url": "",
          "post_date": "06/17/2023 02:15:01",
          "content": "<p>only datset1,only vessel(not unsure)</p>\n<p>gt_wsi2_val_wsi2:map0.5:0.95=0.363<br>\ngt_wsi2_val_wsi1:map0.5:0.95=0.159<br>\ngt_wsi1_val_wsi1:map0.5:0.95=0.475<br>\ngt_wsi1_val_wsi2:map0.5:0.95=0.12</p>\n<p>This may be not surprising since there is very little data, but I have experimented with my current model with an unknown WSI and got these results</p>",
          "votes": null,
          "replies": [
            {
              "id": 2306394,
              "author_name": "ptran1203",
              "author_url": "",
              "post_date": "06/17/2023 09:00:29",
              "content": "<p>Have you try combine dataset 1 and dataset 2? </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2316473,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/25/2023 01:35:37",
              "content": "<p>gt_wsi2_val_wsi2:map0.5:0.95=0.363<br>\ngt_wsi1_val_wsi1:map0.5:0.95=0.475</p>\n<p>my results using yolov7 at 512:<br>\ngt_wsi2_val_wsi2:map0.5:0.95=0.303<br>\ngt_wsi1_val_wsi1:map0.5:0.95=0.480</p>\n<p>my results using yolov7 at 1024:<br>\ngt_wsi2_val_wsi2:map0.5:0.95=0.321<br>\ngt_wsi1_val_wsi1:map0.5:0.95=0.488</p>\n<p>the results are consistent for different fold.<br>\nhow did you end up \"gt_wsi2_val_wsi2:map0.5:0.95=0.363\" ?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2316564,
                  "author_name": "abebe9849",
                  "author_url": "",
                  "post_date": "06/25/2023 05:04:13",
                  "content": "<p><code>yolo segment train model=/home/abe/humap/yolodata/yolov8x-seg.pt epochs=300 imgsz=512 device=1 name=WSI1 flipud=0.5 degrees=90 mixup=0.5</code></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2317168,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "06/25/2023 13:39:33",
                      "content": "<p>in my experiments, yolov7-seg is better then yolov8x-seg.<br>\nIt is the 300 epochs that improve my previous results (100 epoch)</p>\n<p>now:</p>\n<p>my results using yolov7-seg at 512:<br>\ngt_wsi2_val_wsi2:map0.5:0.95=0.331 to 358 (difference fold)</p>\n<p>my results using yolov8x-seg at 512:<br>\ngt_wsi2_val_wsi2:map0.5:0.95=0.331 to 353 (difference fold)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2317979,
                          "author_name": "ptran1203",
                          "author_url": "",
                          "post_date": "06/26/2023 05:45:37",
                          "content": "<p>How did you split K-fold, I think it's better to split by WSI, since private test is completely a difference WSI</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2315012,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/23/2023 18:48:33",
      "content": "<p>you first need to find data smiliarty … (i.e. domain shift)</p>\n<p>you have two choice:</p>\n<ol>\n<li>think of a way to normalise the data such that source_wsi 1,2,3,4, 6,7, …. are the same</li>\n<li>measure the smiliarity of 5 (hidden private test) with respect to 1,2,…..</li>\n</ol>\n<p>you can for example train a classifier to classify if a given tile is of class  1,2,3,4, 6,7, ….<br>\nthen probe the hidden private test to see how smiliar is it.</p>\n<p>or you can embed them to some common feature space using tsne, umap etc …<br>\n(e.g. we find 1,2,3,4, 6,7, … are within xx distance from data center. we probe that hidden 5 is also within xx distance)</p>\n<hr>\n<p>if you know the domain shift, then just need need to train a model that is \"robust within the shift\" (e.g. augmentation, adaption layer, etc)</p>\n<p>or you can use online learning + zeroshot.<br>\n(i.e. your submission code train using hidden test in self/unspervised way like mask auto encoder) ….</p>",
      "votes": null,
      "replies": [
        {
          "id": 2316607,
          "author_name": "abebe9849",
          "author_url": "",
          "post_date": "06/25/2023 05:51:52",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3627759%2F0fdefa2939aad959ad50f10fd48080e1%2Fhubmap_tsne_p30_seed_20230_wsi_type.png?generation=1687672085136194&amp;alt=media\" alt=\"\"></p>\n<p>embed all train images by <a href=\"https://github.com/mahmoodlab/HIPT\" target=\"_blank\">https://github.com/mahmoodlab/HIPT</a> →TSNE</p>\n<p>The color is divided to some extent, and I think that there is likely to be a domain shift.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2316705,
              "author_name": "huyduong7101",
              "author_url": "",
              "post_date": "06/25/2023 07:16:58",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a>, which features correspond to axis x and y?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2316859,
                  "author_name": "abebe9849",
                  "author_url": "",
                  "post_date": "06/25/2023 09:28:14",
                  "content": "<p>512*512*3-&gt; 256*256*3(resize)-&gt;384(embed by ViT-S)-&gt;2(tsne)</p>\n<p>2dim after prereduction by TSNE corresponds to xy</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2316875,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "06/25/2023 09:40:23",
                      "content": "<p>512*512*3-&gt; 256*256*3(resize)-&gt;384(embed by ViT-S)-&gt;2(tsne)</p>\n<pre><code>  slash    show\n</code></pre>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 2321612,
              "author_name": "janglinko2",
              "author_url": "",
              "post_date": "06/28/2023 17:37:19",
              "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> I would like to reproduce your results and check if normalization procedures help to gather dots of different colors in a single group. I want to ask you a few questions.<br>\nHow do You compute embeddings? Do you use only ViT-256 (vit256_small_dino weights) model or maybe ViT-4k too?<br>\nWhat are your TSNE parameters?<br>\nThanks!</p>\n<p>Edit:<br>\nTSNE is very sensitive to initialization and other parameters.<br>\n<code>TSNE(n_components=2, learning_rate='auto', init='random', perplexity=3)</code><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5506850%2F6b23cd24c1b3098c45be4944c7609464%2Fr2.png?generation=1688026623129144&amp;alt=media\" alt=\"\"></p>\n<p><code>TSNE(n_components=2)</code><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5506850%2F63484c4aa1cac0b50acd3a3e4e49a5da%2Fr1.png?generation=1688026648949098&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2324693,
                  "author_name": "abebe9849",
                  "author_url": "",
                  "post_date": "06/30/2023 19:31:13",
                  "content": "<p>only ViT-256<br>\nperplexity 30</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2325497,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "07/01/2023 12:11:23",
                      "content": "<p>check my post.<br>\ni have \"found\" the whole slide for wsi.3.</p>\n<p>see if you can use self-supervised learning (without label) to improve results for private wsi.3.</p>\n<p>i think all test (public and private) are from the 30 slides from previous hubmap competition.<br>\nif so the unlabelled data of dataset.3 should help for private test wsi.5 </p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2334945,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "07/08/2023 06:30:38",
                      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb86588abe96ad3f82e1b6945f141a627%2FSelection_999(2575).png?generation=1688797793176032&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> </p>\n<p>you probably can be better embedding score by usingopen ousrced histotology pretrained model likw thsi</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2302945": "Since the private test WSI is unknown, I would like to separate the WSI used for learning and the WSI not used for local validation, but only two WSIs annotated by experts are provided.\n\n\nIf you have any good ideas to withstand shake, please let me know.",
    "2304699": "I haven't started modelling yet but I think leaving one WSI as validation on each fold makes sense.",
    "2305922": "only datset1,only vessel(not unsure)\n\ngt_wsi2_val_wsi2:map0.5:0.95=0.363\ngt_wsi2_val_wsi1:map0.5:0.95=0.159\ngt_wsi1_val_wsi1:map0.5:0.95=0.475\ngt_wsi1_val_wsi2:map0.5:0.95=0.12\n\nThis may be not surprising since there is very little data, but I have experimented with my current model with an unknown WSI and got these results",
    "2306394": "Have you try combine dataset 1 and dataset 2?",
    "2315012": "you first need to find data smiliarty ... (i.e. domain shift)\n\nyou have two choice:\n1. think of a way to normalise the data such that source\\_wsi 1,2,3,4, 6,7, .... are the same\n2. measure the smiliarity of 5 (hidden private test) with respect to 1,2,.....\n\nyou can for example train a classifier to classify if a given tile is of class  1,2,3,4, 6,7, ....\nthen probe the hidden private test to see how smiliar is it.\n\nor you can embed them to some common feature space using tsne, umap etc ...\n(e.g. we find 1,2,3,4, 6,7, ... are within xx distance from data center. we probe that hidden 5 is also within xx distance)\n\n---\n\nif you know the domain shift, then just need need to train a model that is \"robust within the shift\" (e.g. augmentation, adaption layer, etc)\n\nor you can use online learning + zeroshot.\n(i.e. your submission code train using hidden test in self/unspervised way like mask auto encoder) ....",
    "2316473": "gt_wsi2_val_wsi2:map0.5:0.95=0.363\ngt_wsi1_val_wsi1:map0.5:0.95=0.475\n\nmy results using yolov7 at 512:\ngt_wsi2_val_wsi2:map0.5:0.95=0.303\ngt_wsi1_val_wsi1:map0.5:0.95=0.480\n\nmy results using yolov7 at 1024:\ngt_wsi2_val_wsi2:map0.5:0.95=0.321\ngt_wsi1_val_wsi1:map0.5:0.95=0.488\n\nthe results are consistent for different fold.\nhow did you end up \"gt_wsi2_val_wsi2:map0.5:0.95=0.363\" ?",
    "2316564": "`yolo segment train model=/home/abe/humap/yolodata/yolov8x-seg.pt epochs=300 imgsz=512 device=1 name=WSI1 flipud=0.5 degrees=90 mixup=0.5`",
    "2316607": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3627759%2F0fdefa2939aad959ad50f10fd48080e1%2Fhubmap_tsne_p30_seed_20230_wsi_type.png?generation=1687672085136194&alt=media)\n\nembed all train images by https://github.com/mahmoodlab/HIPT →TSNE\n\nThe color is divided to some extent, and I think that there is likely to be a domain shift.",
    "2316705": "Hi @abebe9849, which features correspond to axis x and y?",
    "2316859": "512\\*512\\*3-> 256\\*256\\*3(resize)->384(embed by ViT-S)->2(tsne)\n\n2dim after prereduction by TSNE corresponds to xy",
    "2316875": "512\\*512\\*3-> 256\\*256\\*3(resize)->384(embed by ViT-S)->2(tsne)\n\n```\nput a slash before \"\\*\" to show\n\n```",
    "2317168": "in my experiments, yolov7-seg is better then yolov8x-seg.\nIt is the 300 epochs that improve my previous results (100 epoch)\n\nnow:\n\nmy results using yolov7-seg at 512:\ngt_wsi2_val_wsi2:map0.5:0.95=0.331 to 358 (difference fold)\n\n\nmy results using yolov8x-seg at 512:\ngt_wsi2_val_wsi2:map0.5:0.95=0.331 to 353 (difference fold)",
    "2317979": "How did you split K-fold, I think it's better to split by WSI, since private test is completely a difference WSI",
    "2321612": "abebe9849 I would like to reproduce your results and check if normalization procedures help to gather dots of different colors in a single group. I want to ask you a few questions.\nHow do You compute embeddings? Do you use only ViT-256 (vit256_small_dino weights) model or maybe ViT-4k too?\nWhat are your TSNE parameters?\nThanks!\n\nEdit:\nTSNE is very sensitive to initialization and other parameters.\n`TSNE(n_components=2, learning_rate='auto', init='random', perplexity=3)`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5506850%2F6b23cd24c1b3098c45be4944c7609464%2Fr2.png?generation=1688026623129144&alt=media)\n\n`TSNE(n_components=2)`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5506850%2F63484c4aa1cac0b50acd3a3e4e49a5da%2Fr1.png?generation=1688026648949098&alt=media)",
    "2324693": "only ViT-256\nperplexity 30",
    "2325497": "check my post.\ni have \"found\" the whole slide for wsi.3.\n\nsee if you can use self-supervised learning (without label) to improve results for private wsi.3.\n\ni think all test (public and private) are from the 30 slides from previous hubmap competition.\nif so the unlabelled data of dataset.3 should help for private test wsi.5",
    "2334945": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb86588abe96ad3f82e1b6945f141a627%2FSelection_999(2575).png?generation=1688797793176032&alt=media)\n\n@abebe9849 \n\nyou probably can be better embedding score by usingopen ousrced histotology pretrained model likw thsi"
  },
  "source": "meta"
}