{
  "id": 342245,
  "title": "HPA unlabeled data (notebook and dataset)",
  "url": "/competitions/hubmap-organ-segmentation/discussion/342245",
  "author_name": "",
  "post_date": "2022-08-06T07:28:38.721442Z",
  "votes": 31,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have downloaded the HPA unlabeled images from its official website using python spider. As the images are grouped by different protein stains, and there are more than 20k protein, so I only downloaded the first 200 proteins' stained images, ~2000 in total. (without lung, I don't think the ground truth of lung segmentation makes sense)</p>\n<p>Here is my spider notebook:<br>\n<a href=\"https://www.kaggle.com/code/carnozhao/hpa-data-download/notebook\" target=\"_blank\">https://www.kaggle.com/code/carnozhao/hpa-data-download/notebook</a></p>\n<p>And saved first 200 protein images dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/carnozhao/hpa-unlabeled-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/carnozhao/hpa-unlabeled-dataset</a></p>\n<p>From my own results, both my CV and LB did not improve after using pseudo-label on this. Hope this can help you if you have other brilliant ideas.</p>\n<blockquote>\n  <p><strong>Please be careful not to use spider too excessively</strong>😃<br>\n  I suggest you can use this 2000 images to do a simple experiment on how to use this to make a considerable improvement on your LB. After then you can decide whether to download more images.</p>\n</blockquote>",
  "messages": [
    {
      "id": "1886832",
      "postDate": "08/06/2022 07:28:38",
      "content": "<p>I have downloaded the HPA unlabeled images from its official website using python spider. As the images are grouped by different protein stains, and there are more than 20k protein, so I only downloaded the first 200 proteins' stained images, ~2000 in total. (without lung, I don't think the ground truth of lung segmentation makes sense)</p>\n<p>Here is my spider notebook:<br>\n<a href=\"https://www.kaggle.com/code/carnozhao/hpa-data-download/notebook\" target=\"_blank\">https://www.kaggle.com/code/carnozhao/hpa-data-download/notebook</a></p>\n<p>And saved first 200 protein images dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/carnozhao/hpa-unlabeled-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/carnozhao/hpa-unlabeled-dataset</a></p>\n<p>From my own results, both my CV and LB did not improve after using pseudo-label on this. Hope this can help you if you have other brilliant ideas.</p>\n<blockquote>\n  <p><strong>Please be careful not to use spider too excessively</strong>😃<br>\n  I suggest you can use this 2000 images to do a simple experiment on how to use this to make a considerable improvement on your LB. After then you can decide whether to download more images.</p>\n</blockquote>",
      "rawMarkdown": "I have downloaded the HPA unlabeled images from its official website using python spider. As the images are grouped by different protein stains, and there are more than 20k protein, so I only downloaded the first 200 proteins' stained images, ~2000 in total. (without lung, I don't think the ground truth of lung segmentation makes sense)\n\nHere is my spider notebook:\nhttps://www.kaggle.com/code/carnozhao/hpa-data-download/notebook\n\nAnd saved first 200 protein images dataset:\nhttps://www.kaggle.com/datasets/carnozhao/hpa-unlabeled-dataset\n\nFrom my own results, both my CV and LB did not improve after using pseudo-label on this. Hope this can help you if you have other brilliant ideas.\n\n> **Please be careful not to use spider too excessively**😃\nI suggest you can use this 2000 images to do a simple experiment on how to use this to make a considerable improvement on your LB. After then you can decide whether to download more images.",
      "votes": null
    },
    {
      "id": "1887546",
      "postDate": "08/06/2022 19:48:24",
      "content": "<p>i strongly suggest download images from GTEX portal.<br>\nI believe it will resembles closely to our Hubmap hidden images.</p>\n<p>I also suggest to probe the stain vector of hidden test images.<br>\nthen you can use this information to create the \"same color\" training images.</p>\n<p>i haven't check the copyright of GTEX portal images</p>\n<p><img src=\"https://i.ibb.co/WGb5X5J/Selection-085.png\" alt=\"https://i.ibb.co/WGb5X5J/Selection-085.png\"></p>",
      "rawMarkdown": "i strongly suggest download images from GTEX portal.\nI believe it will resembles closely to our Hubmap hidden images.\n\nI also suggest to probe the stain vector of hidden test images.\nthen you can use this information to create the \"same color\" training images.\n\ni haven't check the copyright of GTEX portal images\n\n![https://i.ibb.co/WGb5X5J/Selection-085.png](https://i.ibb.co/WGb5X5J/Selection-085.png)",
      "votes": null
    },
    {
      "id": "1887733",
      "postDate": "08/07/2022 01:18:59",
      "content": "<p>Yes GTEX is very useful. I'm working on this currently.</p>",
      "rawMarkdown": "Yes GTEX is very useful. I'm working on this currently.",
      "votes": null
    },
    {
      "id": "1887780",
      "postDate": "08/07/2022 03:11:50",
      "content": "<p>i directly use my kaggle validation images as unlabelled train images in mean teacher pseudo label framework for algorithm debug.<br>\nlocal CV results does not improve or get worse.</p>",
      "rawMarkdown": "i directly use my kaggle validation images as unlabelled train images in mean teacher pseudo label framework for algorithm debug.\nlocal CV results does not improve or get worse.",
      "votes": null
    },
    {
      "id": "1887838",
      "postDate": "08/07/2022 05:21:34",
      "content": "<p>How do you preview whole slide images on GTEX portal? It takes ages for me to load the image and timeouts eventually. QuPath is pretty useful though. It loads whole slide images in couple seconds.</p>",
      "rawMarkdown": "How do you preview whole slide images on GTEX portal? It takes ages for me to load the image and timeouts eventually. QuPath is pretty useful though. It loads whole slide images in couple seconds.",
      "votes": null
    },
    {
      "id": "1887873",
      "postDate": "08/07/2022 06:29:10",
      "content": "<p>i just download the svs file without preview</p>",
      "rawMarkdown": "i just download the svs file without preview",
      "votes": null
    },
    {
      "id": "1887955",
      "postDate": "08/07/2022 08:15:40",
      "content": "<p>i directly use my kaggle validation images as unlabelled train images in mean teacher pseudo label framework for algorithm debug.<br>\nlocal CV results does not improve or get worse.</p>",
      "rawMarkdown": "i directly use my kaggle validation images as unlabelled train images in mean teacher pseudo label framework for algorithm debug.\nlocal CV results does not improve or get worse.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1887546,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/06/2022 19:48:24",
      "content": "<p>i strongly suggest download images from GTEX portal.<br>\nI believe it will resembles closely to our Hubmap hidden images.</p>\n<p>I also suggest to probe the stain vector of hidden test images.<br>\nthen you can use this information to create the \"same color\" training images.</p>\n<p>i haven't check the copyright of GTEX portal images</p>\n<p><img src=\"https://i.ibb.co/WGb5X5J/Selection-085.png\" alt=\"https://i.ibb.co/WGb5X5J/Selection-085.png\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1887733,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "08/07/2022 01:18:59",
          "content": "<p>Yes GTEX is very useful. I'm working on this currently.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1887838,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "08/07/2022 05:21:34",
          "content": "<p>How do you preview whole slide images on GTEX portal? It takes ages for me to load the image and timeouts eventually. QuPath is pretty useful though. It loads whole slide images in couple seconds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1887873,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/07/2022 06:29:10",
          "content": "<p>i just download the svs file without preview</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1887780,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/07/2022 03:11:50",
      "content": "<p>i directly use my kaggle validation images as unlabelled train images in mean teacher pseudo label framework for algorithm debug.<br>\nlocal CV results does not improve or get worse.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1887955,
      "author_name": "bilalsuppal",
      "author_url": "",
      "post_date": "08/07/2022 08:15:40",
      "content": "<p>i directly use my kaggle validation images as unlabelled train images in mean teacher pseudo label framework for algorithm debug.<br>\nlocal CV results does not improve or get worse.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1886832": "I have downloaded the HPA unlabeled images from its official website using python spider. As the images are grouped by different protein stains, and there are more than 20k protein, so I only downloaded the first 200 proteins' stained images, ~2000 in total. (without lung, I don't think the ground truth of lung segmentation makes sense)\n\nHere is my spider notebook:\nhttps://www.kaggle.com/code/carnozhao/hpa-data-download/notebook\n\nAnd saved first 200 protein images dataset:\nhttps://www.kaggle.com/datasets/carnozhao/hpa-unlabeled-dataset\n\nFrom my own results, both my CV and LB did not improve after using pseudo-label on this. Hope this can help you if you have other brilliant ideas.\n\n> **Please be careful not to use spider too excessively**😃\nI suggest you can use this 2000 images to do a simple experiment on how to use this to make a considerable improvement on your LB. After then you can decide whether to download more images.",
    "1887546": "i strongly suggest download images from GTEX portal.\nI believe it will resembles closely to our Hubmap hidden images.\n\nI also suggest to probe the stain vector of hidden test images.\nthen you can use this information to create the \"same color\" training images.\n\ni haven't check the copyright of GTEX portal images\n\n![https://i.ibb.co/WGb5X5J/Selection-085.png](https://i.ibb.co/WGb5X5J/Selection-085.png)",
    "1887733": "Yes GTEX is very useful. I'm working on this currently.",
    "1887780": "i directly use my kaggle validation images as unlabelled train images in mean teacher pseudo label framework for algorithm debug.\nlocal CV results does not improve or get worse.",
    "1887838": "How do you preview whole slide images on GTEX portal? It takes ages for me to load the image and timeouts eventually. QuPath is pretty useful though. It loads whole slide images in couple seconds.",
    "1887873": "i just download the svs file without preview",
    "1887955": "i directly use my kaggle validation images as unlabelled train images in mean teacher pseudo label framework for algorithm debug.\nlocal CV results does not improve or get worse."
  },
  "source": "meta"
}