{
  "id": 338066,
  "title": "Can we use spider to download images from HPA website?",
  "url": "/competitions/hubmap-organ-segmentation/discussion/338066",
  "author_name": "Carno Zhao",
  "post_date": "2022-07-19T02:52:34.519000",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>In this <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/337692\" target=\"_blank\">post</a>, it is mentioned that there are many external images with HPA-source stain.</p>\n<p>However, there is no convenient way to download them from multiple tabs. So, does HPA website allow spider? Or is there any limitation of download frequency?</p>",
  "messages": [
    {
      "id": 1861431,
      "postDate": "2022-07-19T02:52:34.520Z",
      "content": "<p>In this <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/337692\" target=\"_blank\">post</a>, it is mentioned that there are many external images with HPA-source stain.</p>\n<p>However, there is no convenient way to download them from multiple tabs. So, does HPA website allow spider? Or is there any limitation of download frequency?</p>",
      "rawMarkdown": "In this [post](https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/337692), it is mentioned that there are many external images with HPA-source stain.\n\nHowever, there is no convenient way to download them from multiple tabs. So, does HPA website allow spider? Or is there any limitation of download frequency?",
      "votes": 6
    },
    {
      "id": 1861594,
      "postDate": "2022-07-19T05:36:22.510Z",
      "content": "<p>Script which I wrote to download the images. One can loop through multiple base urls as well, this is just for testing</p>\n<pre><code>import os\nimport requests\nfrom PIL import Image\nfrom io import BytesIO\nfrom bs4 import BeautifulSoup\n\norgans = ['colon', 'kidney', 'lung', 'prostate', 'spleen']\nbase_url = 'https://www.proteinatlas.org/ENSG00000121410-A1BG'\nsave_path = '/kaggle/working/unlabelled_images/'\nmetadata = []\n\nfor organ in organs:\n    r = requests.get(f'{base_url}/tissue/{organ}')\n    images = BeautifulSoup(r.text, 'html.parser').findAll('img')\n    links = ['https:'+img['src'].replace('_medium', '') for img in images if img['src'].startswith('//images.proteinatlas.org')]\n    print(f'Found {len(links)} images of {organ}')\n\n    for link in links:\n        img_name = link.split('/')[-1]\n        response = requests.get(link)\n        img = Image.open(BytesIO(response.content))\n        img.save(save_path + img_name)\n        metadata.append([img_name, organ, *img.size])\n</code></pre>",
      "rawMarkdown": "Script which I wrote to download the images. One can loop through multiple base urls as well, this is just for testing\n\n```python\nimport os\nimport requests\nfrom PIL import Image\nfrom io import BytesIO\nfrom bs4 import BeautifulSoup\n\norgans = ['colon', 'kidney', 'lung', 'prostate', 'spleen']\nbase_url = 'https://www.proteinatlas.org/ENSG00000121410-A1BG'\nsave_path = '/kaggle/working/unlabelled_images/'\nmetadata = []\n\nfor organ in organs:\n    r = requests.get(f'{base_url}/tissue/{organ}')\n    images = BeautifulSoup(r.text, 'html.parser').findAll('img')\n    links = ['https:'+img['src'].replace('_medium', '') for img in images if img['src'].startswith('//images.proteinatlas.org')]\n    print(f'Found {len(links)} images of {organ}')\n\n    for link in links:\n        img_name = link.split('/')[-1]\n        response = requests.get(link)\n        img = Image.open(BytesIO(response.content))\n        img.save(save_path + img_name)\n        metadata.append([img_name, organ, *img.size])\n```",
      "votes": 4,
      "replies": [
        {
          "id": 1861596,
          "postDate": "2022-07-19T05:39:58.447Z",
          "content": "<p>how you get \"ENSG00000121410-A1BG\"? i remember there is a site to download HPA meta data.</p>",
          "rawMarkdown": "how you get \"ENSG00000121410-A1BG\"? i remember there is a site to download HPA meta data."
        },
        {
          "id": 1861599,
          "postDate": "2022-07-19T05:42:00.817Z",
          "content": "<p>Go to any row here - <a href=\"https://www.proteinatlas.org/search\" target=\"_blank\">https://www.proteinatlas.org/search</a>. You'll find the base url</p>",
          "rawMarkdown": "Go to any row here - https://www.proteinatlas.org/search. You'll find the base url"
        },
        {
          "id": 1861626,
          "postDate": "2022-07-19T06:09:56.790Z",
          "content": "<p>How many images have you donwloaded? Is there any limitation?</p>",
          "rawMarkdown": "How many images have you donwloaded? Is there any limitation?"
        },
        {
          "id": 1861634,
          "postDate": "2022-07-19T06:20:28.133Z",
          "content": "<p>I tested this script for only 3 urls, should be around 50 images.</p>\n<p>Regarding limitation, I don't have an answer. Perhaps the hosts may know more about it. <a href=\"https://www.kaggle.com/yashvrdnjain\" target=\"_blank\">@yashvrdnjain</a></p>",
          "rawMarkdown": "I tested this script for only 3 urls, should be around 50 images.\n\nRegarding limitation, I don't have an answer. Perhaps the hosts may know more about it. @yashvrdnjain"
        },
        {
          "id": 1861644,
          "postDate": "2022-07-19T06:27:12.237Z",
          "content": "<p>Thats cool! It seems that there are thousands of genes, 3-6 images per gene, which is far larger than given training set.</p>\n<p>I suspect the training data also comes from this site, but the hosts have a better way to access these data. Given HPA <strong>has</strong> such a large amount of data, why don't they use then as training set?🤔🤔🤔</p>",
          "rawMarkdown": "Thats cool! It seems that there are thousands of genes, 3-6 images per gene, which is far larger than given training set.\n\nI suspect the training data also comes from this site, but the hosts have a better way to access these data. Given HPA **has** such a large amount of data, why don't they use then as training set?🤔🤔🤔"
        },
        {
          "id": 1862190,
          "postDate": "2022-07-19T14:01:56.937Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1862267,
          "postDate": "2022-07-19T15:04:51.817Z",
          "content": "<p>any idea to download the WSI H&amp;E slide as well?<br>\n<img src=\"https://i.ibb.co/ZfGx1N4/Selection-058.png\" alt=\"https://i.ibb.co/ZfGx1N4/Selection-058.png\"></p>",
          "rawMarkdown": "any idea to download the WSI H&E slide as well?\n![https://i.ibb.co/ZfGx1N4/Selection-058.png](https://i.ibb.co/ZfGx1N4/Selection-058.png)\n"
        },
        {
          "id": 1862272,
          "postDate": "2022-07-19T15:16:39.197Z",
          "content": "<p>I don’t think there is full resolution of these patient-specific WSI. Maybe the only available WSI on HPA is in dictionary tab.</p>",
          "rawMarkdown": "I don’t think there is full resolution of these patient-specific WSI. Maybe the only available WSI on HPA is in dictionary tab."
        }
      ]
    },
    {
      "id": 1862438,
      "postDate": "2022-07-19T17:29:18.767Z",
      "content": "<p>Please refer to the FAQ section on the Human Protein Atlas website for data download: <a href=\"https://www.proteinatlas.org/about/help#19\" target=\"_blank\">https://www.proteinatlas.org/about/help#19</a> </p>",
      "rawMarkdown": "Please refer to the FAQ section on the Human Protein Atlas website for data download: https://www.proteinatlas.org/about/help#19 ",
      "votes": 1,
      "replies": [
        {
          "id": 1862706,
          "postDate": "2022-07-19T23:07:45.170Z",
          "content": "<p>Thanks for your information！</p>",
          "rawMarkdown": "Thanks for your information！"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1861594,
      "author_name": "Jebastin Nadar",
      "author_url": "",
      "post_date": "2022-07-19T05:36:22.510000",
      "content": "<p>Script which I wrote to download the images. One can loop through multiple base urls as well, this is just for testing</p>\n<pre><code>import os\nimport requests\nfrom PIL import Image\nfrom io import BytesIO\nfrom bs4 import BeautifulSoup\n\norgans = ['colon', 'kidney', 'lung', 'prostate', 'spleen']\nbase_url = 'https://www.proteinatlas.org/ENSG00000121410-A1BG'\nsave_path = '/kaggle/working/unlabelled_images/'\nmetadata = []\n\nfor organ in organs:\n    r = requests.get(f'{base_url}/tissue/{organ}')\n    images = BeautifulSoup(r.text, 'html.parser').findAll('img')\n    links = ['https:'+img['src'].replace('_medium', '') for img in images if img['src'].startswith('//images.proteinatlas.org')]\n    print(f'Found {len(links)} images of {organ}')\n\n    for link in links:\n        img_name = link.split('/')[-1]\n        response = requests.get(link)\n        img = Image.open(BytesIO(response.content))\n        img.save(save_path + img_name)\n        metadata.append([img_name, organ, *img.size])\n</code></pre>",
      "votes": 4,
      "replies": [
        {
          "id": 1861596,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-07-19T05:39:58.447000",
          "content": "<p>how you get \"ENSG00000121410-A1BG\"? i remember there is a site to download HPA meta data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1861599,
          "author_name": "Jebastin Nadar",
          "author_url": "",
          "post_date": "2022-07-19T05:42:00.817000",
          "content": "<p>Go to any row here - <a href=\"https://www.proteinatlas.org/search\" target=\"_blank\">https://www.proteinatlas.org/search</a>. You'll find the base url</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1861626,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2022-07-19T06:09:56.790000",
          "content": "<p>How many images have you donwloaded? Is there any limitation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1861634,
          "author_name": "Jebastin Nadar",
          "author_url": "",
          "post_date": "2022-07-19T06:20:28.133000",
          "content": "<p>I tested this script for only 3 urls, should be around 50 images.</p>\n<p>Regarding limitation, I don't have an answer. Perhaps the hosts may know more about it. <a href=\"https://www.kaggle.com/yashvrdnjain\" target=\"_blank\">@yashvrdnjain</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1861644,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2022-07-19T06:27:12.237000",
          "content": "<p>Thats cool! It seems that there are thousands of genes, 3-6 images per gene, which is far larger than given training set.</p>\n<p>I suspect the training data also comes from this site, but the hosts have a better way to access these data. Given HPA <strong>has</strong> such a large amount of data, why don't they use then as training set?🤔🤔🤔</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1862190,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-19T14:01:56.937000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1862267,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-07-19T15:04:51.817000",
          "content": "<p>any idea to download the WSI H&amp;E slide as well?<br>\n<img src=\"https://i.ibb.co/ZfGx1N4/Selection-058.png\" alt=\"https://i.ibb.co/ZfGx1N4/Selection-058.png\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1862272,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2022-07-19T15:16:39.197000",
          "content": "<p>I don’t think there is full resolution of these patient-specific WSI. Maybe the only available WSI on HPA is in dictionary tab.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1862438,
      "author_name": "Yashvardhan Jain",
      "author_url": "",
      "post_date": "2022-07-19T17:29:18.767000",
      "content": "<p>Please refer to the FAQ section on the Human Protein Atlas website for data download: <a href=\"https://www.proteinatlas.org/about/help#19\" target=\"_blank\">https://www.proteinatlas.org/about/help#19</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1862706,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2022-07-19T23:07:45.170000",
          "content": "<p>Thanks for your information！</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1861431": "In this [post](https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/337692), it is mentioned that there are many external images with HPA-source stain.\n\nHowever, there is no convenient way to download them from multiple tabs. So, does HPA website allow spider? Or is there any limitation of download frequency?",
    "1861594": "Script which I wrote to download the images. One can loop through multiple base urls as well, this is just for testing\n\n```python\nimport os\nimport requests\nfrom PIL import Image\nfrom io import BytesIO\nfrom bs4 import BeautifulSoup\n\norgans = ['colon', 'kidney', 'lung', 'prostate', 'spleen']\nbase_url = 'https://www.proteinatlas.org/ENSG00000121410-A1BG'\nsave_path = '/kaggle/working/unlabelled_images/'\nmetadata = []\n\nfor organ in organs:\n    r = requests.get(f'{base_url}/tissue/{organ}')\n    images = BeautifulSoup(r.text, 'html.parser').findAll('img')\n    links = ['https:'+img['src'].replace('_medium', '') for img in images if img['src'].startswith('//images.proteinatlas.org')]\n    print(f'Found {len(links)} images of {organ}')\n\n    for link in links:\n        img_name = link.split('/')[-1]\n        response = requests.get(link)\n        img = Image.open(BytesIO(response.content))\n        img.save(save_path + img_name)\n        metadata.append([img_name, organ, *img.size])\n```",
    "1862438": "Please refer to the FAQ section on the Human Protein Atlas website for data download: https://www.proteinatlas.org/about/help#19 "
  }
}