{
  "id": 214905,
  "title": "HPA public data: How to speed up downloading?",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/214905",
  "author_name": "Chenglu",
  "post_date": "2021-01-28T04:40:17.687000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have tried the <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">downloading notebook</a>, but it seems like it will takes a million years to download all of the image, even if using multi threading, because the image is downloaded one by one.</p>\n<p>Is there a faster way to download the external dataset?</p>",
  "messages": [
    {
      "id": 1173714,
      "postDate": "2021-01-28T04:40:17.687Z",
      "content": "<p>I have tried the <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">downloading notebook</a>, but it seems like it will takes a million years to download all of the image, even if using multi threading, because the image is downloaded one by one.</p>\n<p>Is there a faster way to download the external dataset?</p>",
      "rawMarkdown": "I have tried the [downloading notebook](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg), but it seems like it will takes a million years to download all of the image, even if using multi threading, because the image is downloaded one by one.\n\nIs there a faster way to download the external dataset?\n\n\n\n",
      "votes": 5
    },
    {
      "id": 1175052,
      "postDate": "2021-01-28T21:49:17.320Z",
      "content": "<p>The images we provided in the notebooks are high resolution 16-bit tiff images, each channel is around 8MB after gzip (sometimes less), and we have 82495*4 images,  so in total it will be 2 to 2.6TB of data. We are currently looking for ways to improve the situation, but in the meantime, you can try to use the jpeg version which is significantly smaller.</p>\n<p>Basically, if you replace <code>.tif.gz</code> to <code>.jpg</code>, then you can get the jpg version. For example:</p>\n<p><a href=\"https://images.proteinatlas.org/10005/921_B9_1_red.tif.gz\" target=\"_blank\">https://images.proteinatlas.org/10005/921_B9_1_red.tif.gz</a> </p>\n<p>will become</p>\n<p><a href=\"https://images.proteinatlas.org/10005/921_B9_1_red.jpg\" target=\"_blank\">https://images.proteinatlas.org/10005/921_B9_1_red.jpg</a></p>\n<p>As pointed out by <a href=\"https://www.kaggle.com/governor\" target=\"_blank\">@governor</a> you can use the multiprocessing code, but please combine that with our latest .tsv file provided in the notebook, and use the url format mentioned above.</p>\n<p>However, if you do use the jpg images, please be aware the following: <br>\n1) while the tif images are single channel grayscale image, the jpg version is in 8-bit RGB format, you will need to convert RGB color image to grayscale<br>\n2) jpg is a lossy compression format, you will get lower image quality; <br>\n3) the segmentation model we provided are not trained on jpg, it might affect the result. You may want to verify that and potentially finetune the segmentation model on jpg images.</p>",
      "rawMarkdown": "The images we provided in the notebooks are high resolution 16-bit tiff images, each channel is around 8MB after gzip (sometimes less), and we have 82495*4 images,  so in total it will be 2 to 2.6TB of data. We are currently looking for ways to improve the situation, but in the meantime, you can try to use the jpeg version which is significantly smaller.\n\nBasically, if you replace `.tif.gz` to `.jpg`, then you can get the jpg version. For example:\n\nhttps://images.proteinatlas.org/10005/921_B9_1_red.tif.gz \n\nwill become\n\n https://images.proteinatlas.org/10005/921_B9_1_red.jpg\n\nAs pointed out by @governor you can use the multiprocessing code, but please combine that with our latest .tsv file provided in the notebook, and use the url format mentioned above.\n\nHowever, if you do use the jpg images, please be aware the following: \n1) while the tif images are single channel grayscale image, the jpg version is in 8-bit RGB format, you will need to convert RGB color image to grayscale\n2) jpg is a lossy compression format, you will get lower image quality; \n3) the segmentation model we provided are not trained on jpg, it might affect the result. You may want to verify that and potentially finetune the segmentation model on jpg images.",
      "votes": 1
    },
    {
      "id": 1175505,
      "postDate": "2021-01-29T07:54:32.990Z",
      "content": "<p>Hi! I have updated cell lines information here as well, so you can consider just downloading 17 cell lines, sampling for rarer classes, or use jpg like <a href=\"https://www.kaggle.com/weiouyang\" target=\"_blank\">@weiouyang</a> pointed out <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a> </p>",
      "rawMarkdown": "Hi! I have updated cell lines information here as well, so you can consider just downloading 17 cell lines, sampling for rarer classes, or use jpg like @weiouyang pointed out https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg "
    },
    {
      "id": 1174469,
      "postDate": "2021-01-28T13:47:20.897Z",
      "content": "<p>From the last competition there were some scripts using multiprocessing to speed it up. You could probably modify them for this dataset. I am just too lazy to do it myself right now 😅. It is probably just extra slow single threaded because the files are converted to another format.</p>\n<p><a href=\"https://github.com/CellProfiling/HPA-competition/blob/master/download_hpa_dataset.py\" target=\"_blank\">https://github.com/CellProfiling/HPA-competition/blob/master/download_hpa_dataset.py</a><br>\n<a href=\"https://github.com/pudae/kaggle-hpa/blob/master/tools/download.py\" target=\"_blank\">https://github.com/pudae/kaggle-hpa/blob/master/tools/download.py</a></p>",
      "rawMarkdown": "From the last competition there were some scripts using multiprocessing to speed it up. You could probably modify them for this dataset. I am just too lazy to do it myself right now 😅. It is probably just extra slow single threaded because the files are converted to another format.\n\nhttps://github.com/CellProfiling/HPA-competition/blob/master/download_hpa_dataset.py\nhttps://github.com/pudae/kaggle-hpa/blob/master/tools/download.py",
      "replies": [
        {
          "id": 1174536,
          "postDate": "2021-01-28T14:50:07.350Z",
          "content": "<p>Already done the mutiprocessing, still, it'll take a thoudand years to finish 😂</p>",
          "rawMarkdown": "Already done the mutiprocessing, still, it'll take a thoudand years to finish 😂"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1175052,
      "author_name": "Wei Ouyang",
      "author_url": "",
      "post_date": "2021-01-28T21:49:17.320000",
      "content": "<p>The images we provided in the notebooks are high resolution 16-bit tiff images, each channel is around 8MB after gzip (sometimes less), and we have 82495*4 images,  so in total it will be 2 to 2.6TB of data. We are currently looking for ways to improve the situation, but in the meantime, you can try to use the jpeg version which is significantly smaller.</p>\n<p>Basically, if you replace <code>.tif.gz</code> to <code>.jpg</code>, then you can get the jpg version. For example:</p>\n<p><a href=\"https://images.proteinatlas.org/10005/921_B9_1_red.tif.gz\" target=\"_blank\">https://images.proteinatlas.org/10005/921_B9_1_red.tif.gz</a> </p>\n<p>will become</p>\n<p><a href=\"https://images.proteinatlas.org/10005/921_B9_1_red.jpg\" target=\"_blank\">https://images.proteinatlas.org/10005/921_B9_1_red.jpg</a></p>\n<p>As pointed out by <a href=\"https://www.kaggle.com/governor\" target=\"_blank\">@governor</a> you can use the multiprocessing code, but please combine that with our latest .tsv file provided in the notebook, and use the url format mentioned above.</p>\n<p>However, if you do use the jpg images, please be aware the following: <br>\n1) while the tif images are single channel grayscale image, the jpg version is in 8-bit RGB format, you will need to convert RGB color image to grayscale<br>\n2) jpg is a lossy compression format, you will get lower image quality; <br>\n3) the segmentation model we provided are not trained on jpg, it might affect the result. You may want to verify that and potentially finetune the segmentation model on jpg images.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1175505,
      "author_name": "Trang Le",
      "author_url": "",
      "post_date": "2021-01-29T07:54:32.990000",
      "content": "<p>Hi! I have updated cell lines information here as well, so you can consider just downloading 17 cell lines, sampling for rarer classes, or use jpg like <a href=\"https://www.kaggle.com/weiouyang\" target=\"_blank\">@weiouyang</a> pointed out <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1174469,
      "author_name": "bright",
      "author_url": "",
      "post_date": "2021-01-28T13:47:20.897000",
      "content": "<p>From the last competition there were some scripts using multiprocessing to speed it up. You could probably modify them for this dataset. I am just too lazy to do it myself right now 😅. It is probably just extra slow single threaded because the files are converted to another format.</p>\n<p><a href=\"https://github.com/CellProfiling/HPA-competition/blob/master/download_hpa_dataset.py\" target=\"_blank\">https://github.com/CellProfiling/HPA-competition/blob/master/download_hpa_dataset.py</a><br>\n<a href=\"https://github.com/pudae/kaggle-hpa/blob/master/tools/download.py\" target=\"_blank\">https://github.com/pudae/kaggle-hpa/blob/master/tools/download.py</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1174536,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2021-01-28T14:50:07.350000",
          "content": "<p>Already done the mutiprocessing, still, it'll take a thoudand years to finish 😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1173714": "I have tried the [downloading notebook](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg), but it seems like it will takes a million years to download all of the image, even if using multi threading, because the image is downloaded one by one.\n\nIs there a faster way to download the external dataset?\n\n\n\n",
    "1175052": "The images we provided in the notebooks are high resolution 16-bit tiff images, each channel is around 8MB after gzip (sometimes less), and we have 82495*4 images,  so in total it will be 2 to 2.6TB of data. We are currently looking for ways to improve the situation, but in the meantime, you can try to use the jpeg version which is significantly smaller.\n\nBasically, if you replace `.tif.gz` to `.jpg`, then you can get the jpg version. For example:\n\nhttps://images.proteinatlas.org/10005/921_B9_1_red.tif.gz \n\nwill become\n\n https://images.proteinatlas.org/10005/921_B9_1_red.jpg\n\nAs pointed out by @governor you can use the multiprocessing code, but please combine that with our latest .tsv file provided in the notebook, and use the url format mentioned above.\n\nHowever, if you do use the jpg images, please be aware the following: \n1) while the tif images are single channel grayscale image, the jpg version is in 8-bit RGB format, you will need to convert RGB color image to grayscale\n2) jpg is a lossy compression format, you will get lower image quality; \n3) the segmentation model we provided are not trained on jpg, it might affect the result. You may want to verify that and potentially finetune the segmentation model on jpg images.",
    "1175505": "Hi! I have updated cell lines information here as well, so you can consider just downloading 17 cell lines, sampling for rarer classes, or use jpg like @weiouyang pointed out https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg ",
    "1174469": "From the last competition there were some scripts using multiprocessing to speed it up. You could probably modify them for this dataset. I am just too lazy to do it myself right now 😅. It is probably just extra slow single threaded because the files are converted to another format.\n\nhttps://github.com/CellProfiling/HPA-competition/blob/master/download_hpa_dataset.py\nhttps://github.com/pudae/kaggle-hpa/blob/master/tools/download.py"
  }
}