{
  "id": 151428,
  "title": "Anyone Tried Google  Colab?",
  "url": "/competitions/alaska2-image-steganalysis/discussion/151428",
  "author_name": "",
  "post_date": "2020-05-15T13:17:55.873463300Z",
  "votes": 7,
  "comment_count": 16,
  "views": 0,
  "content": "<p>i am having some issue with google colab : </p>\n\n<h1>Problem 1 :</h1>\n\n<p>if i do : </p>\n\n<p>!kaggle competitions download   -c alaska2-image-steganalysis \n!unzip /content/alaska2-image-steganalysis.zip -d /content/alaska2-image-steganalysis\n<strong>i get this : content/alaska2-image-steganalysis/Cover/74855.jpg:  write error (disk full?).  Continue? (y/n/^C)</strong> so you can see i am running out of disk space there,but last time i did RSNA competition from colab and that dataset size was 75 gb if i can remember correctly. but with this 30gb data i am running out of disk space,colab disk space is reduced on the other hand it is giving us bit powerful GPU ,previously it offered k80 and now tesla p100 i guess. is there any api command for downloading folder in colab instead of zip file? unzipping costs disk space issue </p>\n\n<h1>problem 2 :</h1>\n\n<p>i downloaded dataset and tried to upload extracted folders in google drive but google drive  crashing even for single folder called \"cover\". so this technique didn't work for me too</p>\n\n<h1>problem 3 :</h1>\n\n<p>tried upload zipped dataset in google drive and it was uploaded successfully there in my google drive (note that i didn't created the zipped dataset,i just simply downloaded the dataset from this competition as a zip file)\nbut from colab when i try to extract this competitions data using this command : \n<strong>!unzip /content/drive/\"My Drive\"/alaska2-image-steganalysis.zip -d /content/alaska2-image-steganalysis</strong>\ni get these errors : </p>\n\n<p>Archive:  /content/drive/My Drive/alaska2-image-steganalysis.zip\nfile #1:  bad zipfile offset (lseek):  0\nfile #2:  bad zipfile offset (lseek):  98304\nfile #3:  bad zipfile offset (lseek):  253952\nfile #4:  bad zipfile offset (lseek):  344064\nfile #5:  bad zipfile offset (lseek):  516096\nfile #6:  bad zipfile offset (lseek):  614400\nfile #7:  bad zipfile offset (lseek):  688128\nfile #8:  bad zipfile offset (lseek):  761856\nfile #9:  bad zipfile offset (lseek):  827392\nfile #10:  bad zipfile offset (lseek):  950272\nfile #11:  bad zipfile offset (lseek):  1015808\nfile #12:  bad zipfile offset (lseek):  1245184\nfile #13:  bad zipfile offset (lseek):  1359872\nfile #14:  bad zipfile offset (lseek):  1499136</p>\n\n<p>i searched in internet but couldn't find useful solutions,anyone help please? thank you</p>",
  "messages": [
    {
      "id": "849053",
      "postDate": "05/15/2020 13:17:55",
      "content": "<p>i am having some issue with google colab : </p>\n\n<h1>Problem 1 :</h1>\n\n<p>if i do : </p>\n\n<p>!kaggle competitions download   -c alaska2-image-steganalysis \n!unzip /content/alaska2-image-steganalysis.zip -d /content/alaska2-image-steganalysis\n<strong>i get this : content/alaska2-image-steganalysis/Cover/74855.jpg:  write error (disk full?).  Continue? (y/n/^C)</strong> so you can see i am running out of disk space there,but last time i did RSNA competition from colab and that dataset size was 75 gb if i can remember correctly. but with this 30gb data i am running out of disk space,colab disk space is reduced on the other hand it is giving us bit powerful GPU ,previously it offered k80 and now tesla p100 i guess. is there any api command for downloading folder in colab instead of zip file? unzipping costs disk space issue </p>\n\n<h1>problem 2 :</h1>\n\n<p>i downloaded dataset and tried to upload extracted folders in google drive but google drive  crashing even for single folder called \"cover\". so this technique didn't work for me too</p>\n\n<h1>problem 3 :</h1>\n\n<p>tried upload zipped dataset in google drive and it was uploaded successfully there in my google drive (note that i didn't created the zipped dataset,i just simply downloaded the dataset from this competition as a zip file)\nbut from colab when i try to extract this competitions data using this command : \n<strong>!unzip /content/drive/\"My Drive\"/alaska2-image-steganalysis.zip -d /content/alaska2-image-steganalysis</strong>\ni get these errors : </p>\n\n<p>Archive:  /content/drive/My Drive/alaska2-image-steganalysis.zip\nfile #1:  bad zipfile offset (lseek):  0\nfile #2:  bad zipfile offset (lseek):  98304\nfile #3:  bad zipfile offset (lseek):  253952\nfile #4:  bad zipfile offset (lseek):  344064\nfile #5:  bad zipfile offset (lseek):  516096\nfile #6:  bad zipfile offset (lseek):  614400\nfile #7:  bad zipfile offset (lseek):  688128\nfile #8:  bad zipfile offset (lseek):  761856\nfile #9:  bad zipfile offset (lseek):  827392\nfile #10:  bad zipfile offset (lseek):  950272\nfile #11:  bad zipfile offset (lseek):  1015808\nfile #12:  bad zipfile offset (lseek):  1245184\nfile #13:  bad zipfile offset (lseek):  1359872\nfile #14:  bad zipfile offset (lseek):  1499136</p>\n\n<p>i searched in internet but couldn't find useful solutions,anyone help please? thank you</p>",
      "rawMarkdown": "i am having some issue with google colab : \n# Problem 1 : \nif i do : \n\n!kaggle competitions download   -c alaska2-image-steganalysis \n!unzip /content/alaska2-image-steganalysis.zip -d /content/alaska2-image-steganalysis\n**i get this : content/alaska2-image-steganalysis/Cover/74855.jpg:  write error (disk full?).  Continue? (y/n/^C)** so you can see i am running out of disk space there,but last time i did RSNA competition from colab and that dataset size was 75 gb if i can remember correctly. but with this 30gb data i am running out of disk space,colab disk space is reduced on the other hand it is giving us bit powerful GPU ,previously it offered k80 and now tesla p100 i guess. is there any api command for downloading folder in colab instead of zip file? unzipping costs disk space issue \n\n# problem 2 : \ni downloaded dataset and tried to upload extracted folders in google drive but google drive  crashing even for single folder called \"cover\". so this technique didn't work for me too\n\n# problem 3 : \ntried upload zipped dataset in google drive and it was uploaded successfully there in my google drive (note that i didn't created the zipped dataset,i just simply downloaded the dataset from this competition as a zip file)\nbut from colab when i try to extract this competitions data using this command : \n**!unzip /content/drive/\"My Drive\"/alaska2-image-steganalysis.zip -d /content/alaska2-image-steganalysis**\ni get these errors : \n\nArchive:  /content/drive/My Drive/alaska2-image-steganalysis.zip\nfile #1:  bad zipfile offset (lseek):  0\nfile #2:  bad zipfile offset (lseek):  98304\nfile #3:  bad zipfile offset (lseek):  253952\nfile #4:  bad zipfile offset (lseek):  344064\nfile #5:  bad zipfile offset (lseek):  516096\nfile #6:  bad zipfile offset (lseek):  614400\nfile #7:  bad zipfile offset (lseek):  688128\nfile #8:  bad zipfile offset (lseek):  761856\nfile #9:  bad zipfile offset (lseek):  827392\nfile #10:  bad zipfile offset (lseek):  950272\nfile #11:  bad zipfile offset (lseek):  1015808\nfile #12:  bad zipfile offset (lseek):  1245184\nfile #13:  bad zipfile offset (lseek):  1359872\nfile #14:  bad zipfile offset (lseek):  1499136\n\ni searched in internet but couldn't find useful solutions,anyone help please? thank you",
      "votes": null
    },
    {
      "id": "849596",
      "postDate": "05/15/2020 23:11:58",
      "content": "<p>Been there before.\nMy opinion:\n1. No solution, you just used up the quote size. I use GCS+TPU instead , kaggle already provided the GCS data: </p>\n\n<p>from kaggle_datasets import KaggleDatasets    </p>\n\n<p>GCS_DS_PATH = KaggleDatasets().get_gcs_path()</p>\n\n<ol>\n<li>Didn't do that , but you should do this in colab console command line mode, I did so , and I met problem 3.</li>\n<li>No solution too. I think may be the size of the google drive is dynamically allocated, and the IO speed of google drive is  lower than the upzip files writing speed , it just can't handle the high speed writing .</li>\n</ol>",
      "rawMarkdown": "Been there before.\nMy opinion:\n1. No solution, you just used up the quote size. I use GCS+TPU instead , kaggle already provided the GCS data: \n\n   from kaggle_datasets import KaggleDatasets    \n\n   GCS_DS_PATH = KaggleDatasets().get_gcs_path()\n\n2. Didn't do that , but you should do this in colab console command line mode, I did so , and I met problem 3.\n3. No solution too. I think may be the size of the google drive is dynamically allocated, and the IO speed of google drive is  lower than the upzip files writing speed , it just can't handle the high speed writing .",
      "votes": null
    },
    {
      "id": "880195",
      "postDate": "06/10/2020 05:01:42",
      "content": "<p>did you solved your issue?</p>",
      "rawMarkdown": "did you solved your issue?",
      "votes": null
    },
    {
      "id": "880965",
      "postDate": "06/10/2020 17:05:28",
      "content": "<p>no <a href=\"/alizasubedi\">@alizasubedi</a> \nno  solution</p>",
      "rawMarkdown": "no @alizasubedi \nno  solution",
      "votes": null
    },
    {
      "id": "881452",
      "postDate": "06/11/2020 03:49:44",
      "content": "<p>I believe this will help you\n<a href=\"https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb\">Steganalysis: ALASKA 2 | JPEG compression and Sampling using TF</a>\n```\ninput_directory = '/content/'\nlocation_dir = {'Cover':&gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'JMiPOD':       &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'JUNIWARD':  &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'UERD':          &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'Test':             &gt;&gt;The URL you get when you click on download&lt;&lt;}</p>\n\n<p>def downloader_unziper(key, url):\n  from tensorflow.keras.utils import get_file\n  get_file(\n        key, origin=url,\n        cache_subdir=input_directory+key, extract=True,\n        archive_format='auto', cache_dir=None\n    )</p>\n\n<p>for key, url in location_dir.items():\n  downloader_unziper(key, url) \n```</p>\n\n<p>Be patient, it Colab get's quite slow while unzipping.</p>",
      "rawMarkdown": "I believe this will help you\n[Steganalysis: ALASKA 2 | JPEG compression and Sampling using TF](https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb)\n```\ninput_directory = '/content/'\nlocation_dir = {'Cover':&gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'JMiPOD':       &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'JUNIWARD':  &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'UERD':          &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'Test':             &gt;&gt;The URL you get when you click on download&lt;&lt;}\n\ndef downloader_unziper(key, url):\n  from tensorflow.keras.utils import get_file\n  get_file(\n        key, origin=url,\n        cache_subdir=input_directory+key, extract=True,\n        archive_format='auto', cache_dir=None\n    )\n\nfor key, url in location_dir.items():\n  downloader_unziper(key, url) \n```\n\nBe patient, it Colab get's quite slow while unzipping.",
      "votes": null
    },
    {
      "id": "882267",
      "postDate": "06/11/2020 17:13:24",
      "content": "<p>Yes,  you can try with my Gist <a href=\"https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb\">https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb</a></p>",
      "rawMarkdown": "Yes,  you can try with my Gist https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb",
      "votes": null
    },
    {
      "id": "882283",
      "postDate": "06/11/2020 17:28:25",
      "content": "<p>thank you <a href=\"/echo9k\">@echo9k</a> \nsorry i missed your comment, thank you a lot for your comment and great work,i have 2 questions : \n1. how long unzipping took? approximately?\n2. location_dir = {'Cover':  'URL you when after clicking on Download.zip',\n               'JMiPOD':  'URL you when after clicking on Download.zip',\n               'JUNIWARD':'URL you when after clicking on Download.zip',\n               'UERD':    'URL you when after clicking on Download.zip',\n               'Test':    'URL you when after clicking on Download.zip'}</p>\n\n<p>here <strong>URL you when after clicking on Download.zip</strong> will be exact url  of google drive where each folder is zipped?</p>\n\n<p>i have whole data in colab as a single zip file and inside that i have each folder, i guess it is not possible to zip entire zip file instead i will have to try folder by folder,am i right?</p>",
      "rawMarkdown": "thank you @echo9k \nsorry i missed your comment, thank you a lot for your comment and great work,i have 2 questions : \n1. how long unzipping took? approximately?\n2. location_dir = {'Cover':  'URL you when after clicking on Download.zip',\n               'JMiPOD':  'URL you when after clicking on Download.zip',\n               'JUNIWARD':'URL you when after clicking on Download.zip',\n               'UERD':    'URL you when after clicking on Download.zip',\n               'Test':    'URL you when after clicking on Download.zip'}\n\n\nhere **URL you when after clicking on Download.zip** will be exact url  of google drive where each folder is zipped?\n\ni have whole data in colab as a single zip file and inside that i have each folder, i guess it is not possible to zip entire zip file instead i will have to try folder by folder,am i right?",
      "votes": null
    },
    {
      "id": "882395",
      "postDate": "06/11/2020 19:07:07",
      "content": "<p>You welcome <a href=\"/mobassir\">@mobassir</a> </p>\n\n<p>No problemo</p>\n\n<ol>\n<li>Around 15 minutes. It's not that much, but enough to be careful with the session. </li>\n<li>The exact URL: What I did was to click on the download button of each file, cancel the download. Then \"copy the link address\"  that shows on your downloads tab (\"chrome://downloads/\")</li>\n</ol>\n\n<p>Yeah, otherwise you don't have enough space to download and unzip.</p>",
      "rawMarkdown": "You welcome @mobassir \n\nNo problemo\n\n1.  Around 15 minutes. It's not that much, but enough to be careful with the session. \n2. The exact URL: What I did was to click on the download button of each file, cancel the download. Then \"copy the link address\"  that shows on your downloads tab (\"chrome://downloads/\")\n\nYeah, otherwise you don't have enough space to download and unzip.",
      "votes": null
    },
    {
      "id": "900193",
      "postDate": "06/24/2020 17:18:23",
      "content": "<p><a href=\"/mobassir\">@mobassir</a> did you solved the  issue? I also meet this probem.</p>",
      "rawMarkdown": "mobassir did you solved the  issue? I also meet this probem.",
      "votes": null
    },
    {
      "id": "911302",
      "postDate": "07/01/2020 16:46:36",
      "content": "<p>Is there any specific reason to stick with colab or kaggle for computing? If it's not sufficient then explore cloud options. I know they can be costly but colab and kaggle won't offer you the desired requirements at your end in terms of storage and even no guarantee of GPUs. </p>\n\n<p>If you are short on budget then I'd recommend exploring this peer to peer computing network Q Blocks for your <a href=\"https://www.qblocks.cloud\">GPU computing</a> needs. Upto 10X less costly GPU instances but pre-configured with AI frameworks like tf and torch and ready to launch jupyter notebooks. Storage upto 150Gb in an instance. Guess that serves the purpose.</p>",
      "rawMarkdown": "Is there any specific reason to stick with colab or kaggle for computing? If it's not sufficient then explore cloud options. I know they can be costly but colab and kaggle won't offer you the desired requirements at your end in terms of storage and even no guarantee of GPUs. \n\nIf you are short on budget then I'd recommend exploring this peer to peer computing network Q Blocks for your [GPU computing](https://www.qblocks.cloud) needs. Upto 10X less costly GPU instances but pre-configured with AI frameworks like tf and torch and ready to launch jupyter notebooks. Storage upto 150Gb in an instance. Guess that serves the purpose.",
      "votes": null
    },
    {
      "id": "912204",
      "postDate": "07/02/2020 09:57:07",
      "content": "<p><a href=\"/dandingclam\">@dandingclam</a> have you solved the problem?\nIf yes, can you share the commands?</p>",
      "rawMarkdown": "dandingclam have you solved the problem?\nIf yes, can you share the commands?",
      "votes": null
    },
    {
      "id": "912206",
      "postDate": "07/02/2020 09:59:26",
      "content": "<p><a href=\"/echo9k\">@echo9k</a> Colab link seems dead</p>",
      "rawMarkdown": "echo9k Colab link seems dead",
      "votes": null
    },
    {
      "id": "912209",
      "postDate": "07/02/2020 10:07:49",
      "content": "<p>But the code works. You can try that out by creating your own colab notebook. </p>",
      "rawMarkdown": "But the code works. You can try that out by creating your own colab notebook.",
      "votes": null
    },
    {
      "id": "912214",
      "postDate": "07/02/2020 10:12:05",
      "content": "<p>No solution to upzip the data. So i try to read the image from .zip now.\n<code>def get_np_array_from_zip_ref(zip_file):\n       return np.asarray(bytearray(zip_file.read()), dtype=np.uint8)</code>\n<code>image=cv2.imdecode(get_np_array_from_zip_ref(zipfile.ZipFile.open(f'{kind}/{image_name}')),cv2.IMREAD_COLOR)</code></p>",
      "rawMarkdown": "No solution to upzip the data. So i try to read the image from .zip now.\n`def get_np_array_from_zip_ref(zip_file):\n       return np.asarray(bytearray(zip_file.read()), dtype=np.uint8)`\n`image=cv2.imdecode(get_np_array_from_zip_ref(zipfile.ZipFile.open(f'{kind}/{image_name}')),cv2.IMREAD_COLOR)`",
      "votes": null
    },
    {
      "id": "912215",
      "postDate": "07/02/2020 10:12:57",
      "content": "<p><a href=\"/urvishp80\">@urvishp80</a> Thanks. \nCan you confirm <code>The URL you get when you click on download</code> link for <code>Cover</code> folder looks like this one ?</p>\n\n<p><a href=\"https://storage.googleapis.com/kaggle-competitions-data/kaggle-v2/19991/1117522/compressed/Cover.zip?GoogleAccessId=web-data@kaggle-161607.iam.gserviceaccount.com&amp;Expires=1593943940&amp;Signature=DPV4ZktF6to9OuFniSZbJ%2BFyW8r0wRIgvBtRrWe9kSFTlSXPcDvdlHeFNtttut6zfCQzceEGs79sYA8DtfP2VJGK%2FosqV0ZdBf3mATVDu9ClO2eYf4Egs6o%2BSFPgjlI7J097%2FaFvGTWXKymso99z5rnGOrnv9PPEPsuC3WRIvqKPemwC3wi8SdErFIw7k38miSKoA%2BH%2FeCcaxKEPBAWoBxspKT73t3WGb1pc3K%2B36jdVuVsKBd9qg%2Bb%2FP4I9KQV9KTRexhzYLiinGzQCdUI2MUiuU2ywpSVwkBs0g%2F7RftUnks3QiISIxZb7q4BdEbai4vtkdTOs96rpl9f9YfXPKQ%3D%3D&amp;response-content-disposition=attachment%3B+filename%3DCover.zip\">https://storage.googleapis.com/kaggle-competitions-data/kaggle-v2/19991/1117522/compressed/Cover.zip?GoogleAccessId=web-data@kaggle-161607.iam.gserviceaccount.com&amp;Expires=1593943940&amp;Signature=DPV4ZktF6to9OuFniSZbJ%2BFyW8r0wRIgvBtRrWe9kSFTlSXPcDvdlHeFNtttut6zfCQzceEGs79sYA8DtfP2VJGK%2FosqV0ZdBf3mATVDu9ClO2eYf4Egs6o%2BSFPgjlI7J097%2FaFvGTWXKymso99z5rnGOrnv9PPEPsuC3WRIvqKPemwC3wi8SdErFIw7k38miSKoA%2BH%2FeCcaxKEPBAWoBxspKT73t3WGb1pc3K%2B36jdVuVsKBd9qg%2Bb%2FP4I9KQV9KTRexhzYLiinGzQCdUI2MUiuU2ywpSVwkBs0g%2F7RftUnks3QiISIxZb7q4BdEbai4vtkdTOs96rpl9f9YfXPKQ%3D%3D&amp;response-content-disposition=attachment%3B+filename%3DCover.zip</a></p>\n\n<p>EDIT: It's working.</p>",
      "rawMarkdown": "urvishp80 Thanks. \nCan you confirm `The URL you get when you click on download` link for `Cover` folder looks like this one ?\n\nhttps://storage.googleapis.com/kaggle-competitions-data/kaggle-v2/19991/1117522/compressed/Cover.zip?GoogleAccessId=web-data@kaggle-161607.iam.gserviceaccount.com&amp;Expires=1593943940&amp;Signature=DPV4ZktF6to9OuFniSZbJ%2BFyW8r0wRIgvBtRrWe9kSFTlSXPcDvdlHeFNtttut6zfCQzceEGs79sYA8DtfP2VJGK%2FosqV0ZdBf3mATVDu9ClO2eYf4Egs6o%2BSFPgjlI7J097%2FaFvGTWXKymso99z5rnGOrnv9PPEPsuC3WRIvqKPemwC3wi8SdErFIw7k38miSKoA%2BH%2FeCcaxKEPBAWoBxspKT73t3WGb1pc3K%2B36jdVuVsKBd9qg%2Bb%2FP4I9KQV9KTRexhzYLiinGzQCdUI2MUiuU2ywpSVwkBs0g%2F7RftUnks3QiISIxZb7q4BdEbai4vtkdTOs96rpl9f9YfXPKQ%3D%3D&amp;response-content-disposition=attachment%3B+filename%3DCover.zip\n\nEDIT: It's working.",
      "votes": null
    },
    {
      "id": "912252",
      "postDate": "07/02/2020 10:57:10",
      "content": "<p>Have anyone used Google colab for this solution?\nI may be missing something but it seems, GPU only have ~70 GB in total(37 GB available) of disk space. Which I don't think is enough for this competition.</p>",
      "rawMarkdown": "Have anyone used Google colab for this solution?\nI may be missing something but it seems, GPU only have ~70 GB in total(37 GB available) of disk space. Which I don't think is enough for this competition.",
      "votes": null
    },
    {
      "id": "912690",
      "postDate": "07/02/2020 16:48:05",
      "content": "<p>Copy the download link of the individual file/folder paste it on tf.keras.utlis.getfile\n<code>\nkeras.utils.get_file(\n    fname = 'name_of_file',\n    origin = 'url_link',\n   cache_dir=\"location where you want to save\",\n    extract=True,\n)\n</code>\nHope this works</p>",
      "rawMarkdown": "Copy the download link of the individual file/folder paste it on tf.keras.utlis.getfile\n```\nkeras.utils.get_file(\n    fname = 'name_of_file',\n    origin = 'url_link',\n   cache_dir=\"location where you want to save\",\n    extract=True,\n)\n```\nHope this works",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 849596,
      "author_name": "zhenglv",
      "author_url": "",
      "post_date": "05/15/2020 23:11:58",
      "content": "<p>Been there before.\nMy opinion:\n1. No solution, you just used up the quote size. I use GCS+TPU instead , kaggle already provided the GCS data: </p>\n\n<p>from kaggle_datasets import KaggleDatasets    </p>\n\n<p>GCS_DS_PATH = KaggleDatasets().get_gcs_path()</p>\n\n<ol>\n<li>Didn't do that , but you should do this in colab console command line mode, I did so , and I met problem 3.</li>\n<li>No solution too. I think may be the size of the google drive is dynamically allocated, and the IO speed of google drive is  lower than the upzip files writing speed , it just can't handle the high speed writing .</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 880195,
      "author_name": "alizasubedi",
      "author_url": "",
      "post_date": "06/10/2020 05:01:42",
      "content": "<p>did you solved your issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 880965,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "06/10/2020 17:05:28",
          "content": "<p>no <a href=\"/alizasubedi\">@alizasubedi</a> \nno  solution</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 882267,
          "author_name": "echo9k",
          "author_url": "",
          "post_date": "06/11/2020 17:13:24",
          "content": "<p>Yes,  you can try with my Gist <a href=\"https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb\">https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 881452,
      "author_name": "echo9k",
      "author_url": "",
      "post_date": "06/11/2020 03:49:44",
      "content": "<p>I believe this will help you\n<a href=\"https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb\">Steganalysis: ALASKA 2 | JPEG compression and Sampling using TF</a>\n```\ninput_directory = '/content/'\nlocation_dir = {'Cover':&gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'JMiPOD':       &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'JUNIWARD':  &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'UERD':          &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'Test':             &gt;&gt;The URL you get when you click on download&lt;&lt;}</p>\n\n<p>def downloader_unziper(key, url):\n  from tensorflow.keras.utils import get_file\n  get_file(\n        key, origin=url,\n        cache_subdir=input_directory+key, extract=True,\n        archive_format='auto', cache_dir=None\n    )</p>\n\n<p>for key, url in location_dir.items():\n  downloader_unziper(key, url) \n```</p>\n\n<p>Be patient, it Colab get's quite slow while unzipping.</p>",
      "votes": null,
      "replies": [
        {
          "id": 882283,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "06/11/2020 17:28:25",
          "content": "<p>thank you <a href=\"/echo9k\">@echo9k</a> \nsorry i missed your comment, thank you a lot for your comment and great work,i have 2 questions : \n1. how long unzipping took? approximately?\n2. location_dir = {'Cover':  'URL you when after clicking on Download.zip',\n               'JMiPOD':  'URL you when after clicking on Download.zip',\n               'JUNIWARD':'URL you when after clicking on Download.zip',\n               'UERD':    'URL you when after clicking on Download.zip',\n               'Test':    'URL you when after clicking on Download.zip'}</p>\n\n<p>here <strong>URL you when after clicking on Download.zip</strong> will be exact url  of google drive where each folder is zipped?</p>\n\n<p>i have whole data in colab as a single zip file and inside that i have each folder, i guess it is not possible to zip entire zip file instead i will have to try folder by folder,am i right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 882395,
          "author_name": "echo9k",
          "author_url": "",
          "post_date": "06/11/2020 19:07:07",
          "content": "<p>You welcome <a href=\"/mobassir\">@mobassir</a> </p>\n\n<p>No problemo</p>\n\n<ol>\n<li>Around 15 minutes. It's not that much, but enough to be careful with the session. </li>\n<li>The exact URL: What I did was to click on the download button of each file, cancel the download. Then \"copy the link address\"  that shows on your downloads tab (\"chrome://downloads/\")</li>\n</ol>\n\n<p>Yeah, otherwise you don't have enough space to download and unzip.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 912206,
          "author_name": "prashantkikani",
          "author_url": "",
          "post_date": "07/02/2020 09:59:26",
          "content": "<p><a href=\"/echo9k\">@echo9k</a> Colab link seems dead</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 912209,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "07/02/2020 10:07:49",
          "content": "<p>But the code works. You can try that out by creating your own colab notebook. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 912215,
          "author_name": "prashantkikani",
          "author_url": "",
          "post_date": "07/02/2020 10:12:57",
          "content": "<p><a href=\"/urvishp80\">@urvishp80</a> Thanks. \nCan you confirm <code>The URL you get when you click on download</code> link for <code>Cover</code> folder looks like this one ?</p>\n\n<p><a href=\"https://storage.googleapis.com/kaggle-competitions-data/kaggle-v2/19991/1117522/compressed/Cover.zip?GoogleAccessId=web-data@kaggle-161607.iam.gserviceaccount.com&amp;Expires=1593943940&amp;Signature=DPV4ZktF6to9OuFniSZbJ%2BFyW8r0wRIgvBtRrWe9kSFTlSXPcDvdlHeFNtttut6zfCQzceEGs79sYA8DtfP2VJGK%2FosqV0ZdBf3mATVDu9ClO2eYf4Egs6o%2BSFPgjlI7J097%2FaFvGTWXKymso99z5rnGOrnv9PPEPsuC3WRIvqKPemwC3wi8SdErFIw7k38miSKoA%2BH%2FeCcaxKEPBAWoBxspKT73t3WGb1pc3K%2B36jdVuVsKBd9qg%2Bb%2FP4I9KQV9KTRexhzYLiinGzQCdUI2MUiuU2ywpSVwkBs0g%2F7RftUnks3QiISIxZb7q4BdEbai4vtkdTOs96rpl9f9YfXPKQ%3D%3D&amp;response-content-disposition=attachment%3B+filename%3DCover.zip\">https://storage.googleapis.com/kaggle-competitions-data/kaggle-v2/19991/1117522/compressed/Cover.zip?GoogleAccessId=web-data@kaggle-161607.iam.gserviceaccount.com&amp;Expires=1593943940&amp;Signature=DPV4ZktF6to9OuFniSZbJ%2BFyW8r0wRIgvBtRrWe9kSFTlSXPcDvdlHeFNtttut6zfCQzceEGs79sYA8DtfP2VJGK%2FosqV0ZdBf3mATVDu9ClO2eYf4Egs6o%2BSFPgjlI7J097%2FaFvGTWXKymso99z5rnGOrnv9PPEPsuC3WRIvqKPemwC3wi8SdErFIw7k38miSKoA%2BH%2FeCcaxKEPBAWoBxspKT73t3WGb1pc3K%2B36jdVuVsKBd9qg%2Bb%2FP4I9KQV9KTRexhzYLiinGzQCdUI2MUiuU2ywpSVwkBs0g%2F7RftUnks3QiISIxZb7q4BdEbai4vtkdTOs96rpl9f9YfXPKQ%3D%3D&amp;response-content-disposition=attachment%3B+filename%3DCover.zip</a></p>\n\n<p>EDIT: It's working.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 900193,
      "author_name": "dandingclam",
      "author_url": "",
      "post_date": "06/24/2020 17:18:23",
      "content": "<p><a href=\"/mobassir\">@mobassir</a> did you solved the  issue? I also meet this probem.</p>",
      "votes": null,
      "replies": [
        {
          "id": 912204,
          "author_name": "prashantkikani",
          "author_url": "",
          "post_date": "07/02/2020 09:57:07",
          "content": "<p><a href=\"/dandingclam\">@dandingclam</a> have you solved the problem?\nIf yes, can you share the commands?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 912214,
          "author_name": "dandingclam",
          "author_url": "",
          "post_date": "07/02/2020 10:12:05",
          "content": "<p>No solution to upzip the data. So i try to read the image from .zip now.\n<code>def get_np_array_from_zip_ref(zip_file):\n       return np.asarray(bytearray(zip_file.read()), dtype=np.uint8)</code>\n<code>image=cv2.imdecode(get_np_array_from_zip_ref(zipfile.ZipFile.open(f'{kind}/{image_name}')),cv2.IMREAD_COLOR)</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 911302,
      "author_name": "genesis96839",
      "author_url": "",
      "post_date": "07/01/2020 16:46:36",
      "content": "<p>Is there any specific reason to stick with colab or kaggle for computing? If it's not sufficient then explore cloud options. I know they can be costly but colab and kaggle won't offer you the desired requirements at your end in terms of storage and even no guarantee of GPUs. </p>\n\n<p>If you are short on budget then I'd recommend exploring this peer to peer computing network Q Blocks for your <a href=\"https://www.qblocks.cloud\">GPU computing</a> needs. Upto 10X less costly GPU instances but pre-configured with AI frameworks like tf and torch and ready to launch jupyter notebooks. Storage upto 150Gb in an instance. Guess that serves the purpose.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 912252,
      "author_name": "prashantkikani",
      "author_url": "",
      "post_date": "07/02/2020 10:57:10",
      "content": "<p>Have anyone used Google colab for this solution?\nI may be missing something but it seems, GPU only have ~70 GB in total(37 GB available) of disk space. Which I don't think is enough for this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 912690,
      "author_name": "mlneo07",
      "author_url": "",
      "post_date": "07/02/2020 16:48:05",
      "content": "<p>Copy the download link of the individual file/folder paste it on tf.keras.utlis.getfile\n<code>\nkeras.utils.get_file(\n    fname = 'name_of_file',\n    origin = 'url_link',\n   cache_dir=\"location where you want to save\",\n    extract=True,\n)\n</code>\nHope this works</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "849053": "i am having some issue with google colab : \n# Problem 1 : \nif i do : \n\n!kaggle competitions download   -c alaska2-image-steganalysis \n!unzip /content/alaska2-image-steganalysis.zip -d /content/alaska2-image-steganalysis\n**i get this : content/alaska2-image-steganalysis/Cover/74855.jpg:  write error (disk full?).  Continue? (y/n/^C)** so you can see i am running out of disk space there,but last time i did RSNA competition from colab and that dataset size was 75 gb if i can remember correctly. but with this 30gb data i am running out of disk space,colab disk space is reduced on the other hand it is giving us bit powerful GPU ,previously it offered k80 and now tesla p100 i guess. is there any api command for downloading folder in colab instead of zip file? unzipping costs disk space issue \n\n# problem 2 : \ni downloaded dataset and tried to upload extracted folders in google drive but google drive  crashing even for single folder called \"cover\". so this technique didn't work for me too\n\n# problem 3 : \ntried upload zipped dataset in google drive and it was uploaded successfully there in my google drive (note that i didn't created the zipped dataset,i just simply downloaded the dataset from this competition as a zip file)\nbut from colab when i try to extract this competitions data using this command : \n**!unzip /content/drive/\"My Drive\"/alaska2-image-steganalysis.zip -d /content/alaska2-image-steganalysis**\ni get these errors : \n\nArchive:  /content/drive/My Drive/alaska2-image-steganalysis.zip\nfile #1:  bad zipfile offset (lseek):  0\nfile #2:  bad zipfile offset (lseek):  98304\nfile #3:  bad zipfile offset (lseek):  253952\nfile #4:  bad zipfile offset (lseek):  344064\nfile #5:  bad zipfile offset (lseek):  516096\nfile #6:  bad zipfile offset (lseek):  614400\nfile #7:  bad zipfile offset (lseek):  688128\nfile #8:  bad zipfile offset (lseek):  761856\nfile #9:  bad zipfile offset (lseek):  827392\nfile #10:  bad zipfile offset (lseek):  950272\nfile #11:  bad zipfile offset (lseek):  1015808\nfile #12:  bad zipfile offset (lseek):  1245184\nfile #13:  bad zipfile offset (lseek):  1359872\nfile #14:  bad zipfile offset (lseek):  1499136\n\ni searched in internet but couldn't find useful solutions,anyone help please? thank you",
    "849596": "Been there before.\nMy opinion:\n1. No solution, you just used up the quote size. I use GCS+TPU instead , kaggle already provided the GCS data: \n\n   from kaggle_datasets import KaggleDatasets    \n\n   GCS_DS_PATH = KaggleDatasets().get_gcs_path()\n\n2. Didn't do that , but you should do this in colab console command line mode, I did so , and I met problem 3.\n3. No solution too. I think may be the size of the google drive is dynamically allocated, and the IO speed of google drive is  lower than the upzip files writing speed , it just can't handle the high speed writing .",
    "880195": "did you solved your issue?",
    "880965": "no @alizasubedi \nno  solution",
    "881452": "I believe this will help you\n[Steganalysis: ALASKA 2 | JPEG compression and Sampling using TF](https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb)\n```\ninput_directory = '/content/'\nlocation_dir = {'Cover':&gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'JMiPOD':       &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'JUNIWARD':  &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'UERD':          &gt;&gt;The URL you get when you click on download&lt;&lt;,\n               'Test':             &gt;&gt;The URL you get when you click on download&lt;&lt;}\n\ndef downloader_unziper(key, url):\n  from tensorflow.keras.utils import get_file\n  get_file(\n        key, origin=url,\n        cache_subdir=input_directory+key, extract=True,\n        archive_format='auto', cache_dir=None\n    )\n\nfor key, url in location_dir.items():\n  downloader_unziper(key, url) \n```\n\nBe patient, it Colab get's quite slow while unzipping.",
    "882267": "Yes,  you can try with my Gist https://colab.research.google.com/gist/Echo9k/7429e1f91b1ee287f243c298fa5340a9/steganalysis-alaska-2-jpeg-compression-and-sampling-using-tf.ipynb",
    "882283": "thank you @echo9k \nsorry i missed your comment, thank you a lot for your comment and great work,i have 2 questions : \n1. how long unzipping took? approximately?\n2. location_dir = {'Cover':  'URL you when after clicking on Download.zip',\n               'JMiPOD':  'URL you when after clicking on Download.zip',\n               'JUNIWARD':'URL you when after clicking on Download.zip',\n               'UERD':    'URL you when after clicking on Download.zip',\n               'Test':    'URL you when after clicking on Download.zip'}\n\n\nhere **URL you when after clicking on Download.zip** will be exact url  of google drive where each folder is zipped?\n\ni have whole data in colab as a single zip file and inside that i have each folder, i guess it is not possible to zip entire zip file instead i will have to try folder by folder,am i right?",
    "882395": "You welcome @mobassir \n\nNo problemo\n\n1.  Around 15 minutes. It's not that much, but enough to be careful with the session. \n2. The exact URL: What I did was to click on the download button of each file, cancel the download. Then \"copy the link address\"  that shows on your downloads tab (\"chrome://downloads/\")\n\nYeah, otherwise you don't have enough space to download and unzip.",
    "900193": "mobassir did you solved the  issue? I also meet this probem.",
    "911302": "Is there any specific reason to stick with colab or kaggle for computing? If it's not sufficient then explore cloud options. I know they can be costly but colab and kaggle won't offer you the desired requirements at your end in terms of storage and even no guarantee of GPUs. \n\nIf you are short on budget then I'd recommend exploring this peer to peer computing network Q Blocks for your [GPU computing](https://www.qblocks.cloud) needs. Upto 10X less costly GPU instances but pre-configured with AI frameworks like tf and torch and ready to launch jupyter notebooks. Storage upto 150Gb in an instance. Guess that serves the purpose.",
    "912204": "dandingclam have you solved the problem?\nIf yes, can you share the commands?",
    "912206": "echo9k Colab link seems dead",
    "912209": "But the code works. You can try that out by creating your own colab notebook.",
    "912214": "No solution to upzip the data. So i try to read the image from .zip now.\n`def get_np_array_from_zip_ref(zip_file):\n       return np.asarray(bytearray(zip_file.read()), dtype=np.uint8)`\n`image=cv2.imdecode(get_np_array_from_zip_ref(zipfile.ZipFile.open(f'{kind}/{image_name}')),cv2.IMREAD_COLOR)`",
    "912215": "urvishp80 Thanks. \nCan you confirm `The URL you get when you click on download` link for `Cover` folder looks like this one ?\n\nhttps://storage.googleapis.com/kaggle-competitions-data/kaggle-v2/19991/1117522/compressed/Cover.zip?GoogleAccessId=web-data@kaggle-161607.iam.gserviceaccount.com&amp;Expires=1593943940&amp;Signature=DPV4ZktF6to9OuFniSZbJ%2BFyW8r0wRIgvBtRrWe9kSFTlSXPcDvdlHeFNtttut6zfCQzceEGs79sYA8DtfP2VJGK%2FosqV0ZdBf3mATVDu9ClO2eYf4Egs6o%2BSFPgjlI7J097%2FaFvGTWXKymso99z5rnGOrnv9PPEPsuC3WRIvqKPemwC3wi8SdErFIw7k38miSKoA%2BH%2FeCcaxKEPBAWoBxspKT73t3WGb1pc3K%2B36jdVuVsKBd9qg%2Bb%2FP4I9KQV9KTRexhzYLiinGzQCdUI2MUiuU2ywpSVwkBs0g%2F7RftUnks3QiISIxZb7q4BdEbai4vtkdTOs96rpl9f9YfXPKQ%3D%3D&amp;response-content-disposition=attachment%3B+filename%3DCover.zip\n\nEDIT: It's working.",
    "912252": "Have anyone used Google colab for this solution?\nI may be missing something but it seems, GPU only have ~70 GB in total(37 GB available) of disk space. Which I don't think is enough for this competition.",
    "912690": "Copy the download link of the individual file/folder paste it on tf.keras.utlis.getfile\n```\nkeras.utils.get_file(\n    fname = 'name_of_file',\n    origin = 'url_link',\n   cache_dir=\"location where you want to save\",\n    extract=True,\n)\n```\nHope this works"
  },
  "source": "meta"
}