{
  "id": 241658,
  "title": "Did anyone tried running experiments on Colab ?",
  "url": "/competitions/seti-breakthrough-listen/discussion/241658",
  "author_name": "Athar Sayed",
  "post_date": "2021-05-25T14:04:48.035000",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I was actually trying to download data using Kaggle API , to run more experiments on colab but the disk space on Google Colab gets Full and the session crashes , has anyone encountered same problem and if yes how have you solve it ?</p>",
  "messages": [
    {
      "id": 1323277,
      "postDate": "2021-05-26T06:01:02.747Z",
      "content": "<p>I divided the whole dataset by sub directory, have a look:</p>\n<p>Training:<br>\n<a href=\"https://www.kaggle.com/snaker/train-0\" target=\"_blank\">https://www.kaggle.com/snaker/train-0</a><br>\n<a href=\"https://www.kaggle.com/snaker/train-1\" target=\"_blank\">https://www.kaggle.com/snaker/train-1</a><br>\n….<br>\n<a href=\"https://www.kaggle.com/snaker/train-f\" target=\"_blank\">https://www.kaggle.com/snaker/train-f</a></p>\n<p>Testing:<br>\n<a href=\"https://www.kaggle.com/snaker/test-0\" target=\"_blank\">https://www.kaggle.com/snaker/test-0</a><br>\n<a href=\"https://www.kaggle.com/snaker/test-1\" target=\"_blank\">https://www.kaggle.com/snaker/test-1</a><br>\n….<br>\n<a href=\"https://www.kaggle.com/snaker/test-f\" target=\"_blank\">https://www.kaggle.com/snaker/test-f</a></p>\n<p>With this, you can extract a subfolder and then delete the zip file, then download another zip and again.</p>",
      "rawMarkdown": "I divided the whole dataset by sub directory, have a look:\n\nTraining:\nhttps://www.kaggle.com/snaker/train-0\nhttps://www.kaggle.com/snaker/train-1\n....\nhttps://www.kaggle.com/snaker/train-f\n\nTesting:\nhttps://www.kaggle.com/snaker/test-0\nhttps://www.kaggle.com/snaker/test-1\n....\nhttps://www.kaggle.com/snaker/test-f\n\nWith this, you can extract a subfolder and then delete the zip file, then download another zip and again.\n",
      "votes": 7,
      "replies": [
        {
          "id": 1324532,
          "postDate": "2021-05-27T03:45:34.690Z",
          "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> thanks but this looks very arduous though :). Wished that Kaggle zip file was sufficiently big enough to be stored in Colab , Bristol Myers Squibb Competition had large number of files but that was zipped and unzipped properly in Colab enviornments using Kaggle Api</p>",
          "rawMarkdown": "@snaker thanks but this looks very arduous though :). Wished that Kaggle zip file was sufficiently big enough to be stored in Colab , Bristol Myers Squibb Competition had large number of files but that was zipped and unzipped properly in Colab enviornments using Kaggle Api"
        },
        {
          "id": 1324700,
          "postDate": "2021-05-27T06:55:50.543Z",
          "content": "<p>If you need both train and test data, then the space of Colab is not enough. But in this way we can save both train and test data</p>",
          "rawMarkdown": "If you need both train and test data, then the space of Colab is not enough. But in this way we can save both train and test data"
        },
        {
          "id": 1326587,
          "postDate": "2021-05-28T14:55:38.647Z",
          "content": "<p>You are a life savior! tnx</p>",
          "rawMarkdown": "You are a life savior! tnx"
        }
      ]
    },
    {
      "id": 1322558,
      "postDate": "2021-05-25T14:04:48.037Z",
      "content": "<p>I was actually trying to download data using Kaggle API , to run more experiments on colab but the disk space on Google Colab gets Full and the session crashes , has anyone encountered same problem and if yes how have you solve it ?</p>",
      "rawMarkdown": "I was actually trying to download data using Kaggle API , to run more experiments on colab but the disk space on Google Colab gets Full and the session crashes , has anyone encountered same problem and if yes how have you solve it ?",
      "votes": 1
    },
    {
      "id": 1322583,
      "postDate": "2021-05-25T14:19:22.177Z",
      "content": "<p>Hey i am doing all my training on google colab. The secret to using on colab it when the test dataset is extracted you have to delete it. you can see this when you use the zip command. when you see that the test dataset is finished extracting you can use the command <code>rm -rf Path_to_your_dataset/test</code>. Keep in mind that because of this you cannot use the test dataset on colab. Furthermore this trick only works on Colab pro because it has more disk space </p>",
      "rawMarkdown": "Hey i am doing all my training on google colab. The secret to using on colab it when the test dataset is extracted you have to delete it. you can see this when you use the zip command. when you see that the test dataset is finished extracting you can use the command `rm -rf Path_to_your_dataset/test`. Keep in mind that because of this you cannot use the test dataset on colab. Furthermore this trick only works on Colab pro because it has more disk space ",
      "votes": 2,
      "replies": [
        {
          "id": 1322609,
          "postDate": "2021-05-25T14:48:15.170Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1322615,
          "postDate": "2021-05-25T14:51:05.663Z",
          "content": "<p>110 - 120 gb </p>",
          "rawMarkdown": "110 - 120 gb "
        }
      ]
    },
    {
      "id": 1325746,
      "postDate": "2021-05-28T01:30:10.097Z",
      "content": "<p>with GPU-high memory option,<br>\n!unzip seti-breakthrough-listen.zip  -x train/* - cut part of data ( make size half ) &amp; load on memory - !sudo rm -r test - np.memmap() &lt;- test_dataset - del test_dataset</p>\n<p>!unzip seti-breakthrough-listen.zip train/[0-9]/*  - cut part of data ( make size half ) &amp; load on memory - !sudo rm -r train - np.memmap() &lt;- train_dataset1 - del train_dataset1 </p>\n<p>!unzip seti-breakthrough-listen.zip train/[^0-9]/* - !sudo rm -r .zip - cut part of data ( make size half ) &amp; load on memory - !sudo rm -r train - np.memmap() &lt;- train_dataset2 - del train_dataset2</p>\n<p>np.memmap() &lt;- np.concatnate([train_dataset1, train_dataset2], 0)</p>\n<p>these will take 1.3 hours (download 0.25h + unzip &amp; process 1.05h ) </p>\n<p>tf.data.Dataset.from_tensor_slices(np.memmap() array)<br>\nwill lead one cycle data iteration in 240s (18GB memmap load minimum time)</p>\n<p>but the most better thing is<br>\ndownload in local -&gt; preprocessing (reduce size almost half) as memmap file -&gt; upload on kaggle (private dataset) -&gt; download in colab with half size!</p>",
      "rawMarkdown": "with GPU-high memory option,\n!unzip seti-breakthrough-listen.zip  -x train/* - cut part of data ( make size half ) & load on memory - !sudo rm -r test - np.memmap() <- test_dataset - del test_dataset\n  \n!unzip seti-breakthrough-listen.zip train/[0-9]/*  - cut part of data ( make size half ) & load on memory - !sudo rm -r train - np.memmap() <- train_dataset1 - del train_dataset1 \n  \n!unzip seti-breakthrough-listen.zip train/[^0-9]/* - !sudo rm -r .zip - cut part of data ( make size half ) & load on memory - !sudo rm -r train - np.memmap() <- train_dataset2 - del train_dataset2\n\nnp.memmap() <- np.concatnate([train_dataset1, train_dataset2], 0)\n\nthese will take 1.3 hours (download 0.25h + unzip & process 1.05h ) \n\ntf.data.Dataset.from_tensor_slices(np.memmap() array)\nwill lead one cycle data iteration in 240s (18GB memmap load minimum time)\n\nbut the most better thing is\ndownload in local -> preprocessing (reduce size almost half) as memmap file -> upload on kaggle (private dataset) -> download in colab with half size!"
    },
    {
      "id": 1323203,
      "postDate": "2021-05-26T04:44:29.013Z",
      "content": "<p>This can be possible. See if you can limit the use of your data to train or the other option would be to go for Colab Pro. Though you need to subscribe for that you will be able to get larger disk space and also your runtime will be much larger.</p>",
      "rawMarkdown": "This can be possible. See if you can limit the use of your data to train or the other option would be to go for Colab Pro. Though you need to subscribe for that you will be able to get larger disk space and also your runtime will be much larger."
    }
  ],
  "comments": [
    {
      "id": 1323277,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2021-05-26T06:01:02.747000",
      "content": "<p>I divided the whole dataset by sub directory, have a look:</p>\n<p>Training:<br>\n<a href=\"https://www.kaggle.com/snaker/train-0\" target=\"_blank\">https://www.kaggle.com/snaker/train-0</a><br>\n<a href=\"https://www.kaggle.com/snaker/train-1\" target=\"_blank\">https://www.kaggle.com/snaker/train-1</a><br>\n….<br>\n<a href=\"https://www.kaggle.com/snaker/train-f\" target=\"_blank\">https://www.kaggle.com/snaker/train-f</a></p>\n<p>Testing:<br>\n<a href=\"https://www.kaggle.com/snaker/test-0\" target=\"_blank\">https://www.kaggle.com/snaker/test-0</a><br>\n<a href=\"https://www.kaggle.com/snaker/test-1\" target=\"_blank\">https://www.kaggle.com/snaker/test-1</a><br>\n….<br>\n<a href=\"https://www.kaggle.com/snaker/test-f\" target=\"_blank\">https://www.kaggle.com/snaker/test-f</a></p>\n<p>With this, you can extract a subfolder and then delete the zip file, then download another zip and again.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1324532,
          "author_name": "Athar Sayed",
          "author_url": "",
          "post_date": "2021-05-27T03:45:34.690000",
          "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> thanks but this looks very arduous though :). Wished that Kaggle zip file was sufficiently big enough to be stored in Colab , Bristol Myers Squibb Competition had large number of files but that was zipped and unzipped properly in Colab enviornments using Kaggle Api</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1324700,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2021-05-27T06:55:50.543000",
          "content": "<p>If you need both train and test data, then the space of Colab is not enough. But in this way we can save both train and test data</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1326587,
          "author_name": "Adriano Passos",
          "author_url": "",
          "post_date": "2021-05-28T14:55:38.647000",
          "content": "<p>You are a life savior! tnx</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1322583,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-05-25T14:19:22.177000",
      "content": "<p>Hey i am doing all my training on google colab. The secret to using on colab it when the test dataset is extracted you have to delete it. you can see this when you use the zip command. when you see that the test dataset is finished extracting you can use the command <code>rm -rf Path_to_your_dataset/test</code>. Keep in mind that because of this you cannot use the test dataset on colab. Furthermore this trick only works on Colab pro because it has more disk space </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1322609,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-25T14:48:15.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1322615,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-05-25T14:51:05.663000",
          "content": "<p>110 - 120 gb </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1325746,
      "author_name": "assign",
      "author_url": "",
      "post_date": "2021-05-28T01:30:10.097000",
      "content": "<p>with GPU-high memory option,<br>\n!unzip seti-breakthrough-listen.zip  -x train/* - cut part of data ( make size half ) &amp; load on memory - !sudo rm -r test - np.memmap() &lt;- test_dataset - del test_dataset</p>\n<p>!unzip seti-breakthrough-listen.zip train/[0-9]/*  - cut part of data ( make size half ) &amp; load on memory - !sudo rm -r train - np.memmap() &lt;- train_dataset1 - del train_dataset1 </p>\n<p>!unzip seti-breakthrough-listen.zip train/[^0-9]/* - !sudo rm -r .zip - cut part of data ( make size half ) &amp; load on memory - !sudo rm -r train - np.memmap() &lt;- train_dataset2 - del train_dataset2</p>\n<p>np.memmap() &lt;- np.concatnate([train_dataset1, train_dataset2], 0)</p>\n<p>these will take 1.3 hours (download 0.25h + unzip &amp; process 1.05h ) </p>\n<p>tf.data.Dataset.from_tensor_slices(np.memmap() array)<br>\nwill lead one cycle data iteration in 240s (18GB memmap load minimum time)</p>\n<p>but the most better thing is<br>\ndownload in local -&gt; preprocessing (reduce size almost half) as memmap file -&gt; upload on kaggle (private dataset) -&gt; download in colab with half size!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1323203,
      "author_name": "Ankit",
      "author_url": "",
      "post_date": "2021-05-26T04:44:29.013000",
      "content": "<p>This can be possible. See if you can limit the use of your data to train or the other option would be to go for Colab Pro. Though you need to subscribe for that you will be able to get larger disk space and also your runtime will be much larger.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1323277": "I divided the whole dataset by sub directory, have a look:\n\nTraining:\nhttps://www.kaggle.com/snaker/train-0\nhttps://www.kaggle.com/snaker/train-1\n....\nhttps://www.kaggle.com/snaker/train-f\n\nTesting:\nhttps://www.kaggle.com/snaker/test-0\nhttps://www.kaggle.com/snaker/test-1\n....\nhttps://www.kaggle.com/snaker/test-f\n\nWith this, you can extract a subfolder and then delete the zip file, then download another zip and again.\n",
    "1322558": "I was actually trying to download data using Kaggle API , to run more experiments on colab but the disk space on Google Colab gets Full and the session crashes , has anyone encountered same problem and if yes how have you solve it ?",
    "1322583": "Hey i am doing all my training on google colab. The secret to using on colab it when the test dataset is extracted you have to delete it. you can see this when you use the zip command. when you see that the test dataset is finished extracting you can use the command `rm -rf Path_to_your_dataset/test`. Keep in mind that because of this you cannot use the test dataset on colab. Furthermore this trick only works on Colab pro because it has more disk space ",
    "1325746": "with GPU-high memory option,\n!unzip seti-breakthrough-listen.zip  -x train/* - cut part of data ( make size half ) & load on memory - !sudo rm -r test - np.memmap() <- test_dataset - del test_dataset\n  \n!unzip seti-breakthrough-listen.zip train/[0-9]/*  - cut part of data ( make size half ) & load on memory - !sudo rm -r train - np.memmap() <- train_dataset1 - del train_dataset1 \n  \n!unzip seti-breakthrough-listen.zip train/[^0-9]/* - !sudo rm -r .zip - cut part of data ( make size half ) & load on memory - !sudo rm -r train - np.memmap() <- train_dataset2 - del train_dataset2\n\nnp.memmap() <- np.concatnate([train_dataset1, train_dataset2], 0)\n\nthese will take 1.3 hours (download 0.25h + unzip & process 1.05h ) \n\ntf.data.Dataset.from_tensor_slices(np.memmap() array)\nwill lead one cycle data iteration in 240s (18GB memmap load minimum time)\n\nbut the most better thing is\ndownload in local -> preprocessing (reduce size almost half) as memmap file -> upload on kaggle (private dataset) -> download in colab with half size!",
    "1323203": "This can be possible. See if you can limit the use of your data to train or the other option would be to go for Colab Pro. Though you need to subscribe for that you will be able to get larger disk space and also your runtime will be much larger."
  }
}