{
  "id": 232619,
  "title": "How to unzip data efficiently ?",
  "url": "/competitions/bms-molecular-translation/discussion/232619",
  "author_name": "",
  "post_date": "2021-04-14T14:51:06.346295700Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This may seem a very trivial question. I work on Vscode connected to google colab via ssh. So, I have to always download the data once I launch a fresh colab instance. Downloading is fine, as it is done in few minutes. But unzipping takes almost one hour and I can't afford to always do that. I know that we have a way to convert image files in .pkl or .npy format, but I can't find an efficient way to do it.<br>\nAny suggestions or resources would be helpful :)</p>",
  "messages": [
    {
      "id": "1273694",
      "postDate": "04/14/2021 14:51:06",
      "content": "<p>This may seem a very trivial question. I work on Vscode connected to google colab via ssh. So, I have to always download the data once I launch a fresh colab instance. Downloading is fine, as it is done in few minutes. But unzipping takes almost one hour and I can't afford to always do that. I know that we have a way to convert image files in .pkl or .npy format, but I can't find an efficient way to do it.<br>\nAny suggestions or resources would be helpful :)</p>",
      "rawMarkdown": "This may seem a very trivial question. I work on Vscode connected to google colab via ssh. So, I have to always download the data once I launch a fresh colab instance. Downloading is fine, as it is done in few minutes. But unzipping takes almost one hour and I can't afford to always do that. I know that we have a way to convert image files in .pkl or .npy format, but I can't find an efficient way to do it.\nAny suggestions or resources would be helpful :)",
      "votes": null
    },
    {
      "id": "1273718",
      "postDate": "04/14/2021 15:11:03",
      "content": "<pre><code>import zipfile\nzip_ref = zipfile.ZipFile(\"train.zip\", 'r')\nzip_ref.extractall(\"/tmp\")\nzip_ref.close()\n!rm train.zip\n</code></pre>\n<p>It takes 6 minutes for this competition dataset</p>",
      "rawMarkdown": "```\nimport zipfile\n\nzip_ref = zipfile.ZipFile(\"train.zip\", 'r')\nzip_ref.extractall(\"/tmp\")\nzip_ref.close()\n!rm train.zip\n```\n\nIt takes 6 minutes for this competition dataset",
      "votes": null
    },
    {
      "id": "1273786",
      "postDate": "04/14/2021 16:40:08",
      "content": "<pre><code>%%time\nDATA_PATH = './bms-molecular-translation/'\n!mkdir {DATA_PATH}\n!unzip -q bms-molecular-translation.zip \"train/**/*\" -d '/content/bms-molecular-translation/'\n!ls -sh {DATA_PATH}\n</code></pre>\n<p>I tried this it takes around 4 minutes 30 seconds to extract all the train data</p>",
      "rawMarkdown": "```\n%%time\nDATA_PATH = './bms-molecular-translation/'\n!mkdir {DATA_PATH}\n!unzip -q bms-molecular-translation.zip \"train/**/*\" -d '/content/bms-molecular-translation/'\n!ls -sh {DATA_PATH}\n```\n\nI tried this it takes around 4 minutes 30 seconds to extract all the train data",
      "votes": null
    },
    {
      "id": "1273849",
      "postDate": "04/14/2021 17:40:23",
      "content": "<p>use commandline</p>",
      "rawMarkdown": "use commandline",
      "votes": null
    },
    {
      "id": "1273917",
      "postDate": "04/14/2021 18:48:52",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/nitindatta\" target=\"_blank\">@nitindatta</a>. I tried all ways of unzipping but this worked fastest for me. Was able to extract all the train, test, csv files in 6-7 mins. Surprised that normal unzip takes 40 mins and unzipping this way <code>unzip -q bms-molecular-translation.zip \"train/**/*\" -d '/content/bms-molecular-translation/</code> takes 7 mins :)</p>",
      "rawMarkdown": "Thanks @nitindatta. I tried all ways of unzipping but this worked fastest for me. Was able to extract all the train, test, csv files in 6-7 mins. Surprised that normal unzip takes 40 mins and unzipping this way `unzip -q bms-molecular-translation.zip \"train/**/*\" -d '/content/bms-molecular-translation/` takes 7 mins :)",
      "votes": null
    },
    {
      "id": "1275603",
      "postDate": "04/16/2021 13:48:20",
      "content": "<p>Hey. I am trying to do the same in Google Colab mounting by google drive. But it is exceeding 12 minutes.</p>",
      "rawMarkdown": "Hey. I am trying to do the same in Google Colab mounting by google drive. But it is exceeding 12 minutes.",
      "votes": null
    },
    {
      "id": "1275731",
      "postDate": "04/16/2021 16:09:47",
      "content": "<p>Try this <a href=\"https://www.kaggle.com/tawheedrony\" target=\"_blank\">@tawheedrony</a>. Takes me 6 mins.</p>\n<p><code>unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"train/**/*\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'</code></p>\n<p><code>echo \"Unzipping test image files\"</code></p>\n<p><code>unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"test/**/*\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'</code></p>\n<p><code>echo \"Unzipping csv files\"</code></p>\n<p><code>unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"*.csv\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'</code></p>",
      "rawMarkdown": "Try this @tawheedrony. Takes me 6 mins.\n\n`unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"train/**/*\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'`\n\n`echo \"Unzipping test image files\"`\n\n`unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"test/**/*\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'`\n\n`echo \"Unzipping csv files\"`\n\n`unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"*.csv\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'`",
      "votes": null
    },
    {
      "id": "1275759",
      "postDate": "04/16/2021 16:59:25",
      "content": "<p><a href=\"https://www.kaggle.com/tawheedrony\" target=\"_blank\">@tawheedrony</a> <br>\nI use it all times. </p>\n<p>And it takes 6 minutes for this competition train dataset.</p>\n<p>I use Colab Pro though. i don't know if this has some effects. </p>",
      "rawMarkdown": "tawheedrony \nI use it all times. \n\nAnd it takes 6 minutes for this competition train dataset.\n\nI use Colab Pro though. i don't know if this has some effects.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1273718,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "04/14/2021 15:11:03",
      "content": "<pre><code>import zipfile\nzip_ref = zipfile.ZipFile(\"train.zip\", 'r')\nzip_ref.extractall(\"/tmp\")\nzip_ref.close()\n!rm train.zip\n</code></pre>\n<p>It takes 6 minutes for this competition dataset</p>",
      "votes": null,
      "replies": [
        {
          "id": 1275603,
          "author_name": "tawheedrony",
          "author_url": "",
          "post_date": "04/16/2021 13:48:20",
          "content": "<p>Hey. I am trying to do the same in Google Colab mounting by google drive. But it is exceeding 12 minutes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1275731,
          "author_name": "atharvaingle",
          "author_url": "",
          "post_date": "04/16/2021 16:09:47",
          "content": "<p>Try this <a href=\"https://www.kaggle.com/tawheedrony\" target=\"_blank\">@tawheedrony</a>. Takes me 6 mins.</p>\n<p><code>unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"train/**/*\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'</code></p>\n<p><code>echo \"Unzipping test image files\"</code></p>\n<p><code>unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"test/**/*\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'</code></p>\n<p><code>echo \"Unzipping csv files\"</code></p>\n<p><code>unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"*.csv\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1275759,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "04/16/2021 16:59:25",
          "content": "<p><a href=\"https://www.kaggle.com/tawheedrony\" target=\"_blank\">@tawheedrony</a> <br>\nI use it all times. </p>\n<p>And it takes 6 minutes for this competition train dataset.</p>\n<p>I use Colab Pro though. i don't know if this has some effects. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1273786,
      "author_name": "nitindatta",
      "author_url": "",
      "post_date": "04/14/2021 16:40:08",
      "content": "<pre><code>%%time\nDATA_PATH = './bms-molecular-translation/'\n!mkdir {DATA_PATH}\n!unzip -q bms-molecular-translation.zip \"train/**/*\" -d '/content/bms-molecular-translation/'\n!ls -sh {DATA_PATH}\n</code></pre>\n<p>I tried this it takes around 4 minutes 30 seconds to extract all the train data</p>",
      "votes": null,
      "replies": [
        {
          "id": 1273917,
          "author_name": "atharvaingle",
          "author_url": "",
          "post_date": "04/14/2021 18:48:52",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/nitindatta\" target=\"_blank\">@nitindatta</a>. I tried all ways of unzipping but this worked fastest for me. Was able to extract all the train, test, csv files in 6-7 mins. Surprised that normal unzip takes 40 mins and unzipping this way <code>unzip -q bms-molecular-translation.zip \"train/**/*\" -d '/content/bms-molecular-translation/</code> takes 7 mins :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1273849,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/14/2021 17:40:23",
      "content": "<p>use commandline</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1273694": "This may seem a very trivial question. I work on Vscode connected to google colab via ssh. So, I have to always download the data once I launch a fresh colab instance. Downloading is fine, as it is done in few minutes. But unzipping takes almost one hour and I can't afford to always do that. I know that we have a way to convert image files in .pkl or .npy format, but I can't find an efficient way to do it.\nAny suggestions or resources would be helpful :)",
    "1273718": "```\nimport zipfile\n\nzip_ref = zipfile.ZipFile(\"train.zip\", 'r')\nzip_ref.extractall(\"/tmp\")\nzip_ref.close()\n!rm train.zip\n```\n\nIt takes 6 minutes for this competition dataset",
    "1273786": "```\n%%time\nDATA_PATH = './bms-molecular-translation/'\n!mkdir {DATA_PATH}\n!unzip -q bms-molecular-translation.zip \"train/**/*\" -d '/content/bms-molecular-translation/'\n!ls -sh {DATA_PATH}\n```\n\nI tried this it takes around 4 minutes 30 seconds to extract all the train data",
    "1273849": "use commandline",
    "1273917": "Thanks @nitindatta. I tried all ways of unzipping but this worked fastest for me. Was able to extract all the train, test, csv files in 6-7 mins. Surprised that normal unzip takes 40 mins and unzipping this way `unzip -q bms-molecular-translation.zip \"train/**/*\" -d '/content/bms-molecular-translation/` takes 7 mins :)",
    "1275603": "Hey. I am trying to do the same in Google Colab mounting by google drive. But it is exceeding 12 minutes.",
    "1275731": "Try this @tawheedrony. Takes me 6 mins.\n\n`unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"train/**/*\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'`\n\n`echo \"Unzipping test image files\"`\n\n`unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"test/**/*\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'`\n\n`echo \"Unzipping csv files\"`\n\n`unzip -q /content/Bristol-Myers-Squibb-Molecular-Translation/input/bms-molecular-translation.zip \"*.csv\" -d '/content/Bristol-Myers-Squibb-Molecular-Translation/input/'`",
    "1275759": "tawheedrony \nI use it all times. \n\nAnd it takes 6 minutes for this competition train dataset.\n\nI use Colab Pro though. i don't know if this has some effects."
  },
  "source": "meta"
}