{
  "id": 146187,
  "title": "Out of Memory on Unzipping",
  "url": "/competitions/quora-insincere-questions-classification/discussion/146187",
  "author_name": "",
  "post_date": "2020-04-26T07:58:37.818989900Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I have been trying to unzip embeddings.zip but it throws out of memory error. Can someone help? It is really frustrating.</p>\n\n<p><code>from zipfile import ZipFile\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as z:\n    z.extract('glove.840B.300d/glove.840B.300d.txt')</code></p>",
  "messages": [
    {
      "id": "821497",
      "postDate": "04/26/2020 07:58:37",
      "content": "<p>I have been trying to unzip embeddings.zip but it throws out of memory error. Can someone help? It is really frustrating.</p>\n\n<p><code>from zipfile import ZipFile\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as z:\n    z.extract('glove.840B.300d/glove.840B.300d.txt')</code></p>",
      "rawMarkdown": "I have been trying to unzip embeddings.zip but it throws out of memory error. Can someone help? It is really frustrating.\n\n`from zipfile import ZipFile\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as z:\n    z.extract('glove.840B.300d/glove.840B.300d.txt')`",
      "votes": null
    },
    {
      "id": "821498",
      "postDate": "04/26/2020 08:00:04",
      "content": "<p>I don't understand how everyone is accessing embeddings without unzipping it. For instance \n`# embdedding setup</p>\n\n<h1>Source <a href=\"https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html\">https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html</a></h1>\n\n<p>embeddings_index = {}\nf = open('/kaggle/input/embeddings/glove.840B.300d/glove.840B.300d.txt')\nfor line in tqdm(f):\n    values = line.split(\" \")\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    embeddings_index[word] = coefs\nf.close()</p>\n\n<p>print('Found %s word vectors.' % len(embeddings_index))`</p>",
      "rawMarkdown": "I don't understand how everyone is accessing embeddings without unzipping it. For instance \n`# embdedding setup\n\n# Source https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html\n\nembeddings_index = {}\nf = open('/kaggle/input/embeddings/glove.840B.300d/glove.840B.300d.txt')\nfor line in tqdm(f):\n    values = line.split(\" \")\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    embeddings_index[word] = coefs\nf.close()\n\nprint('Found %s word vectors.' % len(embeddings_index))`",
      "votes": null
    },
    {
      "id": "821501",
      "postDate": "04/26/2020 08:01:45",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a>  <a href=\"/sudalairajkumar\">@sudalairajkumar</a>  <a href=\"/juliaelliott\">@juliaelliott</a> </p>",
      "rawMarkdown": "christofhenkel  @sudalairajkumar  @juliaelliott",
      "votes": null
    },
    {
      "id": "821520",
      "postDate": "04/26/2020 08:19:26",
      "content": "<p><code>from zipfile import ZipFile\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as z:\n    z.extract('glove.840B.300d/glove.840B.300d.txt')</code></p>",
      "rawMarkdown": "`from zipfile import ZipFile\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as z:\n    z.extract('glove.840B.300d/glove.840B.300d.txt')`",
      "votes": null
    },
    {
      "id": "821581",
      "postDate": "04/26/2020 09:10:33",
      "content": "<p>Found a solution for this problem</p>\n\n<p>`embeddings_index = {}</p>\n\n<p>with ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as myzip:\n    with myzip.open('glove.840B.300d/glove.840B.300d.txt') as myfile:\n        lines = myfile.readlines()\n        for line in tqdm(lines):\n            values = line.decode().split(\" \")\n            word = values[0]\n            coefs = np.asarray(values[1:], dtype='float32')\n            embeddings_index[word] = coefs</p>\n\n<p>print('Found %s word vectors.' % len(embeddings_index))`</p>",
      "rawMarkdown": "Found a solution for this problem\n\n`embeddings_index = {}\n\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as myzip:\n    with myzip.open('glove.840B.300d/glove.840B.300d.txt') as myfile:\n        lines = myfile.readlines()\n        for line in tqdm(lines):\n            values = line.decode().split(\" \")\n            word = values[0]\n            coefs = np.asarray(values[1:], dtype='float32')\n            embeddings_index[word] = coefs\n\nprint('Found %s word vectors.' % len(embeddings_index))`",
      "votes": null
    },
    {
      "id": "867411",
      "postDate": "05/30/2020 08:59:31",
      "content": "<p>Hi, Following your approach, I am able to get the embedding matrix without unzipping. However, it takes up a lot of RAM and later while creating the weights from embedding matrix, script terminates due to lack of RAM space. </p>\n\n<p>Any workaround for that ?</p>",
      "rawMarkdown": "Hi, Following your approach, I am able to get the embedding matrix without unzipping. However, it takes up a lot of RAM and later while creating the weights from embedding matrix, script terminates due to lack of RAM space. \n\nAny workaround for that ?",
      "votes": null
    },
    {
      "id": "868081",
      "postDate": "05/30/2020 22:10:02",
      "content": "<p>Running into the same problem. Frustrating this is.</p>",
      "rawMarkdown": "Running into the same problem. Frustrating this is.",
      "votes": null
    },
    {
      "id": "868244",
      "postDate": "05/31/2020 05:07:39",
      "content": "<p>Hi Keshav,  </p>\n\n<p>You can use the below code.</p>\n\n<p>import io\nimport zipfile</p>\n\n<p>dim=300\nembeddings1_index={}</p>\n\n<p>with zipfile.ZipFile(\"../input/quora-insincere-questions-classification/embeddings.zip\") as zf:\n    with io.TextIOWrapper(zf.open(\"glove.840B.300d/glove.840B.300d.txt\"), encoding=\"utf-8\") as f:\n        for line in tqdm(f):\n            values=line.split(' ') # \".split(' ')\" only for glove-840b-300d; for all other files, \".split()\" works\n            word=values[0]\n            vectors=np.asarray(values[1:],'float32')\n            embeddings1_index[word]=vectors</p>\n\n<p>This works perfectly without using too much RAM</p>",
      "rawMarkdown": "Hi Keshav,  \n\nYou can use the below code.\n\nimport io\nimport zipfile\n\ndim=300\nembeddings1_index={}\n\nwith zipfile.ZipFile(\"../input/quora-insincere-questions-classification/embeddings.zip\") as zf:\n    with io.TextIOWrapper(zf.open(\"glove.840B.300d/glove.840B.300d.txt\"), encoding=\"utf-8\") as f:\n        for line in tqdm(f):\n            values=line.split(' ') # \".split(' ')\" only for glove-840b-300d; for all other files, \".split()\" works\n            word=values[0]\n            vectors=np.asarray(values[1:],'float32')\n            embeddings1_index[word]=vectors\n\nThis works perfectly without using too much RAM",
      "votes": null
    },
    {
      "id": "869027",
      "postDate": "05/31/2020 16:48:19",
      "content": "<p>Hi Nino,\nYes I used a similar code and it worked.  I was running out of memory when trying to convert the text into sequences encoded by the embedding matrix from scratch without using the Tokenizer - I think when we use Tokenizer() to convert the text to sequences, it works just fine. </p>",
      "rawMarkdown": "Hi Nino,\nYes I used a similar code and it worked.  I was running out of memory when trying to convert the text into sequences encoded by the embedding matrix from scratch without using the Tokenizer - I think when we use Tokenizer() to convert the text to sequences, it works just fine.",
      "votes": null
    },
    {
      "id": "1009235",
      "postDate": "09/13/2020 19:02:47",
      "content": "<p>very very helpful</p>",
      "rawMarkdown": "very very helpful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1009235,
      "author_name": "maverickmonk94",
      "author_url": "",
      "post_date": "09/13/2020 19:02:47",
      "content": "<p>very very helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 821498,
      "author_name": "karsal",
      "author_url": "",
      "post_date": "04/26/2020 08:00:04",
      "content": "<p>I don't understand how everyone is accessing embeddings without unzipping it. For instance \n`# embdedding setup</p>\n\n<h1>Source <a href=\"https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html\">https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html</a></h1>\n\n<p>embeddings_index = {}\nf = open('/kaggle/input/embeddings/glove.840B.300d/glove.840B.300d.txt')\nfor line in tqdm(f):\n    values = line.split(\" \")\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    embeddings_index[word] = coefs\nf.close()</p>\n\n<p>print('Found %s word vectors.' % len(embeddings_index))`</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 821501,
      "author_name": "karsal",
      "author_url": "",
      "post_date": "04/26/2020 08:01:45",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a>  <a href=\"/sudalairajkumar\">@sudalairajkumar</a>  <a href=\"/juliaelliott\">@juliaelliott</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 821520,
      "author_name": "karsal",
      "author_url": "",
      "post_date": "04/26/2020 08:19:26",
      "content": "<p><code>from zipfile import ZipFile\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as z:\n    z.extract('glove.840B.300d/glove.840B.300d.txt')</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 821581,
      "author_name": "karsal",
      "author_url": "",
      "post_date": "04/26/2020 09:10:33",
      "content": "<p>Found a solution for this problem</p>\n\n<p>`embeddings_index = {}</p>\n\n<p>with ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as myzip:\n    with myzip.open('glove.840B.300d/glove.840B.300d.txt') as myfile:\n        lines = myfile.readlines()\n        for line in tqdm(lines):\n            values = line.decode().split(\" \")\n            word = values[0]\n            coefs = np.asarray(values[1:], dtype='float32')\n            embeddings_index[word] = coefs</p>\n\n<p>print('Found %s word vectors.' % len(embeddings_index))`</p>",
      "votes": null,
      "replies": [
        {
          "id": 867411,
          "author_name": "fredthered2",
          "author_url": "",
          "post_date": "05/30/2020 08:59:31",
          "content": "<p>Hi, Following your approach, I am able to get the embedding matrix without unzipping. However, it takes up a lot of RAM and later while creating the weights from embedding matrix, script terminates due to lack of RAM space. </p>\n\n<p>Any workaround for that ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868081,
          "author_name": "keshr3106",
          "author_url": "",
          "post_date": "05/30/2020 22:10:02",
          "content": "<p>Running into the same problem. Frustrating this is.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868244,
          "author_name": "fredthered2",
          "author_url": "",
          "post_date": "05/31/2020 05:07:39",
          "content": "<p>Hi Keshav,  </p>\n\n<p>You can use the below code.</p>\n\n<p>import io\nimport zipfile</p>\n\n<p>dim=300\nembeddings1_index={}</p>\n\n<p>with zipfile.ZipFile(\"../input/quora-insincere-questions-classification/embeddings.zip\") as zf:\n    with io.TextIOWrapper(zf.open(\"glove.840B.300d/glove.840B.300d.txt\"), encoding=\"utf-8\") as f:\n        for line in tqdm(f):\n            values=line.split(' ') # \".split(' ')\" only for glove-840b-300d; for all other files, \".split()\" works\n            word=values[0]\n            vectors=np.asarray(values[1:],'float32')\n            embeddings1_index[word]=vectors</p>\n\n<p>This works perfectly without using too much RAM</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869027,
          "author_name": "keshr3106",
          "author_url": "",
          "post_date": "05/31/2020 16:48:19",
          "content": "<p>Hi Nino,\nYes I used a similar code and it worked.  I was running out of memory when trying to convert the text into sequences encoded by the embedding matrix from scratch without using the Tokenizer - I think when we use Tokenizer() to convert the text to sequences, it works just fine. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "821497": "I have been trying to unzip embeddings.zip but it throws out of memory error. Can someone help? It is really frustrating.\n\n`from zipfile import ZipFile\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as z:\n    z.extract('glove.840B.300d/glove.840B.300d.txt')`",
    "821498": "I don't understand how everyone is accessing embeddings without unzipping it. For instance \n`# embdedding setup\n\n# Source https://blog.keras.io/using-pre-trained-word-embeddings-in-a-keras-model.html\n\nembeddings_index = {}\nf = open('/kaggle/input/embeddings/glove.840B.300d/glove.840B.300d.txt')\nfor line in tqdm(f):\n    values = line.split(\" \")\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    embeddings_index[word] = coefs\nf.close()\n\nprint('Found %s word vectors.' % len(embeddings_index))`",
    "821501": "christofhenkel  @sudalairajkumar  @juliaelliott",
    "821520": "`from zipfile import ZipFile\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as z:\n    z.extract('glove.840B.300d/glove.840B.300d.txt')`",
    "821581": "Found a solution for this problem\n\n`embeddings_index = {}\n\nwith ZipFile('/kaggle/input/quora-insincere-questions-classification/embeddings.zip') as myzip:\n    with myzip.open('glove.840B.300d/glove.840B.300d.txt') as myfile:\n        lines = myfile.readlines()\n        for line in tqdm(lines):\n            values = line.decode().split(\" \")\n            word = values[0]\n            coefs = np.asarray(values[1:], dtype='float32')\n            embeddings_index[word] = coefs\n\nprint('Found %s word vectors.' % len(embeddings_index))`",
    "867411": "Hi, Following your approach, I am able to get the embedding matrix without unzipping. However, it takes up a lot of RAM and later while creating the weights from embedding matrix, script terminates due to lack of RAM space. \n\nAny workaround for that ?",
    "868081": "Running into the same problem. Frustrating this is.",
    "868244": "Hi Keshav,  \n\nYou can use the below code.\n\nimport io\nimport zipfile\n\ndim=300\nembeddings1_index={}\n\nwith zipfile.ZipFile(\"../input/quora-insincere-questions-classification/embeddings.zip\") as zf:\n    with io.TextIOWrapper(zf.open(\"glove.840B.300d/glove.840B.300d.txt\"), encoding=\"utf-8\") as f:\n        for line in tqdm(f):\n            values=line.split(' ') # \".split(' ')\" only for glove-840b-300d; for all other files, \".split()\" works\n            word=values[0]\n            vectors=np.asarray(values[1:],'float32')\n            embeddings1_index[word]=vectors\n\nThis works perfectly without using too much RAM",
    "869027": "Hi Nino,\nYes I used a similar code and it worked.  I was running out of memory when trying to convert the text into sequences encoded by the embedding matrix from scratch without using the Tokenizer - I think when we use Tokenizer() to convert the text to sequences, it works just fine.",
    "1009235": "very very helpful"
  },
  "source": "meta"
}