{
  "id": 77968,
  "title": "Load Word2Vec without KeyedVectors ",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77968",
  "author_name": "",
  "post_date": "2019-01-18T06:02:27.675494200Z",
  "votes": 33,
  "comment_count": 4,
  "views": 0,
  "content": "<p>So I modified the gensim's KeyedVectors *load_word2vec_format* method and now I can load the word2vec embeddings within 30 seconds. </p>\n\n<p>Hope it helps. </p>\n\n<pre><code>from gensim import utils\nmax_features = 95000\n\ndef load_word2vec(fname, encoding='utf8', unicode_errors='strict',datatype=np.float32, word_index=None):\n    emb_mean,emb_std = -0.0051106834, 0.18445626\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (max_features, 300))\n    with utils.smart_open(fname) as fin:\n        header = utils.to_unicode(fin.readline(), encoding=encoding)\n        vocab_size, vector_size = (int(x) for x in header.split())\n        binary_len = np.dtype(datatype).itemsize * vector_size\n        for _ in tqdm(range(vocab_size)):\n            # mixed text and binary: read text first, then binary\n            word = []\n            while True:\n                ch = fin.read(1)\n                if ch == b' ':\n                    break\n                if ch == b'':\n                    raise EOFError(\"unexpected end of input\")\n                if ch != b'\\n':\n                    word.append(ch)\n            word = utils.to_unicode(b''.join(word), encoding=encoding, errors=unicode_errors)\n            weights = np.fromstring(fin.read(binary_len), dtype=datatype).astype(datatype)\n            if word not in word_index:\n                continue\n            i = word_index[word]\n            if i &gt;= max_features:\n                continue\n            embedding_matrix[i] = weights\n    return embedding_matrix\n</code></pre>",
  "messages": [
    {
      "id": "457819",
      "postDate": "01/18/2019 06:02:27",
      "content": "<p>So I modified the gensim's KeyedVectors *load_word2vec_format* method and now I can load the word2vec embeddings within 30 seconds. </p>\n\n<p>Hope it helps. </p>\n\n<pre><code>from gensim import utils\nmax_features = 95000\n\ndef load_word2vec(fname, encoding='utf8', unicode_errors='strict',datatype=np.float32, word_index=None):\n    emb_mean,emb_std = -0.0051106834, 0.18445626\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (max_features, 300))\n    with utils.smart_open(fname) as fin:\n        header = utils.to_unicode(fin.readline(), encoding=encoding)\n        vocab_size, vector_size = (int(x) for x in header.split())\n        binary_len = np.dtype(datatype).itemsize * vector_size\n        for _ in tqdm(range(vocab_size)):\n            # mixed text and binary: read text first, then binary\n            word = []\n            while True:\n                ch = fin.read(1)\n                if ch == b' ':\n                    break\n                if ch == b'':\n                    raise EOFError(\"unexpected end of input\")\n                if ch != b'\\n':\n                    word.append(ch)\n            word = utils.to_unicode(b''.join(word), encoding=encoding, errors=unicode_errors)\n            weights = np.fromstring(fin.read(binary_len), dtype=datatype).astype(datatype)\n            if word not in word_index:\n                continue\n            i = word_index[word]\n            if i &gt;= max_features:\n                continue\n            embedding_matrix[i] = weights\n    return embedding_matrix\n</code></pre>",
      "rawMarkdown": "So I modified the gensim's KeyedVectors *load_word2vec_format* method and now I can load the word2vec embeddings within 30 seconds. \n\nHope it helps. \n\n    from gensim import utils\n    max_features = 95000\n\n    def load_word2vec(fname, encoding='utf8', unicode_errors='strict',datatype=np.float32, word_index=None):\n        emb_mean,emb_std = -0.0051106834, 0.18445626\n        embedding_matrix = np.random.normal(emb_mean, emb_std, (max_features, 300))\n        with utils.smart_open(fname) as fin:\n            header = utils.to_unicode(fin.readline(), encoding=encoding)\n            vocab_size, vector_size = (int(x) for x in header.split())\n            binary_len = np.dtype(datatype).itemsize * vector_size\n            for _ in tqdm(range(vocab_size)):\n                # mixed text and binary: read text first, then binary\n                word = []\n                while True:\n                    ch = fin.read(1)\n                    if ch == b' ':\n                        break\n                    if ch == b'':\n                        raise EOFError(\"unexpected end of input\")\n                    if ch != b'\\n':\n                        word.append(ch)\n                word = utils.to_unicode(b''.join(word), encoding=encoding, errors=unicode_errors)\n                weights = np.fromstring(fin.read(binary_len), dtype=datatype).astype(datatype)\n                if word not in word_index:\n                    continue\n                i = word_index[word]\n                if i &gt;= max_features:\n                    continue\n                embedding_matrix[i] = weights\n        return embedding_matrix",
      "votes": null
    },
    {
      "id": "457865",
      "postDate": "01/18/2019 08:29:31",
      "content": "<p>Thank you, exactly what I was looking for</p>",
      "rawMarkdown": "Thank you, exactly what I was looking for",
      "votes": null
    },
    {
      "id": "458800",
      "postDate": "01/20/2019 14:04:23",
      "content": "<p>thanks for ur sharing</p>",
      "rawMarkdown": "thanks for ur sharing",
      "votes": null
    },
    {
      "id": "774945",
      "postDate": "03/16/2020 04:57:25",
      "content": "<p>You reduced the vectors vocab size to top 95000 right ? I suppose if are reducing the vocab size the less loading time of vector makes sense .</p>",
      "rawMarkdown": "You reduced the vectors vocab size to top 95000 right ? I suppose if are reducing the vocab size the less loading time of vector makes sense .",
      "votes": null
    },
    {
      "id": "1077489",
      "postDate": "11/13/2020 16:33:37",
      "content": "<p>Notebook : <a href=\"https://www.kaggle.com/naim99/text-clustering-with-word2vec\" target=\"_blank\">https://www.kaggle.com/naim99/text-clustering-with-word2vec</a> <br>\nMy dataset : <a href=\"https://www.kaggle.com/naim99/ts-naim-mhedhbi\" target=\"_blank\">https://www.kaggle.com/naim99/ts-naim-mhedhbi</a></p>\n<p>Check this kernel for Arabic Text Analysis . I hope you find it useful and informative.</p>",
      "rawMarkdown": "Notebook : https://www.kaggle.com/naim99/text-clustering-with-word2vec \nMy dataset : https://www.kaggle.com/naim99/ts-naim-mhedhbi\n\nCheck this kernel for Arabic Text Analysis . I hope you find it useful and informative.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1077489,
      "author_name": "naim99",
      "author_url": "",
      "post_date": "11/13/2020 16:33:37",
      "content": "<p>Notebook : <a href=\"https://www.kaggle.com/naim99/text-clustering-with-word2vec\" target=\"_blank\">https://www.kaggle.com/naim99/text-clustering-with-word2vec</a> <br>\nMy dataset : <a href=\"https://www.kaggle.com/naim99/ts-naim-mhedhbi\" target=\"_blank\">https://www.kaggle.com/naim99/ts-naim-mhedhbi</a></p>\n<p>Check this kernel for Arabic Text Analysis . I hope you find it useful and informative.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 457865,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "01/18/2019 08:29:31",
      "content": "<p>Thank you, exactly what I was looking for</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 458800,
      "author_name": "karwik",
      "author_url": "",
      "post_date": "01/20/2019 14:04:23",
      "content": "<p>thanks for ur sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 774945,
      "author_name": "akashtyagi08",
      "author_url": "",
      "post_date": "03/16/2020 04:57:25",
      "content": "<p>You reduced the vectors vocab size to top 95000 right ? I suppose if are reducing the vocab size the less loading time of vector makes sense .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "457819": "So I modified the gensim's KeyedVectors *load_word2vec_format* method and now I can load the word2vec embeddings within 30 seconds. \n\nHope it helps. \n\n    from gensim import utils\n    max_features = 95000\n\n    def load_word2vec(fname, encoding='utf8', unicode_errors='strict',datatype=np.float32, word_index=None):\n        emb_mean,emb_std = -0.0051106834, 0.18445626\n        embedding_matrix = np.random.normal(emb_mean, emb_std, (max_features, 300))\n        with utils.smart_open(fname) as fin:\n            header = utils.to_unicode(fin.readline(), encoding=encoding)\n            vocab_size, vector_size = (int(x) for x in header.split())\n            binary_len = np.dtype(datatype).itemsize * vector_size\n            for _ in tqdm(range(vocab_size)):\n                # mixed text and binary: read text first, then binary\n                word = []\n                while True:\n                    ch = fin.read(1)\n                    if ch == b' ':\n                        break\n                    if ch == b'':\n                        raise EOFError(\"unexpected end of input\")\n                    if ch != b'\\n':\n                        word.append(ch)\n                word = utils.to_unicode(b''.join(word), encoding=encoding, errors=unicode_errors)\n                weights = np.fromstring(fin.read(binary_len), dtype=datatype).astype(datatype)\n                if word not in word_index:\n                    continue\n                i = word_index[word]\n                if i &gt;= max_features:\n                    continue\n                embedding_matrix[i] = weights\n        return embedding_matrix",
    "457865": "Thank you, exactly what I was looking for",
    "458800": "thanks for ur sharing",
    "774945": "You reduced the vectors vocab size to top 95000 right ? I suppose if are reducing the vocab size the less loading time of vector makes sense .",
    "1077489": "Notebook : https://www.kaggle.com/naim99/text-clustering-with-word2vec \nMy dataset : https://www.kaggle.com/naim99/ts-naim-mhedhbi\n\nCheck this kernel for Arabic Text Analysis . I hope you find it useful and informative."
  },
  "source": "meta"
}