{
  "id": 70991,
  "title": "How to decode GoogleNews-vectors-negative300.bin ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/70991",
  "author_name": "",
  "post_date": "2018-11-09T04:11:28.559374400Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>This is a very basic question, but it troubles me.</p>\n\n<p>if you read <code>GoogleNews-vectors-negative300.bin</code> as other embeddings like this:</p>\n\n<pre><code>import numpy as np\nembeddings_index = {}\nfilepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\nwith open(filepath, encoding=\"utf-8\", errors='ignore') as f:\n        for line in f:\n                values = line.split(\" \")\n                word = values[0]\n                coefs = np.asarray(values[1:], dtype='float32')\n                embeddings_index[word] = coefs\n</code></pre>\n\n<p>and you will get the following error:</p>\n\n<pre><code>ValueError: could not convert string to float: '\\x00\\x00:\\x00\\x00k\\x00\\x009\\x00\\x00:\\x00\\x00:\\x00\\x00\\x00\\x00\\x00\\x00\\x00ܹ\\x00\\x00\\x17\\x00\\x00:\\x00\\x00\\x00\\x00\"\\x00\\x00F\\x00\\x00:\\x00\\x00\\u05fa\\x00\\x00&amp;amp;\\x00\\x00:\\x00\\x00\\x00\\x00\\x00\\x00+:\\x00\\x00ڹ\\x00\\x00\\x00\\x00:\\x00\\x00\\x00\\x00\\x139\\x00\\x00:\\x00\\x00:\\x00\\x00Z\\x00\\x00\\x00\\x00:\\x00\\x009\\x00\\x00@\\x00\\x00ݸ\\x00\\x00\\x00\\x00:\\x00\\x00,:\\x00\\x00-\\x00\\x00D6\\x00\\x00:\\x00\\x009\\x00\\x00¹\\x00\\x00\\x00\\x00:\\x00\\x00l\\x00\\x009\\x00\\x00ӹ\\x00\\x00\\x149\\x00\\x00\\n'\n</code></pre>\n\n<p>How can we decode this?</p>\n\n<p>Thank you in advance.</p>",
  "messages": [
    {
      "id": "417969",
      "postDate": "11/09/2018 04:11:28",
      "content": "<p>This is a very basic question, but it troubles me.</p>\n\n<p>if you read <code>GoogleNews-vectors-negative300.bin</code> as other embeddings like this:</p>\n\n<pre><code>import numpy as np\nembeddings_index = {}\nfilepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\nwith open(filepath, encoding=\"utf-8\", errors='ignore') as f:\n        for line in f:\n                values = line.split(\" \")\n                word = values[0]\n                coefs = np.asarray(values[1:], dtype='float32')\n                embeddings_index[word] = coefs\n</code></pre>\n\n<p>and you will get the following error:</p>\n\n<pre><code>ValueError: could not convert string to float: '\\x00\\x00:\\x00\\x00k\\x00\\x009\\x00\\x00:\\x00\\x00:\\x00\\x00\\x00\\x00\\x00\\x00\\x00ܹ\\x00\\x00\\x17\\x00\\x00:\\x00\\x00\\x00\\x00\"\\x00\\x00F\\x00\\x00:\\x00\\x00\\u05fa\\x00\\x00&amp;amp;\\x00\\x00:\\x00\\x00\\x00\\x00\\x00\\x00+:\\x00\\x00ڹ\\x00\\x00\\x00\\x00:\\x00\\x00\\x00\\x00\\x139\\x00\\x00:\\x00\\x00:\\x00\\x00Z\\x00\\x00\\x00\\x00:\\x00\\x009\\x00\\x00@\\x00\\x00ݸ\\x00\\x00\\x00\\x00:\\x00\\x00,:\\x00\\x00-\\x00\\x00D6\\x00\\x00:\\x00\\x009\\x00\\x00¹\\x00\\x00\\x00\\x00:\\x00\\x00l\\x00\\x009\\x00\\x00ӹ\\x00\\x00\\x149\\x00\\x00\\n'\n</code></pre>\n\n<p>How can we decode this?</p>\n\n<p>Thank you in advance.</p>",
      "rawMarkdown": "This is a very basic question, but it troubles me.\n\nif you read ```GoogleNews-vectors-negative300.bin``` as other embeddings like this:\n\n    import numpy as np\n    embeddings_index = {}\n    filepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n    with open(filepath, encoding=\"utf-8\", errors='ignore') as f:\n            for line in f:\n                    values = line.split(\" \")\n                    word = values[0]\n                    coefs = np.asarray(values[1:], dtype='float32')\n                    embeddings_index[word] = coefs\n\nand you will get the following error:\n\n    ValueError: could not convert string to float: '\\x00\\x00:\\x00\\x00k\\x00\\x009\\x00\\x00:\\x00\\x00:\\x00\\x00\\x00\\x00\\x00\\x00\\x00ܹ\\x00\\x00\\x17\\x00\\x00:\\x00\\x00\\x00\\x00\"\\x00\\x00F\\x00\\x00:\\x00\\x00\\u05fa\\x00\\x00&amp;\\x00\\x00:\\x00\\x00\\x00\\x00\\x00\\x00+:\\x00\\x00ڹ\\x00\\x00\\x00\\x00:\\x00\\x00\\x00\\x00\\x139\\x00\\x00:\\x00\\x00:\\x00\\x00Z\\x00\\x00\\x00\\x00:\\x00\\x009\\x00\\x00@\\x00\\x00ݸ\\x00\\x00\\x00\\x00:\\x00\\x00,:\\x00\\x00-\\x00\\x00D6\\x00\\x00:\\x00\\x009\\x00\\x00¹\\x00\\x00\\x00\\x00:\\x00\\x00l\\x00\\x009\\x00\\x00ӹ\\x00\\x00\\x149\\x00\\x00\\n'\n\n\nHow can we decode this?\n\nThank you in advance.",
      "votes": null
    },
    {
      "id": "417979",
      "postDate": "11/09/2018 04:23:55",
      "content": "<p>You can use gensim <a href=\"https://radimrehurek.com/gensim/models/keyedvectors.html\">https://radimrehurek.com/gensim/models/keyedvectors.html</a> to load the vectors, I am writing a kernel to give an example</p>",
      "rawMarkdown": "You can use gensim https://radimrehurek.com/gensim/models/keyedvectors.html to load the vectors, I am writing a kernel to give an example",
      "votes": null
    },
    {
      "id": "418006",
      "postDate": "11/09/2018 05:19:41",
      "content": "<p>Here is my kernel <a href=\"https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\">https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go</a></p>",
      "rawMarkdown": "Here is my kernel https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go",
      "votes": null
    },
    {
      "id": "418009",
      "postDate": "11/09/2018 05:23:01",
      "content": "<p>Thanks a lot! I can read it correctly.</p>\n\n<p>The following code may help someone (<a href=\"https://www.kaggle.com/takaishikawa/how-to-decode-googlenews-vectors-negative300-bin\">here</a> is my kernel): </p>\n\n<pre><code>import numpy as np\n\nfilepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n\nembeddings_index = {}\nfrom gensim.models import KeyedVectors\nwv_from_bin = KeyedVectors.load_word2vec_format(filepath, binary=True) \nfor word, vector in zip(wv_from_bin.vocab, wv_from_bin.vectors):\n    coefs = np.asarray(vector, dtype='float32')\n    embeddings_index[word] = coefs\n</code></pre>\n\n<p>And, <a href=\"https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\">your kernel</a> is a good example to use gensim to load word2vec embeddings.</p>",
      "rawMarkdown": "Thanks a lot! I can read it correctly.\n\nThe following code may help someone ([here](https://www.kaggle.com/takaishikawa/how-to-decode-googlenews-vectors-negative300-bin) is my kernel): \n\n    import numpy as np\n\n    filepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n\n    embeddings_index = {}\n    from gensim.models import KeyedVectors\n    wv_from_bin = KeyedVectors.load_word2vec_format(filepath, binary=True) \n    for word, vector in zip(wv_from_bin.vocab, wv_from_bin.vectors):\n        coefs = np.asarray(vector, dtype='float32')\n        embeddings_index[word] = coefs\n\nAnd, [your kernel](https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go) is a good example to use gensim to load word2vec embeddings.",
      "votes": null
    },
    {
      "id": "453030",
      "postDate": "01/09/2019 14:45:36",
      "content": "<p>thanks! it works ok</p>",
      "rawMarkdown": "thanks! it works ok",
      "votes": null
    },
    {
      "id": "453248",
      "postDate": "01/09/2019 23:15:21",
      "content": "<p>I'm glad if it helped you :)</p>",
      "rawMarkdown": "I'm glad if it helped you :)",
      "votes": null
    },
    {
      "id": "518596",
      "postDate": "04/17/2019 13:55:13",
      "content": "<p>how to access data from another kernel, i need GoogleNews-vectors-negative300.bin </p>",
      "rawMarkdown": "how to access data from another kernel, i need GoogleNews-vectors-negative300.bin",
      "votes": null
    },
    {
      "id": "518619",
      "postDate": "04/17/2019 14:26:42",
      "content": "<p>you can add another dataset in your kernel, check the dataset part located in the left of the kernel interface</p>",
      "rawMarkdown": "you can add another dataset in your kernel, check the dataset part located in the left of the kernel interface",
      "votes": null
    },
    {
      "id": "1498202",
      "postDate": "08/31/2021 18:37:30",
      "content": "<p>my question is about clustering vocabolary in 'GoogleNews-vectors-negative300.bin '</p>\n<h2>This is a good code:</h2>\n<p>import numpy as np</p>\n<p>filepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"</p>\n<p>embeddings_index = {}<br>\nfrom gensim.models import KeyedVectors<br>\nwv_from_bin = KeyedVectors.load_word2vec_format(filepath, binary=True) <br>\nfor word, vector in zip(wv_from_bin.vocab, wv_from_bin.vectors):<br>\n    coefs = np.asarray(vector, dtype='float32')</p>\n<h2>    embeddings_index[word] = coefs</h2>\n<p>After getting all words in 'GoogleNews-vectors-negative300.bin'  (about 3 million), what is a optimal cluster number ? is it 1500 or 10000? a logic behind 10000 clusters: there are 3 million words and each word is linked to 300 words, so 3,000,000 / 300 = 10000.</p>\n<p>Thank you in advance<br>\nTursun</p>",
      "rawMarkdown": "my question is about clustering vocabolary in 'GoogleNews-vectors-negative300.bin '\nThis is a good code:\n---------------------------------------------------------------------------------------------------\nimport numpy as np\n\nfilepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n\nembeddings_index = {}\nfrom gensim.models import KeyedVectors\nwv_from_bin = KeyedVectors.load_word2vec_format(filepath, binary=True) \nfor word, vector in zip(wv_from_bin.vocab, wv_from_bin.vectors):\n    coefs = np.asarray(vector, dtype='float32')\n    embeddings_index[word] = coefs\n---------------------------------------------------------------------------------------------------\nAfter getting all words in 'GoogleNews-vectors-negative300.bin'  (about 3 million), what is a optimal cluster number ? is it 1500 or 10000? a logic behind 10000 clusters: there are 3 million words and each word is linked to 300 words, so 3,000,000 / 300 = 10000.\n\nThank you in advance\nTursun",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1498202,
      "author_name": "sayajujur",
      "author_url": "",
      "post_date": "08/31/2021 18:37:30",
      "content": "<p>my question is about clustering vocabolary in 'GoogleNews-vectors-negative300.bin '</p>\n<h2>This is a good code:</h2>\n<p>import numpy as np</p>\n<p>filepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"</p>\n<p>embeddings_index = {}<br>\nfrom gensim.models import KeyedVectors<br>\nwv_from_bin = KeyedVectors.load_word2vec_format(filepath, binary=True) <br>\nfor word, vector in zip(wv_from_bin.vocab, wv_from_bin.vectors):<br>\n    coefs = np.asarray(vector, dtype='float32')</p>\n<h2>    embeddings_index[word] = coefs</h2>\n<p>After getting all words in 'GoogleNews-vectors-negative300.bin'  (about 3 million), what is a optimal cluster number ? is it 1500 or 10000? a logic behind 10000 clusters: there are 3 million words and each word is linked to 300 words, so 3,000,000 / 300 = 10000.</p>\n<p>Thank you in advance<br>\nTursun</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417979,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "11/09/2018 04:23:55",
      "content": "<p>You can use gensim <a href=\"https://radimrehurek.com/gensim/models/keyedvectors.html\">https://radimrehurek.com/gensim/models/keyedvectors.html</a> to load the vectors, I am writing a kernel to give an example</p>",
      "votes": null,
      "replies": [
        {
          "id": 418009,
          "author_name": "takaishikawa",
          "author_url": "",
          "post_date": "11/09/2018 05:23:01",
          "content": "<p>Thanks a lot! I can read it correctly.</p>\n\n<p>The following code may help someone (<a href=\"https://www.kaggle.com/takaishikawa/how-to-decode-googlenews-vectors-negative300-bin\">here</a> is my kernel): </p>\n\n<pre><code>import numpy as np\n\nfilepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n\nembeddings_index = {}\nfrom gensim.models import KeyedVectors\nwv_from_bin = KeyedVectors.load_word2vec_format(filepath, binary=True) \nfor word, vector in zip(wv_from_bin.vocab, wv_from_bin.vectors):\n    coefs = np.asarray(vector, dtype='float32')\n    embeddings_index[word] = coefs\n</code></pre>\n\n<p>And, <a href=\"https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\">your kernel</a> is a good example to use gensim to load word2vec embeddings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 453030,
          "author_name": "galloguille",
          "author_url": "",
          "post_date": "01/09/2019 14:45:36",
          "content": "<p>thanks! it works ok</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 453248,
          "author_name": "takaishikawa",
          "author_url": "",
          "post_date": "01/09/2019 23:15:21",
          "content": "<p>I'm glad if it helped you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418006,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "11/09/2018 05:19:41",
      "content": "<p>Here is my kernel <a href=\"https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\">https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518596,
      "author_name": "mehdihadji1994",
      "author_url": "",
      "post_date": "04/17/2019 13:55:13",
      "content": "<p>how to access data from another kernel, i need GoogleNews-vectors-negative300.bin </p>",
      "votes": null,
      "replies": [
        {
          "id": 518619,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "04/17/2019 14:26:42",
          "content": "<p>you can add another dataset in your kernel, check the dataset part located in the left of the kernel interface</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "417969": "This is a very basic question, but it troubles me.\n\nif you read ```GoogleNews-vectors-negative300.bin``` as other embeddings like this:\n\n    import numpy as np\n    embeddings_index = {}\n    filepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n    with open(filepath, encoding=\"utf-8\", errors='ignore') as f:\n            for line in f:\n                    values = line.split(\" \")\n                    word = values[0]\n                    coefs = np.asarray(values[1:], dtype='float32')\n                    embeddings_index[word] = coefs\n\nand you will get the following error:\n\n    ValueError: could not convert string to float: '\\x00\\x00:\\x00\\x00k\\x00\\x009\\x00\\x00:\\x00\\x00:\\x00\\x00\\x00\\x00\\x00\\x00\\x00ܹ\\x00\\x00\\x17\\x00\\x00:\\x00\\x00\\x00\\x00\"\\x00\\x00F\\x00\\x00:\\x00\\x00\\u05fa\\x00\\x00&amp;\\x00\\x00:\\x00\\x00\\x00\\x00\\x00\\x00+:\\x00\\x00ڹ\\x00\\x00\\x00\\x00:\\x00\\x00\\x00\\x00\\x139\\x00\\x00:\\x00\\x00:\\x00\\x00Z\\x00\\x00\\x00\\x00:\\x00\\x009\\x00\\x00@\\x00\\x00ݸ\\x00\\x00\\x00\\x00:\\x00\\x00,:\\x00\\x00-\\x00\\x00D6\\x00\\x00:\\x00\\x009\\x00\\x00¹\\x00\\x00\\x00\\x00:\\x00\\x00l\\x00\\x009\\x00\\x00ӹ\\x00\\x00\\x149\\x00\\x00\\n'\n\n\nHow can we decode this?\n\nThank you in advance.",
    "417979": "You can use gensim https://radimrehurek.com/gensim/models/keyedvectors.html to load the vectors, I am writing a kernel to give an example",
    "418006": "Here is my kernel https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go",
    "418009": "Thanks a lot! I can read it correctly.\n\nThe following code may help someone ([here](https://www.kaggle.com/takaishikawa/how-to-decode-googlenews-vectors-negative300-bin) is my kernel): \n\n    import numpy as np\n\n    filepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n\n    embeddings_index = {}\n    from gensim.models import KeyedVectors\n    wv_from_bin = KeyedVectors.load_word2vec_format(filepath, binary=True) \n    for word, vector in zip(wv_from_bin.vocab, wv_from_bin.vectors):\n        coefs = np.asarray(vector, dtype='float32')\n        embeddings_index[word] = coefs\n\nAnd, [your kernel](https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go) is a good example to use gensim to load word2vec embeddings.",
    "453030": "thanks! it works ok",
    "453248": "I'm glad if it helped you :)",
    "518596": "how to access data from another kernel, i need GoogleNews-vectors-negative300.bin",
    "518619": "you can add another dataset in your kernel, check the dataset part located in the left of the kernel interface",
    "1498202": "my question is about clustering vocabolary in 'GoogleNews-vectors-negative300.bin '\nThis is a good code:\n---------------------------------------------------------------------------------------------------\nimport numpy as np\n\nfilepath = \"../input/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n\nembeddings_index = {}\nfrom gensim.models import KeyedVectors\nwv_from_bin = KeyedVectors.load_word2vec_format(filepath, binary=True) \nfor word, vector in zip(wv_from_bin.vocab, wv_from_bin.vectors):\n    coefs = np.asarray(vector, dtype='float32')\n    embeddings_index[word] = coefs\n---------------------------------------------------------------------------------------------------\nAfter getting all words in 'GoogleNews-vectors-negative300.bin'  (about 3 million), what is a optimal cluster number ? is it 1500 or 10000? a logic behind 10000 clusters: there are 3 million words and each word is linked to 300 words, so 3,000,000 / 300 = 10000.\n\nThank you in advance\nTursun"
  },
  "source": "meta"
}