{
  "id": 71106,
  "title": "[R] -- Using \"GoogleNews-vectors-negative300.bin\"",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71106",
  "author_name": "",
  "post_date": "2018-11-10T09:33:02.896612Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Has anyone figured out how to successfully read/use GoogleNews-vectors-negative300.bin with R? </p>\n\n<p>What I've tried:\n- converting the binary file to txt first with the 'rword2vec' package \nRuns into disk memory issues on kernels</p>\n\n<ul>\n<li>using the \"reticulate\" package to call the \"gensim\" module\nReturns module not found error. I think the python version and modules available through \"reticulate\" are different from the ones available using a python kernel.  </li>\n</ul>",
  "messages": [
    {
      "id": "418623",
      "postDate": "11/10/2018 09:33:02",
      "content": "<p>Has anyone figured out how to successfully read/use GoogleNews-vectors-negative300.bin with R? </p>\n\n<p>What I've tried:\n- converting the binary file to txt first with the 'rword2vec' package \nRuns into disk memory issues on kernels</p>\n\n<ul>\n<li>using the \"reticulate\" package to call the \"gensim\" module\nReturns module not found error. I think the python version and modules available through \"reticulate\" are different from the ones available using a python kernel.  </li>\n</ul>",
      "rawMarkdown": "Has anyone figured out how to successfully read/use GoogleNews-vectors-negative300.bin with R? \n\nWhat I've tried:\n- converting the binary file to txt first with the 'rword2vec' package \nRuns into disk memory issues on kernels\n\n- using the \"reticulate\" package to call the \"gensim\" module\nReturns module not found error. I think the python version and modules available through \"reticulate\" are different from the ones available using a python kernel.",
      "votes": null
    },
    {
      "id": "424890",
      "postDate": "11/20/2018 20:42:06",
      "content": "<p>You've probably found an answer by now, but below is an approach.</p>\n\n<p>(Slightly adapted from here: <a href=\"https://stackoverflow.com/questions/41903454/importing-and-working-with-word2vec-googlenews-vectors-negative300-bin-gz-into-r\">https://stackoverflow.com/questions/41903454/importing-and-working-with-word2vec-googlenews-vectors-negative300-bin-gz-into-r</a>)</p>\n\n<pre><code>devtools::install_github(\"mukul13/rword2vec\")\nlibrary(rword2vec)\n\n# Extract a word\nfname = \"data_raw/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\ndist=distance(file_name = fname,search_word = \"king\",num = 10)\n\n# convert part (might take a while)\nbin_to_txt(fname,\"googleOut.txt\")\n</code></pre>",
      "rawMarkdown": "You've probably found an answer by now, but below is an approach.\n\n(Slightly adapted from here: https://stackoverflow.com/questions/41903454/importing-and-working-with-word2vec-googlenews-vectors-negative300-bin-gz-into-r)\n\n\n\tdevtools::install_github(\"mukul13/rword2vec\")\n\tlibrary(rword2vec)\n\n\t# Extract a word\n\tfname = \"data_raw/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n\tdist=distance(file_name = fname,search_word = \"king\",num = 10)\n\n\t# convert part (might take a while)\n\tbin_to_txt(fname,\"googleOut.txt\")",
      "votes": null
    },
    {
      "id": "425059",
      "postDate": "11/21/2018 04:22:33",
      "content": "<p>I haven't, Thanks for the response!</p>\n\n<p>I am not sure how this solution is different from the first of the two things I mentioned in the post. I think this does not work in the kernel environment because the outfile (\"googleOut.txt\" in your snippet) size gets too big. Also, one cannot use \"install_github\" here because internet connections are not allowed in the competition...rword2vec is pre-installed though so no need to install either.</p>",
      "rawMarkdown": "I haven't, Thanks for the response!\n\nI am not sure how this solution is different from the first of the two things I mentioned in the post. I think this does not work in the kernel environment because the outfile (\"googleOut.txt\" in your snippet) size gets too big. Also, one cannot use \"install_github\" here because internet connections are not allowed in the competition...rword2vec is pre-installed though so no need to install either.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 424890,
      "author_name": "martinlbarron",
      "author_url": "",
      "post_date": "11/20/2018 20:42:06",
      "content": "<p>You've probably found an answer by now, but below is an approach.</p>\n\n<p>(Slightly adapted from here: <a href=\"https://stackoverflow.com/questions/41903454/importing-and-working-with-word2vec-googlenews-vectors-negative300-bin-gz-into-r\">https://stackoverflow.com/questions/41903454/importing-and-working-with-word2vec-googlenews-vectors-negative300-bin-gz-into-r</a>)</p>\n\n<pre><code>devtools::install_github(\"mukul13/rword2vec\")\nlibrary(rword2vec)\n\n# Extract a word\nfname = \"data_raw/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\ndist=distance(file_name = fname,search_word = \"king\",num = 10)\n\n# convert part (might take a while)\nbin_to_txt(fname,\"googleOut.txt\")\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 425059,
          "author_name": "grmunjal",
          "author_url": "",
          "post_date": "11/21/2018 04:22:33",
          "content": "<p>I haven't, Thanks for the response!</p>\n\n<p>I am not sure how this solution is different from the first of the two things I mentioned in the post. I think this does not work in the kernel environment because the outfile (\"googleOut.txt\" in your snippet) size gets too big. Also, one cannot use \"install_github\" here because internet connections are not allowed in the competition...rword2vec is pre-installed though so no need to install either.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "418623": "Has anyone figured out how to successfully read/use GoogleNews-vectors-negative300.bin with R? \n\nWhat I've tried:\n- converting the binary file to txt first with the 'rword2vec' package \nRuns into disk memory issues on kernels\n\n- using the \"reticulate\" package to call the \"gensim\" module\nReturns module not found error. I think the python version and modules available through \"reticulate\" are different from the ones available using a python kernel.",
    "424890": "You've probably found an answer by now, but below is an approach.\n\n(Slightly adapted from here: https://stackoverflow.com/questions/41903454/importing-and-working-with-word2vec-googlenews-vectors-negative300-bin-gz-into-r)\n\n\n\tdevtools::install_github(\"mukul13/rword2vec\")\n\tlibrary(rword2vec)\n\n\t# Extract a word\n\tfname = \"data_raw/embeddings/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin\"\n\tdist=distance(file_name = fname,search_word = \"king\",num = 10)\n\n\t# convert part (might take a while)\n\tbin_to_txt(fname,\"googleOut.txt\")",
    "425059": "I haven't, Thanks for the response!\n\nI am not sure how this solution is different from the first of the two things I mentioned in the post. I think this does not work in the kernel environment because the outfile (\"googleOut.txt\" in your snippet) size gets too big. Also, one cannot use \"install_github\" here because internet connections are not allowed in the competition...rword2vec is pre-installed though so no need to install either."
  },
  "source": "meta"
}