{
  "id": 72507,
  "title": "Embeddings",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72507",
  "author_name": "",
  "post_date": "2018-11-24T01:45:47.065885500Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<pre><code>1. GoogleNews-vectors-negative300 - https://code.google.com/archive/p/word2vec/\n2. glove.840B.300d - https://nlp.stanford.edu/projects/glove/\n3. paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n4. wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html\n</code></pre>\n\n<p>Why everyone is trying the glove, paragram and wiki not GOOGLE. i have seen majority guys in competition working with 2nd,3rd and 4th but not the 1st one,\nIs there any particular reason for that.</p>",
  "messages": [
    {
      "id": "426843",
      "postDate": "11/24/2018 01:45:47",
      "content": "<pre><code>1. GoogleNews-vectors-negative300 - https://code.google.com/archive/p/word2vec/\n2. glove.840B.300d - https://nlp.stanford.edu/projects/glove/\n3. paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n4. wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html\n</code></pre>\n\n<p>Why everyone is trying the glove, paragram and wiki not GOOGLE. i have seen majority guys in competition working with 2nd,3rd and 4th but not the 1st one,\nIs there any particular reason for that.</p>",
      "rawMarkdown": "1. GoogleNews-vectors-negative300 - https://code.google.com/archive/p/word2vec/\n    2. glove.840B.300d - https://nlp.stanford.edu/projects/glove/\n    3. paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n    4. wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html\n\nWhy everyone is trying the glove, paragram and wiki not GOOGLE. i have seen majority guys in competition working with 2nd,3rd and 4th but not the 1st one,\nIs there any particular reason for that.",
      "votes": null
    },
    {
      "id": "426867",
      "postDate": "11/24/2018 03:21:12",
      "content": "<p>I compared the validation score from these two kernels:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\">https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go</a></li>\n<li><a href=\"https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\">https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings</a></li>\n</ul>\n\n<p>GloVe tends to work best.</p>",
      "rawMarkdown": "I compared the validation score from these two kernels:\n\n* https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\n* https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\n\nGloVe tends to work best.",
      "votes": null
    },
    {
      "id": "428095",
      "postDate": "11/26/2018 18:40:47",
      "content": "<p>thanks finally got the answer </p>",
      "rawMarkdown": "thanks finally got the answer",
      "votes": null
    },
    {
      "id": "428096",
      "postDate": "11/26/2018 18:41:56",
      "content": "<p>I have one more question do pre-processing the data before applying the model will be helpful ??</p>",
      "rawMarkdown": "I have one more question do pre-processing the data before applying the model will be helpful ??",
      "votes": null
    },
    {
      "id": "428099",
      "postDate": "11/26/2018 18:48:38",
      "content": "<p>Not in my case.</p>",
      "rawMarkdown": "Not in my case.",
      "votes": null
    },
    {
      "id": "428284",
      "postDate": "11/27/2018 02:46:23",
      "content": "<p>thankyou @Shujian Liu</p>",
      "rawMarkdown": "thankyou @Shujian Liu",
      "votes": null
    },
    {
      "id": "435344",
      "postDate": "12/07/2018 22:02:47",
      "content": "<p>Personally, I can't even test my code on my own machine if I want to use Google's creation. If there were a .csv version, it'd be fine... but it takes WAY too much memory (I only have 8 gig on my laptop) and I gave up on finding out how much time just to load. For a 'kernels only' competition, just trying to load their darn file would eat up most of the runtime. The .CSV's can be read one line at a time, grabbing words that actually appear in the train and test vocabulary... whereas the \"gensim\" form requires loading the whole darn thing... like I said, I don't know how long that would take because I gave up. So personally I figure the Google creation is a complete non-starter.</p>",
      "rawMarkdown": "Personally, I can't even test my code on my own machine if I want to use Google's creation. If there were a .csv version, it'd be fine... but it takes WAY too much memory (I only have 8 gig on my laptop) and I gave up on finding out how much time just to load. For a 'kernels only' competition, just trying to load their darn file would eat up most of the runtime. The .CSV's can be read one line at a time, grabbing words that actually appear in the train and test vocabulary... whereas the \"gensim\" form requires loading the whole darn thing... like I said, I don't know how long that would take because I gave up. So personally I figure the Google creation is a complete non-starter.",
      "votes": null
    },
    {
      "id": "435749",
      "postDate": "12/08/2018 17:21:24",
      "content": "<p>Got you @Steven \nthankyou</p>",
      "rawMarkdown": "Got you @Steven \nthankyou",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 426867,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/24/2018 03:21:12",
      "content": "<p>I compared the validation score from these two kernels:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\">https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go</a></li>\n<li><a href=\"https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\">https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings</a></li>\n</ul>\n\n<p>GloVe tends to work best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 428095,
          "author_name": "kalyankkr",
          "author_url": "",
          "post_date": "11/26/2018 18:40:47",
          "content": "<p>thanks finally got the answer </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428096,
          "author_name": "kalyankkr",
          "author_url": "",
          "post_date": "11/26/2018 18:41:56",
          "content": "<p>I have one more question do pre-processing the data before applying the model will be helpful ??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428099,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "11/26/2018 18:48:38",
          "content": "<p>Not in my case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428284,
          "author_name": "kalyankkr",
          "author_url": "",
          "post_date": "11/27/2018 02:46:23",
          "content": "<p>thankyou @Shujian Liu</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 435344,
      "author_name": "stevenaleach",
      "author_url": "",
      "post_date": "12/07/2018 22:02:47",
      "content": "<p>Personally, I can't even test my code on my own machine if I want to use Google's creation. If there were a .csv version, it'd be fine... but it takes WAY too much memory (I only have 8 gig on my laptop) and I gave up on finding out how much time just to load. For a 'kernels only' competition, just trying to load their darn file would eat up most of the runtime. The .CSV's can be read one line at a time, grabbing words that actually appear in the train and test vocabulary... whereas the \"gensim\" form requires loading the whole darn thing... like I said, I don't know how long that would take because I gave up. So personally I figure the Google creation is a complete non-starter.</p>",
      "votes": null,
      "replies": [
        {
          "id": 435749,
          "author_name": "kalyankkr",
          "author_url": "",
          "post_date": "12/08/2018 17:21:24",
          "content": "<p>Got you @Steven \nthankyou</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "426843": "1. GoogleNews-vectors-negative300 - https://code.google.com/archive/p/word2vec/\n    2. glove.840B.300d - https://nlp.stanford.edu/projects/glove/\n    3. paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n    4. wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html\n\nWhy everyone is trying the glove, paragram and wiki not GOOGLE. i have seen majority guys in competition working with 2nd,3rd and 4th but not the 1st one,\nIs there any particular reason for that.",
    "426867": "I compared the validation score from these two kernels:\n\n* https://www.kaggle.com/strideradu/word2vec-and-gensim-go-go-go\n* https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\n\nGloVe tends to work best.",
    "428095": "thanks finally got the answer",
    "428096": "I have one more question do pre-processing the data before applying the model will be helpful ??",
    "428099": "Not in my case.",
    "428284": "thankyou @Shujian Liu",
    "435344": "Personally, I can't even test my code on my own machine if I want to use Google's creation. If there were a .csv version, it'd be fine... but it takes WAY too much memory (I only have 8 gig on my laptop) and I gave up on finding out how much time just to load. For a 'kernels only' competition, just trying to load their darn file would eat up most of the runtime. The .CSV's can be read one line at a time, grabbing words that actually appear in the train and test vocabulary... whereas the \"gensim\" form requires loading the whole darn thing... like I said, I don't know how long that would take because I gave up. So personally I figure the Google creation is a complete non-starter.",
    "435749": "Got you @Steven \nthankyou"
  },
  "source": "meta"
}