{
  "id": 78042,
  "title": "Why nobody is using google new vectors",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78042",
  "author_name": "Xingjian",
  "post_date": "2019-01-19T03:11:28.816000",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I notice in most kernels people don't use google new vectors embedding. Is there a reason for this?</p>",
  "messages": [
    {
      "id": 458220,
      "postDate": "2019-01-19T05:32:05.827Z",
      "content": "<p>Maybe the vocabulary coverage of google news vectors is lowest. In my experiment, after cleaning: </p>\n\n<pre><code>Glove : \nFound embeddings for 77.52% of vocab\nFound embeddings for  99.64% of all text\nParagram : \nFound embeddings for 79.69% of vocab\nFound embeddings for  99.68% of all text\nFastText : \nFound embeddings for 70.58% of vocab\nFound embeddings for  99.51% of all text\nGoogle : \nFound embeddings for 62.41% of vocab\nFound embeddings for  81.27% of all text\n</code></pre>",
      "rawMarkdown": "Maybe the vocabulary coverage of google news vectors is lowest. In my experiment, after cleaning: \n\n    Glove : \n    Found embeddings for 77.52% of vocab\n    Found embeddings for  99.64% of all text\n    Paragram : \n    Found embeddings for 79.69% of vocab\n    Found embeddings for  99.68% of all text\n    FastText : \n    Found embeddings for 70.58% of vocab\n    Found embeddings for  99.51% of all text\n    Google : \n    Found embeddings for 62.41% of vocab\n    Found embeddings for  81.27% of all text\n",
      "votes": 14,
      "replies": [
        {
          "id": 459035,
          "postDate": "2019-01-21T04:01:23.680Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 458378,
      "postDate": "2019-01-19T13:27:28.037Z",
      "content": "<p>I have tried to concat 4 embeddings but my public score is worse</p>",
      "rawMarkdown": "I have tried to concat 4 embeddings but my public score is worse",
      "votes": 3,
      "replies": [
        {
          "id": 458385,
          "postDate": "2019-01-19T13:45:53.410Z",
          "content": "<p>Have you tried to average?</p>",
          "rawMarkdown": "Have you tried to average?"
        },
        {
          "id": 458386,
          "postDate": "2019-01-19T13:47:14.240Z",
          "content": "<p>Me too and also I averaged.</p>",
          "rawMarkdown": "Me too and also I averaged.",
          "votes": 2
        },
        {
          "id": 458615,
          "postDate": "2019-01-20T04:00:38.307Z",
          "content": "<p>I've tried concat 3 embeddings except for google news and it's worse than just concat paragram and glove. Concat glove and paragram did improve my LB scores.</p>",
          "rawMarkdown": "I've tried concat 3 embeddings except for google news and it's worse than just concat paragram and glove. Concat glove and paragram did improve my LB scores.",
          "votes": 3
        }
      ]
    },
    {
      "id": 458195,
      "postDate": "2019-01-19T03:11:28.817Z",
      "content": "<p>I notice in most kernels people don't use google new vectors embedding. Is there a reason for this?</p>",
      "rawMarkdown": "I notice in most kernels people don't use google new vectors embedding. Is there a reason for this?",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 458220,
      "author_name": "Qing Liu",
      "author_url": "",
      "post_date": "2019-01-19T05:32:05.827000",
      "content": "<p>Maybe the vocabulary coverage of google news vectors is lowest. In my experiment, after cleaning: </p>\n\n<pre><code>Glove : \nFound embeddings for 77.52% of vocab\nFound embeddings for  99.64% of all text\nParagram : \nFound embeddings for 79.69% of vocab\nFound embeddings for  99.68% of all text\nFastText : \nFound embeddings for 70.58% of vocab\nFound embeddings for  99.51% of all text\nGoogle : \nFound embeddings for 62.41% of vocab\nFound embeddings for  81.27% of all text\n</code></pre>",
      "votes": 14,
      "replies": [
        {
          "id": 459035,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-21T04:01:23.680000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 458378,
      "author_name": "pbcquoc",
      "author_url": "",
      "post_date": "2019-01-19T13:27:28.037000",
      "content": "<p>I have tried to concat 4 embeddings but my public score is worse</p>",
      "votes": 3,
      "replies": [
        {
          "id": 458385,
          "author_name": "ManuelSH",
          "author_url": "",
          "post_date": "2019-01-19T13:45:53.410000",
          "content": "<p>Have you tried to average?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 458386,
          "author_name": "Soonhwan Kwon",
          "author_url": "",
          "post_date": "2019-01-19T13:47:14.240000",
          "content": "<p>Me too and also I averaged.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 458615,
          "author_name": "Xingjian",
          "author_url": "",
          "post_date": "2019-01-20T04:00:38.307000",
          "content": "<p>I've tried concat 3 embeddings except for google news and it's worse than just concat paragram and glove. Concat glove and paragram did improve my LB scores.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "458220": "Maybe the vocabulary coverage of google news vectors is lowest. In my experiment, after cleaning: \n\n    Glove : \n    Found embeddings for 77.52% of vocab\n    Found embeddings for  99.64% of all text\n    Paragram : \n    Found embeddings for 79.69% of vocab\n    Found embeddings for  99.68% of all text\n    FastText : \n    Found embeddings for 70.58% of vocab\n    Found embeddings for  99.51% of all text\n    Google : \n    Found embeddings for 62.41% of vocab\n    Found embeddings for  81.27% of all text\n",
    "458378": "I have tried to concat 4 embeddings but my public score is worse",
    "458195": "I notice in most kernels people don't use google new vectors embedding. Is there a reason for this?"
  }
}