{
  "id": 56897,
  "title": "tf-idf embeddings",
  "url": "/competitions/avito-demand-prediction/discussion/56897",
  "author_name": "",
  "post_date": "2018-05-16T11:00:44.944890800Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>A good number of kernels are going the traditional route with CountVectorizer/TF-IDF, and some brave souls (I say brave because training is slower and the results don't seem as spectacular so far) have been experimenting with embeddings, as per the previous competitions. So I had a showerthought about merging the two in a non-ensembled manner, and... as per usual, of course someone had already beat me to experimenting with that. Take a look at <a href=\"https://github.com/nadbordrozd/blog_stuff/blob/master/classification_w2v/benchmarking_python3.ipynb\">this fascinating notebook</a>.</p>\n\n<p>Joe Eddy (<a href=\"/aquatic\">@aquatic</a>) mentions in <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/56798#329097\">this thread</a> about <em>average the embedding values across a description to get description-level vectors</em>, and that's exactly what this notebook does. The author compares classification results of using some classical nb and svc, fit on countvectorizer, to mean document embeddings generated from glove and w2v. Then, he/she experiment with the same mean embeddings but this time first multiplying each word vec by it's IDF before averaging them.</p>",
  "messages": [
    {
      "id": "329375",
      "postDate": "05/16/2018 11:00:44",
      "content": "<p>A good number of kernels are going the traditional route with CountVectorizer/TF-IDF, and some brave souls (I say brave because training is slower and the results don't seem as spectacular so far) have been experimenting with embeddings, as per the previous competitions. So I had a showerthought about merging the two in a non-ensembled manner, and... as per usual, of course someone had already beat me to experimenting with that. Take a look at <a href=\"https://github.com/nadbordrozd/blog_stuff/blob/master/classification_w2v/benchmarking_python3.ipynb\">this fascinating notebook</a>.</p>\n\n<p>Joe Eddy (<a href=\"/aquatic\">@aquatic</a>) mentions in <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/56798#329097\">this thread</a> about <em>average the embedding values across a description to get description-level vectors</em>, and that's exactly what this notebook does. The author compares classification results of using some classical nb and svc, fit on countvectorizer, to mean document embeddings generated from glove and w2v. Then, he/she experiment with the same mean embeddings but this time first multiplying each word vec by it's IDF before averaging them.</p>",
      "rawMarkdown": "A good number of kernels are going the traditional route with CountVectorizer/TF-IDF, and some brave souls (I say brave because training is slower and the results don't seem as spectacular so far) have been experimenting with embeddings, as per the previous competitions. So I had a showerthought about merging the two in a non-ensembled manner, and... as per usual, of course someone had already beat me to experimenting with that. Take a look at [this fascinating notebook][2].\n\nJoe Eddy (@aquatic) mentions in [this thread][1] about *average the embedding values across a description to get description-level vectors*, and that's exactly what this notebook does. The author compares classification results of using some classical nb and svc, fit on countvectorizer, to mean document embeddings generated from glove and w2v. Then, he/she experiment with the same mean embeddings but this time first multiplying each word vec by it's IDF before averaging them.\n\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/56798#329097\n  [2]: https://github.com/nadbordrozd/blog_stuff/blob/master/classification_w2v/benchmarking_python3.ipynb",
      "votes": null
    },
    {
      "id": "329377",
      "postDate": "05/16/2018 11:03:47",
      "content": "<p>Did you try it on this data? Did it improve the score?</p>",
      "rawMarkdown": "Did you try it on this data? Did it improve the score?",
      "votes": null
    },
    {
      "id": "329381",
      "postDate": "05/16/2018 11:13:36",
      "content": "<p>This week's plan</p>",
      "rawMarkdown": "This week's plan",
      "votes": null
    },
    {
      "id": "329971",
      "postDate": "05/17/2018 17:58:52",
      "content": "<p>Another interesting link you guys might want to check out: <a href=\"http://srome.github.io//Leveraging-Factorization-Machines-for-Sparse-Data-and-Supervised-Visualization/\">Leveraging Factorization Machines for Wide Sparse Data and Supervised Visualization</a></p>",
      "rawMarkdown": "Another interesting link you guys might want to check out: [Leveraging Factorization Machines for Wide Sparse Data and Supervised Visualization][1]\n\n\n  [1]: http://srome.github.io//Leveraging-Factorization-Machines-for-Sparse-Data-and-Supervised-Visualization/",
      "votes": null
    },
    {
      "id": "330038",
      "postDate": "05/17/2018 22:29:29",
      "content": "<p>I'm excited to see the results from this sort of weighted embedding.</p>",
      "rawMarkdown": "I'm excited to see the results from this sort of weighted embedding.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 329377,
      "author_name": "michaelsnell",
      "author_url": "",
      "post_date": "05/16/2018 11:03:47",
      "content": "<p>Did you try it on this data? Did it improve the score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 329381,
          "author_name": "authman",
          "author_url": "",
          "post_date": "05/16/2018 11:13:36",
          "content": "<p>This week's plan</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 329971,
      "author_name": "authman",
      "author_url": "",
      "post_date": "05/17/2018 17:58:52",
      "content": "<p>Another interesting link you guys might want to check out: <a href=\"http://srome.github.io//Leveraging-Factorization-Machines-for-Sparse-Data-and-Supervised-Visualization/\">Leveraging Factorization Machines for Wide Sparse Data and Supervised Visualization</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330038,
      "author_name": "mjs2600",
      "author_url": "",
      "post_date": "05/17/2018 22:29:29",
      "content": "<p>I'm excited to see the results from this sort of weighted embedding.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "329375": "A good number of kernels are going the traditional route with CountVectorizer/TF-IDF, and some brave souls (I say brave because training is slower and the results don't seem as spectacular so far) have been experimenting with embeddings, as per the previous competitions. So I had a showerthought about merging the two in a non-ensembled manner, and... as per usual, of course someone had already beat me to experimenting with that. Take a look at [this fascinating notebook][2].\n\nJoe Eddy (@aquatic) mentions in [this thread][1] about *average the embedding values across a description to get description-level vectors*, and that's exactly what this notebook does. The author compares classification results of using some classical nb and svc, fit on countvectorizer, to mean document embeddings generated from glove and w2v. Then, he/she experiment with the same mean embeddings but this time first multiplying each word vec by it's IDF before averaging them.\n\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/56798#329097\n  [2]: https://github.com/nadbordrozd/blog_stuff/blob/master/classification_w2v/benchmarking_python3.ipynb",
    "329377": "Did you try it on this data? Did it improve the score?",
    "329381": "This week's plan",
    "329971": "Another interesting link you guys might want to check out: [Leveraging Factorization Machines for Wide Sparse Data and Supervised Visualization][1]\n\n\n  [1]: http://srome.github.io//Leveraging-Factorization-Machines-for-Sparse-Data-and-Supervised-Visualization/",
    "330038": "I'm excited to see the results from this sort of weighted embedding."
  },
  "source": "meta"
}