{
  "id": 79162,
  "title": "Reducing embedding dimensions",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79162",
  "author_name": "",
  "post_date": "2019-01-31T17:56:25.354700Z",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Here's a fun trick which I'm still playing with, but you might like: <a href=\"https://www.kaggle.com/hamishdickson/single-rnn-with-4-folds-pca/\">https://www.kaggle.com/hamishdickson/single-rnn-with-4-folds-pca/</a></p>\n\n<p>Basically you can use PCA (or I guess t-SNE etc) on the embeddings to reduce the dimensionality, if you try and keep say 95% of the variance you can reduce the embedding size to around 250</p>\n\n<p>This gives your model less things to learn/speeds it up etc but comes at the cost of maybe your embeddings aren't as good.... having said that I've not done any analysis into how good the embeddings were previously, so they could be they were terrible to start with :)</p>\n\n<p>I'm still playing/confirming things, but this kernel attached seems to get a +0.007LB boost and a bit less in the CV from doing this</p>",
  "messages": [
    {
      "id": "464360",
      "postDate": "01/31/2019 17:56:25",
      "content": "<p>Here's a fun trick which I'm still playing with, but you might like: <a href=\"https://www.kaggle.com/hamishdickson/single-rnn-with-4-folds-pca/\">https://www.kaggle.com/hamishdickson/single-rnn-with-4-folds-pca/</a></p>\n\n<p>Basically you can use PCA (or I guess t-SNE etc) on the embeddings to reduce the dimensionality, if you try and keep say 95% of the variance you can reduce the embedding size to around 250</p>\n\n<p>This gives your model less things to learn/speeds it up etc but comes at the cost of maybe your embeddings aren't as good.... having said that I've not done any analysis into how good the embeddings were previously, so they could be they were terrible to start with :)</p>\n\n<p>I'm still playing/confirming things, but this kernel attached seems to get a +0.007LB boost and a bit less in the CV from doing this</p>",
      "rawMarkdown": "Here's a fun trick which I'm still playing with, but you might like: https://www.kaggle.com/hamishdickson/single-rnn-with-4-folds-pca/\n\nBasically you can use PCA (or I guess t-SNE etc) on the embeddings to reduce the dimensionality, if you try and keep say 95% of the variance you can reduce the embedding size to around 250\n\nThis gives your model less things to learn/speeds it up etc but comes at the cost of maybe your embeddings aren't as good.... having said that I've not done any analysis into how good the embeddings were previously, so they could be they were terrible to start with :)\n\nI'm still playing/confirming things, but this kernel attached seems to get a +0.007LB boost and a bit less in the CV from doing this",
      "votes": null
    },
    {
      "id": "464388",
      "postDate": "01/31/2019 18:51:34",
      "content": "<p>Interesting! I just read a paper about finding optimal embedding dimension. It might help when using PCA <a href=\"https://papers.nips.cc/paper/7368-on-the-dimensionality-of-word-embedding.pdf\">https://papers.nips.cc/paper/7368-on-the-dimensionality-of-word-embedding.pdf</a></p>",
      "rawMarkdown": "Interesting! I just read a paper about finding optimal embedding dimension. It might help when using PCA https://papers.nips.cc/paper/7368-on-the-dimensionality-of-word-embedding.pdf",
      "votes": null
    },
    {
      "id": "464730",
      "postDate": "02/01/2019 11:43:03",
      "content": "<p><a href=\"/chengham\">@chengham</a> I actually found this paper a few weeks ago and did spend a little bit of time on it. </p>\n\n<p>If you keep your embeddings non-trainable the speed-up isn't huge and I found performance dropped about ~0.001-0.002 using embedding dimensions down to 100 but obviously this depends on what else you are doing. </p>\n\n<p>Sadly every time I have tried training embeddings (or a subset) performance has been worse and I don't have time to try anything more funky at the moment, though do have a bundle of ideas.</p>",
      "rawMarkdown": "chengham I actually found this paper a few weeks ago and did spend a little bit of time on it. \n\nIf you keep your embeddings non-trainable the speed-up isn't huge and I found performance dropped about ~0.001-0.002 using embedding dimensions down to 100 but obviously this depends on what else you are doing. \n\nSadly every time I have tried training embeddings (or a subset) performance has been worse and I don't have time to try anything more funky at the moment, though do have a bundle of ideas.",
      "votes": null
    },
    {
      "id": "464740",
      "postDate": "02/01/2019 12:11:59",
      "content": "<blockquote>\n  <p>Sadly every time I have tried training embeddings (or a subset) performance has been worse</p>\n</blockquote>\n\n<p>I had a similar experience, I think that's because there are too many dimensions, from experience this only really works well when you are in ~10-20D</p>\n\n<p>That's interesting you've experienced a performance drop, I've consistently had a small gain - like you say it probably depends what else is going on - yay another hyperparameter to think about 🙃</p>",
      "rawMarkdown": "&gt; Sadly every time I have tried training embeddings (or a subset) performance has been worse\n\nI had a similar experience, I think that's because there are too many dimensions, from experience this only really works well when you are in ~10-20D\n\nThat's interesting you've experienced a performance drop, I've consistently had a small gain - like you say it probably depends what else is going on - yay another hyperparameter to think about 🙃",
      "votes": null
    },
    {
      "id": "464754",
      "postDate": "02/01/2019 12:38:53",
      "content": "<p>You could just have another set of trainable embeddings that you have reduced down to dimension 20 and then combine them later on with the untrained embeddings.</p>\n\n<p>Is your gain on CV or LB?</p>",
      "rawMarkdown": "You could just have another set of trainable embeddings that you have reduced down to dimension 20 and then combine them later on with the untrained embeddings.\n\nIs your gain on CV or LB?",
      "votes": null
    },
    {
      "id": "464757",
      "postDate": "02/01/2019 12:44:53",
      "content": "<p>Oh that's a cool idea - have you tried that?</p>\n\n<p>Both CV and LB - to be honest I'm not sure how much I trust the LB, theres lots of variance in every model I use. At this point I'm not really expecting to do well in this competition, I'm just trying to push up the CV and try new ideas</p>",
      "rawMarkdown": "Oh that's a cool idea - have you tried that?\n\nBoth CV and LB - to be honest I'm not sure how much I trust the LB, theres lots of variance in every model I use. At this point I'm not really expecting to do well in this competition, I'm just trying to push up the CV and try new ideas",
      "votes": null
    },
    {
      "id": "464763",
      "postDate": "02/01/2019 12:59:37",
      "content": "<p>No I haven't tried that yet and not sure if I will - let me know how it goes if you do.</p>\n\n<p>I would also say don't trust public LB at all.</p>",
      "rawMarkdown": "No I haven't tried that yet and not sure if I will - let me know how it goes if you do.\n\nI would also say don't trust public LB at all.",
      "votes": null
    },
    {
      "id": "464767",
      "postDate": "02/01/2019 13:18:14",
      "content": "<p>Good advice and will do!</p>",
      "rawMarkdown": "Good advice and will do!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 464388,
      "author_name": "chengham",
      "author_url": "",
      "post_date": "01/31/2019 18:51:34",
      "content": "<p>Interesting! I just read a paper about finding optimal embedding dimension. It might help when using PCA <a href=\"https://papers.nips.cc/paper/7368-on-the-dimensionality-of-word-embedding.pdf\">https://papers.nips.cc/paper/7368-on-the-dimensionality-of-word-embedding.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 464730,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "02/01/2019 11:43:03",
      "content": "<p><a href=\"/chengham\">@chengham</a> I actually found this paper a few weeks ago and did spend a little bit of time on it. </p>\n\n<p>If you keep your embeddings non-trainable the speed-up isn't huge and I found performance dropped about ~0.001-0.002 using embedding dimensions down to 100 but obviously this depends on what else you are doing. </p>\n\n<p>Sadly every time I have tried training embeddings (or a subset) performance has been worse and I don't have time to try anything more funky at the moment, though do have a bundle of ideas.</p>",
      "votes": null,
      "replies": [
        {
          "id": 464740,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "02/01/2019 12:11:59",
          "content": "<blockquote>\n  <p>Sadly every time I have tried training embeddings (or a subset) performance has been worse</p>\n</blockquote>\n\n<p>I had a similar experience, I think that's because there are too many dimensions, from experience this only really works well when you are in ~10-20D</p>\n\n<p>That's interesting you've experienced a performance drop, I've consistently had a small gain - like you say it probably depends what else is going on - yay another hyperparameter to think about 🙃</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464754,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "02/01/2019 12:38:53",
          "content": "<p>You could just have another set of trainable embeddings that you have reduced down to dimension 20 and then combine them later on with the untrained embeddings.</p>\n\n<p>Is your gain on CV or LB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464757,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "02/01/2019 12:44:53",
          "content": "<p>Oh that's a cool idea - have you tried that?</p>\n\n<p>Both CV and LB - to be honest I'm not sure how much I trust the LB, theres lots of variance in every model I use. At this point I'm not really expecting to do well in this competition, I'm just trying to push up the CV and try new ideas</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464763,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "02/01/2019 12:59:37",
          "content": "<p>No I haven't tried that yet and not sure if I will - let me know how it goes if you do.</p>\n\n<p>I would also say don't trust public LB at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464767,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "02/01/2019 13:18:14",
          "content": "<p>Good advice and will do!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "464360": "Here's a fun trick which I'm still playing with, but you might like: https://www.kaggle.com/hamishdickson/single-rnn-with-4-folds-pca/\n\nBasically you can use PCA (or I guess t-SNE etc) on the embeddings to reduce the dimensionality, if you try and keep say 95% of the variance you can reduce the embedding size to around 250\n\nThis gives your model less things to learn/speeds it up etc but comes at the cost of maybe your embeddings aren't as good.... having said that I've not done any analysis into how good the embeddings were previously, so they could be they were terrible to start with :)\n\nI'm still playing/confirming things, but this kernel attached seems to get a +0.007LB boost and a bit less in the CV from doing this",
    "464388": "Interesting! I just read a paper about finding optimal embedding dimension. It might help when using PCA https://papers.nips.cc/paper/7368-on-the-dimensionality-of-word-embedding.pdf",
    "464730": "chengham I actually found this paper a few weeks ago and did spend a little bit of time on it. \n\nIf you keep your embeddings non-trainable the speed-up isn't huge and I found performance dropped about ~0.001-0.002 using embedding dimensions down to 100 but obviously this depends on what else you are doing. \n\nSadly every time I have tried training embeddings (or a subset) performance has been worse and I don't have time to try anything more funky at the moment, though do have a bundle of ideas.",
    "464740": "&gt; Sadly every time I have tried training embeddings (or a subset) performance has been worse\n\nI had a similar experience, I think that's because there are too many dimensions, from experience this only really works well when you are in ~10-20D\n\nThat's interesting you've experienced a performance drop, I've consistently had a small gain - like you say it probably depends what else is going on - yay another hyperparameter to think about 🙃",
    "464754": "You could just have another set of trainable embeddings that you have reduced down to dimension 20 and then combine them later on with the untrained embeddings.\n\nIs your gain on CV or LB?",
    "464757": "Oh that's a cool idea - have you tried that?\n\nBoth CV and LB - to be honest I'm not sure how much I trust the LB, theres lots of variance in every model I use. At this point I'm not really expecting to do well in this competition, I'm just trying to push up the CV and try new ideas",
    "464763": "No I haven't tried that yet and not sure if I will - let me know how it goes if you do.\n\nI would also say don't trust public LB at all.",
    "464767": "Good advice and will do!"
  },
  "source": "meta"
}