{
  "id": 57086,
  "title": "user_id embedding woes",
  "url": "/competitions/avito-demand-prediction/discussion/57086",
  "author_name": "",
  "post_date": "2018-05-18T23:48:05.697476600Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I did a lot of experimenting today attempting to get user_id to not overfit my NN. No matter what depth I put it in my network, no matter how many (or few) neurons I dedicate to the embedding, it seems to cause the net to drastically overfit each time every time.</p>\n\n<p>The last version I tried had just a 1-d embedding for userid in the penultimate concatenation right before the final sigmoid (191 other neurons feeding into that layer) and it still overfit against that one parameter. I've tried 85% dropout with no luck. Combined 85% dropout with 50% GaussNoise? No luck. Tried replacing both with GaussDropout to no avail. Currently I'm using k=5 fold cv.</p>\n\n<p>What interests me is that I've ran many 1 and 5 fold LGBM kernels that <em>do</em> directly use user_id as a categorical variable. They too overfit ~1930Val 2240LB, but they have reasonable LB. My NNets are getting ~1850Val and 2258LB when I add in user_id inputs, but the same architecture w/o user id scores 2212Val, 2242LB... which is considerably less overfit than even the popular lgbm models.</p>\n\n<p>Have any of you been able to successfully use user_id directly as an embedded variable in your neural networks? I actually started working on putting together a 5-fold validation set were there would only be 22.2% intersection of user_ids between folds just like train vs test; but then I decided not to pursue that because as mentioned, plenty of LGB kernels out there that use 1fold random holdout and are doing fine with it directly without any hard CV lifting.</p>",
  "messages": [
    {
      "id": "330491",
      "postDate": "05/18/2018 23:48:05",
      "content": "<p>I did a lot of experimenting today attempting to get user_id to not overfit my NN. No matter what depth I put it in my network, no matter how many (or few) neurons I dedicate to the embedding, it seems to cause the net to drastically overfit each time every time.</p>\n\n<p>The last version I tried had just a 1-d embedding for userid in the penultimate concatenation right before the final sigmoid (191 other neurons feeding into that layer) and it still overfit against that one parameter. I've tried 85% dropout with no luck. Combined 85% dropout with 50% GaussNoise? No luck. Tried replacing both with GaussDropout to no avail. Currently I'm using k=5 fold cv.</p>\n\n<p>What interests me is that I've ran many 1 and 5 fold LGBM kernels that <em>do</em> directly use user_id as a categorical variable. They too overfit ~1930Val 2240LB, but they have reasonable LB. My NNets are getting ~1850Val and 2258LB when I add in user_id inputs, but the same architecture w/o user id scores 2212Val, 2242LB... which is considerably less overfit than even the popular lgbm models.</p>\n\n<p>Have any of you been able to successfully use user_id directly as an embedded variable in your neural networks? I actually started working on putting together a 5-fold validation set were there would only be 22.2% intersection of user_ids between folds just like train vs test; but then I decided not to pursue that because as mentioned, plenty of LGB kernels out there that use 1fold random holdout and are doing fine with it directly without any hard CV lifting.</p>",
      "rawMarkdown": "I did a lot of experimenting today attempting to get user_id to not overfit my NN. No matter what depth I put it in my network, no matter how many (or few) neurons I dedicate to the embedding, it seems to cause the net to drastically overfit each time every time.\n\nThe last version I tried had just a 1-d embedding for userid in the penultimate concatenation right before the final sigmoid (191 other neurons feeding into that layer) and it still overfit against that one parameter. I've tried 85% dropout with no luck. Combined 85% dropout with 50% GaussNoise? No luck. Tried replacing both with GaussDropout to no avail. Currently I'm using k=5 fold cv.\n\nWhat interests me is that I've ran many 1 and 5 fold LGBM kernels that _do_ directly use user_id as a categorical variable. They too overfit ~1930Val 2240LB, but they have reasonable LB. My NNets are getting ~1850Val and 2258LB when I add in user_id inputs, but the same architecture w/o user id scores 2212Val, 2242LB... which is considerably less overfit than even the popular lgbm models.\n\nHave any of you been able to successfully use user_id directly as an embedded variable in your neural networks? I actually started working on putting together a 5-fold validation set were there would only be 22.2% intersection of user_ids between folds just like train vs test; but then I decided not to pursue that because as mentioned, plenty of LGB kernels out there that use 1fold random holdout and are doing fine with it directly without any hard CV lifting.",
      "votes": null
    },
    {
      "id": "330499",
      "postDate": "05/19/2018 01:07:26",
      "content": "<p>I prefer using user_id to engineer variables and then dropping it as a variable. I don't include user_id in my actual model.</p>",
      "rawMarkdown": "I prefer using user_id to engineer variables and then dropping it as a variable. I don't include user_id in my actual model.",
      "votes": null
    },
    {
      "id": "330697",
      "postDate": "05/19/2018 13:32:47",
      "content": "<p>I've noticed a similar problem-- I've successfully dealt with similar problems in the past by using a dot layer to constrain the number of connections between ID and other variables, but user_id still seems to overfit even with that trick. I think it's the nature of embeddings with a similar order of magnitude to the dataset- if there are only a couple examples for a user, the network will converge to those probabilities during training for that user, even if they have totally different items in the test set. </p>\n\n<p>TL;DR: I'm with Peter- use engineered variables to capture the information in user_id and then drop it. </p>",
      "rawMarkdown": "I've noticed a similar problem-- I've successfully dealt with similar problems in the past by using a dot layer to constrain the number of connections between ID and other variables, but user_id still seems to overfit even with that trick. I think it's the nature of embeddings with a similar order of magnitude to the dataset- if there are only a couple examples for a user, the network will converge to those probabilities during training for that user, even if they have totally different items in the test set. \n\nTL;DR: I'm with Peter- use engineered variables to capture the information in user_id and then drop it.",
      "votes": null
    },
    {
      "id": "330911",
      "postDate": "05/20/2018 02:41:03",
      "content": "<p>l2 regularization on weights helps a bit. </p>\n\n<p>Otherwise as pointed out by others, add more features which capture user related information.</p>",
      "rawMarkdown": "l2 regularization on weights helps a bit. \n\nOtherwise as pointed out by others, add more features which capture user related information.",
      "votes": null
    },
    {
      "id": "331228",
      "postDate": "05/20/2018 17:14:03",
      "content": "<p>Have you tried to group up some user based on a share of data parameter ? For example if user is not present in at least 100 items then put him in a special class ?</p>\n\n<p>Then treat that parameter as a regularization parameter.</p>",
      "rawMarkdown": "Have you tried to group up some user based on a share of data parameter ? For example if user is not present in at least 100 items then put him in a special class ?\n\nThen treat that parameter as a regularization parameter.",
      "votes": null
    },
    {
      "id": "331453",
      "postDate": "05/21/2018 09:18:09",
      "content": "<p>I don't think its wise to be embedding the user ID. Each embedding is trained independently, meaning that if you only have 1 sample of a specific embedding that embedding will (almost certainly) overfit no matter what you do in terms of regularization. </p>\n\n<p>In addition to not embedding user IDs, you should also be replacing almost any infrequent categories with the \"missing\" value. In the worse case scenario, you should have an embedding that isn't even trained at all, or is trained off only one example. Its better to just change that to the \"missing\" value, which can be trained into something the model actually has a chance of utilizing. Literature suggests minimum of 10, but I suggest higher (i use 50) especially if you intend to use more than size 3-5 embeddings. </p>\n\n<p>Affirming: Build features from the information, then drop the column. </p>",
      "rawMarkdown": "I don't think its wise to be embedding the user ID. Each embedding is trained independently, meaning that if you only have 1 sample of a specific embedding that embedding will (almost certainly) overfit no matter what you do in terms of regularization. \n\nIn addition to not embedding user IDs, you should also be replacing almost any infrequent categories with the \"missing\" value. In the worse case scenario, you should have an embedding that isn't even trained at all, or is trained off only one example. Its better to just change that to the \"missing\" value, which can be trained into something the model actually has a chance of utilizing. Literature suggests minimum of 10, but I suggest higher (i use 50) especially if you intend to use more than size 3-5 embeddings. \n\nAffirming: Build features from the information, then drop the column.",
      "votes": null
    },
    {
      "id": "331455",
      "postDate": "05/21/2018 09:42:01",
      "content": "<p>+1.</p>\n\n<p>Do you have the paper title which suggest min of 10? Thanks </p>",
      "rawMarkdown": "1.\n\nDo you have the paper title which suggest min of 10? Thanks",
      "votes": null
    },
    {
      "id": "331476",
      "postDate": "05/21/2018 11:11:36",
      "content": "<p>The minimum of 10 is a heuristic hold over from logistic regression, you can read about it in that context here: <a href=\"https://en.wikipedia.org/wiki/One_in_ten_rule\">https://en.wikipedia.org/wiki/One_in_ten_rule</a></p>",
      "rawMarkdown": "The minimum of 10 is a heuristic hold over from logistic regression, you can read about it in that context here: https://en.wikipedia.org/wiki/One_in_ten_rule",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 330499,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "05/19/2018 01:07:26",
      "content": "<p>I prefer using user_id to engineer variables and then dropping it as a variable. I don't include user_id in my actual model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330697,
      "author_name": "rhgrossm",
      "author_url": "",
      "post_date": "05/19/2018 13:32:47",
      "content": "<p>I've noticed a similar problem-- I've successfully dealt with similar problems in the past by using a dot layer to constrain the number of connections between ID and other variables, but user_id still seems to overfit even with that trick. I think it's the nature of embeddings with a similar order of magnitude to the dataset- if there are only a couple examples for a user, the network will converge to those probabilities during training for that user, even if they have totally different items in the test set. </p>\n\n<p>TL;DR: I'm with Peter- use engineered variables to capture the information in user_id and then drop it. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330911,
      "author_name": "tezdhar",
      "author_url": "",
      "post_date": "05/20/2018 02:41:03",
      "content": "<p>l2 regularization on weights helps a bit. </p>\n\n<p>Otherwise as pointed out by others, add more features which capture user related information.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 331228,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "05/20/2018 17:14:03",
      "content": "<p>Have you tried to group up some user based on a share of data parameter ? For example if user is not present in at least 100 items then put him in a special class ?</p>\n\n<p>Then treat that parameter as a regularization parameter.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 331453,
      "author_name": "vannak",
      "author_url": "",
      "post_date": "05/21/2018 09:18:09",
      "content": "<p>I don't think its wise to be embedding the user ID. Each embedding is trained independently, meaning that if you only have 1 sample of a specific embedding that embedding will (almost certainly) overfit no matter what you do in terms of regularization. </p>\n\n<p>In addition to not embedding user IDs, you should also be replacing almost any infrequent categories with the \"missing\" value. In the worse case scenario, you should have an embedding that isn't even trained at all, or is trained off only one example. Its better to just change that to the \"missing\" value, which can be trained into something the model actually has a chance of utilizing. Literature suggests minimum of 10, but I suggest higher (i use 50) especially if you intend to use more than size 3-5 embeddings. </p>\n\n<p>Affirming: Build features from the information, then drop the column. </p>",
      "votes": null,
      "replies": [
        {
          "id": 331455,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "05/21/2018 09:42:01",
          "content": "<p>+1.</p>\n\n<p>Do you have the paper title which suggest min of 10? Thanks </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 331476,
          "author_name": "vannak",
          "author_url": "",
          "post_date": "05/21/2018 11:11:36",
          "content": "<p>The minimum of 10 is a heuristic hold over from logistic regression, you can read about it in that context here: <a href=\"https://en.wikipedia.org/wiki/One_in_ten_rule\">https://en.wikipedia.org/wiki/One_in_ten_rule</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "330491": "I did a lot of experimenting today attempting to get user_id to not overfit my NN. No matter what depth I put it in my network, no matter how many (or few) neurons I dedicate to the embedding, it seems to cause the net to drastically overfit each time every time.\n\nThe last version I tried had just a 1-d embedding for userid in the penultimate concatenation right before the final sigmoid (191 other neurons feeding into that layer) and it still overfit against that one parameter. I've tried 85% dropout with no luck. Combined 85% dropout with 50% GaussNoise? No luck. Tried replacing both with GaussDropout to no avail. Currently I'm using k=5 fold cv.\n\nWhat interests me is that I've ran many 1 and 5 fold LGBM kernels that _do_ directly use user_id as a categorical variable. They too overfit ~1930Val 2240LB, but they have reasonable LB. My NNets are getting ~1850Val and 2258LB when I add in user_id inputs, but the same architecture w/o user id scores 2212Val, 2242LB... which is considerably less overfit than even the popular lgbm models.\n\nHave any of you been able to successfully use user_id directly as an embedded variable in your neural networks? I actually started working on putting together a 5-fold validation set were there would only be 22.2% intersection of user_ids between folds just like train vs test; but then I decided not to pursue that because as mentioned, plenty of LGB kernels out there that use 1fold random holdout and are doing fine with it directly without any hard CV lifting.",
    "330499": "I prefer using user_id to engineer variables and then dropping it as a variable. I don't include user_id in my actual model.",
    "330697": "I've noticed a similar problem-- I've successfully dealt with similar problems in the past by using a dot layer to constrain the number of connections between ID and other variables, but user_id still seems to overfit even with that trick. I think it's the nature of embeddings with a similar order of magnitude to the dataset- if there are only a couple examples for a user, the network will converge to those probabilities during training for that user, even if they have totally different items in the test set. \n\nTL;DR: I'm with Peter- use engineered variables to capture the information in user_id and then drop it.",
    "330911": "l2 regularization on weights helps a bit. \n\nOtherwise as pointed out by others, add more features which capture user related information.",
    "331228": "Have you tried to group up some user based on a share of data parameter ? For example if user is not present in at least 100 items then put him in a special class ?\n\nThen treat that parameter as a regularization parameter.",
    "331453": "I don't think its wise to be embedding the user ID. Each embedding is trained independently, meaning that if you only have 1 sample of a specific embedding that embedding will (almost certainly) overfit no matter what you do in terms of regularization. \n\nIn addition to not embedding user IDs, you should also be replacing almost any infrequent categories with the \"missing\" value. In the worse case scenario, you should have an embedding that isn't even trained at all, or is trained off only one example. Its better to just change that to the \"missing\" value, which can be trained into something the model actually has a chance of utilizing. Literature suggests minimum of 10, but I suggest higher (i use 50) especially if you intend to use more than size 3-5 embeddings. \n\nAffirming: Build features from the information, then drop the column.",
    "331455": "1.\n\nDo you have the paper title which suggest min of 10? Thanks",
    "331476": "The minimum of 10 is a heuristic hold over from logistic regression, you can read about it in that context here: https://en.wikipedia.org/wiki/One_in_ten_rule"
  },
  "source": "meta"
}