{
  "id": 54367,
  "title": "Questions to all the deep learning pros",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54367",
  "author_name": "",
  "post_date": "2018-04-12T14:58:34.703284400Z",
  "votes": 9,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi guys, I noticed that Andy and Alexander have published a series of kernels which features a NN architecture which is quite different from what I usually see in Kaggle and mainstream places. It basically takes in multiple inputs points and pass them through embeddings then concat them together before proceeding with the usual NN workflow.</p>\n\n<p>Is this technique published from some paper? Or perhaps this is a new trend or something?</p>\n\n<p>Thanks for all the incredible kernels! I have learnt a lot ;)</p>",
  "messages": [
    {
      "id": "312941",
      "postDate": "04/12/2018 14:58:34",
      "content": "<p>Hi guys, I noticed that Andy and Alexander have published a series of kernels which features a NN architecture which is quite different from what I usually see in Kaggle and mainstream places. It basically takes in multiple inputs points and pass them through embeddings then concat them together before proceeding with the usual NN workflow.</p>\n\n<p>Is this technique published from some paper? Or perhaps this is a new trend or something?</p>\n\n<p>Thanks for all the incredible kernels! I have learnt a lot ;)</p>",
      "rawMarkdown": "Hi guys, I noticed that Andy and Alexander have published a series of kernels which features a NN architecture which is quite different from what I usually see in Kaggle and mainstream places. It basically takes in multiple inputs points and pass them through embeddings then concat them together before proceeding with the usual NN workflow.\n\nIs this technique published from some paper? Or perhaps this is a new trend or something?\n\nThanks for all the incredible kernels! I have learnt a lot ;)",
      "votes": null
    },
    {
      "id": "312971",
      "postDate": "04/12/2018 15:47:18",
      "content": "<p>Hi,</p>\n\n<p>Yes, there is a paper about it! <a href=\"https://arxiv.org/abs/1604.06737\">Here you go</a>. The entity embedding paper was actually published after the authors successfully applied this technique to a kaggle competition (Rossman, 3rd place). They also <a href=\"https://github.com/entron/entity-embedding-rossmann\">share their code here</a>.</p>\n\n<p>I wouldn't describe it as being quite different from the basic neural network format. It is a relatively new idea though, since the paper was published just 2 years ago. The simpler approach would be to just OHE all the categorical variables and pass them as features to a fully connected layer all at once. Embedding is actually equivalent to this style of one-hot-encoding and passing to a fully connected layer, with the only difference being that each feature is broken out into independent layers. The idea is to learn separate reduced-dimensional dense representations of each feature then combine this information, instead of trying to learn a messy mixed representation of all the categories at once. Empirically, learning separate dense representations first often works better than mixing everything to start.</p>\n\n<p>If you want to see another example of this approach applied to kaggle, you could checkout my <a href=\"https://www.kaggle.com/aquatic/entity-embedding-neural-net\">seguro kernel</a>. Note that in this example categorical features are embedded individually and the dense category representations are then combined with continuous/dense input variables. When you bring in continuous features to combine with categories, embedding often works really well because everything gets treated together as a homogenous dense mixture instead of a binary (OHE)/dense mix. Neural networks typically like homogeneity :)    </p>",
      "rawMarkdown": "Hi,\n\nYes, there is a paper about it! [Here you go][1]. The entity embedding paper was actually published after the authors successfully applied this technique to a kaggle competition (Rossman, 3rd place). They also [share their code here][2].\n\nI wouldn't describe it as being quite different from the basic neural network format. It is a relatively new idea though, since the paper was published just 2 years ago. The simpler approach would be to just OHE all the categorical variables and pass them as features to a fully connected layer all at once. Embedding is actually equivalent to this style of one-hot-encoding and passing to a fully connected layer, with the only difference being that each feature is broken out into independent layers. The idea is to learn separate reduced-dimensional dense representations of each feature then combine this information, instead of trying to learn a messy mixed representation of all the categories at once. Empirically, learning separate dense representations first often works better than mixing everything to start.\n\nIf you want to see another example of this approach applied to kaggle, you could checkout my [seguro kernel][3]. Note that in this example categorical features are embedded individually and the dense category representations are then combined with continuous/dense input variables. When you bring in continuous features to combine with categories, embedding often works really well because everything gets treated together as a homogenous dense mixture instead of a binary (OHE)/dense mix. Neural networks typically like homogeneity :)    \n\n\n  [1]: https://arxiv.org/abs/1604.06737\n  [2]: https://github.com/entron/entity-embedding-rossmann\n  [3]: https://www.kaggle.com/aquatic/entity-embedding-neural-net",
      "votes": null
    },
    {
      "id": "313051",
      "postDate": "04/12/2018 18:06:02",
      "content": "<p>The advice that you should one hot encode categorical variables is extremely old. What's more recent is using embeddings. Initially, embeddings were used when the category count got absurdly high, like with word vocabulary. More recently embeddings are more popular for even small embedding vocabularies (I think the smallest list of unique IDs is still over 200). This is specifically something I picked up on Kaggle, not something I actually came across in research. </p>\n\n<p>Not every value should be embedded- The basic rule of thumb is: If you rename classes {1, 2, 3} as {A, B, C}, do you lose information? In the case of our values: IP, App, Device, OS, Channel, the answer is yes, we could do that. So embedding should work. </p>",
      "rawMarkdown": "The advice that you should one hot encode categorical variables is extremely old. What's more recent is using embeddings. Initially, embeddings were used when the category count got absurdly high, like with word vocabulary. More recently embeddings are more popular for even small embedding vocabularies (I think the smallest list of unique IDs is still over 200). This is specifically something I picked up on Kaggle, not something I actually came across in research. \n\n\nNot every value should be embedded- The basic rule of thumb is: If you rename classes {1, 2, 3} as {A, B, C}, do you lose information? In the case of our values: IP, App, Device, OS, Channel, the answer is yes, we could do that. So embedding should work.",
      "votes": null
    },
    {
      "id": "313053",
      "postDate": "04/12/2018 18:13:18",
      "content": "<p>This might be a bit pedantic, but I do want to emphasize the point that embeddings still use OHE, so it's certainly not out of fashion. The only difference between embeddings and the traditional method is that in embedding, the OHE features for each category are passed through a category-specific fully connected layer before merging all the category information together in later layers.</p>",
      "rawMarkdown": "This might be a bit pedantic, but I do want to emphasize the point that embeddings still use OHE, so it's certainly not out of fashion. The only difference between embeddings and the traditional method is that in embedding, the OHE features for each category are passed through a category-specific fully connected layer before merging all the category information together in later layers.",
      "votes": null
    },
    {
      "id": "313112",
      "postDate": "04/12/2018 19:40:21",
      "content": "<p>Good night )) my friends, maybe, you have some thoughts about metric for nn? I believe this can help to improve performance.</p>",
      "rawMarkdown": "Good night )) my friends, maybe, you have some thoughts about metric for nn? I believe this can help to improve performance.",
      "votes": null
    },
    {
      "id": "313181",
      "postDate": "04/12/2018 21:44:17",
      "content": "<p>Just wanted to say thanks for the thorough response, much appreciated! Posts like this are the reason I and many other newcomers join competitions.</p>",
      "rawMarkdown": "Just wanted to say thanks for the thorough response, much appreciated! Posts like this are the reason I and many other newcomers join competitions.",
      "votes": null
    },
    {
      "id": "313263",
      "postDate": "04/13/2018 02:27:01",
      "content": "<p>Thanks for the kind words Anthony! I'm glad it was helpful. </p>",
      "rawMarkdown": "Thanks for the kind words Anthony! I'm glad it was helpful.",
      "votes": null
    },
    {
      "id": "313288",
      "postDate": "04/13/2018 03:47:39",
      "content": "<p>@Joe this is quite informative, frankly speaking, I was unaware of anything like this before this great write-up of yours. I just wanted to learn more from your experiences of using embeddings as opposed to OHE on traditional datasets to create features and use them with tree-based or other methods?</p>",
      "rawMarkdown": "Joe this is quite informative, frankly speaking, I was unaware of anything like this before this great write-up of yours. I just wanted to learn more from your experiences of using embeddings as opposed to OHE on traditional datasets to create features and use them with tree-based or other methods?",
      "votes": null
    },
    {
      "id": "313503",
      "postDate": "04/13/2018 11:15:18",
      "content": "<p>Not pedantic at all, perfectly correct and informative!</p>",
      "rawMarkdown": "Not pedantic at all, perfectly correct and informative!",
      "votes": null
    },
    {
      "id": "313508",
      "postDate": "04/13/2018 11:39:03",
      "content": "<p>Thanks shivraj! </p>\n\n<p>The idea you bring up is a very cool technique. You can train a neural network with embeddings, then extract the embedded feature layer as a substitute for the categorical features to feed to a different model. I did this with lightgbm in the seguro competition. The results were comparable to boosting models without the embedding features, but the really interesting thing was that the predictions were different enough to provide a small ensembling gain. I've only tried this approach once, so I can't make any more general claims about how well it tends to work (though the authors of the EE paper also provide some evidence that it works well).</p>",
      "rawMarkdown": "Thanks shivraj! \n\nThe idea you bring up is a very cool technique. You can train a neural network with embeddings, then extract the embedded feature layer as a substitute for the categorical features to feed to a different model. I did this with lightgbm in the seguro competition. The results were comparable to boosting models without the embedding features, but the really interesting thing was that the predictions were different enough to provide a small ensembling gain. I've only tried this approach once, so I can't make any more general claims about how well it tends to work (though the authors of the EE paper also provide some evidence that it works well).",
      "votes": null
    },
    {
      "id": "313511",
      "postDate": "04/13/2018 11:47:35",
      "content": "<p>I also used embeddings with xgb and lgb in Porto Seguro.  The only drawback is that they can lead to a bit of overfit as they encode some of the target information.  Well, if you get these embeddings from a NN that learns to predict the target.  A safer approach was the winner solution one, from Michael Jahrer, where he trained autoencoders and used their weights as features.  This is safer, because the NN producing embeddings do not use the target at all.</p>",
      "rawMarkdown": "I also used embeddings with xgb and lgb in Porto Seguro.  The only drawback is that they can lead to a bit of overfit as they encode some of the target information.  Well, if you get these embeddings from a NN that learns to predict the target.  A safer approach was the winner solution one, from Michael Jahrer, where he trained autoencoders and used their weights as features.  This is safer, because the NN producing embeddings do not use the target at all.",
      "votes": null
    },
    {
      "id": "313655",
      "postDate": "04/13/2018 16:28:47",
      "content": "<p>I used embeddings from a target-trained NN as GB features in the Recruit competition, in several different variations. IIRC it worked best when I added the embeddings to an already otherwise complete set of features, and there was some improvement and definitely some diversity value.</p>",
      "rawMarkdown": "I used embeddings from a target-trained NN as GB features in the Recruit competition, in several different variations. IIRC it worked best when I added the embeddings to an already otherwise complete set of features, and there was some improvement and definitely some diversity value.",
      "votes": null
    },
    {
      "id": "313756",
      "postDate": "04/13/2018 19:40:40",
      "content": "<p>@Joe @CPMP @Andy Thanks for shedding light based on your experience, I'll definitely have a go with this and explore further.</p>",
      "rawMarkdown": "Joe @CPMP @Andy Thanks for shedding light based on your experience, I'll definitely have a go with this and explore further.",
      "votes": null
    },
    {
      "id": "313849",
      "postDate": "04/14/2018 00:23:04",
      "content": "<p>F1 score works pretty well. </p>",
      "rawMarkdown": "F1 score works pretty well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 312971,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "04/12/2018 15:47:18",
      "content": "<p>Hi,</p>\n\n<p>Yes, there is a paper about it! <a href=\"https://arxiv.org/abs/1604.06737\">Here you go</a>. The entity embedding paper was actually published after the authors successfully applied this technique to a kaggle competition (Rossman, 3rd place). They also <a href=\"https://github.com/entron/entity-embedding-rossmann\">share their code here</a>.</p>\n\n<p>I wouldn't describe it as being quite different from the basic neural network format. It is a relatively new idea though, since the paper was published just 2 years ago. The simpler approach would be to just OHE all the categorical variables and pass them as features to a fully connected layer all at once. Embedding is actually equivalent to this style of one-hot-encoding and passing to a fully connected layer, with the only difference being that each feature is broken out into independent layers. The idea is to learn separate reduced-dimensional dense representations of each feature then combine this information, instead of trying to learn a messy mixed representation of all the categories at once. Empirically, learning separate dense representations first often works better than mixing everything to start.</p>\n\n<p>If you want to see another example of this approach applied to kaggle, you could checkout my <a href=\"https://www.kaggle.com/aquatic/entity-embedding-neural-net\">seguro kernel</a>. Note that in this example categorical features are embedded individually and the dense category representations are then combined with continuous/dense input variables. When you bring in continuous features to combine with categories, embedding often works really well because everything gets treated together as a homogenous dense mixture instead of a binary (OHE)/dense mix. Neural networks typically like homogeneity :)    </p>",
      "votes": null,
      "replies": [
        {
          "id": 313181,
          "author_name": "antmarakis",
          "author_url": "",
          "post_date": "04/12/2018 21:44:17",
          "content": "<p>Just wanted to say thanks for the thorough response, much appreciated! Posts like this are the reason I and many other newcomers join competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313263,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/13/2018 02:27:01",
          "content": "<p>Thanks for the kind words Anthony! I'm glad it was helpful. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313288,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "04/13/2018 03:47:39",
          "content": "<p>@Joe this is quite informative, frankly speaking, I was unaware of anything like this before this great write-up of yours. I just wanted to learn more from your experiences of using embeddings as opposed to OHE on traditional datasets to create features and use them with tree-based or other methods?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313508,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/13/2018 11:39:03",
          "content": "<p>Thanks shivraj! </p>\n\n<p>The idea you bring up is a very cool technique. You can train a neural network with embeddings, then extract the embedded feature layer as a substitute for the categorical features to feed to a different model. I did this with lightgbm in the seguro competition. The results were comparable to boosting models without the embedding features, but the really interesting thing was that the predictions were different enough to provide a small ensembling gain. I've only tried this approach once, so I can't make any more general claims about how well it tends to work (though the authors of the EE paper also provide some evidence that it works well).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313511,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/13/2018 11:47:35",
          "content": "<p>I also used embeddings with xgb and lgb in Porto Seguro.  The only drawback is that they can lead to a bit of overfit as they encode some of the target information.  Well, if you get these embeddings from a NN that learns to predict the target.  A safer approach was the winner solution one, from Michael Jahrer, where he trained autoencoders and used their weights as features.  This is safer, because the NN producing embeddings do not use the target at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313655,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "04/13/2018 16:28:47",
          "content": "<p>I used embeddings from a target-trained NN as GB features in the Recruit competition, in several different variations. IIRC it worked best when I added the embeddings to an already otherwise complete set of features, and there was some improvement and definitely some diversity value.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313756,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "04/13/2018 19:40:40",
          "content": "<p>@Joe @CPMP @Andy Thanks for shedding light based on your experience, I'll definitely have a go with this and explore further.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 313051,
      "author_name": "vannak",
      "author_url": "",
      "post_date": "04/12/2018 18:06:02",
      "content": "<p>The advice that you should one hot encode categorical variables is extremely old. What's more recent is using embeddings. Initially, embeddings were used when the category count got absurdly high, like with word vocabulary. More recently embeddings are more popular for even small embedding vocabularies (I think the smallest list of unique IDs is still over 200). This is specifically something I picked up on Kaggle, not something I actually came across in research. </p>\n\n<p>Not every value should be embedded- The basic rule of thumb is: If you rename classes {1, 2, 3} as {A, B, C}, do you lose information? In the case of our values: IP, App, Device, OS, Channel, the answer is yes, we could do that. So embedding should work. </p>",
      "votes": null,
      "replies": [
        {
          "id": 313053,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/12/2018 18:13:18",
          "content": "<p>This might be a bit pedantic, but I do want to emphasize the point that embeddings still use OHE, so it's certainly not out of fashion. The only difference between embeddings and the traditional method is that in embedding, the OHE features for each category are passed through a category-specific fully connected layer before merging all the category information together in later layers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313503,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/13/2018 11:15:18",
          "content": "<p>Not pedantic at all, perfectly correct and informative!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 313112,
      "author_name": "alexanderkireev",
      "author_url": "",
      "post_date": "04/12/2018 19:40:21",
      "content": "<p>Good night )) my friends, maybe, you have some thoughts about metric for nn? I believe this can help to improve performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 313849,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "04/14/2018 00:23:04",
          "content": "<p>F1 score works pretty well. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "312941": "Hi guys, I noticed that Andy and Alexander have published a series of kernels which features a NN architecture which is quite different from what I usually see in Kaggle and mainstream places. It basically takes in multiple inputs points and pass them through embeddings then concat them together before proceeding with the usual NN workflow.\n\nIs this technique published from some paper? Or perhaps this is a new trend or something?\n\nThanks for all the incredible kernels! I have learnt a lot ;)",
    "312971": "Hi,\n\nYes, there is a paper about it! [Here you go][1]. The entity embedding paper was actually published after the authors successfully applied this technique to a kaggle competition (Rossman, 3rd place). They also [share their code here][2].\n\nI wouldn't describe it as being quite different from the basic neural network format. It is a relatively new idea though, since the paper was published just 2 years ago. The simpler approach would be to just OHE all the categorical variables and pass them as features to a fully connected layer all at once. Embedding is actually equivalent to this style of one-hot-encoding and passing to a fully connected layer, with the only difference being that each feature is broken out into independent layers. The idea is to learn separate reduced-dimensional dense representations of each feature then combine this information, instead of trying to learn a messy mixed representation of all the categories at once. Empirically, learning separate dense representations first often works better than mixing everything to start.\n\nIf you want to see another example of this approach applied to kaggle, you could checkout my [seguro kernel][3]. Note that in this example categorical features are embedded individually and the dense category representations are then combined with continuous/dense input variables. When you bring in continuous features to combine with categories, embedding often works really well because everything gets treated together as a homogenous dense mixture instead of a binary (OHE)/dense mix. Neural networks typically like homogeneity :)    \n\n\n  [1]: https://arxiv.org/abs/1604.06737\n  [2]: https://github.com/entron/entity-embedding-rossmann\n  [3]: https://www.kaggle.com/aquatic/entity-embedding-neural-net",
    "313051": "The advice that you should one hot encode categorical variables is extremely old. What's more recent is using embeddings. Initially, embeddings were used when the category count got absurdly high, like with word vocabulary. More recently embeddings are more popular for even small embedding vocabularies (I think the smallest list of unique IDs is still over 200). This is specifically something I picked up on Kaggle, not something I actually came across in research. \n\n\nNot every value should be embedded- The basic rule of thumb is: If you rename classes {1, 2, 3} as {A, B, C}, do you lose information? In the case of our values: IP, App, Device, OS, Channel, the answer is yes, we could do that. So embedding should work.",
    "313053": "This might be a bit pedantic, but I do want to emphasize the point that embeddings still use OHE, so it's certainly not out of fashion. The only difference between embeddings and the traditional method is that in embedding, the OHE features for each category are passed through a category-specific fully connected layer before merging all the category information together in later layers.",
    "313112": "Good night )) my friends, maybe, you have some thoughts about metric for nn? I believe this can help to improve performance.",
    "313181": "Just wanted to say thanks for the thorough response, much appreciated! Posts like this are the reason I and many other newcomers join competitions.",
    "313263": "Thanks for the kind words Anthony! I'm glad it was helpful.",
    "313288": "Joe this is quite informative, frankly speaking, I was unaware of anything like this before this great write-up of yours. I just wanted to learn more from your experiences of using embeddings as opposed to OHE on traditional datasets to create features and use them with tree-based or other methods?",
    "313503": "Not pedantic at all, perfectly correct and informative!",
    "313508": "Thanks shivraj! \n\nThe idea you bring up is a very cool technique. You can train a neural network with embeddings, then extract the embedded feature layer as a substitute for the categorical features to feed to a different model. I did this with lightgbm in the seguro competition. The results were comparable to boosting models without the embedding features, but the really interesting thing was that the predictions were different enough to provide a small ensembling gain. I've only tried this approach once, so I can't make any more general claims about how well it tends to work (though the authors of the EE paper also provide some evidence that it works well).",
    "313511": "I also used embeddings with xgb and lgb in Porto Seguro.  The only drawback is that they can lead to a bit of overfit as they encode some of the target information.  Well, if you get these embeddings from a NN that learns to predict the target.  A safer approach was the winner solution one, from Michael Jahrer, where he trained autoencoders and used their weights as features.  This is safer, because the NN producing embeddings do not use the target at all.",
    "313655": "I used embeddings from a target-trained NN as GB features in the Recruit competition, in several different variations. IIRC it worked best when I added the embeddings to an already otherwise complete set of features, and there was some improvement and definitely some diversity value.",
    "313756": "Joe @CPMP @Andy Thanks for shedding light based on your experience, I'll definitely have a go with this and explore further.",
    "313849": "F1 score works pretty well."
  },
  "source": "meta"
}