{
  "id": 71789,
  "title": "Ading   features ! Does it make sense? And how",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71789",
  "author_name": "AcademycalBastard",
  "post_date": "2018-11-16T14:22:13.343000",
  "votes": 7,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I am thinking of  making use of feature engineering  such number  of punctuation signs , number of  voals  and  stuff like that.\nDoes it make  sense from your  guys experience?\nAnd if so  ,  how or  where  could   I add them to this neural nets?  Could I add them as  a new input  for keras ?</p>",
  "messages": [
    {
      "id": 422628,
      "postDate": "2018-11-16T14:22:13.343Z",
      "content": "<p>I am thinking of  making use of feature engineering  such number  of punctuation signs , number of  voals  and  stuff like that.\nDoes it make  sense from your  guys experience?\nAnd if so  ,  how or  where  could   I add them to this neural nets?  Could I add them as  a new input  for keras ?</p>",
      "rawMarkdown": "I am thinking of  making use of feature engineering  such number  of punctuation signs , number of  voals  and  stuff like that.\nDoes it make  sense from your  guys experience?\nAnd if so  ,  how or  where  could   I add them to this neural nets?  Could I add them as  a new input  for keras ?",
      "votes": 7
    },
    {
      "id": 422639,
      "postDate": "2018-11-16T14:47:46.110Z",
      "content": "<p>It could be useful, I'll have to try that.</p>\n\n<p>The way I'll deal with that is this way:\n- I'll use a 2 elements input, being [word embeddings, features]\n- I will feed the word embeddings in a RNN (or a layer that deals with sequential data)\n- When my network output is not sequential anymore (for example after a GlobalMaxPool, an Attention Layer or a RNN with return Sequence = False), I'll concatenate it with my features.\n- Add a Dense layer with one unit to get a score.</p>",
      "rawMarkdown": "It could be useful, I'll have to try that.\n\nThe way I'll deal with that is this way:\n- I'll use a 2 elements input, being [word embeddings, features]\n- I will feed the word embeddings in a RNN (or a layer that deals with sequential data)\n- When my network output is not sequential anymore (for example after a GlobalMaxPool, an Attention Layer or a RNN with return Sequence = False), I'll concatenate it with my features.\n- Add a Dense layer with one unit to get a score.",
      "votes": 6,
      "replies": [
        {
          "id": 422645,
          "postDate": "2018-11-16T15:03:37.793Z",
          "content": "<p>We could  build  some interesting features  togheter ,   do you have some email  to  discusss more   about this ??</p>",
          "rawMarkdown": "We could  build  some interesting features  togheter ,   do you have some email  to  discusss more   about this ??"
        },
        {
          "id": 422655,
          "postDate": "2018-11-16T15:17:44.290Z",
          "content": "<p>I prefer to work on my own, thanks for the offer anyway !</p>",
          "rawMarkdown": "I prefer to work on my own, thanks for the offer anyway !",
          "votes": 1
        },
        {
          "id": 422935,
          "postDate": "2018-11-17T05:37:02.453Z",
          "content": "<p>@AcademycalBastard Note that it is considered cheating if you discuss privately ideas with someone you are not in a team</p>",
          "rawMarkdown": "@AcademycalBastard Note that it is considered cheating if you discuss privately ideas with someone you are not in a team",
          "votes": 2
        },
        {
          "id": 423014,
          "postDate": "2018-11-17T09:12:13.430Z",
          "content": "<p>Yes, is not the case here ,  because I allready  found some to team </p>",
          "rawMarkdown": "Yes, is not the case here ,  because I allready  found some to team "
        }
      ]
    },
    {
      "id": 422737,
      "postDate": "2018-11-16T17:44:59.897Z",
      "content": "<p>Yup, it actually makes a lot of sense, and has been used widely in the Vietnamese NLP community recently with great results. Since I'm pretty inexperienced working with English I'm not sure how it will play out, but I'm currently working on something similar to yours (maybe a bit more complicated actually) .</p>\n\n<p>I'm open for teaming up but I think it's just too early , so we can lose some advantage by merging, it's better to wait for at least a month I think.</p>",
      "rawMarkdown": "Yup, it actually makes a lot of sense, and has been used widely in the Vietnamese NLP community recently with great results. Since I'm pretty inexperienced working with English I'm not sure how it will play out, but I'm currently working on something similar to yours (maybe a bit more complicated actually) .\n\nI'm open for teaming up but I think it's just too early , so we can lose some advantage by merging, it's better to wait for at least a month I think.",
      "votes": 2,
      "replies": [
        {
          "id": 422750,
          "postDate": "2018-11-16T18:18:43.147Z",
          "content": "<p>I  will  be  delighted   to  team  up with   you . I have strong feature engineering  ideeas  only thing  is  to  put  it all togheter in the  NN.   Do you have an email  so I  could send  you  one  of  my   features  I have in mind that  I  think will boost up score  ? Tnks</p>",
          "rawMarkdown": "I  will  be  delighted   to  team  up with   you . I have strong feature engineering  ideeas  only thing  is  to  put  it all togheter in the  NN.   Do you have an email  so I  could send  you  one  of  my   features  I have in mind that  I  think will boost up score  ? Tnks",
          "votes": -2
        },
        {
          "id": 422936,
          "postDate": "2018-11-17T05:37:32.007Z",
          "content": "<p>@AcademycalBastard Note that it is considered cheating if you discuss privately ideas with someone you are not in a team</p>",
          "rawMarkdown": "@AcademycalBastard Note that it is considered cheating if you discuss privately ideas with someone you are not in a team",
          "votes": 3
        }
      ]
    },
    {
      "id": 428047,
      "postDate": "2018-11-26T16:57:00.190Z",
      "content": "<p>Here is an example where I used sentence length as an extra feature (not helpful): <a href=\"https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7596269\">https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7596269</a></p>",
      "rawMarkdown": "Here is an example where I used sentence length as an extra feature (not helpful): https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7596269"
    },
    {
      "id": 422975,
      "postDate": "2018-11-17T06:51:22.857Z",
      "content": "<p>I was toying myself with this idea, and i think i found some features with a decent correlation, but when I concat them to the features found after the Embedding layers is done proccesing, my results are ussualy worse. Now here are two otpion: 1. I am doing it wrong, concatentating in a wrong place ... etc. or my features are garbage..    Does anybody have a good source (paper) about combining engineering features + word embeddings? Thanks </p>",
      "rawMarkdown": "I was toying myself with this idea, and i think i found some features with a decent correlation, but when I concat them to the features found after the Embedding layers is done proccesing, my results are ussualy worse. Now here are two otpion: 1. I am doing it wrong, concatentating in a wrong place ... etc. or my features are garbage..    Does anybody have a good source (paper) about combining engineering features + word embeddings? Thanks ",
      "replies": [
        {
          "id": 423011,
          "postDate": "2018-11-17T08:53:53.013Z",
          "content": "<p>You can check my model of avito demand prediction competition, where I combined all sorts of different input\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59902\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59902</a></p>",
          "rawMarkdown": "You can check my model of avito demand prediction competition, where I combined all sorts of different input\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59902",
          "votes": 1
        },
        {
          "id": 423013,
          "postDate": "2018-11-17T08:59:32.723Z",
          "content": "<p>In general it does not work out of the box, there are lots of small things to take into consideration when combining embeddings with additional features. Do you normalize your features?  </p>",
          "rawMarkdown": "In general it does not work out of the box, there are lots of small things to take into consideration when combining embeddings with additional features. Do you normalize your features?  "
        },
        {
          "id": 423241,
          "postDate": "2018-11-17T18:52:08.750Z",
          "content": "<p>I do batch normlazation before comibing features. <br>\nfor example the output of embedding features are normlaized + output of hand picked features is normlaized too.     </p>",
          "rawMarkdown": "I do batch normlazation before comibing features.  \nfor example the output of embedding features are normlaized + output of hand picked features is normlaized too.     "
        }
      ]
    },
    {
      "id": 423210,
      "postDate": "2018-11-17T17:58:29.540Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 422639,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2018-11-16T14:47:46.110000",
      "content": "<p>It could be useful, I'll have to try that.</p>\n\n<p>The way I'll deal with that is this way:\n- I'll use a 2 elements input, being [word embeddings, features]\n- I will feed the word embeddings in a RNN (or a layer that deals with sequential data)\n- When my network output is not sequential anymore (for example after a GlobalMaxPool, an Attention Layer or a RNN with return Sequence = False), I'll concatenate it with my features.\n- Add a Dense layer with one unit to get a score.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 422645,
          "author_name": "AcademycalBastard",
          "author_url": "",
          "post_date": "2018-11-16T15:03:37.793000",
          "content": "<p>We could  build  some interesting features  togheter ,   do you have some email  to  discusss more   about this ??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422655,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2018-11-16T15:17:44.290000",
          "content": "<p>I prefer to work on my own, thanks for the offer anyway !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422935,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-17T05:37:02.453000",
          "content": "<p>@AcademycalBastard Note that it is considered cheating if you discuss privately ideas with someone you are not in a team</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423014,
          "author_name": "AcademycalBastard",
          "author_url": "",
          "post_date": "2018-11-17T09:12:13.430000",
          "content": "<p>Yes, is not the case here ,  because I allready  found some to team </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422737,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2018-11-16T17:44:59.897000",
      "content": "<p>Yup, it actually makes a lot of sense, and has been used widely in the Vietnamese NLP community recently with great results. Since I'm pretty inexperienced working with English I'm not sure how it will play out, but I'm currently working on something similar to yours (maybe a bit more complicated actually) .</p>\n\n<p>I'm open for teaming up but I think it's just too early , so we can lose some advantage by merging, it's better to wait for at least a month I think.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 422750,
          "author_name": "AcademycalBastard",
          "author_url": "",
          "post_date": "2018-11-16T18:18:43.147000",
          "content": "<p>I  will  be  delighted   to  team  up with   you . I have strong feature engineering  ideeas  only thing  is  to  put  it all togheter in the  NN.   Do you have an email  so I  could send  you  one  of  my   features  I have in mind that  I  think will boost up score  ? Tnks</p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 422936,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-17T05:37:32.007000",
          "content": "<p>@AcademycalBastard Note that it is considered cheating if you discuss privately ideas with someone you are not in a team</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 428047,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2018-11-26T16:57:00.190000",
      "content": "<p>Here is an example where I used sentence length as an extra feature (not helpful): <a href=\"https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7596269\">https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7596269</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422975,
      "author_name": "Artiom Zayats",
      "author_url": "",
      "post_date": "2018-11-17T06:51:22.857000",
      "content": "<p>I was toying myself with this idea, and i think i found some features with a decent correlation, but when I concat them to the features found after the Embedding layers is done proccesing, my results are ussualy worse. Now here are two otpion: 1. I am doing it wrong, concatentating in a wrong place ... etc. or my features are garbage..    Does anybody have a good source (paper) about combining engineering features + word embeddings? Thanks </p>",
      "votes": 0,
      "replies": [
        {
          "id": 423011,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-17T08:53:53.013000",
          "content": "<p>You can check my model of avito demand prediction competition, where I combined all sorts of different input\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59902\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59902</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423013,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-17T08:59:32.723000",
          "content": "<p>In general it does not work out of the box, there are lots of small things to take into consideration when combining embeddings with additional features. Do you normalize your features?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423241,
          "author_name": "Artiom Zayats",
          "author_url": "",
          "post_date": "2018-11-17T18:52:08.750000",
          "content": "<p>I do batch normlazation before comibing features. <br>\nfor example the output of embedding features are normlaized + output of hand picked features is normlaized too.     </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423210,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-17T17:58:29.540000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422628": "I am thinking of  making use of feature engineering  such number  of punctuation signs , number of  voals  and  stuff like that.\nDoes it make  sense from your  guys experience?\nAnd if so  ,  how or  where  could   I add them to this neural nets?  Could I add them as  a new input  for keras ?",
    "422639": "It could be useful, I'll have to try that.\n\nThe way I'll deal with that is this way:\n- I'll use a 2 elements input, being [word embeddings, features]\n- I will feed the word embeddings in a RNN (or a layer that deals with sequential data)\n- When my network output is not sequential anymore (for example after a GlobalMaxPool, an Attention Layer or a RNN with return Sequence = False), I'll concatenate it with my features.\n- Add a Dense layer with one unit to get a score.",
    "422737": "Yup, it actually makes a lot of sense, and has been used widely in the Vietnamese NLP community recently with great results. Since I'm pretty inexperienced working with English I'm not sure how it will play out, but I'm currently working on something similar to yours (maybe a bit more complicated actually) .\n\nI'm open for teaming up but I think it's just too early , so we can lose some advantage by merging, it's better to wait for at least a month I think.",
    "428047": "Here is an example where I used sentence length as an extra feature (not helpful): https://www.kaggle.com/shujian/single-rnn-model-with-meta-features?scriptVersionId=7596269",
    "422975": "I was toying myself with this idea, and i think i found some features with a decent correlation, but when I concat them to the features found after the Embedding layers is done proccesing, my results are ussualy worse. Now here are two otpion: 1. I am doing it wrong, concatentating in a wrong place ... etc. or my features are garbage..    Does anybody have a good source (paper) about combining engineering features + word embeddings? Thanks ",
    "423210": ""
  }
}