{
  "id": 77736,
  "title": "How to use extra features ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77736",
  "author_name": "",
  "post_date": "2019-01-16T07:10:58.324502500Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi this is my first competition for medals.\nI have extracted simple  extra features from comments like word density,length,capital letters etc.\nBut I don't understand where to use them in models?\nAfter reading some solutions I found 2 ways,are they correct ?</p>\n\n<ol>\n<li>1.Concat them with maxpool,avgpool layers before final dense layer.</li>\n<li>Put them in embedding.(If it's a word based feature)</li>\n</ol>\n\n<p>What are other ways of using them better ?</p>",
  "messages": [
    {
      "id": "456604",
      "postDate": "01/16/2019 07:10:58",
      "content": "<p>Hi this is my first competition for medals.\nI have extracted simple  extra features from comments like word density,length,capital letters etc.\nBut I don't understand where to use them in models?\nAfter reading some solutions I found 2 ways,are they correct ?</p>\n\n<ol>\n<li>1.Concat them with maxpool,avgpool layers before final dense layer.</li>\n<li>Put them in embedding.(If it's a word based feature)</li>\n</ol>\n\n<p>What are other ways of using them better ?</p>",
      "rawMarkdown": "Hi this is my first competition for medals.\nI have extracted simple  extra features from comments like word density,length,capital letters etc.\nBut I don't understand where to use them in models?\nAfter reading some solutions I found 2 ways,are they correct ?\n\n 1. 1.Concat them with maxpool,avgpool layers before final dense layer.\n 2. Put them in embedding.(If it's a word based feature)\n\nWhat are other ways of using them better ?",
      "votes": null
    },
    {
      "id": "456713",
      "postDate": "01/16/2019 11:01:12",
      "content": "<p>Use 1st approach if it is some statistic of a question and 2nd if this feature is for each word. </p>",
      "rawMarkdown": "Use 1st approach if it is some statistic of a question and 2nd if this feature is for each word.",
      "votes": null
    },
    {
      "id": "456903",
      "postDate": "01/16/2019 18:12:05",
      "content": "<p>Multi-stream network. This paper discusses it, but on a different type of NLP problem..\n<a href=\"https://www.isca-speech.org/archive/interspeech_2015/papers/i15_1413.pdf\">https://www.isca-speech.org/archive/interspeech_2015/papers/i15_1413.pdf</a></p>",
      "rawMarkdown": "Multi-stream network. This paper discusses it, but on a different type of NLP problem..\nhttps://www.isca-speech.org/archive/interspeech_2015/papers/i15_1413.pdf",
      "votes": null
    },
    {
      "id": "457124",
      "postDate": "01/17/2019 01:29:48",
      "content": "<p>Has someone get extra feature working in the right way? I have tried concat them with maxpool,avgpool layers before final dense layer, it does not give better performance.</p>",
      "rawMarkdown": "Has someone get extra feature working in the right way? I have tried concat them with maxpool,avgpool layers before final dense layer, it does not give better performance.",
      "votes": null
    },
    {
      "id": "457582",
      "postDate": "01/17/2019 18:51:39",
      "content": "<p>As for additional features that describes word ex. is the word written in all capital letters you need to pass this as additional embedding dimension.      </p>\n\n<p><code>cap_flag = word.isupper() and word!='I'. <br>\ncap = np.where(cap_flag,np.array([1]),np.array([0])). <br>\n..... <br>\nwv_matrix[i] = np.hstack((np.mean([second_embed[word.lower()],first_embed[word.lower()]], axis = 0),cap))\n</code></p>\n\n<p>As for whole sentence features ex. how long is the sentence, I'm concatynating it before dense layer so I have ex.:    </p>\n\n<p><code>\nconc = torch.cat((h_gru_atten,h_lstm_atten, avg_pool, max_pool, x_a), 1)\n</code></p>\n\n<p>where x_a is my additional features.  </p>\n\n<p>Here is my kernel: <br>\n<a href=\"https://www.kaggle.com/nicke1/fine-text-preproc-concat-embedding-lstm-gru-att\">https://www.kaggle.com/nicke1/fine-text-preproc-concat-embedding-lstm-gru-att</a></p>",
      "rawMarkdown": "As for additional features that describes word ex. is the word written in all capital letters you need to pass this as additional embedding dimension.      \n\n```cap_flag = word.isupper() and word!='I'.    \ncap = np.where(cap_flag,np.array([1]),np.array([0])).        \n.....   \nwv_matrix[i] = np.hstack((np.mean([second_embed[word.lower()],first_embed[word.lower()]], axis = 0),cap))\n```\n\nAs for whole sentence features ex. how long is the sentence, I'm concatynating it before dense layer so I have ex.:    \n\n```\nconc = torch.cat((h_gru_atten,h_lstm_atten, avg_pool, max_pool, x_a), 1)\n```\n\nwhere x_a is my additional features.  \n\nHere is my kernel:      \nhttps://www.kaggle.com/nicke1/fine-text-preproc-concat-embedding-lstm-gru-att",
      "votes": null
    },
    {
      "id": "458882",
      "postDate": "01/20/2019 18:12:04",
      "content": "<p>Do have an example of a kernel that relies heavily on feature engineering? I haven't seen much of it</p>",
      "rawMarkdown": "Do have an example of a kernel that relies heavily on feature engineering? I haven't seen much of it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 456713,
      "author_name": "",
      "author_url": "",
      "post_date": "01/16/2019 11:01:12",
      "content": "<p>Use 1st approach if it is some statistic of a question and 2nd if this feature is for each word. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 456903,
      "author_name": "julius6",
      "author_url": "",
      "post_date": "01/16/2019 18:12:05",
      "content": "<p>Multi-stream network. This paper discusses it, but on a different type of NLP problem..\n<a href=\"https://www.isca-speech.org/archive/interspeech_2015/papers/i15_1413.pdf\">https://www.isca-speech.org/archive/interspeech_2015/papers/i15_1413.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 457124,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "01/17/2019 01:29:48",
      "content": "<p>Has someone get extra feature working in the right way? I have tried concat them with maxpool,avgpool layers before final dense layer, it does not give better performance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 457582,
      "author_name": "nicke1",
      "author_url": "",
      "post_date": "01/17/2019 18:51:39",
      "content": "<p>As for additional features that describes word ex. is the word written in all capital letters you need to pass this as additional embedding dimension.      </p>\n\n<p><code>cap_flag = word.isupper() and word!='I'. <br>\ncap = np.where(cap_flag,np.array([1]),np.array([0])). <br>\n..... <br>\nwv_matrix[i] = np.hstack((np.mean([second_embed[word.lower()],first_embed[word.lower()]], axis = 0),cap))\n</code></p>\n\n<p>As for whole sentence features ex. how long is the sentence, I'm concatynating it before dense layer so I have ex.:    </p>\n\n<p><code>\nconc = torch.cat((h_gru_atten,h_lstm_atten, avg_pool, max_pool, x_a), 1)\n</code></p>\n\n<p>where x_a is my additional features.  </p>\n\n<p>Here is my kernel: <br>\n<a href=\"https://www.kaggle.com/nicke1/fine-text-preproc-concat-embedding-lstm-gru-att\">https://www.kaggle.com/nicke1/fine-text-preproc-concat-embedding-lstm-gru-att</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 458882,
      "author_name": "joaopduro",
      "author_url": "",
      "post_date": "01/20/2019 18:12:04",
      "content": "<p>Do have an example of a kernel that relies heavily on feature engineering? I haven't seen much of it</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "456604": "Hi this is my first competition for medals.\nI have extracted simple  extra features from comments like word density,length,capital letters etc.\nBut I don't understand where to use them in models?\nAfter reading some solutions I found 2 ways,are they correct ?\n\n 1. 1.Concat them with maxpool,avgpool layers before final dense layer.\n 2. Put them in embedding.(If it's a word based feature)\n\nWhat are other ways of using them better ?",
    "456713": "Use 1st approach if it is some statistic of a question and 2nd if this feature is for each word.",
    "456903": "Multi-stream network. This paper discusses it, but on a different type of NLP problem..\nhttps://www.isca-speech.org/archive/interspeech_2015/papers/i15_1413.pdf",
    "457124": "Has someone get extra feature working in the right way? I have tried concat them with maxpool,avgpool layers before final dense layer, it does not give better performance.",
    "457582": "As for additional features that describes word ex. is the word written in all capital letters you need to pass this as additional embedding dimension.      \n\n```cap_flag = word.isupper() and word!='I'.    \ncap = np.where(cap_flag,np.array([1]),np.array([0])).        \n.....   \nwv_matrix[i] = np.hstack((np.mean([second_embed[word.lower()],first_embed[word.lower()]], axis = 0),cap))\n```\n\nAs for whole sentence features ex. how long is the sentence, I'm concatynating it before dense layer so I have ex.:    \n\n```\nconc = torch.cat((h_gru_atten,h_lstm_atten, avg_pool, max_pool, x_a), 1)\n```\n\nwhere x_a is my additional features.  \n\nHere is my kernel:      \nhttps://www.kaggle.com/nicke1/fine-text-preproc-concat-embedding-lstm-gru-att",
    "458882": "Do have an example of a kernel that relies heavily on feature engineering? I haven't seen much of it"
  },
  "source": "meta"
}