{
  "id": 73069,
  "title": "XGBoost Potential And Features",
  "url": "/competitions/quora-insincere-questions-classification/discussion/73069",
  "author_name": "",
  "post_date": "2018-11-29T13:27:53.230507200Z",
  "votes": 10,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all xgboost results are almost robust, it generates the same range of results each run. In addition, it can be enhanced with well-selected features.</p>\n\n<p>The first xgboost model published by me scored 0.614 then by adding 2 extra features​,​ it reached to <strong>0.623</strong> in version 3 <a href=\"https://www.kaggle.com/jaguar00/xgboost-baseline\">here</a>. Which shows the potential to increase the score with more well selecte​​​d features or params tuning. <em>\"2 features increased = 0.09\"</em>\n<img src=\"https://i.ibb.co/Lv5Yn1s/score.png\" alt=\"enter image description here\"></p>\n\n<p>Following features importance by xgboost model:\n<img src=\"https://i.imgur.com/SpcFb19.png\" alt=\"features\"></p>\n\n<p>If you suggest more features to test please add as comments</p>",
  "messages": [
    {
      "id": "429863",
      "postDate": "11/29/2018 13:27:53",
      "content": "<p>First of all xgboost results are almost robust, it generates the same range of results each run. In addition, it can be enhanced with well-selected features.</p>\n\n<p>The first xgboost model published by me scored 0.614 then by adding 2 extra features​,​ it reached to <strong>0.623</strong> in version 3 <a href=\"https://www.kaggle.com/jaguar00/xgboost-baseline\">here</a>. Which shows the potential to increase the score with more well selecte​​​d features or params tuning. <em>\"2 features increased = 0.09\"</em>\n<img src=\"https://i.ibb.co/Lv5Yn1s/score.png\" alt=\"enter image description here\"></p>\n\n<p>Following features importance by xgboost model:\n<img src=\"https://i.imgur.com/SpcFb19.png\" alt=\"features\"></p>\n\n<p>If you suggest more features to test please add as comments</p>",
      "rawMarkdown": "First of all xgboost results are almost robust, it generates the same range of results each run. In addition, it can be enhanced with well-selected features.\n\nThe first xgboost model published by me scored 0.614 then by adding 2 extra features​,​ it reached to **0.623** in version 3 [here][1]. Which shows the potential to increase the score with more well selecte​​​d features or params tuning. *\"2 features increased = 0.09\"*\n![enter image description here][2]\n\nFollowing features importance by xgboost model:\n![features][3]\n\nIf you suggest more features to test please add as comments\n\n  [1]: https://www.kaggle.com/jaguar00/xgboost-baseline\n  [2]: https://i.ibb.co/Lv5Yn1s/score.png\n  [3]: https://i.imgur.com/SpcFb19.png",
      "votes": null
    },
    {
      "id": "435672",
      "postDate": "12/08/2018 14:13:51",
      "content": "<p>You could try counting the number of some swear words in each sentence. Also, people tend to say 'you' before swearing or mention someone's race. Also, you might get a boost from checking if a question has a link, email, date or time.\nHope this helps!</p>",
      "rawMarkdown": "You could try counting the number of some swear words in each sentence. Also, people tend to say 'you' before swearing or mention someone's race. Also, you might get a boost from checking if a question has a link, email, date or time.\nHope this helps!",
      "votes": null
    },
    {
      "id": "435753",
      "postDate": "12/08/2018 17:30:34",
      "content": "<p>Nice ideas, going to test it and update you with results if any.</p>",
      "rawMarkdown": "Nice ideas, going to test it and update you with results if any.",
      "votes": null
    },
    {
      "id": "444051",
      "postDate": "12/23/2018 02:43:11",
      "content": "<p>Can such hand feature improve the NN score in your model?</p>",
      "rawMarkdown": "Can such hand feature improve the NN score in your model?",
      "votes": null
    },
    {
      "id": "445244",
      "postDate": "12/26/2018 03:41:39",
      "content": "<p>Not all the features maybe one or two, the 3rd place solution for similar competition \"Toxic\" used two features: <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644</a></p>",
      "rawMarkdown": "Not all the features maybe one or two, the 3rd place solution for similar competition \"Toxic\" used two features: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644",
      "votes": null
    },
    {
      "id": "457118",
      "postDate": "01/17/2019 01:13:39",
      "content": "<p>Hi, did you get performance improvement after adding extra features? For me, I was unable to get extra feature working....</p>",
      "rawMarkdown": "Hi, did you get performance improvement after adding extra features? For me, I was unable to get extra feature working....",
      "votes": null
    },
    {
      "id": "457167",
      "postDate": "01/17/2019 03:27:35",
      "content": "<p>I tried other features such as the number of verbs and nouns however did not help to improve. Even tried it with RNN as an extra input and no enhancement. </p>",
      "rawMarkdown": "I tried other features such as the number of verbs and nouns however did not help to improve. Even tried it with RNN as an extra input and no enhancement.",
      "votes": null
    },
    {
      "id": "458739",
      "postDate": "01/20/2019 12:03:41",
      "content": "<p>Hi, I'm new kaggler. I use 300dims word embedding as input of first layer, and how to use such these statistics features in rnn? For example, when i get the number of verbs, should I take it after the word embedding and make it become 301 dims? Thank u so much!</p>",
      "rawMarkdown": "Hi, I'm new kaggler. I use 300dims word embedding as input of first layer, and how to use such these statistics features in rnn? For example, when i get the number of verbs, should I take it after the word embedding and make it become 301 dims? Thank u so much!",
      "votes": null
    },
    {
      "id": "459580",
      "postDate": "01/22/2019 03:10:18",
      "content": "<p>Have you tried POS, NER, syllables per word, readability indexes, sentiments, bad word list features?</p>",
      "rawMarkdown": "Have you tried POS, NER, syllables per word, readability indexes, sentiments, bad word list features?",
      "votes": null
    },
    {
      "id": "459647",
      "postDate": "01/22/2019 06:04:31",
      "content": "<p>Yes, not a much enhancement ... with RNN it showed a little enhancement although my best RNN model till now does not use it as input. \nCapital rate and unique words have a good impact on both XGboost and RNN based on mine testing.</p>",
      "rawMarkdown": "Yes, not a much enhancement ... with RNN it showed a little enhancement although my best RNN model till now does not use it as input. \nCapital rate and unique words have a good impact on both XGboost and RNN based on mine testing.",
      "votes": null
    },
    {
      "id": "459649",
      "postDate": "01/22/2019 06:07:18",
      "content": "<p>Thanks for your comment, following public kernel use multi-input for RNN (<a href=\"https://www.kaggle.com/jannen/reaching-0-7-fork-from-bilstm-attention-kfold\">https://www.kaggle.com/jannen/reaching-0-7-fork-from-bilstm-attention-kfold</a>) the code not much organized you have to read it carefully to avoid lines that consume time with no reason in this kernel.</p>",
      "rawMarkdown": "Thanks for your comment, following public kernel use multi-input for RNN (https://www.kaggle.com/jannen/reaching-0-7-fork-from-bilstm-attention-kfold) the code not much organized you have to read it carefully to avoid lines that consume time with no reason in this kernel.",
      "votes": null
    },
    {
      "id": "461704",
      "postDate": "01/26/2019 19:34:09",
      "content": "<p>I added some of these features my Local CV has improved while on LB it is performing worse. </p>",
      "rawMarkdown": "I added some of these features my Local CV has improved while on LB it is performing worse.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 435672,
      "author_name": "craq21",
      "author_url": "",
      "post_date": "12/08/2018 14:13:51",
      "content": "<p>You could try counting the number of some swear words in each sentence. Also, people tend to say 'you' before swearing or mention someone's race. Also, you might get a boost from checking if a question has a link, email, date or time.\nHope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 435753,
          "author_name": "jaguar00",
          "author_url": "",
          "post_date": "12/08/2018 17:30:34",
          "content": "<p>Nice ideas, going to test it and update you with results if any.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 444051,
      "author_name": "codebb",
      "author_url": "",
      "post_date": "12/23/2018 02:43:11",
      "content": "<p>Can such hand feature improve the NN score in your model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 445244,
          "author_name": "jaguar00",
          "author_url": "",
          "post_date": "12/26/2018 03:41:39",
          "content": "<p>Not all the features maybe one or two, the 3rd place solution for similar competition \"Toxic\" used two features: <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 457118,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "01/17/2019 01:13:39",
      "content": "<p>Hi, did you get performance improvement after adding extra features? For me, I was unable to get extra feature working....</p>",
      "votes": null,
      "replies": [
        {
          "id": 457167,
          "author_name": "jaguar00",
          "author_url": "",
          "post_date": "01/17/2019 03:27:35",
          "content": "<p>I tried other features such as the number of verbs and nouns however did not help to improve. Even tried it with RNN as an extra input and no enhancement. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461704,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "01/26/2019 19:34:09",
          "content": "<p>I added some of these features my Local CV has improved while on LB it is performing worse. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 458739,
      "author_name": "karwik",
      "author_url": "",
      "post_date": "01/20/2019 12:03:41",
      "content": "<p>Hi, I'm new kaggler. I use 300dims word embedding as input of first layer, and how to use such these statistics features in rnn? For example, when i get the number of verbs, should I take it after the word embedding and make it become 301 dims? Thank u so much!</p>",
      "votes": null,
      "replies": [
        {
          "id": 459649,
          "author_name": "jaguar00",
          "author_url": "",
          "post_date": "01/22/2019 06:07:18",
          "content": "<p>Thanks for your comment, following public kernel use multi-input for RNN (<a href=\"https://www.kaggle.com/jannen/reaching-0-7-fork-from-bilstm-attention-kfold\">https://www.kaggle.com/jannen/reaching-0-7-fork-from-bilstm-attention-kfold</a>) the code not much organized you have to read it carefully to avoid lines that consume time with no reason in this kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 459580,
      "author_name": "kwigan",
      "author_url": "",
      "post_date": "01/22/2019 03:10:18",
      "content": "<p>Have you tried POS, NER, syllables per word, readability indexes, sentiments, bad word list features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 459647,
          "author_name": "jaguar00",
          "author_url": "",
          "post_date": "01/22/2019 06:04:31",
          "content": "<p>Yes, not a much enhancement ... with RNN it showed a little enhancement although my best RNN model till now does not use it as input. \nCapital rate and unique words have a good impact on both XGboost and RNN based on mine testing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "429863": "First of all xgboost results are almost robust, it generates the same range of results each run. In addition, it can be enhanced with well-selected features.\n\nThe first xgboost model published by me scored 0.614 then by adding 2 extra features​,​ it reached to **0.623** in version 3 [here][1]. Which shows the potential to increase the score with more well selecte​​​d features or params tuning. *\"2 features increased = 0.09\"*\n![enter image description here][2]\n\nFollowing features importance by xgboost model:\n![features][3]\n\nIf you suggest more features to test please add as comments\n\n  [1]: https://www.kaggle.com/jaguar00/xgboost-baseline\n  [2]: https://i.ibb.co/Lv5Yn1s/score.png\n  [3]: https://i.imgur.com/SpcFb19.png",
    "435672": "You could try counting the number of some swear words in each sentence. Also, people tend to say 'you' before swearing or mention someone's race. Also, you might get a boost from checking if a question has a link, email, date or time.\nHope this helps!",
    "435753": "Nice ideas, going to test it and update you with results if any.",
    "444051": "Can such hand feature improve the NN score in your model?",
    "445244": "Not all the features maybe one or two, the 3rd place solution for similar competition \"Toxic\" used two features: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644",
    "457118": "Hi, did you get performance improvement after adding extra features? For me, I was unable to get extra feature working....",
    "457167": "I tried other features such as the number of verbs and nouns however did not help to improve. Even tried it with RNN as an extra input and no enhancement.",
    "458739": "Hi, I'm new kaggler. I use 300dims word embedding as input of first layer, and how to use such these statistics features in rnn? For example, when i get the number of verbs, should I take it after the word embedding and make it become 301 dims? Thank u so much!",
    "459580": "Have you tried POS, NER, syllables per word, readability indexes, sentiments, bad word list features?",
    "459647": "Yes, not a much enhancement ... with RNN it showed a little enhancement although my best RNN model till now does not use it as input. \nCapital rate and unique words have a good impact on both XGboost and RNN based on mine testing.",
    "459649": "Thanks for your comment, following public kernel use multi-input for RNN (https://www.kaggle.com/jannen/reaching-0-7-fork-from-bilstm-attention-kfold) the code not much organized you have to read it carefully to avoid lines that consume time with no reason in this kernel.",
    "461704": "I added some of these features my Local CV has improved while on LB it is performing worse."
  },
  "source": "meta"
}