{
  "id": 59415,
  "title": "Boosting the scores beyond 0.221X",
  "url": "/competitions/avito-demand-prediction/discussion/59415",
  "author_name": "",
  "post_date": "2018-06-22T00:45:54.066621Z",
  "votes": 9,
  "comment_count": 14,
  "views": 0,
  "content": "<p>It may be little bit late to discuss this, since we don't have much dates left, but I want to discuss some of the features / process to push scores forward. </p>\n\n<p>I'm kind of stuck around 0.222X. With a little trick - blending within my results (<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/58377#339448\">see this</a>) I boosted my score toward 0.221X. However, I'm kind of stuck here, so I want to discuss on what else could I try, and what else other people are trying so we could figure out some of the significant processing process / features. </p>\n\n<p>What I've tried for basic feature processing is:\n- Fill NAs\n- Labeling categorical feature\n- Log processing price feature\n- description and title features : Countvec (Count number of words / unique words, punctuations, etx)\n- Tfidf vectorization</p>\n\n<p>Also here are some of extended things I've tried so far \n- Converting Russians into English and then process <strong><em>- Seems like working as Russians itself gives better result</em></strong></p>\n\n<ul>\n<li><p>Scraping Regional information from Wikipedia and adding them as features ([See this][2 - This one does not work on kaggle kernal for some reasons]) <strong><em>- Seems like it did not work. I believe it's because those regional information (Time zone, population, etc) are too dependent on 'region' feature.</em></strong> </p></li>\n<li><p>Time Features:\nI used extension of date features from <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">this kernal</a>  <strong><em>- I believe this works, thanks for original kernal author :)</em></strong></p></li>\n<li><p>Image blurrness score - <strong>*Seems like adding image blurrness itself did not improve the result that much. *</strong></p></li>\n</ul>\n\n<p>What could I try more? And what others are trying to improve results?\n(I am using LGBM model) What could jump the score beyond 0.221X?</p>",
  "messages": [
    {
      "id": "346565",
      "postDate": "06/22/2018 00:45:54",
      "content": "<p>It may be little bit late to discuss this, since we don't have much dates left, but I want to discuss some of the features / process to push scores forward. </p>\n\n<p>I'm kind of stuck around 0.222X. With a little trick - blending within my results (<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/58377#339448\">see this</a>) I boosted my score toward 0.221X. However, I'm kind of stuck here, so I want to discuss on what else could I try, and what else other people are trying so we could figure out some of the significant processing process / features. </p>\n\n<p>What I've tried for basic feature processing is:\n- Fill NAs\n- Labeling categorical feature\n- Log processing price feature\n- description and title features : Countvec (Count number of words / unique words, punctuations, etx)\n- Tfidf vectorization</p>\n\n<p>Also here are some of extended things I've tried so far \n- Converting Russians into English and then process <strong><em>- Seems like working as Russians itself gives better result</em></strong></p>\n\n<ul>\n<li><p>Scraping Regional information from Wikipedia and adding them as features ([See this][2 - This one does not work on kaggle kernal for some reasons]) <strong><em>- Seems like it did not work. I believe it's because those regional information (Time zone, population, etc) are too dependent on 'region' feature.</em></strong> </p></li>\n<li><p>Time Features:\nI used extension of date features from <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">this kernal</a>  <strong><em>- I believe this works, thanks for original kernal author :)</em></strong></p></li>\n<li><p>Image blurrness score - <strong>*Seems like adding image blurrness itself did not improve the result that much. *</strong></p></li>\n</ul>\n\n<p>What could I try more? And what others are trying to improve results?\n(I am using LGBM model) What could jump the score beyond 0.221X?</p>",
      "rawMarkdown": "It may be little bit late to discuss this, since we don't have much dates left, but I want to discuss some of the features / process to push scores forward. \n\nI'm kind of stuck around 0.222X. With a little trick - blending within my results ([see this][1]) I boosted my score toward 0.221X. However, I'm kind of stuck here, so I want to discuss on what else could I try, and what else other people are trying so we could figure out some of the significant processing process / features. \n\nWhat I've tried for basic feature processing is:\n- Fill NAs\n- Labeling categorical feature\n- Log processing price feature\n- description and title features : Countvec (Count number of words / unique words, punctuations, etx)\n- Tfidf vectorization\n\nAlso here are some of extended things I've tried so far \n- Converting Russians into English and then process ***- Seems like working as Russians itself gives better result***\n\n- Scraping Regional information from Wikipedia and adding them as features ([See this][2 - This one does not work on kaggle kernal for some reasons]) ***- Seems like it did not work. I believe it's because those regional information (Time zone, population, etc) are too dependent on 'region' feature.*** \n\n- Time Features:\nI used extension of date features from [this kernal][2]  ***- I believe this works, thanks for original kernal author :)***\n\n- Image blurrness score - ***Seems like adding image blurrness itself did not improve the result that much. ***\n\n\nWhat could I try more? And what others are trying to improve results?\n(I am using LGBM model) What could jump the score beyond 0.221X?\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/58377#339448\n  [2]: https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm",
      "votes": null
    },
    {
      "id": "346607",
      "postDate": "06/22/2018 03:29:36",
      "content": "<p>Maybe it will help you but I got a 0.001 improvement when I optimized hyper parameters for bigger trees which are not possible in kernels due to memory issues.</p>",
      "rawMarkdown": "Maybe it will help you but I got a 0.001 improvement when I optimized hyper parameters for bigger trees which are not possible in kernels due to memory issues.",
      "votes": null
    },
    {
      "id": "346618",
      "postDate": "06/22/2018 04:02:15",
      "content": "<p>Oh the memory issue :(. Thanks for the advice!</p>",
      "rawMarkdown": "Oh the memory issue :(. Thanks for the advice!",
      "votes": null
    },
    {
      "id": "346646",
      "postDate": "06/22/2018 05:22:47",
      "content": "<p>You are going great btw. Impressive work and Best of Luck buddy.\nAs for the improvements, let me share a few pointers that helped me getting along;</p>\n\n<ul>\n<li><p>Try adding more image meta features. size (x,y) / Blurness / Brightness / No. of key points in an image. (This gave me max 0.0004 of improvement both on LB/CV.</p></li>\n<li><p>Try normalizing the text. There is too much noise in param/description and title. (Gave me a little improvement, not much)</p></li>\n<li><p>Try filling in the top features. Region/Param/image_top_1 are one of those that can be taken care of. By filling, I mean with a smarter way rather than a static value. How about a model that predict those missing values? (Gave me major improvement)</p></li>\n<li><p>You should go through all the EDAs we have. Look closely, and try to get something from there. I handcrafted a few features using price, image_top_1 and user. (It might be luck) :)</p></li>\n<li><p>Try Catboost and XGB as well. They are memory intensive but building them shallow would still help.</p></li>\n<li><p>One last kinda greedy approach, which I usually try at the end. I start sub-setting the data and change hyper-params for the best performing models. Here we have a ton of features so it seems to be working for me. (Killing time)</p></li>\n</ul>\n\n<p>Though the time is short but you can try any of them if you want to - as always, learning is fun and is even more fun with sharing. Also, I'm not a guru and learning with you guys so I might end up over-fitting real bad. You have been warned :D</p>\n\n<p>Best of luck!</p>",
      "rawMarkdown": "You are going great btw. Impressive work and Best of Luck buddy.\nAs for the improvements, let me share a few pointers that helped me getting along;\n\n - Try adding more image meta features. size (x,y) / Blurness / Brightness / No. of key points in an image. (This gave me max 0.0004 of improvement both on LB/CV.\n\n - Try normalizing the text. There is too much noise in param/description and title. (Gave me a little improvement, not much)\n\n - Try filling in the top features. Region/Param/image_top_1 are one of those that can be taken care of. By filling, I mean with a smarter way rather than a static value. How about a model that predict those missing values? (Gave me major improvement)\n\n - You should go through all the EDAs we have. Look closely, and try to get something from there. I handcrafted a few features using price, image_top_1 and user. (It might be luck) :)\n\n - Try Catboost and XGB as well. They are memory intensive but building them shallow would still help.\n\n - One last kinda greedy approach, which I usually try at the end. I start sub-setting the data and change hyper-params for the best performing models. Here we have a ton of features so it seems to be working for me. (Killing time)\n\n\nThough the time is short but you can try any of them if you want to - as always, learning is fun and is even more fun with sharing. Also, I'm not a guru and learning with you guys so I might end up over-fitting real bad. You have been warned :D\n\nBest of luck!",
      "votes": null
    },
    {
      "id": "346658",
      "postDate": "06/22/2018 06:12:00",
      "content": "<p>Thank you for so much great tips! I think I'm going to try xBoost too. Also trying smarter ways to fill three major features! :)</p>",
      "rawMarkdown": "Thank you for so much great tips! I think I'm going to try xBoost too. Also trying smarter ways to fill three major features! :)",
      "votes": null
    },
    {
      "id": "346721",
      "postDate": "06/22/2018 08:38:47",
      "content": "<p>+1 for image features. they give a good boost, especially # of key points and brightness. \nyou can try the <a href=\"https://docs.opencv.org/3.1.0/da/df5/tutorial_py_sift_intro.html\">opencv SIFT implementation</a></p>",
      "rawMarkdown": "1 for image features. they give a good boost, especially # of key points and brightness. \nyou can try the [opencv SIFT implementation][1]\n\n\n  [1]: https://docs.opencv.org/3.1.0/da/df5/tutorial_py_sift_intro.html",
      "votes": null
    },
    {
      "id": "346724",
      "postDate": "06/22/2018 08:49:09",
      "content": "<p>I believe there are many more thing to try on image, and thanks for the comment!</p>",
      "rawMarkdown": "I believe there are many more thing to try on image, and thanks for the comment!",
      "votes": null
    },
    {
      "id": "347133",
      "postDate": "06/23/2018 10:39:18",
      "content": "<p>Clever Neural nets structures can give you 0.220X and even 0.219x without lot of features engineering (Text features are too important in this competition) </p>\n\n<p>But will of course require some power (GPU). </p>",
      "rawMarkdown": "Clever Neural nets structures can give you 0.220X and even 0.219x without lot of features engineering (Text features are too important in this competition) \n\nBut will of course require some power (GPU).",
      "votes": null
    },
    {
      "id": "347612",
      "postDate": "06/24/2018 20:43:11",
      "content": "<p>What is the meaning of bigger tree？ adding more leaves？</p>",
      "rawMarkdown": "What is the meaning of bigger tree？ adding more leaves？",
      "votes": null
    },
    {
      "id": "347636",
      "postDate": "06/24/2018 23:10:26",
      "content": "<p>Depth and leaves yes. But I mostly mean bigger compared to kernels.</p>",
      "rawMarkdown": "Depth and leaves yes. But I mostly mean bigger compared to kernels.",
      "votes": null
    },
    {
      "id": "347639",
      "postDate": "06/24/2018 23:29:50",
      "content": "<p>about 300-400？</p>",
      "rawMarkdown": "about 300-400？",
      "votes": null
    },
    {
      "id": "347643",
      "postDate": "06/25/2018 00:17:39",
      "content": "<p>Is there any references or kernals regarding Clever Neural nets structure which I can look for?</p>",
      "rawMarkdown": "Is there any references or kernals regarding Clever Neural nets structure which I can look for?",
      "votes": null
    },
    {
      "id": "348833",
      "postDate": "06/27/2018 11:15:28",
      "content": "<blockquote>\n  <p><strong>Ethan Sukhyun Hong wrote</strong></p>\n  \n  <blockquote>\n    <p>Is there any references or kernals regarding Clever Neural nets structure which I can look for?</p>\n  </blockquote>\n</blockquote>\n\n<p>Really sorry for answering too late. I was too busy in the last days to focus on discussions here. \nI don't about this competition, but there are many good NN kernels on Toxic competition. They may serve as reference/baseline to improve. </p>\n\n<p>Not a kernel, but I found this <a href=\"https://explosion.ai/blog/deep-learning-formula-nlp\">article</a> very useful to build NN structure for text/cat features</p>",
      "rawMarkdown": "&gt; **Ethan Sukhyun Hong wrote**\n&gt; \n&gt; &gt; Is there any references or kernals regarding Clever Neural nets structure which I can look for?\n\nReally sorry for answering too late. I was too busy in the last days to focus on discussions here. \nI don't about this competition, but there are many good NN kernels on Toxic competition. They may serve as reference/baseline to improve. \n\nNot a kernel, but I found this [article][1] very useful to build NN structure for text/cat features\n\n\n  [1]: https://explosion.ai/blog/deep-learning-formula-nlp",
      "votes": null
    },
    {
      "id": "349322",
      "postDate": "06/28/2018 02:28:20",
      "content": "<p>Hey, my friend could  you share some of your solutions like how to predict the missing value? That is a very interesting point!</p>",
      "rawMarkdown": "Hey, my friend could  you share some of your solutions like how to predict the missing value? That is a very interesting point!",
      "votes": null
    },
    {
      "id": "349378",
      "postDate": "06/28/2018 03:39:35",
      "content": "<p>I personally used a neural network on text + categories to predict missings like image_top, param and price.</p>",
      "rawMarkdown": "I personally used a neural network on text + categories to predict missings like image_top, param and price.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 346607,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "06/22/2018 03:29:36",
      "content": "<p>Maybe it will help you but I got a 0.001 improvement when I optimized hyper parameters for bigger trees which are not possible in kernels due to memory issues.</p>",
      "votes": null,
      "replies": [
        {
          "id": 346618,
          "author_name": "sukhyun9673",
          "author_url": "",
          "post_date": "06/22/2018 04:02:15",
          "content": "<p>Oh the memory issue :(. Thanks for the advice!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347612,
          "author_name": "sheboke93",
          "author_url": "",
          "post_date": "06/24/2018 20:43:11",
          "content": "<p>What is the meaning of bigger tree？ adding more leaves？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347636,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "06/24/2018 23:10:26",
          "content": "<p>Depth and leaves yes. But I mostly mean bigger compared to kernels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347639,
          "author_name": "sheboke93",
          "author_url": "",
          "post_date": "06/24/2018 23:29:50",
          "content": "<p>about 300-400？</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 346646,
      "author_name": "nuhsikander",
      "author_url": "",
      "post_date": "06/22/2018 05:22:47",
      "content": "<p>You are going great btw. Impressive work and Best of Luck buddy.\nAs for the improvements, let me share a few pointers that helped me getting along;</p>\n\n<ul>\n<li><p>Try adding more image meta features. size (x,y) / Blurness / Brightness / No. of key points in an image. (This gave me max 0.0004 of improvement both on LB/CV.</p></li>\n<li><p>Try normalizing the text. There is too much noise in param/description and title. (Gave me a little improvement, not much)</p></li>\n<li><p>Try filling in the top features. Region/Param/image_top_1 are one of those that can be taken care of. By filling, I mean with a smarter way rather than a static value. How about a model that predict those missing values? (Gave me major improvement)</p></li>\n<li><p>You should go through all the EDAs we have. Look closely, and try to get something from there. I handcrafted a few features using price, image_top_1 and user. (It might be luck) :)</p></li>\n<li><p>Try Catboost and XGB as well. They are memory intensive but building them shallow would still help.</p></li>\n<li><p>One last kinda greedy approach, which I usually try at the end. I start sub-setting the data and change hyper-params for the best performing models. Here we have a ton of features so it seems to be working for me. (Killing time)</p></li>\n</ul>\n\n<p>Though the time is short but you can try any of them if you want to - as always, learning is fun and is even more fun with sharing. Also, I'm not a guru and learning with you guys so I might end up over-fitting real bad. You have been warned :D</p>\n\n<p>Best of luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 346658,
          "author_name": "sukhyun9673",
          "author_url": "",
          "post_date": "06/22/2018 06:12:00",
          "content": "<p>Thank you for so much great tips! I think I'm going to try xBoost too. Also trying smarter ways to fill three major features! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346721,
          "author_name": "yk1598",
          "author_url": "",
          "post_date": "06/22/2018 08:38:47",
          "content": "<p>+1 for image features. they give a good boost, especially # of key points and brightness. \nyou can try the <a href=\"https://docs.opencv.org/3.1.0/da/df5/tutorial_py_sift_intro.html\">opencv SIFT implementation</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346724,
          "author_name": "sukhyun9673",
          "author_url": "",
          "post_date": "06/22/2018 08:49:09",
          "content": "<p>I believe there are many more thing to try on image, and thanks for the comment!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 349322,
          "author_name": "sheboke93",
          "author_url": "",
          "post_date": "06/28/2018 02:28:20",
          "content": "<p>Hey, my friend could  you share some of your solutions like how to predict the missing value? That is a very interesting point!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 349378,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "06/28/2018 03:39:35",
          "content": "<p>I personally used a neural network on text + categories to predict missings like image_top, param and price.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 347133,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "06/23/2018 10:39:18",
      "content": "<p>Clever Neural nets structures can give you 0.220X and even 0.219x without lot of features engineering (Text features are too important in this competition) </p>\n\n<p>But will of course require some power (GPU). </p>",
      "votes": null,
      "replies": [
        {
          "id": 347643,
          "author_name": "sukhyun9673",
          "author_url": "",
          "post_date": "06/25/2018 00:17:39",
          "content": "<p>Is there any references or kernals regarding Clever Neural nets structure which I can look for?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 348833,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "06/27/2018 11:15:28",
          "content": "<blockquote>\n  <p><strong>Ethan Sukhyun Hong wrote</strong></p>\n  \n  <blockquote>\n    <p>Is there any references or kernals regarding Clever Neural nets structure which I can look for?</p>\n  </blockquote>\n</blockquote>\n\n<p>Really sorry for answering too late. I was too busy in the last days to focus on discussions here. \nI don't about this competition, but there are many good NN kernels on Toxic competition. They may serve as reference/baseline to improve. </p>\n\n<p>Not a kernel, but I found this <a href=\"https://explosion.ai/blog/deep-learning-formula-nlp\">article</a> very useful to build NN structure for text/cat features</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "346565": "It may be little bit late to discuss this, since we don't have much dates left, but I want to discuss some of the features / process to push scores forward. \n\nI'm kind of stuck around 0.222X. With a little trick - blending within my results ([see this][1]) I boosted my score toward 0.221X. However, I'm kind of stuck here, so I want to discuss on what else could I try, and what else other people are trying so we could figure out some of the significant processing process / features. \n\nWhat I've tried for basic feature processing is:\n- Fill NAs\n- Labeling categorical feature\n- Log processing price feature\n- description and title features : Countvec (Count number of words / unique words, punctuations, etx)\n- Tfidf vectorization\n\nAlso here are some of extended things I've tried so far \n- Converting Russians into English and then process ***- Seems like working as Russians itself gives better result***\n\n- Scraping Regional information from Wikipedia and adding them as features ([See this][2 - This one does not work on kaggle kernal for some reasons]) ***- Seems like it did not work. I believe it's because those regional information (Time zone, population, etc) are too dependent on 'region' feature.*** \n\n- Time Features:\nI used extension of date features from [this kernal][2]  ***- I believe this works, thanks for original kernal author :)***\n\n- Image blurrness score - ***Seems like adding image blurrness itself did not improve the result that much. ***\n\n\nWhat could I try more? And what others are trying to improve results?\n(I am using LGBM model) What could jump the score beyond 0.221X?\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/58377#339448\n  [2]: https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm",
    "346607": "Maybe it will help you but I got a 0.001 improvement when I optimized hyper parameters for bigger trees which are not possible in kernels due to memory issues.",
    "346618": "Oh the memory issue :(. Thanks for the advice!",
    "346646": "You are going great btw. Impressive work and Best of Luck buddy.\nAs for the improvements, let me share a few pointers that helped me getting along;\n\n - Try adding more image meta features. size (x,y) / Blurness / Brightness / No. of key points in an image. (This gave me max 0.0004 of improvement both on LB/CV.\n\n - Try normalizing the text. There is too much noise in param/description and title. (Gave me a little improvement, not much)\n\n - Try filling in the top features. Region/Param/image_top_1 are one of those that can be taken care of. By filling, I mean with a smarter way rather than a static value. How about a model that predict those missing values? (Gave me major improvement)\n\n - You should go through all the EDAs we have. Look closely, and try to get something from there. I handcrafted a few features using price, image_top_1 and user. (It might be luck) :)\n\n - Try Catboost and XGB as well. They are memory intensive but building them shallow would still help.\n\n - One last kinda greedy approach, which I usually try at the end. I start sub-setting the data and change hyper-params for the best performing models. Here we have a ton of features so it seems to be working for me. (Killing time)\n\n\nThough the time is short but you can try any of them if you want to - as always, learning is fun and is even more fun with sharing. Also, I'm not a guru and learning with you guys so I might end up over-fitting real bad. You have been warned :D\n\nBest of luck!",
    "346658": "Thank you for so much great tips! I think I'm going to try xBoost too. Also trying smarter ways to fill three major features! :)",
    "346721": "1 for image features. they give a good boost, especially # of key points and brightness. \nyou can try the [opencv SIFT implementation][1]\n\n\n  [1]: https://docs.opencv.org/3.1.0/da/df5/tutorial_py_sift_intro.html",
    "346724": "I believe there are many more thing to try on image, and thanks for the comment!",
    "347133": "Clever Neural nets structures can give you 0.220X and even 0.219x without lot of features engineering (Text features are too important in this competition) \n\nBut will of course require some power (GPU).",
    "347612": "What is the meaning of bigger tree？ adding more leaves？",
    "347636": "Depth and leaves yes. But I mostly mean bigger compared to kernels.",
    "347639": "about 300-400？",
    "347643": "Is there any references or kernals regarding Clever Neural nets structure which I can look for?",
    "348833": "&gt; **Ethan Sukhyun Hong wrote**\n&gt; \n&gt; &gt; Is there any references or kernals regarding Clever Neural nets structure which I can look for?\n\nReally sorry for answering too late. I was too busy in the last days to focus on discussions here. \nI don't about this competition, but there are many good NN kernels on Toxic competition. They may serve as reference/baseline to improve. \n\nNot a kernel, but I found this [article][1] very useful to build NN structure for text/cat features\n\n\n  [1]: https://explosion.ai/blog/deep-learning-formula-nlp",
    "349322": "Hey, my friend could  you share some of your solutions like how to predict the missing value? That is a very interesting point!",
    "349378": "I personally used a neural network on text + categories to predict missings like image_top, param and price."
  },
  "source": "meta"
}