{
  "id": 59922,
  "title": "The way to the 18th place, just for rookies.",
  "url": "/competitions/avito-demand-prediction/writeups/we-had-great-fun-the-way-to-the-18th-place-just-fo",
  "author_name": "",
  "post_date": "2018-06-28T11:38:01.797Z",
  "votes": 38,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, congratulations to all of you guys who put great effort and learned a lot of new things from this competition. I had a great fun and meet new friends here, so lucky.</p>\n\n<p>Now comes the solution of our team, which should be simple without any magic. The steps are:</p>\n\n<h2>Best single model part (LGB):</h2>\n\n<ol>\n<li>Folked the public kernel and did some research such as parameter tuning which gave us a slightly better result than all the public kernels. </li>\n<li>Doing more feature engineering such as keywords count, text hash,  categorical feature hash. </li>\n<li>We found the price and item_seq_number features were interesting. We applied exponential and linear segmentation of them and adding the segmented data as new categorical features.</li>\n<li>We added \"char\" level analyzer for TFID process.</li>\n<li>Target encoding.</li>\n<li>Applying external data. we found the \"income\", \"population\" features are helpful. The Location Information did not help us.</li>\n<li>Adding basic image features based on public kernels. </li>\n<li>Final parameter tunning.\nThis model helped us to score 0.2184 after doing a 5-folds average. But it needs 3 hours to run on a 20 core server for each fold.</li>\n</ol>\n\n<h2>NN part:</h2>\n\n<p>We used most the similar features for NN, which gave us a 0.2205 score on LB. By a simple stacking of the LGB and NN models, we got 0.2179 and after some blending, we got 0.2176 on LB.</p>\n\n<h2>Team up part:</h2>\n\n<p>It was so lucky for us that @Yuki.O joined us. He is a genius guy who did massive feature engineering which makes the big diversity between our models. A simple stack gave us 0.2169 on the public LB.</p>\n\n<h2>Stacking part:</h2>\n\n<p>In the last week, we start to stack all the models. We created aprrox. 40 models with catboost, FM, mlp, LSTM. I think because it was limited by our best single model, the can just reach 0.2156 on the public LB. We fight till the last minutes though. :P</p>\n\n<p>Ok, this is all about what we did in this competition, great fun! Many thanks to all of you guys <a href=\"/huiqin\">@huiqin</a>, @Steeve Huang, <a href=\"/steinhafen\">@steinhafen</a>, and <a href=\"/extremin\">@extremin</a>. Special thanks to my new friend @Yuki.O, who put lots of effort into this competition and I learned a lot from him, not only the knowledge but also the professional attitude.</p>",
  "messages": [
    {
      "id": "349604",
      "postDate": "06/28/2018 10:57:09",
      "content": "<p>First of all, congratulations to all of you guys who put great effort and learned a lot of new things from this competition. I had a great fun and meet new friends here, so lucky.</p>\n\n<p>Now comes the solution of our team, which should be simple without any magic. The steps are:</p>\n\n<h2>Best single model part (LGB):</h2>\n\n<ol>\n<li>Folked the public kernel and did some research such as parameter tuning which gave us a slightly better result than all the public kernels. </li>\n<li>Doing more feature engineering such as keywords count, text hash,  categorical feature hash. </li>\n<li>We found the price and item_seq_number features were interesting. We applied exponential and linear segmentation of them and adding the segmented data as new categorical features.</li>\n<li>We added \"char\" level analyzer for TFID process.</li>\n<li>Target encoding.</li>\n<li>Applying external data. we found the \"income\", \"population\" features are helpful. The Location Information did not help us.</li>\n<li>Adding basic image features based on public kernels. </li>\n<li>Final parameter tunning.\nThis model helped us to score 0.2184 after doing a 5-folds average. But it needs 3 hours to run on a 20 core server for each fold.</li>\n</ol>\n\n<h2>NN part:</h2>\n\n<p>We used most the similar features for NN, which gave us a 0.2205 score on LB. By a simple stacking of the LGB and NN models, we got 0.2179 and after some blending, we got 0.2176 on LB.</p>\n\n<h2>Team up part:</h2>\n\n<p>It was so lucky for us that @Yuki.O joined us. He is a genius guy who did massive feature engineering which makes the big diversity between our models. A simple stack gave us 0.2169 on the public LB.</p>\n\n<h2>Stacking part:</h2>\n\n<p>In the last week, we start to stack all the models. We created aprrox. 40 models with catboost, FM, mlp, LSTM. I think because it was limited by our best single model, the can just reach 0.2156 on the public LB. We fight till the last minutes though. :P</p>\n\n<p>Ok, this is all about what we did in this competition, great fun! Many thanks to all of you guys <a href=\"/huiqin\">@huiqin</a>, @Steeve Huang, <a href=\"/steinhafen\">@steinhafen</a>, and <a href=\"/extremin\">@extremin</a>. Special thanks to my new friend @Yuki.O, who put lots of effort into this competition and I learned a lot from him, not only the knowledge but also the professional attitude.</p>",
      "rawMarkdown": "First of all, congratulations to all of you guys who put great effort and learned a lot of new things from this competition. I had a great fun and meet new friends here, so lucky.\n\nNow comes the solution of our team, which should be simple without any magic. The steps are:\n\nBest single model part (LGB):\n-------------------------------------------------------------------------------\n1. Folked the public kernel and did some research such as parameter tuning which gave us a slightly better result than all the public kernels. \n2. Doing more feature engineering such as keywords count, text hash,  categorical feature hash. \n3. We found the price and item_seq_number features were interesting. We applied exponential and linear segmentation of them and adding the segmented data as new categorical features.\n4. We added \"char\" level analyzer for TFID process.\n5. Target encoding.\n6. Applying external data. we found the \"income\", \"population\" features are helpful. The Location Information did not help us.\n7. Adding basic image features based on public kernels. \n8. Final parameter tunning.\nThis model helped us to score 0.2184 after doing a 5-folds average. But it needs 3 hours to run on a 20 core server for each fold.\n\nNN part:\n-----------------------------------------------------------------------------\nWe used most the similar features for NN, which gave us a 0.2205 score on LB. By a simple stacking of the LGB and NN models, we got 0.2179 and after some blending, we got 0.2176 on LB.\n\nTeam up part:\n-------------------------------------------------------------------------------\nIt was so lucky for us that @Yuki.O joined us. He is a genius guy who did massive feature engineering which makes the big diversity between our models. A simple stack gave us 0.2169 on the public LB.\n\nStacking part:\n-------------------------------------------------------------------------------\nIn the last week, we start to stack all the models. We created aprrox. 40 models with catboost, FM, mlp, LSTM. I think because it was limited by our best single model, the can just reach 0.2156 on the public LB. We fight till the last minutes though. :P\n\nOk, this is all about what we did in this competition, great fun! Many thanks to all of you guys @huiqin, @Steeve Huang, @steinhafen, and @extremin. Special thanks to my new friend @Yuki.O, who put lots of effort into this competition and I learned a lot from him, not only the knowledge but also the professional attitude.",
      "votes": null
    },
    {
      "id": "349630",
      "postDate": "06/28/2018 12:06:51",
      "content": "<p>Thank you for the great writeup <a href=\"/meli19\">@meli19</a> and the efforts of the team members! It was a fun journey of learning while support each other.  </p>",
      "rawMarkdown": "Thank you for the great writeup @meli19 and the efforts of the team members! It was a fun journey of learning while support each other.",
      "votes": null
    },
    {
      "id": "349903",
      "postDate": "06/28/2018 20:24:18",
      "content": "<p>Congrats Mengfei and to your team!</p>\n\n<p>Could you please detail what you meant by segmentation:</p>\n\n<p>\"We found the price and item_seq_number features were interesting. We applied exponential and linear segmentation of them and adding the segmented data as new categorical features.\" </p>",
      "rawMarkdown": "Congrats Mengfei and to your team!\n\nCould you please detail what you meant by segmentation:\n\n \"We found the price and item_seq_number features were interesting. We applied exponential and linear segmentation of them and adding the segmented data as new categorical features.\"",
      "votes": null
    },
    {
      "id": "350012",
      "postDate": "06/29/2018 03:01:47",
      "content": "<p>Congratulations @Mengfei Li and team on a strong finish. Thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congratulations @Mengfei Li and team on a strong finish. Thanks for sharing your solution overview.",
      "votes": null
    },
    {
      "id": "350030",
      "postDate": "06/29/2018 03:32:24",
      "content": "<p>good job! Your mentor also enjoy this game : )</p>",
      "rawMarkdown": "good job! Your mentor also enjoy this game : )",
      "votes": null
    },
    {
      "id": "350093",
      "postDate": "06/29/2018 06:48:57",
      "content": "<p>Thanks man. </p>\n\n<p>long for short, things will be explained in the script as follow: </p>\n\n<p>df[\"price_resampled\"] = np.round(np.log1p(df[\"price\"]*res_coef)).astype(np.int16) </p>",
      "rawMarkdown": "Thanks man. \n\nlong for short, things will be explained in the script as follow: \n\ndf[\"price_resampled\"] = np.round(np.log1p(df[\"price\"]*res_coef)).astype(np.int16)",
      "votes": null
    },
    {
      "id": "350094",
      "postDate": "06/29/2018 06:50:30",
      "content": "<p>Thanks bud, congratulations to you and Shannon.</p>",
      "rawMarkdown": "Thanks bud, congratulations to you and Shannon.",
      "votes": null
    },
    {
      "id": "350095",
      "postDate": "06/29/2018 06:50:55",
      "content": "<p>Thanks man, happy to see you again here :P</p>",
      "rawMarkdown": "Thanks man, happy to see you again here :P",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 349630,
      "author_name": "steinhafen",
      "author_url": "",
      "post_date": "06/28/2018 12:06:51",
      "content": "<p>Thank you for the great writeup <a href=\"/meli19\">@meli19</a> and the efforts of the team members! It was a fun journey of learning while support each other.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 349903,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "06/28/2018 20:24:18",
      "content": "<p>Congrats Mengfei and to your team!</p>\n\n<p>Could you please detail what you meant by segmentation:</p>\n\n<p>\"We found the price and item_seq_number features were interesting. We applied exponential and linear segmentation of them and adding the segmented data as new categorical features.\" </p>",
      "votes": null,
      "replies": [
        {
          "id": 350093,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "06/29/2018 06:48:57",
          "content": "<p>Thanks man. </p>\n\n<p>long for short, things will be explained in the script as follow: </p>\n\n<p>df[\"price_resampled\"] = np.round(np.log1p(df[\"price\"]*res_coef)).astype(np.int16) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350012,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/29/2018 03:01:47",
      "content": "<p>Congratulations @Mengfei Li and team on a strong finish. Thanks for sharing your solution overview.</p>",
      "votes": null,
      "replies": [
        {
          "id": 350095,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "06/29/2018 06:50:55",
          "content": "<p>Thanks man, happy to see you again here :P</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 350030,
      "author_name": "fengari",
      "author_url": "",
      "post_date": "06/29/2018 03:32:24",
      "content": "<p>good job! Your mentor also enjoy this game : )</p>",
      "votes": null,
      "replies": [
        {
          "id": 350094,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "06/29/2018 06:50:30",
          "content": "<p>Thanks bud, congratulations to you and Shannon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "349604": "First of all, congratulations to all of you guys who put great effort and learned a lot of new things from this competition. I had a great fun and meet new friends here, so lucky.\n\nNow comes the solution of our team, which should be simple without any magic. The steps are:\n\nBest single model part (LGB):\n-------------------------------------------------------------------------------\n1. Folked the public kernel and did some research such as parameter tuning which gave us a slightly better result than all the public kernels. \n2. Doing more feature engineering such as keywords count, text hash,  categorical feature hash. \n3. We found the price and item_seq_number features were interesting. We applied exponential and linear segmentation of them and adding the segmented data as new categorical features.\n4. We added \"char\" level analyzer for TFID process.\n5. Target encoding.\n6. Applying external data. we found the \"income\", \"population\" features are helpful. The Location Information did not help us.\n7. Adding basic image features based on public kernels. \n8. Final parameter tunning.\nThis model helped us to score 0.2184 after doing a 5-folds average. But it needs 3 hours to run on a 20 core server for each fold.\n\nNN part:\n-----------------------------------------------------------------------------\nWe used most the similar features for NN, which gave us a 0.2205 score on LB. By a simple stacking of the LGB and NN models, we got 0.2179 and after some blending, we got 0.2176 on LB.\n\nTeam up part:\n-------------------------------------------------------------------------------\nIt was so lucky for us that @Yuki.O joined us. He is a genius guy who did massive feature engineering which makes the big diversity between our models. A simple stack gave us 0.2169 on the public LB.\n\nStacking part:\n-------------------------------------------------------------------------------\nIn the last week, we start to stack all the models. We created aprrox. 40 models with catboost, FM, mlp, LSTM. I think because it was limited by our best single model, the can just reach 0.2156 on the public LB. We fight till the last minutes though. :P\n\nOk, this is all about what we did in this competition, great fun! Many thanks to all of you guys @huiqin, @Steeve Huang, @steinhafen, and @extremin. Special thanks to my new friend @Yuki.O, who put lots of effort into this competition and I learned a lot from him, not only the knowledge but also the professional attitude.",
    "349630": "Thank you for the great writeup @meli19 and the efforts of the team members! It was a fun journey of learning while support each other.",
    "349903": "Congrats Mengfei and to your team!\n\nCould you please detail what you meant by segmentation:\n\n \"We found the price and item_seq_number features were interesting. We applied exponential and linear segmentation of them and adding the segmented data as new categorical features.\"",
    "350012": "Congratulations @Mengfei Li and team on a strong finish. Thanks for sharing your solution overview.",
    "350030": "good job! Your mentor also enjoy this game : )",
    "350093": "Thanks man. \n\nlong for short, things will be explained in the script as follow: \n\ndf[\"price_resampled\"] = np.round(np.log1p(df[\"price\"]*res_coef)).astype(np.int16)",
    "350094": "Thanks bud, congratulations to you and Shannon.",
    "350095": "Thanks man, happy to see you again here :P"
  },
  "source": "meta"
}