{
  "id": 60102,
  "title": "22nd place solution [Team NoVices]",
  "url": "/competitions/avito-demand-prediction/writeups/no-vices-22nd-place-solution-team-novices",
  "author_name": "",
  "post_date": "2018-06-30T06:32:20.812092500Z",
  "votes": 21,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to Kaggle and Avito for hosting this competition. This was a great dataset - so much to do.</p>\n\n<p>We had 3 main models - LGB, XGB and NN. LGB Public score 0.2186, XGB Public score 0.2185, NN Public score 0.2186.  We also did stacking and Yi Tang can talk more about it. \nThis was quite a team effort for us and I really enjoyed working with team on this project. \nAbhimanyu and YiTang were feature engineers and I was responsible for modeling LGB, XGB. Yiang Zheng handled NN. I had few features of my own before we formed the team. Our score dramatically improved after I merged everyone’s feature into one model. \nI also did similar to what Joe Eddy’s post mentions. I was merging or concatenating (or pickled objects) different feature groups - sort of relational matching. I could quickly build the model and see what is working or not. </p>\n\n<p>For some reason we couldn’t use LGB/XGB features in NN and vice versa so Yiang Zheng ended up creating his own feature set. </p>\n\n<p>Image features - We had lot of image features - all mentioned in public kernels or discussion forums. I used dask to generate many image features. It worked really well. Sample code on GitHub. - <a href=\"https://github.com/rashmibanthia/Avito\">https://github.com/rashmibanthia/Avito</a> </p>\n\n<p>We had TFIDF features for title and description with Russian stemmer- code on GitHub </p>\n\n<p>Ridge features - Up until last day we were using features as is extracted from public kernel. Then YiTang had the feature generated with similar folds as our other OOF models. So we used YiTang’s ridge features to avoid any leakage. </p>\n\n<p>Aggregate features - again as is from public kernel. \nI’ll let my team elaborate more on the features:</p>\n\n<p><strong>Yiang Zheng</strong> - I'll briefly talk about my nn model (0.2186 public, 0.2225 private), since @Little Boat and @Liu Jilong have already shared their nn architectures and their amazing work. It turns out all nns overall architectures are quite similar, only differs in details. In short, there are four branches in the whole model dealing with categorical, numerical, image and text separately and concatenate into one vector, followed by several dense layer. The thing I found useful is resizing all image to small scale, I used 32 by 32 and trained a CNN on it as a part of it. It gave me significant boost. Little Boat used middle layer of Resnet and Jilong used 64 by 64 here, I think potential boost lies on using both. <br>\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59880</a>\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59917\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59917</a>\nFor simplicity, I only trained one LSTM followed by [max,ave] for title_description as text info, as littleboat mentioned that might lose roughly 0.0002 boost, though increased the speed a bit. There is not too much needed to mention in categorical embedding, only tuning the embed size based on understanding. Last, for numerical data, I did find category based features are quite useful, like mean price on all categories. I didn't think of adding more interaction between categories, which I really should do. A more clever way might be creating more categories which mentioned by Webber in here: <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59886\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59886</a>        Thanks for my teammates' hard work and everyone's sharing and I indeed enjoy this wonderful Kaggle journey.</p>",
  "messages": [
    {
      "id": "350645",
      "postDate": "06/30/2018 06:32:20",
      "content": "<p>Thanks to Kaggle and Avito for hosting this competition. This was a great dataset - so much to do.</p>\n\n<p>We had 3 main models - LGB, XGB and NN. LGB Public score 0.2186, XGB Public score 0.2185, NN Public score 0.2186.  We also did stacking and Yi Tang can talk more about it. \nThis was quite a team effort for us and I really enjoyed working with team on this project. \nAbhimanyu and YiTang were feature engineers and I was responsible for modeling LGB, XGB. Yiang Zheng handled NN. I had few features of my own before we formed the team. Our score dramatically improved after I merged everyone’s feature into one model. \nI also did similar to what Joe Eddy’s post mentions. I was merging or concatenating (or pickled objects) different feature groups - sort of relational matching. I could quickly build the model and see what is working or not. </p>\n\n<p>For some reason we couldn’t use LGB/XGB features in NN and vice versa so Yiang Zheng ended up creating his own feature set. </p>\n\n<p>Image features - We had lot of image features - all mentioned in public kernels or discussion forums. I used dask to generate many image features. It worked really well. Sample code on GitHub. - <a href=\"https://github.com/rashmibanthia/Avito\">https://github.com/rashmibanthia/Avito</a> </p>\n\n<p>We had TFIDF features for title and description with Russian stemmer- code on GitHub </p>\n\n<p>Ridge features - Up until last day we were using features as is extracted from public kernel. Then YiTang had the feature generated with similar folds as our other OOF models. So we used YiTang’s ridge features to avoid any leakage. </p>\n\n<p>Aggregate features - again as is from public kernel. \nI’ll let my team elaborate more on the features:</p>\n\n<p><strong>Yiang Zheng</strong> - I'll briefly talk about my nn model (0.2186 public, 0.2225 private), since @Little Boat and @Liu Jilong have already shared their nn architectures and their amazing work. It turns out all nns overall architectures are quite similar, only differs in details. In short, there are four branches in the whole model dealing with categorical, numerical, image and text separately and concatenate into one vector, followed by several dense layer. The thing I found useful is resizing all image to small scale, I used 32 by 32 and trained a CNN on it as a part of it. It gave me significant boost. Little Boat used middle layer of Resnet and Jilong used 64 by 64 here, I think potential boost lies on using both. <br>\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59880</a>\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59917\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59917</a>\nFor simplicity, I only trained one LSTM followed by [max,ave] for title_description as text info, as littleboat mentioned that might lose roughly 0.0002 boost, though increased the speed a bit. There is not too much needed to mention in categorical embedding, only tuning the embed size based on understanding. Last, for numerical data, I did find category based features are quite useful, like mean price on all categories. I didn't think of adding more interaction between categories, which I really should do. A more clever way might be creating more categories which mentioned by Webber in here: <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59886\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59886</a>        Thanks for my teammates' hard work and everyone's sharing and I indeed enjoy this wonderful Kaggle journey.</p>",
      "rawMarkdown": "Thanks to Kaggle and Avito for hosting this competition. This was a great dataset - so much to do.\n\nWe had 3 main models - LGB, XGB and NN. LGB Public score 0.2186, XGB Public score 0.2185, NN Public score 0.2186.  We also did stacking and Yi Tang can talk more about it. \nThis was quite a team effort for us and I really enjoyed working with team on this project. \nAbhimanyu and YiTang were feature engineers and I was responsible for modeling LGB, XGB. Yiang Zheng handled NN. I had few features of my own before we formed the team. Our score dramatically improved after I merged everyone’s feature into one model. \nI also did similar to what Joe Eddy’s post mentions. I was merging or concatenating (or pickled objects) different feature groups - sort of relational matching. I could quickly build the model and see what is working or not. \n\nFor some reason we couldn’t use LGB/XGB features in NN and vice versa so Yiang Zheng ended up creating his own feature set. \n\nImage features - We had lot of image features - all mentioned in public kernels or discussion forums. I used dask to generate many image features. It worked really well. Sample code on GitHub. - [https://github.com/rashmibanthia/Avito][1] \n\nWe had TFIDF features for title and description with Russian stemmer- code on GitHub \n\nRidge features - Up until last day we were using features as is extracted from public kernel. Then YiTang had the feature generated with similar folds as our other OOF models. So we used YiTang’s ridge features to avoid any leakage. \n\nAggregate features - again as is from public kernel. \nI’ll let my team elaborate more on the features:\n\n**Yiang Zheng** - I'll briefly talk about my nn model (0.2186 public, 0.2225 private), since @Little Boat and @Liu Jilong have already shared their nn architectures and their amazing work. It turns out all nns overall architectures are quite similar, only differs in details. In short, there are four branches in the whole model dealing with categorical, numerical, image and text separately and concatenate into one vector, followed by several dense layer. The thing I found useful is resizing all image to small scale, I used 32 by 32 and trained a CNN on it as a part of it. It gave me significant boost. Little Boat used middle layer of Resnet and Jilong used 64 by 64 here, I think potential boost lies on using both.  \nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59880\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59917\nFor simplicity, I only trained one LSTM followed by [max,ave] for title_description as text info, as littleboat mentioned that might lose roughly 0.0002 boost, though increased the speed a bit. There is not too much needed to mention in categorical embedding, only tuning the embed size based on understanding. Last, for numerical data, I did find category based features are quite useful, like mean price on all categories. I didn't think of adding more interaction between categories, which I really should do. A more clever way might be creating more categories which mentioned by Webber in here: https://www.kaggle.com/c/avito-demand-prediction/discussion/59886        Thanks for my teammates' hard work and everyone's sharing and I indeed enjoy this wonderful Kaggle journey.\n\n\n  [1]: https://github.com/rashmibanthia/Avito",
      "votes": null
    },
    {
      "id": "351113",
      "postDate": "07/01/2018 11:08:27",
      "content": "<p>This is an interesting competition for me. I decided to quit this competition and Kaggle because of other commitments in life/work. just one day before team merge deadline, Rashmi asked me to join, at the time, my position is 880th, about 50%, and Rashmis' team is about 82. so I decided to join and finish this competition which I already about many many hours.</p>\n\n<p>as part of this team, my work mostly on ensemble workflow:\n   1. make sure everyone uses the same agreed cross validation schema,\n   2. provide model_zoo.md to keep track of all level 1 models, their\n      train/valid/lb scores, feature used, and file path to their oof/test prediction.\n   3. write merge_oof.py to combine all oof/test predictions together.\n   4. write R scripts for glmnet ensemble\n   5. write python scripts for LightGBM ensemble</p>\n\n<p>Once the new models are built, other team member updates the model_zoo.md and upload the data to a private github repo. then I update the merge_oof.py to include new models' result and run glmnet and LightGBM ensemble. we had this ensemble workflow automated so it takes little effort.</p>\n\n<p>I spent some times analyzing the coefficients/weights of L1 model and tried to exclude models with negative and lower weights, but it doesn't help at all. the final submission is a glmnet ensemble with 41 models (lgb + xgb + NN).</p>\n\n<p>Also, lgb ensemble has much better cv score but the LB score is worse. i suspect it is because there are leakage in L1 models and glmnet is more robust to leakage since it's linear model. unfortunately, there's no enough time to identify which models have leakage.</p>\n\n<p>This is my 2nd-time working in a team for a Kaggle competition. Although there's a lot of room to improve in collaborating when compared with a professional data scientist team but as night/weekend project, we have done a really good job as a team.</p>\n\n<p>The setup for collaboration:\n   1. slack for discussion. we have channels for general, final_ensemble, random for cat photos etc.\n   2. we also used slack for sharing features which I personally don't like.\n   3. private github repo for sharing code and oof/test predictions.\n   4. Monday.com for managing tasks. it gives a nice overview of what everyone's up to.</p>\n\n<p>we tried very hard to get a gold, but other teams work even harder. At one point we were at 17, and finished at 22. It's not gold but very close :) </p>\n\n<p>finally, when we waited 1 hour for the final score, we had a lovely discussion, including our past disqualification experience. we were all shocked when we were on the different team but actually team up with the same person.</p>",
      "rawMarkdown": "This is an interesting competition for me. I decided to quit this competition and Kaggle because of other commitments in life/work. just one day before team merge deadline, Rashmi asked me to join, at the time, my position is 880th, about 50%, and Rashmis' team is about 82. so I decided to join and finish this competition which I already about many many hours.\n\nas part of this team, my work mostly on ensemble workflow:\n   1. make sure everyone uses the same agreed cross validation schema,\n   2. provide model_zoo.md to keep track of all level 1 models, their\n      train/valid/lb scores, feature used, and file path to their oof/test prediction.\n   3. write merge_oof.py to combine all oof/test predictions together.\n   4. write R scripts for glmnet ensemble\n   5. write python scripts for LightGBM ensemble\n\n      \nOnce the new models are built, other team member updates the model_zoo.md and upload the data to a private github repo. then I update the merge_oof.py to include new models' result and run glmnet and LightGBM ensemble. we had this ensemble workflow automated so it takes little effort.\n\nI spent some times analyzing the coefficients/weights of L1 model and tried to exclude models with negative and lower weights, but it doesn't help at all. the final submission is a glmnet ensemble with 41 models (lgb + xgb + NN).\n\nAlso, lgb ensemble has much better cv score but the LB score is worse. i suspect it is because there are leakage in L1 models and glmnet is more robust to leakage since it's linear model. unfortunately, there's no enough time to identify which models have leakage.\n   \nThis is my 2nd-time working in a team for a Kaggle competition. Although there's a lot of room to improve in collaborating when compared with a professional data scientist team but as night/weekend project, we have done a really good job as a team.\n\nThe setup for collaboration:\n   1. slack for discussion. we have channels for general, final_ensemble, random for cat photos etc.\n   2. we also used slack for sharing features which I personally don't like.\n   3. private github repo for sharing code and oof/test predictions.\n   4. Monday.com for managing tasks. it gives a nice overview of what everyone's up to.\n\nwe tried very hard to get a gold, but other teams work even harder. At one point we were at 17, and finished at 22. It's not gold but very close :) \n   \nfinally, when we waited 1 hour for the final score, we had a lovely discussion, including our past disqualification experience. we were all shocked when we were on the different team but actually team up with the same person.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 351113,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "07/01/2018 11:08:27",
      "content": "<p>This is an interesting competition for me. I decided to quit this competition and Kaggle because of other commitments in life/work. just one day before team merge deadline, Rashmi asked me to join, at the time, my position is 880th, about 50%, and Rashmis' team is about 82. so I decided to join and finish this competition which I already about many many hours.</p>\n\n<p>as part of this team, my work mostly on ensemble workflow:\n   1. make sure everyone uses the same agreed cross validation schema,\n   2. provide model_zoo.md to keep track of all level 1 models, their\n      train/valid/lb scores, feature used, and file path to their oof/test prediction.\n   3. write merge_oof.py to combine all oof/test predictions together.\n   4. write R scripts for glmnet ensemble\n   5. write python scripts for LightGBM ensemble</p>\n\n<p>Once the new models are built, other team member updates the model_zoo.md and upload the data to a private github repo. then I update the merge_oof.py to include new models' result and run glmnet and LightGBM ensemble. we had this ensemble workflow automated so it takes little effort.</p>\n\n<p>I spent some times analyzing the coefficients/weights of L1 model and tried to exclude models with negative and lower weights, but it doesn't help at all. the final submission is a glmnet ensemble with 41 models (lgb + xgb + NN).</p>\n\n<p>Also, lgb ensemble has much better cv score but the LB score is worse. i suspect it is because there are leakage in L1 models and glmnet is more robust to leakage since it's linear model. unfortunately, there's no enough time to identify which models have leakage.</p>\n\n<p>This is my 2nd-time working in a team for a Kaggle competition. Although there's a lot of room to improve in collaborating when compared with a professional data scientist team but as night/weekend project, we have done a really good job as a team.</p>\n\n<p>The setup for collaboration:\n   1. slack for discussion. we have channels for general, final_ensemble, random for cat photos etc.\n   2. we also used slack for sharing features which I personally don't like.\n   3. private github repo for sharing code and oof/test predictions.\n   4. Monday.com for managing tasks. it gives a nice overview of what everyone's up to.</p>\n\n<p>we tried very hard to get a gold, but other teams work even harder. At one point we were at 17, and finished at 22. It's not gold but very close :) </p>\n\n<p>finally, when we waited 1 hour for the final score, we had a lovely discussion, including our past disqualification experience. we were all shocked when we were on the different team but actually team up with the same person.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "350645": "Thanks to Kaggle and Avito for hosting this competition. This was a great dataset - so much to do.\n\nWe had 3 main models - LGB, XGB and NN. LGB Public score 0.2186, XGB Public score 0.2185, NN Public score 0.2186.  We also did stacking and Yi Tang can talk more about it. \nThis was quite a team effort for us and I really enjoyed working with team on this project. \nAbhimanyu and YiTang were feature engineers and I was responsible for modeling LGB, XGB. Yiang Zheng handled NN. I had few features of my own before we formed the team. Our score dramatically improved after I merged everyone’s feature into one model. \nI also did similar to what Joe Eddy’s post mentions. I was merging or concatenating (or pickled objects) different feature groups - sort of relational matching. I could quickly build the model and see what is working or not. \n\nFor some reason we couldn’t use LGB/XGB features in NN and vice versa so Yiang Zheng ended up creating his own feature set. \n\nImage features - We had lot of image features - all mentioned in public kernels or discussion forums. I used dask to generate many image features. It worked really well. Sample code on GitHub. - [https://github.com/rashmibanthia/Avito][1] \n\nWe had TFIDF features for title and description with Russian stemmer- code on GitHub \n\nRidge features - Up until last day we were using features as is extracted from public kernel. Then YiTang had the feature generated with similar folds as our other OOF models. So we used YiTang’s ridge features to avoid any leakage. \n\nAggregate features - again as is from public kernel. \nI’ll let my team elaborate more on the features:\n\n**Yiang Zheng** - I'll briefly talk about my nn model (0.2186 public, 0.2225 private), since @Little Boat and @Liu Jilong have already shared their nn architectures and their amazing work. It turns out all nns overall architectures are quite similar, only differs in details. In short, there are four branches in the whole model dealing with categorical, numerical, image and text separately and concatenate into one vector, followed by several dense layer. The thing I found useful is resizing all image to small scale, I used 32 by 32 and trained a CNN on it as a part of it. It gave me significant boost. Little Boat used middle layer of Resnet and Jilong used 64 by 64 here, I think potential boost lies on using both.  \nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59880\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59917\nFor simplicity, I only trained one LSTM followed by [max,ave] for title_description as text info, as littleboat mentioned that might lose roughly 0.0002 boost, though increased the speed a bit. There is not too much needed to mention in categorical embedding, only tuning the embed size based on understanding. Last, for numerical data, I did find category based features are quite useful, like mean price on all categories. I didn't think of adding more interaction between categories, which I really should do. A more clever way might be creating more categories which mentioned by Webber in here: https://www.kaggle.com/c/avito-demand-prediction/discussion/59886        Thanks for my teammates' hard work and everyone's sharing and I indeed enjoy this wonderful Kaggle journey.\n\n\n  [1]: https://github.com/rashmibanthia/Avito",
    "351113": "This is an interesting competition for me. I decided to quit this competition and Kaggle because of other commitments in life/work. just one day before team merge deadline, Rashmi asked me to join, at the time, my position is 880th, about 50%, and Rashmis' team is about 82. so I decided to join and finish this competition which I already about many many hours.\n\nas part of this team, my work mostly on ensemble workflow:\n   1. make sure everyone uses the same agreed cross validation schema,\n   2. provide model_zoo.md to keep track of all level 1 models, their\n      train/valid/lb scores, feature used, and file path to their oof/test prediction.\n   3. write merge_oof.py to combine all oof/test predictions together.\n   4. write R scripts for glmnet ensemble\n   5. write python scripts for LightGBM ensemble\n\n      \nOnce the new models are built, other team member updates the model_zoo.md and upload the data to a private github repo. then I update the merge_oof.py to include new models' result and run glmnet and LightGBM ensemble. we had this ensemble workflow automated so it takes little effort.\n\nI spent some times analyzing the coefficients/weights of L1 model and tried to exclude models with negative and lower weights, but it doesn't help at all. the final submission is a glmnet ensemble with 41 models (lgb + xgb + NN).\n\nAlso, lgb ensemble has much better cv score but the LB score is worse. i suspect it is because there are leakage in L1 models and glmnet is more robust to leakage since it's linear model. unfortunately, there's no enough time to identify which models have leakage.\n   \nThis is my 2nd-time working in a team for a Kaggle competition. Although there's a lot of room to improve in collaborating when compared with a professional data scientist team but as night/weekend project, we have done a really good job as a team.\n\nThe setup for collaboration:\n   1. slack for discussion. we have channels for general, final_ensemble, random for cat photos etc.\n   2. we also used slack for sharing features which I personally don't like.\n   3. private github repo for sharing code and oof/test predictions.\n   4. Monday.com for managing tasks. it gives a nice overview of what everyone's up to.\n\nwe tried very hard to get a gold, but other teams work even harder. At one point we were at 17, and finished at 22. It's not gold but very close :) \n   \nfinally, when we waited 1 hour for the final score, we had a lovely discussion, including our past disqualification experience. we were all shocked when we were on the different team but actually team up with the same person."
  },
  "source": "meta"
}