{
  "id": 59880,
  "title": "\"Dance with Ensemble\" Sharing Thread",
  "url": "/competitions/avito-demand-prediction/discussion/59880",
  "author_name": "Little Boat",
  "post_date": "2018-06-28T01:17:21.266000",
  "votes": 257,
  "comment_count": 92,
  "views": 0,
  "content": "<p><strong>Preface</strong></p>\n\n<p>First of all, I would like to thank Avito and Kaggle for hosting such an interesting competition! Lots of ways to do feature engineering, data quality is not bad, and data size is arguably accessible to everyone. </p>\n\n<p>I would also like to thank my three awesome teammates, Arsenal, Georgiy Danshchin and thousandvoices (ordered alphabetically) ! You all are truly amazing feature engineering masters! </p>\n\n<p>Arsenal has been a long-time friend of mine and we decided to go with “Dance with Ensemble” again in honor of our last loss under the same team name (which was 3 years ago already!) I believe that time we dropped out of top 3 mainly because we didn’t use NN at all. And the funny thing is, this time (R)NN is one of the main reasons that we were always staying ahead of other teams.</p>\n\n<p><strong>Summary</strong></p>\n\n<p>So our approach is probably not different than what other top teams did in a high level. We have some lgb models, some NN models, some xgb models as first layer, and some lgb models, some xgb models and some NN models as second layer, and one NN as the final layer. But honestly, the complicated structure (3 layer) probably gave us about 0.0002 - 0.0004 improvement. Just several models with a simple linear stacking should be able to achieve not exactly the same, but quite similar score.</p>\n\n<p>A few days ago our best single models were both 215X (on public LB) for NN and lgb. And then amazingly Georgiy Danshchin discovered a few features based on active train+test and immediately it boosted the best single lgb to 213X! Purely including them to my NN didn't help much but I couldn't find much time tweaking it (and honestly I was lazy at that point). I think in the end, it had 0.0007 ish improvement on our final score. I will leave this black magic to Georgiy Danshchin to disclose (Hint: RNN is involved there).</p>\n\n<p>Stacking is extremely important here, which means that building diversified models are extremely important. I remember when four of us merged, just a linear blending of our models could get to 0.2133. </p>\n\n<p><strong>NN</strong></p>\n\n<p>I exclusively worked on NNs for this one and didn’t do much feature engineering otherwise. So I would like to share how you can achieve 0.215X with a single NN, and leave the rest (truly amazing stuff) to my awesome teammates.</p>\n\n<p>All features matter here. Text, categorical, numerical, images (and probably in this order).  And to my best memory, here is how I did it:</p>\n\n<ul>\n<li>I got 0.227X with numerical features and categorical embedding</li>\n<li>And then I included titile and description with 2 RNNs, with fastText pretrained embedding, with some tuning, the score dropped to 0.221X.</li>\n<li>Played with self training fastText embedding on train+test, and also train active, test active. It turned out that self training on train+test was the best. Score got to 0.220X.</li>\n<li>Added VGG16 top layer with average pooling. It made my score worse. Did some tuning, specifically, had a separate layer before merging text, image, categorical, numerical features together, and started to see the improvement . Got to about 0.219X.</li>\n<li>Tried to tweak text models, with CNN or Attention etc. None worked. In the end, went with 2 layer LSTM followed by a dense layer. Probably 0.0003 improvement here.</li>\n<li>Tried different CNN models for images. None of the \"fine tuning\" models worked (and GOD it was slow). But fixed ResNet50 middle layer helped by probably another 0.0005. Now the score became 0.218X.</li>\n<li>Started doing all sorts of tuning (based on intuition mostly). And found that adding spatial dropout between text and LSTM helped quite a bit, probably 0.0007 - 0.001. And fine tuned dropout ratio overall helped too. In the end, about 0.001 - 0.0015 improvement here. So now the score was around 0.2165 - 0.217.</li>\n<li>Started including all engineered features from teammates. Lots of engineered features from them (the ones based on text) didn't help but others did. So in the end, a NN with 0.215X!</li>\n<li>If you kept saving models along the way (models with fewer features, models with more features that got worse result, etc.), you could train a fully connected NN on top of them and for me it was around 0.008 improvement in addition. In other words, you can easily get into top 10 with only NN!</li>\n</ul>\n\n<p>I also attached a simple sketch of what the model architecture looks like. </p>\n\n<p>And this thread is <strong>To Be Continued</strong> by my amazing teammates!</p>\n\n<p>Jump to: </p>\n\n<p><strong>Arsenal's approach:</strong>\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349563\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349563</a></p>\n\n<p><strong>thousandvoices's approach:</strong> \n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349386\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349386</a></p>\n\n<p><strong>Georgiy Danshchin:</strong> \n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349710\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349710</a></p>\n\n<p><img src=\"https://pbs.twimg.com/media/DgvX3pWUYAAlCAK.jpg:large\" alt=\"NN Model\"></p>\n\n<p>Once again, thank you all for this amazing experience!</p>",
  "messages": [
    {
      "id": 349269,
      "postDate": "2018-06-28T01:17:21.267Z",
      "content": "<p><strong>Preface</strong></p>\n\n<p>First of all, I would like to thank Avito and Kaggle for hosting such an interesting competition! Lots of ways to do feature engineering, data quality is not bad, and data size is arguably accessible to everyone. </p>\n\n<p>I would also like to thank my three awesome teammates, Arsenal, Georgiy Danshchin and thousandvoices (ordered alphabetically) ! You all are truly amazing feature engineering masters! </p>\n\n<p>Arsenal has been a long-time friend of mine and we decided to go with “Dance with Ensemble” again in honor of our last loss under the same team name (which was 3 years ago already!) I believe that time we dropped out of top 3 mainly because we didn’t use NN at all. And the funny thing is, this time (R)NN is one of the main reasons that we were always staying ahead of other teams.</p>\n\n<p><strong>Summary</strong></p>\n\n<p>So our approach is probably not different than what other top teams did in a high level. We have some lgb models, some NN models, some xgb models as first layer, and some lgb models, some xgb models and some NN models as second layer, and one NN as the final layer. But honestly, the complicated structure (3 layer) probably gave us about 0.0002 - 0.0004 improvement. Just several models with a simple linear stacking should be able to achieve not exactly the same, but quite similar score.</p>\n\n<p>A few days ago our best single models were both 215X (on public LB) for NN and lgb. And then amazingly Georgiy Danshchin discovered a few features based on active train+test and immediately it boosted the best single lgb to 213X! Purely including them to my NN didn't help much but I couldn't find much time tweaking it (and honestly I was lazy at that point). I think in the end, it had 0.0007 ish improvement on our final score. I will leave this black magic to Georgiy Danshchin to disclose (Hint: RNN is involved there).</p>\n\n<p>Stacking is extremely important here, which means that building diversified models are extremely important. I remember when four of us merged, just a linear blending of our models could get to 0.2133. </p>\n\n<p><strong>NN</strong></p>\n\n<p>I exclusively worked on NNs for this one and didn’t do much feature engineering otherwise. So I would like to share how you can achieve 0.215X with a single NN, and leave the rest (truly amazing stuff) to my awesome teammates.</p>\n\n<p>All features matter here. Text, categorical, numerical, images (and probably in this order).  And to my best memory, here is how I did it:</p>\n\n<ul>\n<li>I got 0.227X with numerical features and categorical embedding</li>\n<li>And then I included titile and description with 2 RNNs, with fastText pretrained embedding, with some tuning, the score dropped to 0.221X.</li>\n<li>Played with self training fastText embedding on train+test, and also train active, test active. It turned out that self training on train+test was the best. Score got to 0.220X.</li>\n<li>Added VGG16 top layer with average pooling. It made my score worse. Did some tuning, specifically, had a separate layer before merging text, image, categorical, numerical features together, and started to see the improvement . Got to about 0.219X.</li>\n<li>Tried to tweak text models, with CNN or Attention etc. None worked. In the end, went with 2 layer LSTM followed by a dense layer. Probably 0.0003 improvement here.</li>\n<li>Tried different CNN models for images. None of the \"fine tuning\" models worked (and GOD it was slow). But fixed ResNet50 middle layer helped by probably another 0.0005. Now the score became 0.218X.</li>\n<li>Started doing all sorts of tuning (based on intuition mostly). And found that adding spatial dropout between text and LSTM helped quite a bit, probably 0.0007 - 0.001. And fine tuned dropout ratio overall helped too. In the end, about 0.001 - 0.0015 improvement here. So now the score was around 0.2165 - 0.217.</li>\n<li>Started including all engineered features from teammates. Lots of engineered features from them (the ones based on text) didn't help but others did. So in the end, a NN with 0.215X!</li>\n<li>If you kept saving models along the way (models with fewer features, models with more features that got worse result, etc.), you could train a fully connected NN on top of them and for me it was around 0.008 improvement in addition. In other words, you can easily get into top 10 with only NN!</li>\n</ul>\n\n<p>I also attached a simple sketch of what the model architecture looks like. </p>\n\n<p>And this thread is <strong>To Be Continued</strong> by my amazing teammates!</p>\n\n<p>Jump to: </p>\n\n<p><strong>Arsenal's approach:</strong>\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349563\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349563</a></p>\n\n<p><strong>thousandvoices's approach:</strong> \n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349386\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349386</a></p>\n\n<p><strong>Georgiy Danshchin:</strong> \n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349710\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349710</a></p>\n\n<p><img src=\"https://pbs.twimg.com/media/DgvX3pWUYAAlCAK.jpg:large\" alt=\"NN Model\"></p>\n\n<p>Once again, thank you all for this amazing experience!</p>",
      "rawMarkdown": "**Preface**\n\nFirst of all, I would like to thank Avito and Kaggle for hosting such an interesting competition! Lots of ways to do feature engineering, data quality is not bad, and data size is arguably accessible to everyone. \n\nI would also like to thank my three awesome teammates, Arsenal, Georgiy Danshchin and thousandvoices (ordered alphabetically) ! You all are truly amazing feature engineering masters! \n\nArsenal has been a long-time friend of mine and we decided to go with “Dance with Ensemble” again in honor of our last loss under the same team name (which was 3 years ago already!) I believe that time we dropped out of top 3 mainly because we didn’t use NN at all. And the funny thing is, this time (R)NN is one of the main reasons that we were always staying ahead of other teams.\n\n**Summary**\n\nSo our approach is probably not different than what other top teams did in a high level. We have some lgb models, some NN models, some xgb models as first layer, and some lgb models, some xgb models and some NN models as second layer, and one NN as the final layer. But honestly, the complicated structure (3 layer) probably gave us about 0.0002 - 0.0004 improvement. Just several models with a simple linear stacking should be able to achieve not exactly the same, but quite similar score.\n\nA few days ago our best single models were both 215X (on public LB) for NN and lgb. And then amazingly Georgiy Danshchin discovered a few features based on active train+test and immediately it boosted the best single lgb to 213X! Purely including them to my NN didn't help much but I couldn't find much time tweaking it (and honestly I was lazy at that point). I think in the end, it had 0.0007 ish improvement on our final score. I will leave this black magic to Georgiy Danshchin to disclose (Hint: RNN is involved there).\n\nStacking is extremely important here, which means that building diversified models are extremely important. I remember when four of us merged, just a linear blending of our models could get to 0.2133. \n\n**NN**\n\nI exclusively worked on NNs for this one and didn’t do much feature engineering otherwise. So I would like to share how you can achieve 0.215X with a single NN, and leave the rest (truly amazing stuff) to my awesome teammates.\n\nAll features matter here. Text, categorical, numerical, images (and probably in this order).  And to my best memory, here is how I did it:\n\n - I got 0.227X with numerical features and categorical embedding\n - And then I included titile and description with 2 RNNs, with fastText pretrained embedding, with some tuning, the score dropped to 0.221X.\n - Played with self training fastText embedding on train+test, and also train active, test active. It turned out that self training on train+test was the best. Score got to 0.220X.\n - Added VGG16 top layer with average pooling. It made my score worse. Did some tuning, specifically, had a separate layer before merging text, image, categorical, numerical features together, and started to see the improvement . Got to about 0.219X.\n - Tried to tweak text models, with CNN or Attention etc. None worked. In the end, went with 2 layer LSTM followed by a dense layer. Probably 0.0003 improvement here.\n - Tried different CNN models for images. None of the \"fine tuning\" models worked (and GOD it was slow). But fixed ResNet50 middle layer helped by probably another 0.0005. Now the score became 0.218X.\n - Started doing all sorts of tuning (based on intuition mostly). And found that adding spatial dropout between text and LSTM helped quite a bit, probably 0.0007 - 0.001. And fine tuned dropout ratio overall helped too. In the end, about 0.001 - 0.0015 improvement here. So now the score was around 0.2165 - 0.217.\n - Started including all engineered features from teammates. Lots of engineered features from them (the ones based on text) didn't help but others did. So in the end, a NN with 0.215X!\n - If you kept saving models along the way (models with fewer features, models with more features that got worse result, etc.), you could train a fully connected NN on top of them and for me it was around 0.008 improvement in addition. In other words, you can easily get into top 10 with only NN!\n\n\nI also attached a simple sketch of what the model architecture looks like. \n\n\nAnd this thread is **To Be Continued** by my amazing teammates!\n\nJump to: \n\n**Arsenal's approach:**\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349563\n\n**thousandvoices's approach:** \nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349386\n\n**Georgiy Danshchin:** \nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349710\n\n![NN Model][1]\n\n\n  [1]: https://pbs.twimg.com/media/DgvX3pWUYAAlCAK.jpg:large\n\n\nOnce again, thank you all for this amazing experience!",
      "votes": 257
    },
    {
      "id": 349710,
      "postDate": "2018-06-28T14:18:45.603Z",
      "content": "<h2>The last part of Dance with Ensemble solution</h2>\n\n<p>First of all, I want to thank my great teammates for this awesome experience and Avito for organizing very interesting competition!</p>\n\n<p>From my point of view main things for this competition were strong feature engineering and stacking of diverse models. So let’s dive into the details.</p>\n\n<h2>Validation</h2>\n\n<p>Before joining the team I've used a time rolling window validation split with 4 folds (8 days in train 3 days in test for each fold). It was very consistent with public lb score, gap between cv and lb was between 0.002 and 0.003. After joining the team I've regenerated all my features and models for common kfold split, and noticed that cv gap increased by 0.003. I figured out that this was mainly due to user_id based target encoding features, so I put them aside and used them only for second level models later.</p>\n\n<h2>Features</h2>\n\n<p><strong>Categories</strong></p>\n\n<ul>\n<li>Percent encoding of categories computed on train+test parts of data</li>\n<li>Percent encoding of user_id category computed on train_active+test_active.</li>\n<li>Mean target encoding of different categories, which were computed separately for each fold in the oof manner using inner folds.</li>\n<li>Mean log1p(price) by different categories and differences between these features and actual log1p(price) as well. For most categories these features were computed using train+test, for user_id and title I used train_active+test_active.</li>\n<li>Mean 'renewed' value by different categories (train_active + test_active)</li>\n<li>Mean 'total_duration' value by different categories (train_active + test_active)</li>\n<li>Sparse one hot encoded categories for Lightgbm models</li>\n</ul>\n\n<p>I’ve used only single categories for all these features, didn’t see any consistent improvement with category interactions. In addition to category_name/region/param_* I treated stemmed title as category too. The 'renewed' value is just a binary value which indicates whether an item has more than one period in periods_train/periods_test data, 'total_duration' is the sum of lengths of all periods for particular item.\nAll 'mean' features were postprocessed with additive smoothing. The coefficients were different for different groups of features and were tuned manually with help of cv feedback.</p>\n\n<p><strong>Text</strong></p>\n\n<ul>\n<li>Sparse tfidf 1-2grams vectors for title and joined title+description+params (used only in Lightgbm models)</li>\n<li>14 different out of fold ridge predictions over different sparse tfidf vectors. Different combinations of fields, and their concatenations, stemmed/raw tokens etc. Computed separately for each fold using inner folds as mean target features.</li>\n<li>Various text statistics: caps quotient, alfanum quotient, char counts, word counts etc.</li>\n<li>Fasttext embeddings trained on train_active + test_active were used as initial embedding matrix for rnn models.</li>\n</ul>\n\n<p><strong>Numeric</strong></p>\n\n<ul>\n<li>Just log1p(price) and log1p(item_seq_number)</li>\n</ul>\n\n<p><strong>Images</strong></p>\n\n<ul>\n<li>Channel mean values</li>\n<li>Image width and height</li>\n<li>Pretrained resnet152 last hidden layer output (only for nn models)</li>\n<li>100 components of svd over previous feature</li>\n<li>Out of fold ridge predictions trained on pretrained resnet101 and resnet152 last hidden layers.</li>\n<li>Finetuned resnet34 out of fold predictions gave a significant boost to my models but took a loooong time to train :)</li>\n</ul>\n\n<p><strong>Active data</strong></p>\n\n<p>I’ve designed 3 nn models trained on active data to predict log1p(price), renewed and log1p(total_duration). These models had 2 rnn branches for title and description and also used category embeddings. The difference between actual log1p(price) and predicted one was an extremely important feature.</p>\n\n<h2>Models</h2>\n\n<p>I had 25 base models (15 lgbs and 10 nns) trained on different feature subsets and with different parameters. My best lgb reaches 0.2155/0.2196 public/private score, best nn is 0.2160/0.2200. Stacking of these models + finetuned resnet34 oof + user_id mean target with one lgb gives 0.2132 on public and 0.2174 on private.</p>\n\n<p>Lgb models could be divided into two groups. One of them uses small set of dense features, it’s trained for 700 rounds with num_leaves=31. Second group uses all features from the first group with addition of sparse tfidf/categories and svd decomposition of pretrained resnet152. It’s trained for 2300 rounds with num_leaves=384, it takes about 1 hour to train on one fold.</p>\n\n<p>NN models have separate branches for title and description: 2-level bidirectional lstm/gru with attention. They also contain separate branches for category embeddings/numeric features and pretrained resnet152 vectors. Every branch is followed by a separate dense layer and finally everything is concatenated and followed by another common dense layer. Text embedding layer is shared for title/description, it’s initialized with fasttext trained on train_active + test_active and frozen for the first epoch, but finetuned later.</p>",
      "rawMarkdown": "## The last part of Dance with Ensemble solution ##\n\nFirst of all, I want to thank my great teammates for this awesome experience and Avito for organizing very interesting competition!\n\nFrom my point of view main things for this competition were strong feature engineering and stacking of diverse models. So let’s dive into the details.\n\nValidation\n----------\n\nBefore joining the team I've used a time rolling window validation split with 4 folds (8 days in train 3 days in test for each fold). It was very consistent with public lb score, gap between cv and lb was between 0.002 and 0.003. After joining the team I've regenerated all my features and models for common kfold split, and noticed that cv gap increased by 0.003. I figured out that this was mainly due to user_id based target encoding features, so I put them aside and used them only for second level models later.\n\nFeatures\n--------\n\n**Categories**\n\n - Percent encoding of categories computed on train+test parts of data\n - Percent encoding of user_id category computed on train_active+test_active.\n - Mean target encoding of different categories, which were computed separately for each fold in the oof manner using inner folds.\n - Mean log1p(price) by different categories and differences between these features and actual log1p(price) as well. For most categories these features were computed using train+test, for user_id and title I used train_active+test_active.\n - Mean 'renewed' value by different categories (train_active + test_active)\n - Mean 'total_duration' value by different categories (train_active + test_active)\n - Sparse one hot encoded categories for Lightgbm models\n\nI’ve used only single categories for all these features, didn’t see any consistent improvement with category interactions. In addition to category_name/region/param_* I treated stemmed title as category too. The 'renewed' value is just a binary value which indicates whether an item has more than one period in periods_train/periods_test data, 'total_duration' is the sum of lengths of all periods for particular item.\nAll 'mean' features were postprocessed with additive smoothing. The coefficients were different for different groups of features and were tuned manually with help of cv feedback.\n\n**Text**\n\n - Sparse tfidf 1-2grams vectors for title and joined title+description+params (used only in Lightgbm models)\n - 14 different out of fold ridge predictions over different sparse tfidf vectors. Different combinations of fields, and their concatenations, stemmed/raw tokens etc. Computed separately for each fold using inner folds as mean target features.\n - Various text statistics: caps quotient, alfanum quotient, char counts, word counts etc.\n - Fasttext embeddings trained on train_active + test_active were used as initial embedding matrix for rnn models.\n\n**Numeric**\n\n -  Just log1p(price) and log1p(item_seq_number)\n\n**Images**\n\n - Channel mean values\n - Image width and height\n - Pretrained resnet152 last hidden layer output (only for nn models)\n - 100 components of svd over previous feature\n - Out of fold ridge predictions trained on pretrained resnet101 and resnet152 last hidden layers.\n - Finetuned resnet34 out of fold predictions gave a significant boost to my models but took a loooong time to train :)\n\n**Active data**\n\nI’ve designed 3 nn models trained on active data to predict log1p(price), renewed and log1p(total_duration). These models had 2 rnn branches for title and description and also used category embeddings. The difference between actual log1p(price) and predicted one was an extremely important feature.\n\nModels\n------\n\nI had 25 base models (15 lgbs and 10 nns) trained on different feature subsets and with different parameters. My best lgb reaches 0.2155/0.2196 public/private score, best nn is 0.2160/0.2200. Stacking of these models + finetuned resnet34 oof + user_id mean target with one lgb gives 0.2132 on public and 0.2174 on private.\n\nLgb models could be divided into two groups. One of them uses small set of dense features, it’s trained for 700 rounds with num_leaves=31. Second group uses all features from the first group with addition of sparse tfidf/categories and svd decomposition of pretrained resnet152. It’s trained for 2300 rounds with num_leaves=384, it takes about 1 hour to train on one fold.\n\nNN models have separate branches for title and description: 2-level bidirectional lstm/gru with attention. They also contain separate branches for category embeddings/numeric features and pretrained resnet152 vectors. Every branch is followed by a separate dense layer and finally everything is concatenated and followed by another common dense layer. Text embedding layer is shared for title/description, it’s initialized with fasttext trained on train_active + test_active and frozen for the first epoch, but finetuned later.\n",
      "votes": 53,
      "replies": [
        {
          "id": 349829,
          "postDate": "2018-06-28T17:34:11.503Z",
          "content": "<p>Congratulations and thanks for sharing.</p>",
          "rawMarkdown": "Congratulations and thanks for sharing."
        },
        {
          "id": 349832,
          "postDate": "2018-06-28T17:39:58.140Z",
          "content": "<p>Excellent, very nice approach on <strong>Active data</strong> I guess this was the magic features, is that correct ? </p>",
          "rawMarkdown": "Excellent, very nice approach on **Active data** I guess this was the magic features, is that correct ? "
        },
        {
          "id": 349918,
          "postDate": "2018-06-28T21:15:55.453Z",
          "content": "<p>Hi Georgiy ,</p>\n\n<p>Thanks a lot for the explanation...</p>\n\n<p>It seems that you used mean target encoding for user_id for lgb. There is probably 2/3 of test data where the user_id is not in train data, so I was scared that this would overfit ? Can you explain why you still choose to include it ?</p>\n\n<p>Also what was the reason behind percent encoding of categories ?</p>\n\n<p>Thanks a lot,\nAntoine.</p>",
          "rawMarkdown": "Hi Georgiy ,\n\nThanks a lot for the explanation...\n\nIt seems that you used mean target encoding for user_id for lgb. There is probably 2/3 of test data where the user_id is not in train data, so I was scared that this would overfit ? Can you explain why you still choose to include it ?\n\nAlso what was the reason behind percent encoding of categories ?\n\nThanks a lot,\nAntoine."
        },
        {
          "id": 349964,
          "postDate": "2018-06-29T00:31:36.227Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 349966,
          "postDate": "2018-06-29T00:33:46.197Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 349970,
          "postDate": "2018-06-29T01:00:12.607Z",
          "content": "<p>&gt; It seems that you used mean target encoding for user_id for lgb. There is probably 2/3 of test data where the user_id is not in train data, so I was scared that this would overfit ? Can you explain why you still choose to include it ?</p>\n\n<p>Actually it was not just user_id, but common_user_id, which means that the mean target was computed only for those users which were shared between train and test parts of each fold, so I had different sets of common_users for different folds and another set of common_users for the train+test data used for submissions. This approach worked really well with time based split versus kfold split, maybe just the rate of shared users matters (0.06 with time based cv versus 0.15 with plain kfold, and 0.067 for train/test), or maybe there is some complicated time dependency for this kind of information, I don't know for sure.</p>\n\n<p>&gt; Also what was the reason behind percent encoding of categories ?</p>\n\n<p>It's just a first type of categorical encoding which comes to my mind :) I suppose that the frequency of each category value across the training data could give a relevant information (at lest more than raw category value :)). After the percent encoding I usually try mean targets (which are more vulnerable to overfitting) and other approaches.</p>",
          "rawMarkdown": "&gt; It seems that you used mean target encoding for user_id for lgb. There is probably 2/3 of test data where the user_id is not in train data, so I was scared that this would overfit ? Can you explain why you still choose to include it ?\n\nActually it was not just user_id, but common_user_id, which means that the mean target was computed only for those users which were shared between train and test parts of each fold, so I had different sets of common_users for different folds and another set of common_users for the train+test data used for submissions. This approach worked really well with time based split versus kfold split, maybe just the rate of shared users matters (0.06 with time based cv versus 0.15 with plain kfold, and 0.067 for train/test), or maybe there is some complicated time dependency for this kind of information, I don't know for sure.\n\n&gt; Also what was the reason behind percent encoding of categories ?\n\nIt's just a first type of categorical encoding which comes to my mind :) I suppose that the frequency of each category value across the training data could give a relevant information (at lest more than raw category value :)). After the percent encoding I usually try mean targets (which are more vulnerable to overfitting) and other approaches.",
          "votes": 2
        },
        {
          "id": 350136,
          "postDate": "2018-06-29T08:48:50.033Z",
          "content": "<p>Thanks for the sharing. I'm not clear about the category embedding, could you explain more details how to do it ? with NN  or some sklearn lib ?</p>",
          "rawMarkdown": "Thanks for the sharing. I'm not clear about the category embedding, could you explain more details how to do it ? with NN  or some sklearn lib ?"
        },
        {
          "id": 350314,
          "postDate": "2018-06-29T14:39:46.453Z",
          "content": "<p>I See.. thanks a lot for clarifying this... I was thinking about doing a similar approach but I wasn't sure if it was the proper way to do it... so instead I choose to generate as many feature as possible around user_id. Thanks a lot for your detailed explanation, very helpful !</p>",
          "rawMarkdown": "I See.. thanks a lot for clarifying this... I was thinking about doing a similar approach but I wasn't sure if it was the proper way to do it... so instead I choose to generate as many feature as possible around user_id. Thanks a lot for your detailed explanation, very helpful !"
        },
        {
          "id": 353007,
          "postDate": "2018-07-05T16:31:00.470Z",
          "content": "<p>Thanks for sharing.</p>",
          "rawMarkdown": "Thanks for sharing."
        },
        {
          "id": 353300,
          "postDate": "2018-07-06T11:34:16.397Z",
          "content": "<blockquote>\n  <p><strong>Darragh wrote</strong></p>\n  \n  <blockquote>\n    <p>Excellent, very nice approach on <strong>Active data</strong> I guess this was the magic features, is that correct ? </p>\n  </blockquote>\n</blockquote>\n\n<p>Yes, these are last strong feature I've generated.</p>\n\n<blockquote>\n  <p><strong>yyqing wrote</strong></p>\n  \n  <blockquote>\n    <p>Thanks for the sharing. I'm not clear about the category embedding, could you explain more details how to do it ? with NN  or some sklearn lib ?</p>\n  </blockquote>\n</blockquote>\n\n<p>It's just an embedding layer which acts like a lookup dictionary between specific category value and an n-dimensional vector. You can check a pytorch implementation: <a href=\"https://pytorch.org/docs/stable/nn.html#embedding\">https://pytorch.org/docs/stable/nn.html#embedding</a>.</p>",
          "rawMarkdown": "\n&gt; **Darragh wrote**\n&gt; \n&gt; &gt; Excellent, very nice approach on **Active data** I guess this was the magic features, is that correct ? \n\nYes, these are last strong feature I've generated.\n\n\n&gt; **yyqing wrote**\n&gt; \n&gt; &gt; Thanks for the sharing. I'm not clear about the category embedding, could you explain more details how to do it ? with NN  or some sklearn lib ?\n\nIt's just an embedding layer which acts like a lookup dictionary between specific category value and an n-dimensional vector. You can check a pytorch implementation: https://pytorch.org/docs/stable/nn.html#embedding."
        }
      ]
    },
    {
      "id": 349386,
      "postDate": "2018-06-28T03:57:29.850Z",
      "content": "<p>First of all, I would like to thank my teammates for their excellent performance. Your discussions and feature engineering insights were incredibly helpful. I learnt a lot from you.</p>\n\n<p>Here is a short summary of my approach. I trained 10 first level models (5 lightgbms, 3 neural nets, 1 catboost and 1 ridge). Let me share some ideas behind my best single lightgbm:</p>\n\n<h3>Categorical features interactions</h3>\n\n<p>Concatenating categorical features is a well-known method for taking this into account. Rather surprisingly, I wasn't able to come up with an automated pipeline that outperformed a list of pairs of features to merge that I created based simply on common sense. Here it is:</p>\n\n<ul>\n<li>region, parent category name</li>\n<li>region, category name</li>\n<li>region, image top-1</li>\n<li>region, param 1</li>\n<li>region, param 2</li>\n<li>region, param 3</li>\n<li>city, parent category name</li>\n<li>city, category name</li>\n<li>city, param 1</li>\n<li>city, param 2</li>\n<li>city, param 3</li>\n<li>category name, image top-1</li>\n<li>parent category name, image top-1</li>\n<li>param 1, param 2</li>\n<li>param 1, image top-1</li>\n</ul>\n\n<h3>Categorical features encoding</h3>\n\n<p>For each categorical variable I added the following features to the model based on the value it took for a particular advertisement:</p>\n\n<ul>\n<li>Mean target of previous ads</li>\n<li>Total occurence count</li>\n<li>Median price</li>\n</ul>\n\n<h3>Text features</h3>\n\n<p>Raw word counts extracted by CountVectorizer for title, description and concatenated titles of all user's ads worked best for me. My efforts to design dense representations by applying svd or training embeddings with supplementary data only decreased the score. My model also included several engineered text features like word and char length, amount of caps, digits, punctuation, cyrillic and latin letters, and unique words.</p>\n\n<h3>Image features</h3>\n\n<p>I didn't focus on this part of the challenge much, that's why the only image features I used were image width and height.</p>\n\n<h3>Aggregated user features</h3>\n\n<p>Minimum, median and maximum price, total amount of ads and number of ads with unique titles.</p>\n\n<h3>Using supplementary data</h3>\n\n<p>I trained 2 neural networks that predicted price and length of active period for the ad on supplementary data provided in train_active and test_active. Both of this features turned out to be very strong, making it into 5 most important features of my model. I also used features produced by <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">this kernel</a> without any changes.</p>\n\n<h3>Hyperparameters</h3>\n\n<p>This competition favored VERY deep trees, and that was the point that most public scripts have missed. Every time I increased num_leaves my score improved significantly. I ended up using trees with 1024 leaves because it was the maximum that fitted into my memory.</p>\n\n<h3>Parting words</h3>\n\n<p>Overall, this was an exciting competition with lots of raw data and feature engineering opportunities (this is probably the reason why team merges worked so well here :) ), and I'd love to thank Avito for providing such a challenging dataset.</p>",
      "rawMarkdown": "First of all, I would like to thank my teammates for their excellent performance. Your discussions and feature engineering insights were incredibly helpful. I learnt a lot from you.\n\nHere is a short summary of my approach. I trained 10 first level models (5 lightgbms, 3 neural nets, 1 catboost and 1 ridge). Let me share some ideas behind my best single lightgbm:\n### Categorical features interactions ###\nConcatenating categorical features is a well-known method for taking this into account. Rather surprisingly, I wasn't able to come up with an automated pipeline that outperformed a list of pairs of features to merge that I created based simply on common sense. Here it is:\n\n - region, parent category name\n - region, category name\n - region, image top-1\n - region, param 1\n - region, param 2\n - region, param 3\n - city, parent category name\n - city, category name\n - city, param 1\n - city, param 2\n - city, param 3\n - category name, image top-1\n - parent category name, image top-1\n - param 1, param 2\n - param 1, image top-1\n\n### Categorical features encoding ###\nFor each categorical variable I added the following features to the model based on the value it took for a particular advertisement:\n\n - Mean target of previous ads\n - Total occurence count\n - Median price\n\n### Text features ###\nRaw word counts extracted by CountVectorizer for title, description and concatenated titles of all user's ads worked best for me. My efforts to design dense representations by applying svd or training embeddings with supplementary data only decreased the score. My model also included several engineered text features like word and char length, amount of caps, digits, punctuation, cyrillic and latin letters, and unique words.\n### Image features ###\nI didn't focus on this part of the challenge much, that's why the only image features I used were image width and height.\n### Aggregated user features ###\nMinimum, median and maximum price, total amount of ads and number of ads with unique titles.\n### Using supplementary data ###\nI trained 2 neural networks that predicted price and length of active period for the ad on supplementary data provided in train_active and test_active. Both of this features turned out to be very strong, making it into 5 most important features of my model. I also used features produced by [this kernel][1] without any changes.\n### Hyperparameters ###\nThis competition favored VERY deep trees, and that was the point that most public scripts have missed. Every time I increased num_leaves my score improved significantly. I ended up using trees with 1024 leaves because it was the maximum that fitted into my memory.\n### Parting words ###\nOverall, this was an exciting competition with lots of raw data and feature engineering opportunities (this is probably the reason why team merges worked so well here :) ), and I'd love to thank Avito for providing such a challenging dataset.\n\n\n  [1]: https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm",
      "votes": 43,
      "replies": [
        {
          "id": 349390,
          "postDate": "2018-06-28T04:10:18.933Z",
          "content": "<p>Thanks ! Great work !\nWhat is mean target of previous ads ?</p>",
          "rawMarkdown": "Thanks ! Great work !\nWhat is mean target of previous ads ?",
          "votes": 2
        },
        {
          "id": 349424,
          "postDate": "2018-06-28T05:05:19.497Z",
          "content": "<p>Thanks for sharing and congrats on your great result!</p>\n\n<blockquote>\n  <p>Every time I increased num_leaves my score improved significantly.</p>\n</blockquote>\n\n<p>I confirm :) how much RAM did you use in order to fit num_leaves=1024? </p>",
          "rawMarkdown": "Thanks for sharing and congrats on your great result!\n\n&gt; Every time I increased num_leaves my score improved significantly.\n\nI confirm :) how much RAM did you use in order to fit num_leaves=1024? ",
          "votes": 1
        },
        {
          "id": 349731,
          "postDate": "2018-06-28T14:47:32.180Z",
          "content": "<p>Hi thousandvoices, congratulations on your 1st place!<br>\nI have two questions about num leaves. <br>\n1) Is there any reason you can think of, that this competition favored larger num leaves? <br>\n2) In general, increasing num leaves tends to overfit your model (at least that is what the LGBM tuning guide says). Why were you confident that your model wasn't overfitting? Did you trust your CVscore?<br>\n<br>\nI also had the same experience about increasing num leaves, but wasn't bold enough to increase it to 1024 :( <br></p>",
          "rawMarkdown": "Hi thousandvoices, congratulations on your 1st place!<br>\nI have two questions about num leaves. <br>\n1) Is there any reason you can think of, that this competition favored larger num leaves? <br>\n2) In general, increasing num leaves tends to overfit your model (at least that is what the LGBM tuning guide says). Why were you confident that your model wasn't overfitting? Did you trust your CVscore?<br>\n<br>\nI also had the same experience about increasing num leaves, but wasn't bold enough to increase it to 1024 :( <br>\n",
          "votes": 1
        },
        {
          "id": 349745,
          "postDate": "2018-06-28T15:12:28.920Z",
          "content": "<p>Congratulations and thanks for sharing <a href=\"/thousandvoices\">@thousandvoices</a>. </p>\n\n<p>I discovered that increasing LGBM's number of leaves improved my results even better than some of my features in the last 3 days of the contest. I was concerned about over-fitting so did not go beyond 370 as I have never experienced that before in practice.</p>",
          "rawMarkdown": "Congratulations and thanks for sharing @thousandvoices. \n\nI discovered that increasing LGBM's number of leaves improved my results even better than some of my features in the last 3 days of the contest. I was concerned about over-fitting so did not go beyond 370 as I have never experienced that before in practice.",
          "votes": 1
        },
        {
          "id": 349831,
          "postDate": "2018-06-28T17:36:59.927Z",
          "content": "<blockquote>\n  <p><strong>Arnaud Roussel wrote</strong></p>\n  \n  <p>Thanks ! Great work !\n  What is mean target of previous ads ?</p>\n</blockquote>\n\n<p>Well, it is literally deal probability averaged among all ads that precede the ad being encoded and have the same value of the categorical feature.</p>\n\n<blockquote>\n  <p><strong>A^b wrote</strong></p>\n  \n  <p>Thanks for sharing and congrats on your great result!</p>\n  \n  <blockquote>\n    <p>Every time I increased num_leaves my score improved significantly.</p>\n  </blockquote>\n  \n  <p>I confirm :) how much RAM did you use in order to fit num_leaves=1024? </p>\n</blockquote>\n\n<p>Thank you for your kind words! I used a desktop with 16 gigabytes of RAM.</p>\n\n<blockquote>\n  <p><strong>pocket wrote</strong></p>\n  \n  <p>Hi thousandvoices, congratulations on your 1st place!\n  I have two questions about num leaves.\n  1) Is there any reason you can think of, that this competition favored larger num leaves?\n  2) In general, increasing num leaves tends to overfit your model (at least that is what the LGBM tuning guide says). Why were you confident that your model wasn't overfitting? Did you trust your CVscore?</p>\n  \n  <p>I also had the same experience about increasing num leaves, but wasn't bold enough to increase it to 1024 :(</p>\n</blockquote>\n\n<p>CV score is the only thing I actually trust. This approach always worked for me in the past and this fact along with consistent cv/leaderboard difference made rather confident. Regarding your first question... I don't know and would love to hear if anyone has a good explanation for that.</p>",
          "rawMarkdown": "\n&gt; **Arnaud Roussel wrote**\n&gt; \n&gt; Thanks ! Great work !\n&gt; What is mean target of previous ads ?\n\nWell, it is literally deal probability averaged among all ads that precede the ad being encoded and have the same value of the categorical feature.\n\n&gt; **A^b wrote**\n&gt; \n&gt; Thanks for sharing and congrats on your great result!\n&gt; \n&gt; &gt; Every time I increased num_leaves my score improved significantly.\n&gt; \n&gt; I confirm :) how much RAM did you use in order to fit num_leaves=1024? \n\nThank you for your kind words! I used a desktop with 16 gigabytes of RAM.\n\n&gt; **pocket wrote**\n&gt; \n&gt; Hi thousandvoices, congratulations on your 1st place!\n&gt; I have two questions about num leaves.\n&gt; 1) Is there any reason you can think of, that this competition favored larger num leaves?\n&gt; 2) In general, increasing num leaves tends to overfit your model (at least that is what the LGBM tuning guide says). Why were you confident that your model wasn't overfitting? Did you trust your CVscore?\n&gt; \n&gt; I also had the same experience about increasing num leaves, but wasn't bold enough to increase it to 1024 :(\n&gt; \n\nCV score is the only thing I actually trust. This approach always worked for me in the past and this fact along with consistent cv/leaderboard difference made rather confident. Regarding your first question... I don't know and would love to hear if anyone has a good explanation for that.",
          "votes": 2
        },
        {
          "id": 350423,
          "postDate": "2018-06-29T18:23:17.763Z",
          "content": "<p>Thank you for your detailed explanation! Congrats for your 1st place!</p>",
          "rawMarkdown": "Thank you for your detailed explanation! Congrats for your 1st place!",
          "votes": 1
        },
        {
          "id": 350784,
          "postDate": "2018-06-30T13:12:00.137Z",
          "content": "<p>It was the first time I used LGBM, but I couldn't believe it: validation went up until 3000-4000 num leaves and begun to overfit at 5000 (128 GB RAM on a desktop machine).</p>",
          "rawMarkdown": "It was the first time I used LGBM, but I couldn't believe it: validation went up until 3000-4000 num leaves and begun to overfit at 5000 (128 GB RAM on a desktop machine).",
          "votes": 2
        }
      ]
    },
    {
      "id": 349563,
      "postDate": "2018-06-28T09:23:30.837Z",
      "content": "<h1>Preface</h1>\n\n<p>I need to thank my teammates again for this amazing team up experience on such a challenging problem!</p>\n\n<p>I formed a team with Little Boat early, and we soon decided to focus exclusively on tree-based models (xgb/lgb) and NN models, respectively.</p>\n\n<p>I spent 50% of my time training my original xgb/lgb, and 50% of my time consolidating all features shared among the 4 of us. I trained 6 base xgb models, 7 level2 xgb models, 13 base lgb models, and 12 level2 lgb models. The best base lgb got 0.2138 on public LB, and the best level2 lgb got 0.2110 on public LB.</p>\n\n<h1>Feature Engineering</h1>\n\n<p>I will share my part of feature engineering, please note that it's a joint effort of all of my teammates to achieve a 0.213X single lgb. Particularly, Georgiy Danshchin found 4 amazing features that boosted our best base lgb from 0.215X to 0.213X in the last couple of days.</p>\n\n<p><strong>1. text features</strong></p>\n\n<ul>\n<li><p>tfidf on title, description, title+description, titile+description+param_1, etc. I kept these features sparse for all my xgb models; and I used svd and oof ridge features for all my lgb models to keep it diversified.</p></li>\n<li><p>text statistics, such as words length (#characters/#words) at various levels (title, description, etc.), unique words in title additional to description, etc.</p></li>\n</ul>\n\n<p><strong>2. image features</strong></p>\n\n<ul>\n<li><p>image statistics, reference: <a href=\"https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\">https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality</a></p></li>\n<li><p>features from 3 pretrained nn models, reference: <a href=\"https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf\">https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf</a>. I extracted predicted scores and categories for top-5-category exactly from ResNet50, InceptionV3 and Xception. The predicted categories were treated as new categorical variables, in addition to parent_category_name and category_name.</p></li>\n<li><p>features from vgg16 pretrained network, 512 raw features + first 15 PCA features. (extracted by Little Boat)</p></li>\n<li><p>keypoint feature, reference: <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59414\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59414</a></p></li>\n</ul>\n\n<p><strong>3. categorical features</strong></p>\n\n<ul>\n<li><p>count/count unique features at various levels. These features were generated for both train+test, and train+test+train_active+test_active. For example, number of ads in (parent_category_name, category_name, param_1), number of unique user_ids in (region, city).</p></li>\n<li><p>target encoding at various levels. These include mean deal_probability for categories with count &gt;= 5000; mean predicted deal_probability (train+test) for categories with count &gt;= 5000 (Note that threshold was chosen such that cv/lb gap remained roughly the same.); OOF mean deal_probability encoding; OOF mean deal_probability * min(1, log(count)/log(10000)) encoding.</p></li>\n</ul>\n\n<p><strong>4. predicted independent variables features</strong> (price, image_top_1, item_seq_number, day_diff)</p>\n\n<ul>\n<li><p>xgb predicted price, image_top_1 (train on train+test, predict on train+test); lgb predicted price, image_top_1, item_seq_number (oof prediction); rnn predicted day_diff = day_to - day_from (trained on active and period, shared by Little Boat)</p></li>\n<li><p>mean predicted price, image_top_1, item_seq_number, day_diff at various categorical levels.</p></li>\n<li><p>diff features at various levels, for example, (price-xgb_price)/price at (category_name) level, log(image_top_1) - log(lgb_image_top_1) at (resnet50_category1) level, etc.</p></li>\n<li><p>I tried predicting parent_category_name and category_name (multicalss classification) from text features, but including them made cv slightly worse.</p></li>\n</ul>\n\n<p><strong>5. user_id features</strong></p>\n\n<p>We were aware that our team was slightly overfitting cv because of my excessive user_id features. We chose to work on fixed 5 folds for stacking on team merge. It would be a tremendous effort if we retrained our models using a different validation strategy.</p>\n\n<p>Georgiy Danshchin was working on time-based 4 folds before he joined us. User_id features worked fine in his setting because the shared user_id percentage between time-based folds was ~6%, roughly the same as the shared user_id percentage between train and test.</p>\n\n<p>It's another story for the fixed 5 folds, where the shared user_id percentage between folds increased to ~15%.\nThis enlarged our cv/lb gap to ~0.006, but fortunately we could get really low cv and consistent cv/lb gap ~0.006 going forward.</p>\n\n<ul>\n<li><p>journey features, first order by (user_id, item_seq_number, activation_date) for both train+test, and train+test+train_active+test_active, then caculate journey number (and percentage), reversed journey number (and percentage) at various levels, e.g. (user_id, parent_category_name), (user_id, parent_category_name, category_name, activation_date), etc.</p></li>\n<li><p>categorical features, treat user_id as categorical variable, generate features at (user_id) or (user_id, other categorical variables) level. e.g. unique item_seq_number, price range = log1p(max(price)) - log1p(min(price))).</p></li>\n<li><p>periods features, generated from periods_train and periods_test.</p></li>\n</ul>\n\n<p><strong>First calculate</strong></p>\n\n<p>for each row in periods_train+periods_test (periods table), activation_time = date_from - activation_date, activation_len = date_to - date_from, etc.;</p>\n\n<p>for each item_id, lag_activation = activation_date - lag(activation_date), toact_diff = activation_date - lag(date_to), etc.</p>\n\n<p><strong>Second aggregate</strong></p>\n\n<p>periods table to (item_id) level, activation_count = .N, and calculate min, max, mean, sd of the above numerical features.</p>\n\n<p><strong>Third aggregate</strong></p>\n\n<p>item_id table to (user_id) level, again calculate min, max, mean, sd.</p>\n\n<p>(BTW I also aggregated the item_id table from second step to other categorical levels as well.)</p>\n\n<p><strong>For example</strong>, one of the important features is al2max_mean_uid, this means:</p>\n\n<p>activation_len2 = date_to - activation_date</p>\n\n<p>activation_len2_max = max(activation_len2) at (item_id) level</p>\n\n<p>al2max_mean = mean(activation_len2_max) at (user_id) level</p>\n\n<h1>Other comments</h1>\n\n<p>I've used three sets of parameters through my series of lgb models (be it base models or level2 models). The parameters were found by baysian optimization. This added some diversity.</p>\n\n<p>I also only kept features that constitute 99% of gain of the previous models. This helped me remain 700+ features for most of my base lgb models, and 300+ for most of my level2 lgb models. (Roughly 3000 sparse features for xgb models)</p>\n\n<p>I've also trained models predicting modified labels to add diversities:</p>\n\n<ul>\n<li>use mean deal_probability instead at (title, description) level, etc.</li>\n<li>train a multiclass classification model, for example, 5-class classification for class 0.005, class 0.125, class 0.765, class 0.805, class 0.865.</li>\n</ul>",
      "rawMarkdown": "Preface\n=======\n\nI need to thank my teammates again for this amazing team up experience on such a challenging problem!\n\nI formed a team with Little Boat early, and we soon decided to focus exclusively on tree-based models (xgb/lgb) and NN models, respectively.\n\nI spent 50% of my time training my original xgb/lgb, and 50% of my time consolidating all features shared among the 4 of us. I trained 6 base xgb models, 7 level2 xgb models, 13 base lgb models, and 12 level2 lgb models. The best base lgb got 0.2138 on public LB, and the best level2 lgb got 0.2110 on public LB.\n\n\nFeature Engineering\n=========================\n\nI will share my part of feature engineering, please note that it's a joint effort of all of my teammates to achieve a 0.213X single lgb. Particularly, Georgiy Danshchin found 4 amazing features that boosted our best base lgb from 0.215X to 0.213X in the last couple of days.\n\n**1. text features**\n\n- tfidf on title, description, title+description, titile+description+param_1, etc. I kept these features sparse for all my xgb models; and I used svd and oof ridge features for all my lgb models to keep it diversified.\n\n- text statistics, such as words length (#characters/#words) at various levels (title, description, etc.), unique words in title additional to description, etc.\n\n\n**2. image features**\n\n- image statistics, reference: https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\n\n- features from 3 pretrained nn models, reference: https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf. I extracted predicted scores and categories for top-5-category exactly from ResNet50, InceptionV3 and Xception. The predicted categories were treated as new categorical variables, in addition to parent_category_name and category_name.\n\n- features from vgg16 pretrained network, 512 raw features + first 15 PCA features. (extracted by Little Boat)\n\n- keypoint feature, reference: https://www.kaggle.com/c/avito-demand-prediction/discussion/59414\n\n\n**3. categorical features**\n\n- count/count unique features at various levels. These features were generated for both train+test, and train+test+train_active+test_active. For example, number of ads in (parent_category_name, category_name, param_1), number of unique user_ids in (region, city).\n\n- target encoding at various levels. These include mean deal_probability for categories with count &gt;= 5000; mean predicted deal_probability (train+test) for categories with count &gt;= 5000 (Note that threshold was chosen such that cv/lb gap remained roughly the same.); OOF mean deal_probability encoding; OOF mean deal_probability * min(1, log(count)/log(10000)) encoding.\n\n\n**4. predicted independent variables features** (price, image_top_1, item_seq_number, day_diff)\n\n- xgb predicted price, image_top_1 (train on train+test, predict on train+test); lgb predicted price, image_top_1, item_seq_number (oof prediction); rnn predicted day_diff = day_to - day_from (trained on active and period, shared by Little Boat)\n\n- mean predicted price, image_top_1, item_seq_number, day_diff at various categorical levels.\n\n- diff features at various levels, for example, (price-xgb_price)/price at (category_name) level, log(image_top_1) - log(lgb_image_top_1) at (resnet50_category1) level, etc.\n\n- I tried predicting parent_category_name and category_name (multicalss classification) from text features, but including them made cv slightly worse.\n\n\n**5. user_id features**\n\nWe were aware that our team was slightly overfitting cv because of my excessive user_id features. We chose to work on fixed 5 folds for stacking on team merge. It would be a tremendous effort if we retrained our models using a different validation strategy.\n\nGeorgiy Danshchin was working on time-based 4 folds before he joined us. User_id features worked fine in his setting because the shared user_id percentage between time-based folds was ~6%, roughly the same as the shared user_id percentage between train and test.\n\nIt's another story for the fixed 5 folds, where the shared user_id percentage between folds increased to ~15%.\nThis enlarged our cv/lb gap to ~0.006, but fortunately we could get really low cv and consistent cv/lb gap ~0.006 going forward.\n\n- journey features, first order by (user_id, item_seq_number, activation_date) for both train+test, and train+test+train_active+test_active, then caculate journey number (and percentage), reversed journey number (and percentage) at various levels, e.g. (user_id, parent_category_name), (user_id, parent_category_name, category_name, activation_date), etc.\n\n- categorical features, treat user_id as categorical variable, generate features at (user_id) or (user_id, other categorical variables) level. e.g. unique item_seq_number, price range = log1p(max(price)) - log1p(min(price))).\n\n- periods features, generated from periods_train and periods_test.\n\n**First calculate**\n\nfor each row in periods_train+periods_test (periods table), activation_time = date_from - activation_date, activation_len = date_to - date_from, etc.;\n\nfor each item_id, lag_activation = activation_date - lag(activation_date), toact_diff = activation_date - lag(date_to), etc.\n\n**Second aggregate**\n\nperiods table to (item_id) level, activation_count = .N, and calculate min, max, mean, sd of the above numerical features.\n\n**Third aggregate**\n\nitem_id table to (user_id) level, again calculate min, max, mean, sd.\n\n(BTW I also aggregated the item_id table from second step to other categorical levels as well.)\n\n**For example**, one of the important features is al2max_mean_uid, this means:\n\nactivation_len2 = date_to - activation_date\n\nactivation_len2_max = max(activation_len2) at (item_id) level\n\nal2max_mean = mean(activation_len2_max) at (user_id) level\n\n\nOther comments\n====================\n\nI've used three sets of parameters through my series of lgb models (be it base models or level2 models). The parameters were found by baysian optimization. This added some diversity.\n\nI also only kept features that constitute 99% of gain of the previous models. This helped me remain 700+ features for most of my base lgb models, and 300+ for most of my level2 lgb models. (Roughly 3000 sparse features for xgb models)\n\nI've also trained models predicting modified labels to add diversities:\n\n- use mean deal_probability instead at (title, description) level, etc.\n- train a multiclass classification model, for example, 5-class classification for class 0.005, class 0.125, class 0.765, class 0.805, class 0.865.",
      "votes": 38,
      "replies": [
        {
          "id": 349746,
          "postDate": "2018-06-28T15:13:59.037Z",
          "content": "<p>Congratulations and thanks for sharing.</p>",
          "rawMarkdown": "Congratulations and thanks for sharing."
        },
        {
          "id": 349867,
          "postDate": "2018-06-28T19:25:48.113Z",
          "content": "<p>Congratulations on this exceptionally amazing achievement! Your team well deserve this!</p>",
          "rawMarkdown": "Congratulations on this exceptionally amazing achievement! Your team well deserve this!"
        },
        {
          "id": 349915,
          "postDate": "2018-06-28T20:59:04.080Z",
          "content": "<p>Congratulation and thanks for sharing! This one has caught my attention: \"I also only kept features that constitute 99% of gain of the previous models. \". Do I understand this correctly that you looked at feature importance and kept only first 99% of features? Or was it something more complicated?</p>",
          "rawMarkdown": "Congratulation and thanks for sharing! This one has caught my attention: \"I also only kept features that constitute 99% of gain of the previous models. \". Do I understand this correctly that you looked at feature importance and kept only first 99% of features? Or was it something more complicated?"
        },
        {
          "id": 349927,
          "postDate": "2018-06-28T21:42:24.437Z",
          "content": "<p>Yes you got the idea, more precisely, I kept minimum number of features such that sum of gain &gt;= 0.99. Since most of the features can be calculated for different categorical levels, some of them might be meaningless, developing models this way served as an automatic way of feature selection.</p>",
          "rawMarkdown": "Yes you got the idea, more precisely, I kept minimum number of features such that sum of gain &gt;= 0.99. Since most of the features can be calculated for different categorical levels, some of them might be meaningless, developing models this way served as an automatic way of feature selection.",
          "votes": 1
        },
        {
          "id": 350111,
          "postDate": "2018-06-29T07:40:50.913Z",
          "content": "<p>Thank you for sharing !   I have questions about how do you orgnize a few thousand features?   Saved to many files or a huge DB ?  how to load features programmaticly ?  Do you have any suggestions or what's your trick to play with 40 models and thousands features ?</p>",
          "rawMarkdown": "Thank you for sharing !   I have questions about how do you orgnize a few thousand features?   Saved to many files or a huge DB ?  how to load features programmaticly ?  Do you have any suggestions or what's your trick to play with 40 models and thousands features ?"
        },
        {
          "id": 350126,
          "postDate": "2018-06-29T08:13:02.237Z",
          "content": "<p>I only have thousands or tens of thousands of sparse features and I saved them in .RData format, which only takes ~10G space for one model. I have hundreds of dense features and I saved them as csv files, which are roughly ~10G as well. If I run out of space, I would probably delete saved features for most of the models but only keep the milestone ones. Anyways, I consider this dataset only medium size, so this probably won't be a big issue anyways :)</p>",
          "rawMarkdown": "I only have thousands or tens of thousands of sparse features and I saved them in .RData format, which only takes ~10G space for one model. I have hundreds of dense features and I saved them as csv files, which are roughly ~10G as well. If I run out of space, I would probably delete saved features for most of the models but only keep the milestone ones. Anyways, I consider this dataset only medium size, so this probably won't be a big issue anyways :)",
          "votes": 1
        },
        {
          "id": 350134,
          "postDate": "2018-06-29T08:40:55.463Z",
          "content": "<p>Thank you very much. One more question:  I don't understand the No.4 of FeatureEngineering. What's the predicted features like price?  Price is a feature already in train.csv?  how to get your \"predicted price\"?</p>",
          "rawMarkdown": "Thank you very much. One more question:  I don't understand the No.4 of FeatureEngineering. What's the predicted features like price?  Price is a feature already in train.csv?  how to get your \"predicted price\"?"
        },
        {
          "id": 350418,
          "postDate": "2018-06-29T18:18:06.350Z",
          "content": "<p>Basically you can train a model using every features except price related features (or some specific subsets of features) to predict price. The predicted price alone can be a single feature, but then you can aggregate the predicted price to different categorical levels, and compare the true price vs predicted price in various ways to generate more features. Hope it clarifies.</p>",
          "rawMarkdown": "Basically you can train a model using every features except price related features (or some specific subsets of features) to predict price. The predicted price alone can be a single feature, but then you can aggregate the predicted price to different categorical levels, and compare the true price vs predicted price in various ways to generate more features. Hope it clarifies."
        },
        {
          "id": 350452,
          "postDate": "2018-06-29T19:33:11.753Z",
          "content": "<p>Got it, thanks. I'll try this solution on my next competition. Do you have any idea this method will work on which kind of features ? ( price ? or some continues numbers ? or something else ? )</p>",
          "rawMarkdown": "Got it, thanks. I'll try this solution on my next competition. Do you have any idea this method will work on which kind of features ? ( price ? or some continues numbers ? or something else ? )"
        },
        {
          "id": 350627,
          "postDate": "2018-06-30T05:20:12.467Z",
          "content": "<p>In this dataset, we know the meaning of each raw feature. You can always try something that intuitively makes sense. A good example would be Georgiy's predictions on log1p(price), renewed and log1p(total_duration) using active data. Those are extremely strong features. You can tell by CV if it works or not.</p>",
          "rawMarkdown": "In this dataset, we know the meaning of each raw feature. You can always try something that intuitively makes sense. A good example would be Georgiy's predictions on log1p(price), renewed and log1p(total_duration) using active data. Those are extremely strong features. You can tell by CV if it works or not.",
          "votes": 1
        }
      ]
    },
    {
      "id": 352479,
      "postDate": "2018-07-04T13:07:53.253Z",
      "content": "<p>Study</p>",
      "rawMarkdown": "Study",
      "votes": 1
    },
    {
      "id": 349674,
      "postDate": "2018-06-28T13:20:26.397Z",
      "content": "<p>Congrats~~~\nHope I can stay close to you guys. : )\nAttention really helps me a lot, I use several attention models with rnn and cnn structure. included self attention, content attention(text, image and  handcrafted features interaction)\nText translation augmentation and deep oof image features boost a lot to us, however, we just find it in last days.</p>",
      "rawMarkdown": "Congrats~~~\nHope I can stay close to you guys. : )\nAttention really helps me a lot, I use several attention models with rnn and cnn structure. included self attention, content attention(text, image and  handcrafted features interaction)\nText translation augmentation and deep oof image features boost a lot to us, however, we just find it in last days.",
      "votes": 1
    },
    {
      "id": 349603,
      "postDate": "2018-06-28T10:51:34.110Z",
      "content": "<p>Congratulations!  Thanks for sharing your massive and impressive amount of work !</p>",
      "rawMarkdown": "Congratulations!  Thanks for sharing your massive and impressive amount of work !",
      "votes": 1
    },
    {
      "id": 349524,
      "postDate": "2018-06-28T08:22:16.473Z",
      "content": "<p>Great work! </p>\n\n<p>Have you tried to finetune the image NNs instead of use single layers?</p>\n\n<p>Thanks for sharing!</p>",
      "rawMarkdown": "Great work! \n\nHave you tried to finetune the image NNs instead of use single layers?\n\nThanks for sharing!",
      "votes": 1
    },
    {
      "id": 349340,
      "postDate": "2018-06-28T02:40:51.617Z",
      "content": "<p>Congrats on your first place.... Thanks for your solution details.... </p>",
      "rawMarkdown": "Congrats on your first place.... Thanks for your solution details.... ",
      "votes": 1
    },
    {
      "id": 349285,
      "postDate": "2018-06-28T01:36:51.763Z",
      "content": "<p>Brilliant step-by-step solution.</p>\n\n<p>I could barely get a strong nn model. A lot to try in another competition.</p>\n\n<p>Thanks and congrats.</p>",
      "rawMarkdown": "Brilliant step-by-step solution.\n\nI could barely get a strong nn model. A lot to try in another competition.\n\nThanks and congrats.",
      "votes": 1
    },
    {
      "id": 351355,
      "postDate": "2018-07-02T01:45:11.103Z",
      "content": "<p>Congratulations！Your NN solution is awesome and well worth learning. Could you share more NN model code? Many Thanks! </p>",
      "rawMarkdown": "Congratulations！Your NN solution is awesome and well worth learning. Could you share more NN model code? Many Thanks! ",
      "votes": 2
    },
    {
      "id": 3361834,
      "postDate": "2025-12-03T21:28:29.243Z",
      "content": "<p>Great Learning! Have a look at my notebook <a href=\"https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble\" target=\"_blank\">https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble</a></p>",
      "rawMarkdown": "Great Learning! Have a look at my notebook https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble"
    },
    {
      "id": 1000233,
      "postDate": "2020-09-06T12:03:11.833Z",
      "content": "<p><a href=\"https://www.kaggle.com/gdanschin\" target=\"_blank\">@gdanschin</a> Can you please share a notebook / code example of the create of the NN itself? including the part that merges the different types of data (Categorical/Numerical/Image/Text)</p>",
      "rawMarkdown": "@gdanschin Can you please share a notebook / code example of the create of the NN itself? including the part that merges the different types of data (Categorical/Numerical/Image/Text)"
    },
    {
      "id": 383189,
      "postDate": "2018-09-08T01:05:09.457Z",
      "content": "<p>Congratulations! And thanks for sharing such an inspiring solution!</p>",
      "rawMarkdown": "Congratulations! And thanks for sharing such an inspiring solution!"
    },
    {
      "id": 382836,
      "postDate": "2018-09-07T07:40:52.983Z",
      "content": "<p>Congratulations！</p>",
      "rawMarkdown": "Congratulations！"
    },
    {
      "id": 352218,
      "postDate": "2018-07-03T23:17:01.473Z",
      "content": "<p>Great Kernel. Congrats!</p>",
      "rawMarkdown": "Great Kernel. Congrats!"
    },
    {
      "id": 351825,
      "postDate": "2018-07-03T05:36:37.167Z",
      "content": "<p>Nice Ensemble solution!</p>",
      "rawMarkdown": "Nice Ensemble solution!"
    },
    {
      "id": 350984,
      "postDate": "2018-07-01T00:31:08.887Z",
      "content": "<p>nice!</p>",
      "rawMarkdown": "nice!"
    },
    {
      "id": 350186,
      "postDate": "2018-06-29T11:15:31.747Z",
      "content": "<p>Great work!\nCongratulations!</p>",
      "rawMarkdown": "Great work!\nCongratulations!"
    },
    {
      "id": 349841,
      "postDate": "2018-06-28T17:57:19.980Z",
      "content": "<p>Congratulation! Your NN solution is awesome! would u mind share some NN details like Dense Layer size, Dropout rate and Learning Rate? Did u tune them one by one through CV?  Too many parameters to tune in NN. Feel overwhelmed</p>",
      "rawMarkdown": "Congratulation! Your NN solution is awesome! would u mind share some NN details like Dense Layer size, Dropout rate and Learning Rate? Did u tune them one by one through CV?  Too many parameters to tune in NN. Feel overwhelmed"
    },
    {
      "id": 349822,
      "postDate": "2018-06-28T17:27:20.350Z",
      "content": "<p>Huge win. Congrats!!!</p>",
      "rawMarkdown": "Huge win. Congrats!!!"
    },
    {
      "id": 349788,
      "postDate": "2018-06-28T16:40:39.893Z",
      "content": "<p>Wow.</p>",
      "rawMarkdown": "Wow."
    },
    {
      "id": 349785,
      "postDate": "2018-06-28T16:36:57.467Z",
      "content": "<p>Thanks, Thousand voices and Little Boat for amazing achievement and also sharing the details. Learned a lot from your NN architecture. Did you try GRU at all? If yes then can you mind sharing why/how LSTM scored over GRU and what approach you took? Thanks in advance.</p>",
      "rawMarkdown": "Thanks, Thousand voices and Little Boat for amazing achievement and also sharing the details. Learned a lot from your NN architecture. Did you try GRU at all? If yes then can you mind sharing why/how LSTM scored over GRU and what approach you took? Thanks in advance."
    },
    {
      "id": 349712,
      "postDate": "2018-06-28T14:24:39.797Z",
      "content": "<p>Congratulations @LittleBoat, @Arsenal', and Thousandvoices for a well deserved win. Thanks for sharing your solutions. </p>\n\n<p>I can see why I did not go far in as I focused more on FE because I wanted to sharpen my skills there. Felt like this is the right competition for feature engineering. Leaving NN as the last in my list to tackle turned out to be a big mistake as I could not take enough time from my day job to do much in the last 2 weeks. I found progress was slow and incremental in this competition. I certainly learnt a bunch of lessons for the future.</p>\n\n<p>Thanks to all the contributors in kernels as well as discussions. There are many people but special shout out to @Benjamin Minixhofer,  @kxx, @SRK, @SamratP, <a href=\"/peterhurford\">@peterhurford</a>, @Shanth,  @Derek, @Serigne, @Ethan, @Nooh, and many more.</p>\n\n<p>See you in the next one and Happy Kaggling !!!</p>",
      "rawMarkdown": "Congratulations @LittleBoat, @Arsenal', and Thousandvoices for a well deserved win. Thanks for sharing your solutions. \n\nI can see why I did not go far in as I focused more on FE because I wanted to sharpen my skills there. Felt like this is the right competition for feature engineering. Leaving NN as the last in my list to tackle turned out to be a big mistake as I could not take enough time from my day job to do much in the last 2 weeks. I found progress was slow and incremental in this competition. I certainly learnt a bunch of lessons for the future.\n\nThanks to all the contributors in kernels as well as discussions. There are many people but special shout out to @Benjamin Minixhofer,  @kxx, @SRK, @SamratP, @peterhurford, @Shanth,  @Derek, @Serigne, @Ethan, @Nooh, and many more.\n\nSee you in the next one and Happy Kaggling !!!"
    },
    {
      "id": 349637,
      "postDate": "2018-06-28T12:20:56.273Z",
      "content": "<p>Congrats~~ What an impressive solution! And thx for sharing:) </p>",
      "rawMarkdown": "Congrats~~ What an impressive solution! And thx for sharing:) "
    },
    {
      "id": 349570,
      "postDate": "2018-06-28T09:33:17.223Z",
      "content": "<p>Looking forward to see you code;) Thank you for the post</p>",
      "rawMarkdown": "Looking forward to see you code;) Thank you for the post"
    },
    {
      "id": 349538,
      "postDate": "2018-06-28T08:45:00.107Z",
      "content": "<p>Congrats Little Boat and the team! Would you mind to post code for your RNN architecture to process text? Have you done any text cleanup?</p>",
      "rawMarkdown": "Congrats Little Boat and the team! Would you mind to post code for your RNN architecture to process text? Have you done any text cleanup?"
    },
    {
      "id": 349484,
      "postDate": "2018-06-28T07:13:07.567Z",
      "content": "<p>Congrats you &amp; members,\nIt was quite impressive to keep leading position during competition.\nI am really thankful to kind explanation too~!!</p>",
      "rawMarkdown": "Congrats you &amp; members,\nIt was quite impressive to keep leading position during competition.\nI am really thankful to kind explanation too~!!"
    },
    {
      "id": 349427,
      "postDate": "2018-06-28T05:07:29.893Z",
      "content": "<p>Very nice NN approach and congratulations on your consistent leading position over time! How long did it take to train your NN architecture and how much RAM did you need? - Cheers </p>",
      "rawMarkdown": "Very nice NN approach and congratulations on your consistent leading position over time! How long did it take to train your NN architecture and how much RAM did you need? - Cheers ",
      "replies": [
        {
          "id": 349672,
          "postDate": "2018-06-28T13:19:23.923Z",
          "content": "<p>1 epoch took about 3 mins for the final model. And each model is about 8 epochs. I have a machine with 64GB RAM, but you can do it with 32 or 48GB.</p>",
          "rawMarkdown": "1 epoch took about 3 mins for the final model. And each model is about 8 epochs. I have a machine with 64GB RAM, but you can do it with 32 or 48GB.",
          "votes": 3
        },
        {
          "id": 349934,
          "postDate": "2018-06-28T22:02:36.417Z",
          "content": "<p>@LittleBoat, great job on improving the RNN over time. Any chance to see the code on the RNN at any stage would be great. I struggled with RNN runtimes, got 5 mins per epoch - best model was 26 epochs bagged twice with 5mins per epoch - but over 5 folds + test - this adds up to over 24 hours. \nDid you bag the models; or just create a lot of different ones and add the individuals to the stack; did you bag epochs, or just choose the last one ? Would love to learn from you how to improve in RNNs.</p>",
          "rawMarkdown": "@LittleBoat, great job on improving the RNN over time. Any chance to see the code on the RNN at any stage would be great. I struggled with RNN runtimes, got 5 mins per epoch - best model was 26 epochs bagged twice with 5mins per epoch - but over 5 folds + test - this adds up to over 24 hours. \nDid you bag the models; or just create a lot of different ones and add the individuals to the stack; did you bag epochs, or just choose the last one ? Would love to learn from you how to improve in RNNs."
        },
        {
          "id": 349950,
          "postDate": "2018-06-28T23:35:21.100Z",
          "content": "<p>The RNN code is surprisingly easy here</p>\n\n<pre><code>seq_title_description = Input(shape=[max_seq_description_length], name=\"seq_description\")\nemb_seq_title_description = Embedding(vocab_size, EMBEDDING_DIM1, weights=[embedding_matrix1], trainable=False)(\n    seq_title_description)\nemb_seq_title_description = SpatialDropout1D(0.25)(emb_seq_title_description)\nrnn_layer1 = CuDNNLSTM(128, return_sequences=True)(emb_seq_title_description)\nrnn_layer1 = CuDNNLSTM(128, return_sequences=False)(rnn_layer1)\nfc_rnn = Dense(256, init='he_normal')(rnn_layer1)\nfc_rnn = PReLU()(fc_rnn)\nfc_rnn = BatchNormalization()(fc_rnn)\nfc_rnn = Dropout(0.25)(fc_rnn)\n</code></pre>\n\n<p>If you are using Keras you should use CuDNN implementation which would be much faster. It is surprising that you needed 26 epochs to converge. For me it converged after 6 or so and I ran it for 8 because 8 seemed to give slightly better average with different seeds. For each fold, I ran 8 epochs (some fold actually only needed 5, but I just let it \"overfit\" a bit), with 5 fold, so that is 40 epochs per seed. And I ran 4 seeds and then average them. So in total, it is 40 * 4 = 160 epochs. Before adding all the engineered features I remember each epoch took about 100 seconds so in total it is only 16000 seconds so less than 5 hours.</p>\n\n<p>Hope it helps.</p>",
          "rawMarkdown": "The RNN code is surprisingly easy here\n\n    seq_title_description = Input(shape=[max_seq_description_length], name=\"seq_description\")\n    emb_seq_title_description = Embedding(vocab_size, EMBEDDING_DIM1, weights=[embedding_matrix1], trainable=False)(\n        seq_title_description)\n    emb_seq_title_description = SpatialDropout1D(0.25)(emb_seq_title_description)\n    rnn_layer1 = CuDNNLSTM(128, return_sequences=True)(emb_seq_title_description)\n    rnn_layer1 = CuDNNLSTM(128, return_sequences=False)(rnn_layer1)\n    fc_rnn = Dense(256, init='he_normal')(rnn_layer1)\n    fc_rnn = PReLU()(fc_rnn)\n    fc_rnn = BatchNormalization()(fc_rnn)\n    fc_rnn = Dropout(0.25)(fc_rnn)\n\nIf you are using Keras you should use CuDNN implementation which would be much faster. It is surprising that you needed 26 epochs to converge. For me it converged after 6 or so and I ran it for 8 because 8 seemed to give slightly better average with different seeds. For each fold, I ran 8 epochs (some fold actually only needed 5, but I just let it \"overfit\" a bit), with 5 fold, so that is 40 epochs per seed. And I ran 4 seeds and then average them. So in total, it is 40 * 4 = 160 epochs. Before adding all the engineered features I remember each epoch took about 100 seconds so in total it is only 16000 seconds so less than 5 hours.\n\nHope it helps.",
          "votes": 5
        },
        {
          "id": 350080,
          "postDate": "2018-06-29T06:19:54.240Z",
          "content": "<p>Super, thanks, yes it helps. I think ours converged sooner but we ran for more epochs and predicted and bagged each epoch. \nOk, I see now you read the whole dscr and title as one string... Great thanks for sharing.</p>",
          "rawMarkdown": "Super, thanks, yes it helps. I think ours converged sooner but we ran for more epochs and predicted and bagged each epoch. \nOk, I see now you read the whole dscr and title as one string... Great thanks for sharing."
        },
        {
          "id": 350329,
          "postDate": "2018-06-29T15:06:31.217Z",
          "content": "<p>@Darragh,\nActually, title and description were read separately :( it was just dummy name really.... Reading them separately gave about 0.0002-0.0003 improvement.</p>",
          "rawMarkdown": "@Darragh,\nActually, title and description were read separately :( it was just dummy name really.... Reading them separately gave about 0.0002-0.0003 improvement."
        },
        {
          "id": 350341,
          "postDate": "2018-06-29T15:22:37.027Z",
          "content": "<p>Thanks, one more question if its ok; we struggled with Batchnorm after the PRelu, it just destroyed the weights. On retrospect, I suspect its because we concatenated the output of the text RNN and the descr RNN; then passed to a dense and applied batchnorm; whereas maybe we should have applied batchnorm indvidually to each output. Did you face this problem with batchnorm ?</p>",
          "rawMarkdown": "Thanks, one more question if its ok; we struggled with Batchnorm after the PRelu, it just destroyed the weights. On retrospect, I suspect its because we concatenated the output of the text RNN and the descr RNN; then passed to a dense and applied batchnorm; whereas maybe we should have applied batchnorm indvidually to each output. Did you face this problem with batchnorm ?"
        },
        {
          "id": 350345,
          "postDate": "2018-06-29T15:27:41.440Z",
          "content": "<p>That is interesting. I didn't investigate if it was batchnorm or other causes but when I started adding image features it made my model worse and only after using separate layers the model got better. You might want to try that out on your features to see if that would help. My intuition of separate dense layer for each input is that they act like a block/gate to force them to be transformed before concatenating. So you can also in theory tweak the unit size to assign different \"weights\" to them.</p>",
          "rawMarkdown": "That is interesting. I didn't investigate if it was batchnorm or other causes but when I started adding image features it made my model worse and only after using separate layers the model got better. You might want to try that out on your features to see if that would help. My intuition of separate dense layer for each input is that they act like a block/gate to force them to be transformed before concatenating. So you can also in theory tweak the unit size to assign different \"weights\" to them.",
          "votes": 5
        }
      ]
    },
    {
      "id": 349408,
      "postDate": "2018-06-28T04:40:22.170Z",
      "content": "<p>Congrats! And thanks for sharing best solution!</p>",
      "rawMarkdown": "Congrats! And thanks for sharing best solution!"
    },
    {
      "id": 349374,
      "postDate": "2018-06-28T03:31:28.230Z",
      "content": "<p>Congratulations! It's impressive. </p>\n\n<p>Thank you for your sharing.</p>",
      "rawMarkdown": "Congratulations! It's impressive. \n\nThank you for your sharing."
    },
    {
      "id": 349371,
      "postDate": "2018-06-28T03:27:09.227Z",
      "content": "<p>your nn arch is great! thanks for your graph!</p>",
      "rawMarkdown": "your nn arch is great! thanks for your graph!"
    },
    {
      "id": 349353,
      "postDate": "2018-06-28T02:57:28.143Z",
      "content": "<p>Congratulations! And thank you for sharing how you built your NN.</p>",
      "rawMarkdown": "Congratulations! And thank you for sharing how you built your NN."
    },
    {
      "id": 349330,
      "postDate": "2018-06-28T02:33:43.727Z",
      "content": "<p>Congratulations! And thanks for the detailed step-by-step explanation. </p>\n\n<p>May I ask what kind of computing resources did you use for the NN? </p>",
      "rawMarkdown": "Congratulations! And thanks for the detailed step-by-step explanation. \n\nMay I ask what kind of computing resources did you use for the NN? ",
      "replies": [
        {
          "id": 349360,
          "postDate": "2018-06-28T03:06:39.197Z",
          "content": "<p>one GTX 1080ti</p>",
          "rawMarkdown": "one GTX 1080ti",
          "votes": 3
        }
      ]
    },
    {
      "id": 349327,
      "postDate": "2018-06-28T02:31:04.260Z",
      "content": "<p>Amazing scores and amazing feature engineer,  thank you for you detailed NN , it is very helpful. And can't wait to see feature engineer part.</p>",
      "rawMarkdown": "Amazing scores and amazing feature engineer,  thank you for you detailed NN , it is very helpful. And can't wait to see feature engineer part."
    },
    {
      "id": 349295,
      "postDate": "2018-06-28T01:51:31.650Z",
      "content": "<p>Congrats. I can't wait to see your teammates talking the feature engineering things.\nI have two questions please.\n1. Have you tried other CNN pre-trained models other than VGG and Resnet50?\n2. Have you tried to train your own word2vec model on the competition text data in train and train_active other than fasttext pre-trained?\nThank you for sharing such a brilliant solution. It's super detailed.</p>",
      "rawMarkdown": "Congrats. I can't wait to see your teammates talking the feature engineering things.\nI have two questions please.\n1. Have you tried other CNN pre-trained models other than VGG and Resnet50?\n2. Have you tried to train your own word2vec model on the competition text data in train and train_active other than fasttext pre-trained?\nThank you for sharing such a brilliant solution. It's super detailed.",
      "replies": [
        {
          "id": 349299,
          "postDate": "2018-06-28T01:57:12.050Z",
          "content": "<ol>\n<li>I had. But with VGG16 and ResNet50 included, I couldn't get any improvements with others added.</li>\n<li>I didn't try anything other than fastText. Maybe I should have...</li>\n</ol>",
          "rawMarkdown": "1. I had. But with VGG16 and ResNet50 included, I couldn't get any improvements with others added.\n2. I didn't try anything other than fastText. Maybe I should have...",
          "votes": 2
        },
        {
          "id": 349598,
          "postDate": "2018-06-28T10:40:21.060Z",
          "content": "<p>Your fasttext vectors were self trained right?</p>",
          "rawMarkdown": "Your fasttext vectors were self trained right?"
        },
        {
          "id": 349670,
          "postDate": "2018-06-28T13:15:53.200Z",
          "content": "<p>yes.</p>",
          "rawMarkdown": "yes."
        }
      ]
    },
    {
      "id": 349292,
      "postDate": "2018-06-28T01:49:47.970Z",
      "content": "<p>Exhaustive kernel. Excellent work </p>",
      "rawMarkdown": "Exhaustive kernel. Excellent work "
    },
    {
      "id": 349291,
      "postDate": "2018-06-28T01:47:54.537Z",
      "content": "<p>Congrats  Little Boat and you team for the amazing winning and thanks for sharing.</p>",
      "rawMarkdown": "Congrats  Little Boat and you team for the amazing winning and thanks for sharing."
    },
    {
      "id": 349286,
      "postDate": "2018-06-28T01:39:52.180Z",
      "content": "<p>WOW! Mind blowing! Congrats God Boat!</p>",
      "rawMarkdown": "WOW! Mind blowing! Congrats God Boat!"
    },
    {
      "id": 349273,
      "postDate": "2018-06-28T01:20:38.077Z",
      "content": "<p>Amazing！！！ Congrats for your another first place solution.  Thanks for sharing your wisdom.</p>\n\n<p>Plus, Can I ask did you accelerate the process of loading image? For me it takes extremely long for loading images for one epoch. (1 million images here)</p>",
      "rawMarkdown": "Amazing！！！ Congrats for your another first place solution.  Thanks for sharing your wisdom.\n\nPlus, Can I ask did you accelerate the process of loading image? For me it takes extremely long for loading images for one epoch. (1 million images here)",
      "replies": [
        {
          "id": 349298,
          "postDate": "2018-06-28T01:55:29.173Z",
          "content": "<p>For the VGG16 and ResNet50 features I just did the predictions with the pretrained model and saved them for late use. </p>\n\n<p>When I was fine tuning them, I had to reduce the batch size and it was still very slow for 224 * 224... Didn't really know a way to accelerate that... So I didn't do much search there which means maybe there is a big potential there that I didn't have the computing resources to explore.</p>",
          "rawMarkdown": "For the VGG16 and ResNet50 features I just did the predictions with the pretrained model and saved them for late use. \n\nWhen I was fine tuning them, I had to reduce the batch size and it was still very slow for 224 * 224... Didn't really know a way to accelerate that... So I didn't do much search there which means maybe there is a big potential there that I didn't have the computing resources to explore.",
          "votes": 1
        },
        {
          "id": 349977,
          "postDate": "2018-06-29T01:21:45.317Z",
          "content": "<p>I don't have enough resource either, but I did try resizing all images to 32 by 32 and it did help.\nDidn't think of using mid layer of resnet.. \nAnother thing I noticed is, on sketch you trained 2 two-layer LSTM for title and description seperately. Did you use Fasttext embedding size of 300 for both ?  but when I saw your code posted above, it shows emb_seq_title_description, it's like you already concatenate title and description? Thanks!</p>",
          "rawMarkdown": "I don't have enough resource either, but I did try resizing all images to 32 by 32 and it did help.\nDidn't think of using mid layer of resnet.. \nAnother thing I noticed is, on sketch you trained 2 two-layer LSTM for title and description seperately. Did you use Fasttext embedding size of 300 for both ?  but when I saw your code posted above, it shows emb_seq_title_description, it's like you already concatenate title and description? Thanks!"
        },
        {
          "id": 349991,
          "postDate": "2018-06-29T01:45:11.590Z",
          "content": "<p>oh yeah that was due to a silly name that I forget to change. I trained LSTM on title and description separately. Separating them probably has about 0.0002 improvement.</p>",
          "rawMarkdown": "oh yeah that was due to a silly name that I forget to change. I trained LSTM on title and description separately. Separating them probably has about 0.0002 improvement.",
          "votes": 1
        }
      ]
    },
    {
      "id": 349683,
      "postDate": "2018-06-28T13:29:38.763Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 349561,
      "postDate": "2018-06-28T09:22:14.133Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 349426,
      "postDate": "2018-06-28T05:06:12.707Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 349290,
      "postDate": "2018-06-28T01:47:29.193Z",
      "content": "<p>Thank you for the sharing!</p>",
      "rawMarkdown": "Thank you for the sharing!"
    },
    {
      "id": 706978,
      "postDate": "2019-12-31T05:30:50.403Z",
      "content": "<p>Nice Post! Thanks for sharing!</p>",
      "rawMarkdown": "Nice Post! Thanks for sharing!"
    },
    {
      "id": 450523,
      "postDate": "2019-01-05T06:59:39.940Z",
      "content": "<p>Thanks or sharing your work.</p>",
      "rawMarkdown": "Thanks or sharing your work."
    },
    {
      "id": 352560,
      "postDate": "2018-07-04T16:18:00.400Z",
      "content": "<p>Thank you for sharing your insights!</p>",
      "rawMarkdown": "Thank you for sharing your insights!"
    },
    {
      "id": 352377,
      "postDate": "2018-07-04T07:50:59.180Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 351974,
      "postDate": "2018-07-03T12:38:45.943Z",
      "content": "<p>Really helpful! Thank you so much.</p>",
      "rawMarkdown": "Really helpful! Thank you so much."
    },
    {
      "id": 351455,
      "postDate": "2018-07-02T08:20:30.100Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing"
    },
    {
      "id": 349564,
      "postDate": "2018-06-28T09:25:03.457Z",
      "content": "<p>Congrats and thanks for sharing!!!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!!!"
    },
    {
      "id": 349406,
      "postDate": "2018-06-28T04:39:47.673Z",
      "content": "<p>Well done and thanks for sharing.</p>",
      "rawMarkdown": "Well done and thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 349710,
      "author_name": "Georgiy Danshchin",
      "author_url": "",
      "post_date": "2018-06-28T14:18:45.603000",
      "content": "<h2>The last part of Dance with Ensemble solution</h2>\n\n<p>First of all, I want to thank my great teammates for this awesome experience and Avito for organizing very interesting competition!</p>\n\n<p>From my point of view main things for this competition were strong feature engineering and stacking of diverse models. So let’s dive into the details.</p>\n\n<h2>Validation</h2>\n\n<p>Before joining the team I've used a time rolling window validation split with 4 folds (8 days in train 3 days in test for each fold). It was very consistent with public lb score, gap between cv and lb was between 0.002 and 0.003. After joining the team I've regenerated all my features and models for common kfold split, and noticed that cv gap increased by 0.003. I figured out that this was mainly due to user_id based target encoding features, so I put them aside and used them only for second level models later.</p>\n\n<h2>Features</h2>\n\n<p><strong>Categories</strong></p>\n\n<ul>\n<li>Percent encoding of categories computed on train+test parts of data</li>\n<li>Percent encoding of user_id category computed on train_active+test_active.</li>\n<li>Mean target encoding of different categories, which were computed separately for each fold in the oof manner using inner folds.</li>\n<li>Mean log1p(price) by different categories and differences between these features and actual log1p(price) as well. For most categories these features were computed using train+test, for user_id and title I used train_active+test_active.</li>\n<li>Mean 'renewed' value by different categories (train_active + test_active)</li>\n<li>Mean 'total_duration' value by different categories (train_active + test_active)</li>\n<li>Sparse one hot encoded categories for Lightgbm models</li>\n</ul>\n\n<p>I’ve used only single categories for all these features, didn’t see any consistent improvement with category interactions. In addition to category_name/region/param_* I treated stemmed title as category too. The 'renewed' value is just a binary value which indicates whether an item has more than one period in periods_train/periods_test data, 'total_duration' is the sum of lengths of all periods for particular item.\nAll 'mean' features were postprocessed with additive smoothing. The coefficients were different for different groups of features and were tuned manually with help of cv feedback.</p>\n\n<p><strong>Text</strong></p>\n\n<ul>\n<li>Sparse tfidf 1-2grams vectors for title and joined title+description+params (used only in Lightgbm models)</li>\n<li>14 different out of fold ridge predictions over different sparse tfidf vectors. Different combinations of fields, and their concatenations, stemmed/raw tokens etc. Computed separately for each fold using inner folds as mean target features.</li>\n<li>Various text statistics: caps quotient, alfanum quotient, char counts, word counts etc.</li>\n<li>Fasttext embeddings trained on train_active + test_active were used as initial embedding matrix for rnn models.</li>\n</ul>\n\n<p><strong>Numeric</strong></p>\n\n<ul>\n<li>Just log1p(price) and log1p(item_seq_number)</li>\n</ul>\n\n<p><strong>Images</strong></p>\n\n<ul>\n<li>Channel mean values</li>\n<li>Image width and height</li>\n<li>Pretrained resnet152 last hidden layer output (only for nn models)</li>\n<li>100 components of svd over previous feature</li>\n<li>Out of fold ridge predictions trained on pretrained resnet101 and resnet152 last hidden layers.</li>\n<li>Finetuned resnet34 out of fold predictions gave a significant boost to my models but took a loooong time to train :)</li>\n</ul>\n\n<p><strong>Active data</strong></p>\n\n<p>I’ve designed 3 nn models trained on active data to predict log1p(price), renewed and log1p(total_duration). These models had 2 rnn branches for title and description and also used category embeddings. The difference between actual log1p(price) and predicted one was an extremely important feature.</p>\n\n<h2>Models</h2>\n\n<p>I had 25 base models (15 lgbs and 10 nns) trained on different feature subsets and with different parameters. My best lgb reaches 0.2155/0.2196 public/private score, best nn is 0.2160/0.2200. Stacking of these models + finetuned resnet34 oof + user_id mean target with one lgb gives 0.2132 on public and 0.2174 on private.</p>\n\n<p>Lgb models could be divided into two groups. One of them uses small set of dense features, it’s trained for 700 rounds with num_leaves=31. Second group uses all features from the first group with addition of sparse tfidf/categories and svd decomposition of pretrained resnet152. It’s trained for 2300 rounds with num_leaves=384, it takes about 1 hour to train on one fold.</p>\n\n<p>NN models have separate branches for title and description: 2-level bidirectional lstm/gru with attention. They also contain separate branches for category embeddings/numeric features and pretrained resnet152 vectors. Every branch is followed by a separate dense layer and finally everything is concatenated and followed by another common dense layer. Text embedding layer is shared for title/description, it’s initialized with fasttext trained on train_active + test_active and frozen for the first epoch, but finetuned later.</p>",
      "votes": 53,
      "replies": [
        {
          "id": 349829,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-28T17:34:11.503000",
          "content": "<p>Congratulations and thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349832,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T17:39:58.140000",
          "content": "<p>Excellent, very nice approach on <strong>Active data</strong> I guess this was the magic features, is that correct ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349918,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-06-28T21:15:55.453000",
          "content": "<p>Hi Georgiy ,</p>\n\n<p>Thanks a lot for the explanation...</p>\n\n<p>It seems that you used mean target encoding for user_id for lgb. There is probably 2/3 of test data where the user_id is not in train data, so I was scared that this would overfit ? Can you explain why you still choose to include it ?</p>\n\n<p>Also what was the reason behind percent encoding of categories ?</p>\n\n<p>Thanks a lot,\nAntoine.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349964,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-29T00:31:36.227000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349966,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-29T00:33:46.197000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349970,
          "author_name": "Georgiy Danshchin",
          "author_url": "",
          "post_date": "2018-06-29T01:00:12.607000",
          "content": "<p>&gt; It seems that you used mean target encoding for user_id for lgb. There is probably 2/3 of test data where the user_id is not in train data, so I was scared that this would overfit ? Can you explain why you still choose to include it ?</p>\n\n<p>Actually it was not just user_id, but common_user_id, which means that the mean target was computed only for those users which were shared between train and test parts of each fold, so I had different sets of common_users for different folds and another set of common_users for the train+test data used for submissions. This approach worked really well with time based split versus kfold split, maybe just the rate of shared users matters (0.06 with time based cv versus 0.15 with plain kfold, and 0.067 for train/test), or maybe there is some complicated time dependency for this kind of information, I don't know for sure.</p>\n\n<p>&gt; Also what was the reason behind percent encoding of categories ?</p>\n\n<p>It's just a first type of categorical encoding which comes to my mind :) I suppose that the frequency of each category value across the training data could give a relevant information (at lest more than raw category value :)). After the percent encoding I usually try mean targets (which are more vulnerable to overfitting) and other approaches.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 350136,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-06-29T08:48:50.033000",
          "content": "<p>Thanks for the sharing. I'm not clear about the category embedding, could you explain more details how to do it ? with NN  or some sklearn lib ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350314,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-06-29T14:39:46.453000",
          "content": "<p>I See.. thanks a lot for clarifying this... I was thinking about doing a similar approach but I wasn't sure if it was the proper way to do it... so instead I choose to generate as many feature as possible around user_id. Thanks a lot for your detailed explanation, very helpful !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 353007,
          "author_name": "Sérgio Passos",
          "author_url": "",
          "post_date": "2018-07-05T16:31:00.470000",
          "content": "<p>Thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 353300,
          "author_name": "Georgiy Danshchin",
          "author_url": "",
          "post_date": "2018-07-06T11:34:16.397000",
          "content": "<blockquote>\n  <p><strong>Darragh wrote</strong></p>\n  \n  <blockquote>\n    <p>Excellent, very nice approach on <strong>Active data</strong> I guess this was the magic features, is that correct ? </p>\n  </blockquote>\n</blockquote>\n\n<p>Yes, these are last strong feature I've generated.</p>\n\n<blockquote>\n  <p><strong>yyqing wrote</strong></p>\n  \n  <blockquote>\n    <p>Thanks for the sharing. I'm not clear about the category embedding, could you explain more details how to do it ? with NN  or some sklearn lib ?</p>\n  </blockquote>\n</blockquote>\n\n<p>It's just an embedding layer which acts like a lookup dictionary between specific category value and an n-dimensional vector. You can check a pytorch implementation: <a href=\"https://pytorch.org/docs/stable/nn.html#embedding\">https://pytorch.org/docs/stable/nn.html#embedding</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349386,
      "author_name": "thousandvoices",
      "author_url": "",
      "post_date": "2018-06-28T03:57:29.850000",
      "content": "<p>First of all, I would like to thank my teammates for their excellent performance. Your discussions and feature engineering insights were incredibly helpful. I learnt a lot from you.</p>\n\n<p>Here is a short summary of my approach. I trained 10 first level models (5 lightgbms, 3 neural nets, 1 catboost and 1 ridge). Let me share some ideas behind my best single lightgbm:</p>\n\n<h3>Categorical features interactions</h3>\n\n<p>Concatenating categorical features is a well-known method for taking this into account. Rather surprisingly, I wasn't able to come up with an automated pipeline that outperformed a list of pairs of features to merge that I created based simply on common sense. Here it is:</p>\n\n<ul>\n<li>region, parent category name</li>\n<li>region, category name</li>\n<li>region, image top-1</li>\n<li>region, param 1</li>\n<li>region, param 2</li>\n<li>region, param 3</li>\n<li>city, parent category name</li>\n<li>city, category name</li>\n<li>city, param 1</li>\n<li>city, param 2</li>\n<li>city, param 3</li>\n<li>category name, image top-1</li>\n<li>parent category name, image top-1</li>\n<li>param 1, param 2</li>\n<li>param 1, image top-1</li>\n</ul>\n\n<h3>Categorical features encoding</h3>\n\n<p>For each categorical variable I added the following features to the model based on the value it took for a particular advertisement:</p>\n\n<ul>\n<li>Mean target of previous ads</li>\n<li>Total occurence count</li>\n<li>Median price</li>\n</ul>\n\n<h3>Text features</h3>\n\n<p>Raw word counts extracted by CountVectorizer for title, description and concatenated titles of all user's ads worked best for me. My efforts to design dense representations by applying svd or training embeddings with supplementary data only decreased the score. My model also included several engineered text features like word and char length, amount of caps, digits, punctuation, cyrillic and latin letters, and unique words.</p>\n\n<h3>Image features</h3>\n\n<p>I didn't focus on this part of the challenge much, that's why the only image features I used were image width and height.</p>\n\n<h3>Aggregated user features</h3>\n\n<p>Minimum, median and maximum price, total amount of ads and number of ads with unique titles.</p>\n\n<h3>Using supplementary data</h3>\n\n<p>I trained 2 neural networks that predicted price and length of active period for the ad on supplementary data provided in train_active and test_active. Both of this features turned out to be very strong, making it into 5 most important features of my model. I also used features produced by <a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">this kernel</a> without any changes.</p>\n\n<h3>Hyperparameters</h3>\n\n<p>This competition favored VERY deep trees, and that was the point that most public scripts have missed. Every time I increased num_leaves my score improved significantly. I ended up using trees with 1024 leaves because it was the maximum that fitted into my memory.</p>\n\n<h3>Parting words</h3>\n\n<p>Overall, this was an exciting competition with lots of raw data and feature engineering opportunities (this is probably the reason why team merges worked so well here :) ), and I'd love to thank Avito for providing such a challenging dataset.</p>",
      "votes": 43,
      "replies": [
        {
          "id": 349390,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2018-06-28T04:10:18.933000",
          "content": "<p>Thanks ! Great work !\nWhat is mean target of previous ads ?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 349424,
          "author_name": "A_b",
          "author_url": "",
          "post_date": "2018-06-28T05:05:19.497000",
          "content": "<p>Thanks for sharing and congrats on your great result!</p>\n\n<blockquote>\n  <p>Every time I increased num_leaves my score improved significantly.</p>\n</blockquote>\n\n<p>I confirm :) how much RAM did you use in order to fit num_leaves=1024? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349731,
          "author_name": "pocket",
          "author_url": "",
          "post_date": "2018-06-28T14:47:32.180000",
          "content": "<p>Hi thousandvoices, congratulations on your 1st place!<br>\nI have two questions about num leaves. <br>\n1) Is there any reason you can think of, that this competition favored larger num leaves? <br>\n2) In general, increasing num leaves tends to overfit your model (at least that is what the LGBM tuning guide says). Why were you confident that your model wasn't overfitting? Did you trust your CVscore?<br>\n<br>\nI also had the same experience about increasing num leaves, but wasn't bold enough to increase it to 1024 :( <br></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349745,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-28T15:12:28.920000",
          "content": "<p>Congratulations and thanks for sharing <a href=\"/thousandvoices\">@thousandvoices</a>. </p>\n\n<p>I discovered that increasing LGBM's number of leaves improved my results even better than some of my features in the last 3 days of the contest. I was concerned about over-fitting so did not go beyond 370 as I have never experienced that before in practice.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349831,
          "author_name": "thousandvoices",
          "author_url": "",
          "post_date": "2018-06-28T17:36:59.927000",
          "content": "<blockquote>\n  <p><strong>Arnaud Roussel wrote</strong></p>\n  \n  <p>Thanks ! Great work !\n  What is mean target of previous ads ?</p>\n</blockquote>\n\n<p>Well, it is literally deal probability averaged among all ads that precede the ad being encoded and have the same value of the categorical feature.</p>\n\n<blockquote>\n  <p><strong>A^b wrote</strong></p>\n  \n  <p>Thanks for sharing and congrats on your great result!</p>\n  \n  <blockquote>\n    <p>Every time I increased num_leaves my score improved significantly.</p>\n  </blockquote>\n  \n  <p>I confirm :) how much RAM did you use in order to fit num_leaves=1024? </p>\n</blockquote>\n\n<p>Thank you for your kind words! I used a desktop with 16 gigabytes of RAM.</p>\n\n<blockquote>\n  <p><strong>pocket wrote</strong></p>\n  \n  <p>Hi thousandvoices, congratulations on your 1st place!\n  I have two questions about num leaves.\n  1) Is there any reason you can think of, that this competition favored larger num leaves?\n  2) In general, increasing num leaves tends to overfit your model (at least that is what the LGBM tuning guide says). Why were you confident that your model wasn't overfitting? Did you trust your CVscore?</p>\n  \n  <p>I also had the same experience about increasing num leaves, but wasn't bold enough to increase it to 1024 :(</p>\n</blockquote>\n\n<p>CV score is the only thing I actually trust. This approach always worked for me in the past and this fact along with consistent cv/leaderboard difference made rather confident. Regarding your first question... I don't know and would love to hear if anyone has a good explanation for that.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 350423,
          "author_name": "Eric Park",
          "author_url": "",
          "post_date": "2018-06-29T18:23:17.763000",
          "content": "<p>Thank you for your detailed explanation! Congrats for your 1st place!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 350784,
          "author_name": "Andrea Rapuzzi",
          "author_url": "",
          "post_date": "2018-06-30T13:12:00.137000",
          "content": "<p>It was the first time I used LGBM, but I couldn't believe it: validation went up until 3000-4000 num leaves and begun to overfit at 5000 (128 GB RAM on a desktop machine).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 349563,
      "author_name": "Arsenal",
      "author_url": "",
      "post_date": "2018-06-28T09:23:30.837000",
      "content": "<h1>Preface</h1>\n\n<p>I need to thank my teammates again for this amazing team up experience on such a challenging problem!</p>\n\n<p>I formed a team with Little Boat early, and we soon decided to focus exclusively on tree-based models (xgb/lgb) and NN models, respectively.</p>\n\n<p>I spent 50% of my time training my original xgb/lgb, and 50% of my time consolidating all features shared among the 4 of us. I trained 6 base xgb models, 7 level2 xgb models, 13 base lgb models, and 12 level2 lgb models. The best base lgb got 0.2138 on public LB, and the best level2 lgb got 0.2110 on public LB.</p>\n\n<h1>Feature Engineering</h1>\n\n<p>I will share my part of feature engineering, please note that it's a joint effort of all of my teammates to achieve a 0.213X single lgb. Particularly, Georgiy Danshchin found 4 amazing features that boosted our best base lgb from 0.215X to 0.213X in the last couple of days.</p>\n\n<p><strong>1. text features</strong></p>\n\n<ul>\n<li><p>tfidf on title, description, title+description, titile+description+param_1, etc. I kept these features sparse for all my xgb models; and I used svd and oof ridge features for all my lgb models to keep it diversified.</p></li>\n<li><p>text statistics, such as words length (#characters/#words) at various levels (title, description, etc.), unique words in title additional to description, etc.</p></li>\n</ul>\n\n<p><strong>2. image features</strong></p>\n\n<ul>\n<li><p>image statistics, reference: <a href=\"https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\">https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality</a></p></li>\n<li><p>features from 3 pretrained nn models, reference: <a href=\"https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf\">https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf</a>. I extracted predicted scores and categories for top-5-category exactly from ResNet50, InceptionV3 and Xception. The predicted categories were treated as new categorical variables, in addition to parent_category_name and category_name.</p></li>\n<li><p>features from vgg16 pretrained network, 512 raw features + first 15 PCA features. (extracted by Little Boat)</p></li>\n<li><p>keypoint feature, reference: <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59414\">https://www.kaggle.com/c/avito-demand-prediction/discussion/59414</a></p></li>\n</ul>\n\n<p><strong>3. categorical features</strong></p>\n\n<ul>\n<li><p>count/count unique features at various levels. These features were generated for both train+test, and train+test+train_active+test_active. For example, number of ads in (parent_category_name, category_name, param_1), number of unique user_ids in (region, city).</p></li>\n<li><p>target encoding at various levels. These include mean deal_probability for categories with count &gt;= 5000; mean predicted deal_probability (train+test) for categories with count &gt;= 5000 (Note that threshold was chosen such that cv/lb gap remained roughly the same.); OOF mean deal_probability encoding; OOF mean deal_probability * min(1, log(count)/log(10000)) encoding.</p></li>\n</ul>\n\n<p><strong>4. predicted independent variables features</strong> (price, image_top_1, item_seq_number, day_diff)</p>\n\n<ul>\n<li><p>xgb predicted price, image_top_1 (train on train+test, predict on train+test); lgb predicted price, image_top_1, item_seq_number (oof prediction); rnn predicted day_diff = day_to - day_from (trained on active and period, shared by Little Boat)</p></li>\n<li><p>mean predicted price, image_top_1, item_seq_number, day_diff at various categorical levels.</p></li>\n<li><p>diff features at various levels, for example, (price-xgb_price)/price at (category_name) level, log(image_top_1) - log(lgb_image_top_1) at (resnet50_category1) level, etc.</p></li>\n<li><p>I tried predicting parent_category_name and category_name (multicalss classification) from text features, but including them made cv slightly worse.</p></li>\n</ul>\n\n<p><strong>5. user_id features</strong></p>\n\n<p>We were aware that our team was slightly overfitting cv because of my excessive user_id features. We chose to work on fixed 5 folds for stacking on team merge. It would be a tremendous effort if we retrained our models using a different validation strategy.</p>\n\n<p>Georgiy Danshchin was working on time-based 4 folds before he joined us. User_id features worked fine in his setting because the shared user_id percentage between time-based folds was ~6%, roughly the same as the shared user_id percentage between train and test.</p>\n\n<p>It's another story for the fixed 5 folds, where the shared user_id percentage between folds increased to ~15%.\nThis enlarged our cv/lb gap to ~0.006, but fortunately we could get really low cv and consistent cv/lb gap ~0.006 going forward.</p>\n\n<ul>\n<li><p>journey features, first order by (user_id, item_seq_number, activation_date) for both train+test, and train+test+train_active+test_active, then caculate journey number (and percentage), reversed journey number (and percentage) at various levels, e.g. (user_id, parent_category_name), (user_id, parent_category_name, category_name, activation_date), etc.</p></li>\n<li><p>categorical features, treat user_id as categorical variable, generate features at (user_id) or (user_id, other categorical variables) level. e.g. unique item_seq_number, price range = log1p(max(price)) - log1p(min(price))).</p></li>\n<li><p>periods features, generated from periods_train and periods_test.</p></li>\n</ul>\n\n<p><strong>First calculate</strong></p>\n\n<p>for each row in periods_train+periods_test (periods table), activation_time = date_from - activation_date, activation_len = date_to - date_from, etc.;</p>\n\n<p>for each item_id, lag_activation = activation_date - lag(activation_date), toact_diff = activation_date - lag(date_to), etc.</p>\n\n<p><strong>Second aggregate</strong></p>\n\n<p>periods table to (item_id) level, activation_count = .N, and calculate min, max, mean, sd of the above numerical features.</p>\n\n<p><strong>Third aggregate</strong></p>\n\n<p>item_id table to (user_id) level, again calculate min, max, mean, sd.</p>\n\n<p>(BTW I also aggregated the item_id table from second step to other categorical levels as well.)</p>\n\n<p><strong>For example</strong>, one of the important features is al2max_mean_uid, this means:</p>\n\n<p>activation_len2 = date_to - activation_date</p>\n\n<p>activation_len2_max = max(activation_len2) at (item_id) level</p>\n\n<p>al2max_mean = mean(activation_len2_max) at (user_id) level</p>\n\n<h1>Other comments</h1>\n\n<p>I've used three sets of parameters through my series of lgb models (be it base models or level2 models). The parameters were found by baysian optimization. This added some diversity.</p>\n\n<p>I also only kept features that constitute 99% of gain of the previous models. This helped me remain 700+ features for most of my base lgb models, and 300+ for most of my level2 lgb models. (Roughly 3000 sparse features for xgb models)</p>\n\n<p>I've also trained models predicting modified labels to add diversities:</p>\n\n<ul>\n<li>use mean deal_probability instead at (title, description) level, etc.</li>\n<li>train a multiclass classification model, for example, 5-class classification for class 0.005, class 0.125, class 0.765, class 0.805, class 0.865.</li>\n</ul>",
      "votes": 38,
      "replies": [
        {
          "id": 349746,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-06-28T15:13:59.037000",
          "content": "<p>Congratulations and thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349867,
          "author_name": "yi",
          "author_url": "",
          "post_date": "2018-06-28T19:25:48.113000",
          "content": "<p>Congratulations on this exceptionally amazing achievement! Your team well deserve this!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349915,
          "author_name": "Andrey Vykhodtsev",
          "author_url": "",
          "post_date": "2018-06-28T20:59:04.080000",
          "content": "<p>Congratulation and thanks for sharing! This one has caught my attention: \"I also only kept features that constitute 99% of gain of the previous models. \". Do I understand this correctly that you looked at feature importance and kept only first 99% of features? Or was it something more complicated?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349927,
          "author_name": "Arsenal",
          "author_url": "",
          "post_date": "2018-06-28T21:42:24.437000",
          "content": "<p>Yes you got the idea, more precisely, I kept minimum number of features such that sum of gain &gt;= 0.99. Since most of the features can be calculated for different categorical levels, some of them might be meaningless, developing models this way served as an automatic way of feature selection.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 350111,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-06-29T07:40:50.913000",
          "content": "<p>Thank you for sharing !   I have questions about how do you orgnize a few thousand features?   Saved to many files or a huge DB ?  how to load features programmaticly ?  Do you have any suggestions or what's your trick to play with 40 models and thousands features ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350126,
          "author_name": "Arsenal",
          "author_url": "",
          "post_date": "2018-06-29T08:13:02.237000",
          "content": "<p>I only have thousands or tens of thousands of sparse features and I saved them in .RData format, which only takes ~10G space for one model. I have hundreds of dense features and I saved them as csv files, which are roughly ~10G as well. If I run out of space, I would probably delete saved features for most of the models but only keep the milestone ones. Anyways, I consider this dataset only medium size, so this probably won't be a big issue anyways :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 350134,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-06-29T08:40:55.463000",
          "content": "<p>Thank you very much. One more question:  I don't understand the No.4 of FeatureEngineering. What's the predicted features like price?  Price is a feature already in train.csv?  how to get your \"predicted price\"?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350418,
          "author_name": "Arsenal",
          "author_url": "",
          "post_date": "2018-06-29T18:18:06.350000",
          "content": "<p>Basically you can train a model using every features except price related features (or some specific subsets of features) to predict price. The predicted price alone can be a single feature, but then you can aggregate the predicted price to different categorical levels, and compare the true price vs predicted price in various ways to generate more features. Hope it clarifies.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350452,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-06-29T19:33:11.753000",
          "content": "<p>Got it, thanks. I'll try this solution on my next competition. Do you have any idea this method will work on which kind of features ? ( price ? or some continues numbers ? or something else ? )</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350627,
          "author_name": "Arsenal",
          "author_url": "",
          "post_date": "2018-06-30T05:20:12.467000",
          "content": "<p>In this dataset, we know the meaning of each raw feature. You can always try something that intuitively makes sense. A good example would be Georgiy's predictions on log1p(price), renewed and log1p(total_duration) using active data. Those are extremely strong features. You can tell by CV if it works or not.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 352479,
      "author_name": "StellaXu",
      "author_url": "",
      "post_date": "2018-07-04T13:07:53.253000",
      "content": "<p>Study</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349674,
      "author_name": "Fengari",
      "author_url": "",
      "post_date": "2018-06-28T13:20:26.397000",
      "content": "<p>Congrats~~~\nHope I can stay close to you guys. : )\nAttention really helps me a lot, I use several attention models with rnn and cnn structure. included self attention, content attention(text, image and  handcrafted features interaction)\nText translation augmentation and deep oof image features boost a lot to us, however, we just find it in last days.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349603,
      "author_name": "Atanas Atanasov",
      "author_url": "",
      "post_date": "2018-06-28T10:51:34.110000",
      "content": "<p>Congratulations!  Thanks for sharing your massive and impressive amount of work !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349524,
      "author_name": "PLA",
      "author_url": "",
      "post_date": "2018-06-28T08:22:16.473000",
      "content": "<p>Great work! </p>\n\n<p>Have you tried to finetune the image NNs instead of use single layers?</p>\n\n<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349340,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-06-28T02:40:51.617000",
      "content": "<p>Congrats on your first place.... Thanks for your solution details.... </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349285,
      "author_name": "Ahmed Alesh",
      "author_url": "",
      "post_date": "2018-06-28T01:36:51.763000",
      "content": "<p>Brilliant step-by-step solution.</p>\n\n<p>I could barely get a strong nn model. A lot to try in another competition.</p>\n\n<p>Thanks and congrats.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 351355,
      "author_name": "zuri",
      "author_url": "",
      "post_date": "2018-07-02T01:45:11.103000",
      "content": "<p>Congratulations！Your NN solution is awesome and well worth learning. Could you share more NN model code? Many Thanks! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3361834,
      "author_name": "Shivani Sharma",
      "author_url": "",
      "post_date": "2025-12-03T21:28:29.243000",
      "content": "<p>Great Learning! Have a look at my notebook <a href=\"https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble\" target=\"_blank\">https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1000233,
      "author_name": "Okoub1998",
      "author_url": "",
      "post_date": "2020-09-06T12:03:11.833000",
      "content": "<p><a href=\"https://www.kaggle.com/gdanschin\" target=\"_blank\">@gdanschin</a> Can you please share a notebook / code example of the create of the NN itself? including the part that merges the different types of data (Categorical/Numerical/Image/Text)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 383189,
      "author_name": "DeepLyle",
      "author_url": "",
      "post_date": "2018-09-08T01:05:09.457000",
      "content": "<p>Congratulations! And thanks for sharing such an inspiring solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 382836,
      "author_name": "agnostic",
      "author_url": "",
      "post_date": "2018-09-07T07:40:52.983000",
      "content": "<p>Congratulations！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 352218,
      "author_name": "AI GO",
      "author_url": "",
      "post_date": "2018-07-03T23:17:01.473000",
      "content": "<p>Great Kernel. Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 351825,
      "author_name": "Jalindar Karande",
      "author_url": "",
      "post_date": "2018-07-03T05:36:37.167000",
      "content": "<p>Nice Ensemble solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 350984,
      "author_name": "Matthew Frost",
      "author_url": "",
      "post_date": "2018-07-01T00:31:08.887000",
      "content": "<p>nice!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 350186,
      "author_name": "Sunny Ghosh",
      "author_url": "",
      "post_date": "2018-06-29T11:15:31.747000",
      "content": "<p>Great work!\nCongratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349841,
      "author_name": "Webber",
      "author_url": "",
      "post_date": "2018-06-28T17:57:19.980000",
      "content": "<p>Congratulation! Your NN solution is awesome! would u mind share some NN details like Dense Layer size, Dropout rate and Learning Rate? Did u tune them one by one through CV?  Too many parameters to tune in NN. Feel overwhelmed</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349822,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-06-28T17:27:20.350000",
      "content": "<p>Huge win. Congrats!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349788,
      "author_name": "nicapotato",
      "author_url": "",
      "post_date": "2018-06-28T16:40:39.893000",
      "content": "<p>Wow.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349785,
      "author_name": "PrasunMishra",
      "author_url": "",
      "post_date": "2018-06-28T16:36:57.467000",
      "content": "<p>Thanks, Thousand voices and Little Boat for amazing achievement and also sharing the details. Learned a lot from your NN architecture. Did you try GRU at all? If yes then can you mind sharing why/how LSTM scored over GRU and what approach you took? Thanks in advance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349712,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-06-28T14:24:39.797000",
      "content": "<p>Congratulations @LittleBoat, @Arsenal', and Thousandvoices for a well deserved win. Thanks for sharing your solutions. </p>\n\n<p>I can see why I did not go far in as I focused more on FE because I wanted to sharpen my skills there. Felt like this is the right competition for feature engineering. Leaving NN as the last in my list to tackle turned out to be a big mistake as I could not take enough time from my day job to do much in the last 2 weeks. I found progress was slow and incremental in this competition. I certainly learnt a bunch of lessons for the future.</p>\n\n<p>Thanks to all the contributors in kernels as well as discussions. There are many people but special shout out to @Benjamin Minixhofer,  @kxx, @SRK, @SamratP, <a href=\"/peterhurford\">@peterhurford</a>, @Shanth,  @Derek, @Serigne, @Ethan, @Nooh, and many more.</p>\n\n<p>See you in the next one and Happy Kaggling !!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349637,
      "author_name": "JamesYuan",
      "author_url": "",
      "post_date": "2018-06-28T12:20:56.273000",
      "content": "<p>Congrats~~ What an impressive solution! And thx for sharing:) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349570,
      "author_name": "Insaf Ashrapov",
      "author_url": "",
      "post_date": "2018-06-28T09:33:17.223000",
      "content": "<p>Looking forward to see you code;) Thank you for the post</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349538,
      "author_name": "Andrii Sydorchuk",
      "author_url": "",
      "post_date": "2018-06-28T08:45:00.107000",
      "content": "<p>Congrats Little Boat and the team! Would you mind to post code for your RNN architecture to process text? Have you done any text cleanup?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349484,
      "author_name": "Stephan Jo",
      "author_url": "",
      "post_date": "2018-06-28T07:13:07.567000",
      "content": "<p>Congrats you &amp; members,\nIt was quite impressive to keep leading position during competition.\nI am really thankful to kind explanation too~!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349427,
      "author_name": "A_b",
      "author_url": "",
      "post_date": "2018-06-28T05:07:29.893000",
      "content": "<p>Very nice NN approach and congratulations on your consistent leading position over time! How long did it take to train your NN architecture and how much RAM did you need? - Cheers </p>",
      "votes": 0,
      "replies": [
        {
          "id": 349672,
          "author_name": "Little Boat",
          "author_url": "",
          "post_date": "2018-06-28T13:19:23.923000",
          "content": "<p>1 epoch took about 3 mins for the final model. And each model is about 8 epochs. I have a machine with 64GB RAM, but you can do it with 32 or 48GB.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 349934,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T22:02:36.417000",
          "content": "<p>@LittleBoat, great job on improving the RNN over time. Any chance to see the code on the RNN at any stage would be great. I struggled with RNN runtimes, got 5 mins per epoch - best model was 26 epochs bagged twice with 5mins per epoch - but over 5 folds + test - this adds up to over 24 hours. \nDid you bag the models; or just create a lot of different ones and add the individuals to the stack; did you bag epochs, or just choose the last one ? Would love to learn from you how to improve in RNNs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349950,
          "author_name": "Little Boat",
          "author_url": "",
          "post_date": "2018-06-28T23:35:21.100000",
          "content": "<p>The RNN code is surprisingly easy here</p>\n\n<pre><code>seq_title_description = Input(shape=[max_seq_description_length], name=\"seq_description\")\nemb_seq_title_description = Embedding(vocab_size, EMBEDDING_DIM1, weights=[embedding_matrix1], trainable=False)(\n    seq_title_description)\nemb_seq_title_description = SpatialDropout1D(0.25)(emb_seq_title_description)\nrnn_layer1 = CuDNNLSTM(128, return_sequences=True)(emb_seq_title_description)\nrnn_layer1 = CuDNNLSTM(128, return_sequences=False)(rnn_layer1)\nfc_rnn = Dense(256, init='he_normal')(rnn_layer1)\nfc_rnn = PReLU()(fc_rnn)\nfc_rnn = BatchNormalization()(fc_rnn)\nfc_rnn = Dropout(0.25)(fc_rnn)\n</code></pre>\n\n<p>If you are using Keras you should use CuDNN implementation which would be much faster. It is surprising that you needed 26 epochs to converge. For me it converged after 6 or so and I ran it for 8 because 8 seemed to give slightly better average with different seeds. For each fold, I ran 8 epochs (some fold actually only needed 5, but I just let it \"overfit\" a bit), with 5 fold, so that is 40 epochs per seed. And I ran 4 seeds and then average them. So in total, it is 40 * 4 = 160 epochs. Before adding all the engineered features I remember each epoch took about 100 seconds so in total it is only 16000 seconds so less than 5 hours.</p>\n\n<p>Hope it helps.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 350080,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-29T06:19:54.240000",
          "content": "<p>Super, thanks, yes it helps. I think ours converged sooner but we ran for more epochs and predicted and bagged each epoch. \nOk, I see now you read the whole dscr and title as one string... Great thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350329,
          "author_name": "Little Boat",
          "author_url": "",
          "post_date": "2018-06-29T15:06:31.217000",
          "content": "<p>@Darragh,\nActually, title and description were read separately :( it was just dummy name really.... Reading them separately gave about 0.0002-0.0003 improvement.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350341,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-29T15:22:37.027000",
          "content": "<p>Thanks, one more question if its ok; we struggled with Batchnorm after the PRelu, it just destroyed the weights. On retrospect, I suspect its because we concatenated the output of the text RNN and the descr RNN; then passed to a dense and applied batchnorm; whereas maybe we should have applied batchnorm indvidually to each output. Did you face this problem with batchnorm ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350345,
          "author_name": "Little Boat",
          "author_url": "",
          "post_date": "2018-06-29T15:27:41.440000",
          "content": "<p>That is interesting. I didn't investigate if it was batchnorm or other causes but when I started adding image features it made my model worse and only after using separate layers the model got better. You might want to try that out on your features to see if that would help. My intuition of separate dense layer for each input is that they act like a block/gate to force them to be transformed before concatenating. So you can also in theory tweak the unit size to assign different \"weights\" to them.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 349408,
      "author_name": "Mishunyayev Nikita",
      "author_url": "",
      "post_date": "2018-06-28T04:40:22.170000",
      "content": "<p>Congrats! And thanks for sharing best solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349374,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T03:31:28.230000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349371,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T03:27:09.227000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349353,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T02:57:28.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349330,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T02:33:43.727000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 349360,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-28T03:06:39.197000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 349327,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T02:31:04.260000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349295,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T01:51:31.650000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 349299,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-28T01:57:12.050000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 349598,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-28T10:40:21.060000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349670,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-28T13:15:53.200000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349292,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T01:49:47.970000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349291,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T01:47:54.537000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349286,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T01:39:52.180000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349273,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T01:20:38.077000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 349298,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-28T01:55:29.173000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349977,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-29T01:21:45.317000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349991,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-29T01:45:11.590000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 349683,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T13:29:38.763000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349561,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T09:22:14.133000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349426,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T05:06:12.707000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349290,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T01:47:29.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706978,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-31T05:30:50.403000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450523,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-05T06:59:39.940000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 352560,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-04T16:18:00.400000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 352377,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-04T07:50:59.180000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 351974,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-03T12:38:45.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 351455,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-02T08:20:30.100000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349564,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T09:25:03.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349406,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T04:39:47.673000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349269": "**Preface**\n\nFirst of all, I would like to thank Avito and Kaggle for hosting such an interesting competition! Lots of ways to do feature engineering, data quality is not bad, and data size is arguably accessible to everyone. \n\nI would also like to thank my three awesome teammates, Arsenal, Georgiy Danshchin and thousandvoices (ordered alphabetically) ! You all are truly amazing feature engineering masters! \n\nArsenal has been a long-time friend of mine and we decided to go with “Dance with Ensemble” again in honor of our last loss under the same team name (which was 3 years ago already!) I believe that time we dropped out of top 3 mainly because we didn’t use NN at all. And the funny thing is, this time (R)NN is one of the main reasons that we were always staying ahead of other teams.\n\n**Summary**\n\nSo our approach is probably not different than what other top teams did in a high level. We have some lgb models, some NN models, some xgb models as first layer, and some lgb models, some xgb models and some NN models as second layer, and one NN as the final layer. But honestly, the complicated structure (3 layer) probably gave us about 0.0002 - 0.0004 improvement. Just several models with a simple linear stacking should be able to achieve not exactly the same, but quite similar score.\n\nA few days ago our best single models were both 215X (on public LB) for NN and lgb. And then amazingly Georgiy Danshchin discovered a few features based on active train+test and immediately it boosted the best single lgb to 213X! Purely including them to my NN didn't help much but I couldn't find much time tweaking it (and honestly I was lazy at that point). I think in the end, it had 0.0007 ish improvement on our final score. I will leave this black magic to Georgiy Danshchin to disclose (Hint: RNN is involved there).\n\nStacking is extremely important here, which means that building diversified models are extremely important. I remember when four of us merged, just a linear blending of our models could get to 0.2133. \n\n**NN**\n\nI exclusively worked on NNs for this one and didn’t do much feature engineering otherwise. So I would like to share how you can achieve 0.215X with a single NN, and leave the rest (truly amazing stuff) to my awesome teammates.\n\nAll features matter here. Text, categorical, numerical, images (and probably in this order).  And to my best memory, here is how I did it:\n\n - I got 0.227X with numerical features and categorical embedding\n - And then I included titile and description with 2 RNNs, with fastText pretrained embedding, with some tuning, the score dropped to 0.221X.\n - Played with self training fastText embedding on train+test, and also train active, test active. It turned out that self training on train+test was the best. Score got to 0.220X.\n - Added VGG16 top layer with average pooling. It made my score worse. Did some tuning, specifically, had a separate layer before merging text, image, categorical, numerical features together, and started to see the improvement . Got to about 0.219X.\n - Tried to tweak text models, with CNN or Attention etc. None worked. In the end, went with 2 layer LSTM followed by a dense layer. Probably 0.0003 improvement here.\n - Tried different CNN models for images. None of the \"fine tuning\" models worked (and GOD it was slow). But fixed ResNet50 middle layer helped by probably another 0.0005. Now the score became 0.218X.\n - Started doing all sorts of tuning (based on intuition mostly). And found that adding spatial dropout between text and LSTM helped quite a bit, probably 0.0007 - 0.001. And fine tuned dropout ratio overall helped too. In the end, about 0.001 - 0.0015 improvement here. So now the score was around 0.2165 - 0.217.\n - Started including all engineered features from teammates. Lots of engineered features from them (the ones based on text) didn't help but others did. So in the end, a NN with 0.215X!\n - If you kept saving models along the way (models with fewer features, models with more features that got worse result, etc.), you could train a fully connected NN on top of them and for me it was around 0.008 improvement in addition. In other words, you can easily get into top 10 with only NN!\n\n\nI also attached a simple sketch of what the model architecture looks like. \n\n\nAnd this thread is **To Be Continued** by my amazing teammates!\n\nJump to: \n\n**Arsenal's approach:**\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349563\n\n**thousandvoices's approach:** \nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349386\n\n**Georgiy Danshchin:** \nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/59880#349710\n\n![NN Model][1]\n\n\n  [1]: https://pbs.twimg.com/media/DgvX3pWUYAAlCAK.jpg:large\n\n\nOnce again, thank you all for this amazing experience!",
    "349710": "## The last part of Dance with Ensemble solution ##\n\nFirst of all, I want to thank my great teammates for this awesome experience and Avito for organizing very interesting competition!\n\nFrom my point of view main things for this competition were strong feature engineering and stacking of diverse models. So let’s dive into the details.\n\nValidation\n----------\n\nBefore joining the team I've used a time rolling window validation split with 4 folds (8 days in train 3 days in test for each fold). It was very consistent with public lb score, gap between cv and lb was between 0.002 and 0.003. After joining the team I've regenerated all my features and models for common kfold split, and noticed that cv gap increased by 0.003. I figured out that this was mainly due to user_id based target encoding features, so I put them aside and used them only for second level models later.\n\nFeatures\n--------\n\n**Categories**\n\n - Percent encoding of categories computed on train+test parts of data\n - Percent encoding of user_id category computed on train_active+test_active.\n - Mean target encoding of different categories, which were computed separately for each fold in the oof manner using inner folds.\n - Mean log1p(price) by different categories and differences between these features and actual log1p(price) as well. For most categories these features were computed using train+test, for user_id and title I used train_active+test_active.\n - Mean 'renewed' value by different categories (train_active + test_active)\n - Mean 'total_duration' value by different categories (train_active + test_active)\n - Sparse one hot encoded categories for Lightgbm models\n\nI’ve used only single categories for all these features, didn’t see any consistent improvement with category interactions. In addition to category_name/region/param_* I treated stemmed title as category too. The 'renewed' value is just a binary value which indicates whether an item has more than one period in periods_train/periods_test data, 'total_duration' is the sum of lengths of all periods for particular item.\nAll 'mean' features were postprocessed with additive smoothing. The coefficients were different for different groups of features and were tuned manually with help of cv feedback.\n\n**Text**\n\n - Sparse tfidf 1-2grams vectors for title and joined title+description+params (used only in Lightgbm models)\n - 14 different out of fold ridge predictions over different sparse tfidf vectors. Different combinations of fields, and their concatenations, stemmed/raw tokens etc. Computed separately for each fold using inner folds as mean target features.\n - Various text statistics: caps quotient, alfanum quotient, char counts, word counts etc.\n - Fasttext embeddings trained on train_active + test_active were used as initial embedding matrix for rnn models.\n\n**Numeric**\n\n -  Just log1p(price) and log1p(item_seq_number)\n\n**Images**\n\n - Channel mean values\n - Image width and height\n - Pretrained resnet152 last hidden layer output (only for nn models)\n - 100 components of svd over previous feature\n - Out of fold ridge predictions trained on pretrained resnet101 and resnet152 last hidden layers.\n - Finetuned resnet34 out of fold predictions gave a significant boost to my models but took a loooong time to train :)\n\n**Active data**\n\nI’ve designed 3 nn models trained on active data to predict log1p(price), renewed and log1p(total_duration). These models had 2 rnn branches for title and description and also used category embeddings. The difference between actual log1p(price) and predicted one was an extremely important feature.\n\nModels\n------\n\nI had 25 base models (15 lgbs and 10 nns) trained on different feature subsets and with different parameters. My best lgb reaches 0.2155/0.2196 public/private score, best nn is 0.2160/0.2200. Stacking of these models + finetuned resnet34 oof + user_id mean target with one lgb gives 0.2132 on public and 0.2174 on private.\n\nLgb models could be divided into two groups. One of them uses small set of dense features, it’s trained for 700 rounds with num_leaves=31. Second group uses all features from the first group with addition of sparse tfidf/categories and svd decomposition of pretrained resnet152. It’s trained for 2300 rounds with num_leaves=384, it takes about 1 hour to train on one fold.\n\nNN models have separate branches for title and description: 2-level bidirectional lstm/gru with attention. They also contain separate branches for category embeddings/numeric features and pretrained resnet152 vectors. Every branch is followed by a separate dense layer and finally everything is concatenated and followed by another common dense layer. Text embedding layer is shared for title/description, it’s initialized with fasttext trained on train_active + test_active and frozen for the first epoch, but finetuned later.\n",
    "349386": "First of all, I would like to thank my teammates for their excellent performance. Your discussions and feature engineering insights were incredibly helpful. I learnt a lot from you.\n\nHere is a short summary of my approach. I trained 10 first level models (5 lightgbms, 3 neural nets, 1 catboost and 1 ridge). Let me share some ideas behind my best single lightgbm:\n### Categorical features interactions ###\nConcatenating categorical features is a well-known method for taking this into account. Rather surprisingly, I wasn't able to come up with an automated pipeline that outperformed a list of pairs of features to merge that I created based simply on common sense. Here it is:\n\n - region, parent category name\n - region, category name\n - region, image top-1\n - region, param 1\n - region, param 2\n - region, param 3\n - city, parent category name\n - city, category name\n - city, param 1\n - city, param 2\n - city, param 3\n - category name, image top-1\n - parent category name, image top-1\n - param 1, param 2\n - param 1, image top-1\n\n### Categorical features encoding ###\nFor each categorical variable I added the following features to the model based on the value it took for a particular advertisement:\n\n - Mean target of previous ads\n - Total occurence count\n - Median price\n\n### Text features ###\nRaw word counts extracted by CountVectorizer for title, description and concatenated titles of all user's ads worked best for me. My efforts to design dense representations by applying svd or training embeddings with supplementary data only decreased the score. My model also included several engineered text features like word and char length, amount of caps, digits, punctuation, cyrillic and latin letters, and unique words.\n### Image features ###\nI didn't focus on this part of the challenge much, that's why the only image features I used were image width and height.\n### Aggregated user features ###\nMinimum, median and maximum price, total amount of ads and number of ads with unique titles.\n### Using supplementary data ###\nI trained 2 neural networks that predicted price and length of active period for the ad on supplementary data provided in train_active and test_active. Both of this features turned out to be very strong, making it into 5 most important features of my model. I also used features produced by [this kernel][1] without any changes.\n### Hyperparameters ###\nThis competition favored VERY deep trees, and that was the point that most public scripts have missed. Every time I increased num_leaves my score improved significantly. I ended up using trees with 1024 leaves because it was the maximum that fitted into my memory.\n### Parting words ###\nOverall, this was an exciting competition with lots of raw data and feature engineering opportunities (this is probably the reason why team merges worked so well here :) ), and I'd love to thank Avito for providing such a challenging dataset.\n\n\n  [1]: https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm",
    "349563": "Preface\n=======\n\nI need to thank my teammates again for this amazing team up experience on such a challenging problem!\n\nI formed a team with Little Boat early, and we soon decided to focus exclusively on tree-based models (xgb/lgb) and NN models, respectively.\n\nI spent 50% of my time training my original xgb/lgb, and 50% of my time consolidating all features shared among the 4 of us. I trained 6 base xgb models, 7 level2 xgb models, 13 base lgb models, and 12 level2 lgb models. The best base lgb got 0.2138 on public LB, and the best level2 lgb got 0.2110 on public LB.\n\n\nFeature Engineering\n=========================\n\nI will share my part of feature engineering, please note that it's a joint effort of all of my teammates to achieve a 0.213X single lgb. Particularly, Georgiy Danshchin found 4 amazing features that boosted our best base lgb from 0.215X to 0.213X in the last couple of days.\n\n**1. text features**\n\n- tfidf on title, description, title+description, titile+description+param_1, etc. I kept these features sparse for all my xgb models; and I used svd and oof ridge features for all my lgb models to keep it diversified.\n\n- text statistics, such as words length (#characters/#words) at various levels (title, description, etc.), unique words in title additional to description, etc.\n\n\n**2. image features**\n\n- image statistics, reference: https://www.kaggle.com/shivamb/ideas-for-image-features-and-image-quality\n\n- features from 3 pretrained nn models, reference: https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf. I extracted predicted scores and categories for top-5-category exactly from ResNet50, InceptionV3 and Xception. The predicted categories were treated as new categorical variables, in addition to parent_category_name and category_name.\n\n- features from vgg16 pretrained network, 512 raw features + first 15 PCA features. (extracted by Little Boat)\n\n- keypoint feature, reference: https://www.kaggle.com/c/avito-demand-prediction/discussion/59414\n\n\n**3. categorical features**\n\n- count/count unique features at various levels. These features were generated for both train+test, and train+test+train_active+test_active. For example, number of ads in (parent_category_name, category_name, param_1), number of unique user_ids in (region, city).\n\n- target encoding at various levels. These include mean deal_probability for categories with count &gt;= 5000; mean predicted deal_probability (train+test) for categories with count &gt;= 5000 (Note that threshold was chosen such that cv/lb gap remained roughly the same.); OOF mean deal_probability encoding; OOF mean deal_probability * min(1, log(count)/log(10000)) encoding.\n\n\n**4. predicted independent variables features** (price, image_top_1, item_seq_number, day_diff)\n\n- xgb predicted price, image_top_1 (train on train+test, predict on train+test); lgb predicted price, image_top_1, item_seq_number (oof prediction); rnn predicted day_diff = day_to - day_from (trained on active and period, shared by Little Boat)\n\n- mean predicted price, image_top_1, item_seq_number, day_diff at various categorical levels.\n\n- diff features at various levels, for example, (price-xgb_price)/price at (category_name) level, log(image_top_1) - log(lgb_image_top_1) at (resnet50_category1) level, etc.\n\n- I tried predicting parent_category_name and category_name (multicalss classification) from text features, but including them made cv slightly worse.\n\n\n**5. user_id features**\n\nWe were aware that our team was slightly overfitting cv because of my excessive user_id features. We chose to work on fixed 5 folds for stacking on team merge. It would be a tremendous effort if we retrained our models using a different validation strategy.\n\nGeorgiy Danshchin was working on time-based 4 folds before he joined us. User_id features worked fine in his setting because the shared user_id percentage between time-based folds was ~6%, roughly the same as the shared user_id percentage between train and test.\n\nIt's another story for the fixed 5 folds, where the shared user_id percentage between folds increased to ~15%.\nThis enlarged our cv/lb gap to ~0.006, but fortunately we could get really low cv and consistent cv/lb gap ~0.006 going forward.\n\n- journey features, first order by (user_id, item_seq_number, activation_date) for both train+test, and train+test+train_active+test_active, then caculate journey number (and percentage), reversed journey number (and percentage) at various levels, e.g. (user_id, parent_category_name), (user_id, parent_category_name, category_name, activation_date), etc.\n\n- categorical features, treat user_id as categorical variable, generate features at (user_id) or (user_id, other categorical variables) level. e.g. unique item_seq_number, price range = log1p(max(price)) - log1p(min(price))).\n\n- periods features, generated from periods_train and periods_test.\n\n**First calculate**\n\nfor each row in periods_train+periods_test (periods table), activation_time = date_from - activation_date, activation_len = date_to - date_from, etc.;\n\nfor each item_id, lag_activation = activation_date - lag(activation_date), toact_diff = activation_date - lag(date_to), etc.\n\n**Second aggregate**\n\nperiods table to (item_id) level, activation_count = .N, and calculate min, max, mean, sd of the above numerical features.\n\n**Third aggregate**\n\nitem_id table to (user_id) level, again calculate min, max, mean, sd.\n\n(BTW I also aggregated the item_id table from second step to other categorical levels as well.)\n\n**For example**, one of the important features is al2max_mean_uid, this means:\n\nactivation_len2 = date_to - activation_date\n\nactivation_len2_max = max(activation_len2) at (item_id) level\n\nal2max_mean = mean(activation_len2_max) at (user_id) level\n\n\nOther comments\n====================\n\nI've used three sets of parameters through my series of lgb models (be it base models or level2 models). The parameters were found by baysian optimization. This added some diversity.\n\nI also only kept features that constitute 99% of gain of the previous models. This helped me remain 700+ features for most of my base lgb models, and 300+ for most of my level2 lgb models. (Roughly 3000 sparse features for xgb models)\n\nI've also trained models predicting modified labels to add diversities:\n\n- use mean deal_probability instead at (title, description) level, etc.\n- train a multiclass classification model, for example, 5-class classification for class 0.005, class 0.125, class 0.765, class 0.805, class 0.865.",
    "352479": "Study",
    "349674": "Congrats~~~\nHope I can stay close to you guys. : )\nAttention really helps me a lot, I use several attention models with rnn and cnn structure. included self attention, content attention(text, image and  handcrafted features interaction)\nText translation augmentation and deep oof image features boost a lot to us, however, we just find it in last days.",
    "349603": "Congratulations!  Thanks for sharing your massive and impressive amount of work !",
    "349524": "Great work! \n\nHave you tried to finetune the image NNs instead of use single layers?\n\nThanks for sharing!",
    "349340": "Congrats on your first place.... Thanks for your solution details.... ",
    "349285": "Brilliant step-by-step solution.\n\nI could barely get a strong nn model. A lot to try in another competition.\n\nThanks and congrats.",
    "351355": "Congratulations！Your NN solution is awesome and well worth learning. Could you share more NN model code? Many Thanks! ",
    "3361834": "Great Learning! Have a look at my notebook https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble",
    "1000233": "@gdanschin Can you please share a notebook / code example of the create of the NN itself? including the part that merges the different types of data (Categorical/Numerical/Image/Text)",
    "383189": "Congratulations! And thanks for sharing such an inspiring solution!",
    "382836": "Congratulations！",
    "352218": "Great Kernel. Congrats!",
    "351825": "Nice Ensemble solution!",
    "350984": "nice!",
    "350186": "Great work!\nCongratulations!",
    "349841": "Congratulation! Your NN solution is awesome! would u mind share some NN details like Dense Layer size, Dropout rate and Learning Rate? Did u tune them one by one through CV?  Too many parameters to tune in NN. Feel overwhelmed",
    "349822": "Huge win. Congrats!!!",
    "349788": "Wow.",
    "349785": "Thanks, Thousand voices and Little Boat for amazing achievement and also sharing the details. Learned a lot from your NN architecture. Did you try GRU at all? If yes then can you mind sharing why/how LSTM scored over GRU and what approach you took? Thanks in advance.",
    "349712": "Congratulations @LittleBoat, @Arsenal', and Thousandvoices for a well deserved win. Thanks for sharing your solutions. \n\nI can see why I did not go far in as I focused more on FE because I wanted to sharpen my skills there. Felt like this is the right competition for feature engineering. Leaving NN as the last in my list to tackle turned out to be a big mistake as I could not take enough time from my day job to do much in the last 2 weeks. I found progress was slow and incremental in this competition. I certainly learnt a bunch of lessons for the future.\n\nThanks to all the contributors in kernels as well as discussions. There are many people but special shout out to @Benjamin Minixhofer,  @kxx, @SRK, @SamratP, @peterhurford, @Shanth,  @Derek, @Serigne, @Ethan, @Nooh, and many more.\n\nSee you in the next one and Happy Kaggling !!!",
    "349637": "Congrats~~ What an impressive solution! And thx for sharing:) ",
    "349570": "Looking forward to see you code;) Thank you for the post",
    "349538": "Congrats Little Boat and the team! Would you mind to post code for your RNN architecture to process text? Have you done any text cleanup?",
    "349484": "Congrats you &amp; members,\nIt was quite impressive to keep leading position during competition.\nI am really thankful to kind explanation too~!!",
    "349427": "Very nice NN approach and congratulations on your consistent leading position over time! How long did it take to train your NN architecture and how much RAM did you need? - Cheers ",
    "349408": "Congrats! And thanks for sharing best solution!",
    "349374": "Congratulations! It's impressive. \n\nThank you for your sharing.",
    "349371": "your nn arch is great! thanks for your graph!",
    "349353": "Congratulations! And thank you for sharing how you built your NN.",
    "349330": "Congratulations! And thanks for the detailed step-by-step explanation. \n\nMay I ask what kind of computing resources did you use for the NN? ",
    "349327": "Amazing scores and amazing feature engineer,  thank you for you detailed NN , it is very helpful. And can't wait to see feature engineer part.",
    "349295": "Congrats. I can't wait to see your teammates talking the feature engineering things.\nI have two questions please.\n1. Have you tried other CNN pre-trained models other than VGG and Resnet50?\n2. Have you tried to train your own word2vec model on the competition text data in train and train_active other than fasttext pre-trained?\nThank you for sharing such a brilliant solution. It's super detailed.",
    "349292": "Exhaustive kernel. Excellent work ",
    "349291": "Congrats  Little Boat and you team for the amazing winning and thanks for sharing.",
    "349286": "WOW! Mind blowing! Congrats God Boat!",
    "349273": "Amazing！！！ Congrats for your another first place solution.  Thanks for sharing your wisdom.\n\nPlus, Can I ask did you accelerate the process of loading image? For me it takes extremely long for loading images for one epoch. (1 million images here)",
    "349683": "",
    "349561": "",
    "349426": "",
    "349290": "Thank you for the sharing!",
    "706978": "Nice Post! Thanks for sharing!",
    "450523": "Thanks or sharing your work.",
    "352560": "Thank you for sharing your insights!",
    "352377": "Thanks for sharing",
    "351974": "Really helpful! Thank you so much.",
    "351455": "thanks for sharing",
    "349564": "Congrats and thanks for sharing!!!",
    "349406": "Well done and thanks for sharing."
  }
}