{
  "id": 22205,
  "title": "2nd Place Solution: TheQuants",
  "url": "/competitions/avito-duplicate-ads-detection/writeups/thequants-2nd-place-solution-thequants",
  "author_name": "",
  "post_date": "2016-07-12T18:57:08.277Z",
  "votes": 48,
  "comment_count": 22,
  "views": 3260,
  "content": "<p>Hi everyone!</p>\n\n<p>Here's a description of our team's efforts for this competition. I think this competition was very much driven on feature engineering, with meta-modelling more as a finishing touch.</p>\n\n<p><strong>Our Strategy</strong></p>\n\n<ul>\n<li>Understanding the Data, developing features until team merger deadline, metamodelling in the last week.</li>\n<li>Merging early based on standing of the leaderboard at the time</li>\n<li>Each individually building features to try and branch out as much as possible, with constant discussion about where to look next</li>\n<li>Onboarding an experienced w&#822;i&#822;z&#822;a&#822;r&#822;d&#822; Kaggler for the final stage. </li>\n</ul>\n\n<hr>\n\n<p><strong>Data Cleaning:</strong></p>\n\n<p>In order to clean the text, we applied stemming using the NLTK Snowball Stemmer, and removed stopwords/punctuation as well as transforming to lowercase.</p>\n\n<p><strong>Validation Strategy:</strong></p>\n\n<p>Initially, we were using a random validation set before switching to a set of non-overlapping items, where none of the items in the valset appeared in the train set. This performed somewhat better, however we had failed to notice that the training set was ordered based on time! We later noticed this and switched to using last 33% as a valset. This set correlated relatively well with the leaderboard until the last week, when we were doing meta-modelling and it fell apart - at a point where it would be too much work to switch to a better set. This hurt us a lot towards the end of the competition.</p>\n\n<h2>Features:</h2>\n\n<p>In order to pre-emptively find over-fitting features, we built a script that looks at the changes in the properties (histograms and split purity) of a feature over time, which allowed us to quickly (200ms/feature) identify overfitting features without having to run overnight xgboost jobs. If there is interest, I would be willing to put it on github. </p>\n\n<p>After removing overfitting features, our final feature space had 587 features in it. Here&#8217;s a summary:</p>\n\n<p><strong>General:</strong></p>\n\n<ul>\n<li>CategoryID, parentCategoryID raw CategoryID, parentCategoryID one-hot\n(except overfitting ones) </li>\n<li>Price difference / mean</li>\n<li>Generation3probability (output from model trained to detect <br>\ngenerationmethod=3)</li>\n</ul>\n\n<p><strong>Location:</strong></p>\n\n<ul>\n<li>LocationID &amp; RegionID raw</li>\n<li>Total latitude/longtitude</li>\n<li>SameMetro, samelocation, same region etc.</li>\n<li>Distance from city centres  (kalingrad, moscow, petersburg, krasnodar, makhachkala, murmansk, perm, omsk, khabarovsk, kluichi, norilsk)</li>\n</ul>\n\n<p>Gaussian noise was added to the location features to prevent overfitting to specific locations, whilst allowing xgboost to create its own regions.</p>\n\n<p><strong>All Text:</strong></p>\n\n<ul>\n<li>Length / difference in length</li>\n<li>nGrams Features (n = 1,2,3) for title and description (Both Words and Characters)\n\n<ul><li>Count of Ngrams (#, Sum, Diff, Max, Min)</li>\n<li>Count of Unique Ngrams</li>\n<li>Ratio of Intersect Ngrams</li>\n<li>Ratio of Unique Intersect Ngrams</li></ul></li>\n<li>Distance Features:\n\n<ul><li>Jaccard , Cosine, Levenshtein and Hamming Distance between the titles and descriptions</li></ul></li>\n<li>Special Character Counting &amp; Ratio Features:\n\n<ul><li>Counting &amp; Ratio features of Capital Letters in title and description</li>\n<li>Counting &amp; Ratio features of Special Letters (digits, punctuations, etc.) in title and description</li></ul></li>\n<li>Similarity between sets of words/characters</li>\n<li>Fuzzywuzzy/jellyfish distances</li>\n<li>Number of overlapping sets of n words (n=1,2,3)</li>\n<li>Matching moving windows of strings</li>\n<li>Cross-matching columns (eg. title1 with description2)</li>\n</ul>\n\n<p><strong>Bag of words:</strong></p>\n\n<p>For each of the text columns, we created a bag of words for both the intersection of words and the difference in words and encoded these in a sparse format resulting in ~80,000 columns each. We then used this to build Naive Bayes, SGD and similar models to be used as features.</p>\n\n<p><strong>Price Features:</strong></p>\n\n<ul>\n<li>Price Ratio</li>\n<li>Is both/one price NaN</li>\n<li>Total Price</li>\n</ul>\n\n<p><strong>JSON Features:</strong></p>\n\n<ul>\n<li>Attribute Counting Features:\n\n<ul><li>Sum, diff, max, min</li></ul></li>\n<li>Count of Common Attributes Names</li>\n<li>Count of Common Attributes Values</li>\n<li>Weights of Evidence model on keys/values\nXGBoost model on sparse encoded attributes</li>\n</ul>\n\n<p><strong>Image Features:</strong></p>\n\n<ul>\n<li># of Images in each Set</li>\n<li>Difference Hashing of images</li>\n<li>Hamming distance between each pair of images</li>\n<li>Pairwise comparison of file size of each image</li>\n<li>Image dimension Features:</li>\n<li>Pairwise comparison of dimension of each image</li>\n<li>BRISK keypoint/descriptor matching</li>\n<li>Image histogram comparisons</li>\n<li>Dominant colour analysis</li>\n<li>Uniqueness of images (how many other items have the same images)</li>\n<li>Difference in number of images</li>\n</ul>\n\n<p>I found a possible image metadata leak (the creation dates of the images were embedded in the zip files) but didn&#8217;t try to use these for features - gotta play fair after all!</p>\n\n<p><strong>Clusters:</strong></p>\n\n<p>We found clusters of rows by grouping rows which contain the same items (eg. if row1 has items 123, 456 and row2 has items 456, 789 they are in the same cluster). We discovered that the size of these clusters was a very good feature (larger clusters were more likely to be non-duplicates), as well as the fact that clusters always the same generationMethod. Adding cluster-size features gave us a 0.003 to 0.004 improvement. It would be interesting to know if anyone else found these features :)</p>\n\n<h3>Feature Graveyard:</h3>\n\n<p>Overfitting was probably the biggest problem throughout the competition, and lots of features which destroyed in validation didn&#8217;t do so well on the leaderboard. This is likely because the very powerful features learn to recognise specific products or sellers that do not appear in the test set. Hence, our feature graveyard:</p>\n\n<p><strong>TF-IDF:</strong> This was something we tried very early into the competition, adapting our code from the Home Depot competition. Unfortunately, it overfitted very strongly, netting us 0.98 val-auc and only 0.89 on LB. We tried adding noise, reducing complexity etc. but in the end we gave up.</p>\n\n<p><strong>Word2vec:</strong> We tried both training a model on our cleaned data and using the pretrained model posted in the forums. We tried using word-mover distance from our model as features, but they were rather weak (0.70AUC) so in the end we decided to drop these for simplicity. Using the pre-trained model did not help, as the authors used MyStem for stemming (which is not open-source) so we could not replicate their data cleaning. After doing some transformations on the pre-trained model to try and make it work with our stemming (we got it down to about 20% missing words), it scored the same as our custom word2vec model.</p>\n\n<p><strong>Advanced cluster features:</strong> We tried to expand the gain from our cluster features in several ways. I found that taking the mean prediction for the cluster as well as cluster_size * (1-cluster_mean) provided excellent features in validation (50% of gain in xgb importance), however these overfits for reasons I still have to investigate. We also tried taking features such as the stdev of locations of items in a cluster, but these also overfitted. I suspect that Avito used slightly different methods for generating the test set.</p>\n\n<p><strong>Grammar features:</strong> We tried building features to &#8216;fingerprint&#8217; different types of sellers, such as usage of capital letters, special characters, newlines, punctuation etc. However while these helped a lot in CV, they overfitted on the leaderboard.</p>\n\n<p><strong>Brand violations:</strong> We built some features based around words that could never appear together in duplicate listings. (For example, if one item wrote &#8216;iPhone 4s&#8217; but the other one wrote &#8216;iPhone 5s&#8217;, they could not be duplicates). While they worked well at finding non-duplicates, there were just too few cases where these violations occurred to make a difference to the score.</p>\n\n<h2>Meta-model:</h2>\n\n<p>Using a meta-model improved our score by roughly 0.0015 versus our best XGBoost. In the end, a single model would have been enough to net us second place, but you can never be too careful!</p>\n\n<p>Below is an overview of what our final meta looked like. Each generation is different set of features, with later generations having more features. Note that all scores are private LB scores:</p>\n\n<p><img src=\"https://files.slack.com/files-pri/T18TSG1C3-F1QTWCZ5G/avito-meta.png?pub_secret=14957b8a3e\" alt=\"Meta-model schematic\" title></p>\n\n<p>We also tried using Extra Trees, Random Forest, Adaboost &amp; approximate kNN as base models, however these overfit to our validation-set and so they couldn&#8217;t be used in the final meta-model.</p>\n\n<hr>\n\n<p>If anyone has any questions about our solution, we would be happy to go into more depth!</p>\n\n<p><strong>Finally, a great big thank you to my teammates Peter, Sonny and Marios for making TheQuants a brilliant success!</strong> We couldn&#8217;t have done it without every one of you.</p>\n\n<p>And thanks to everyone else for making this such an enjoyable competition</p>\n\n<p>- Mikel, <br>\nTheQuants</p>",
  "messages": [
    {
      "id": "126869",
      "postDate": "07/12/2016 18:57:08",
      "content": "<p>Hi everyone!</p>\n\n<p>Here's a description of our team's efforts for this competition. I think this competition was very much driven on feature engineering, with meta-modelling more as a finishing touch.</p>\n\n<p><strong>Our Strategy</strong></p>\n\n<ul>\n<li>Understanding the Data, developing features until team merger deadline, metamodelling in the last week.</li>\n<li>Merging early based on standing of the leaderboard at the time</li>\n<li>Each individually building features to try and branch out as much as possible, with constant discussion about where to look next</li>\n<li>Onboarding an experienced w&#822;i&#822;z&#822;a&#822;r&#822;d&#822; Kaggler for the final stage. </li>\n</ul>\n\n<hr>\n\n<p><strong>Data Cleaning:</strong></p>\n\n<p>In order to clean the text, we applied stemming using the NLTK Snowball Stemmer, and removed stopwords/punctuation as well as transforming to lowercase.</p>\n\n<p><strong>Validation Strategy:</strong></p>\n\n<p>Initially, we were using a random validation set before switching to a set of non-overlapping items, where none of the items in the valset appeared in the train set. This performed somewhat better, however we had failed to notice that the training set was ordered based on time! We later noticed this and switched to using last 33% as a valset. This set correlated relatively well with the leaderboard until the last week, when we were doing meta-modelling and it fell apart - at a point where it would be too much work to switch to a better set. This hurt us a lot towards the end of the competition.</p>\n\n<h2>Features:</h2>\n\n<p>In order to pre-emptively find over-fitting features, we built a script that looks at the changes in the properties (histograms and split purity) of a feature over time, which allowed us to quickly (200ms/feature) identify overfitting features without having to run overnight xgboost jobs. If there is interest, I would be willing to put it on github. </p>\n\n<p>After removing overfitting features, our final feature space had 587 features in it. Here&#8217;s a summary:</p>\n\n<p><strong>General:</strong></p>\n\n<ul>\n<li>CategoryID, parentCategoryID raw CategoryID, parentCategoryID one-hot\n(except overfitting ones) </li>\n<li>Price difference / mean</li>\n<li>Generation3probability (output from model trained to detect <br>\ngenerationmethod=3)</li>\n</ul>\n\n<p><strong>Location:</strong></p>\n\n<ul>\n<li>LocationID &amp; RegionID raw</li>\n<li>Total latitude/longtitude</li>\n<li>SameMetro, samelocation, same region etc.</li>\n<li>Distance from city centres  (kalingrad, moscow, petersburg, krasnodar, makhachkala, murmansk, perm, omsk, khabarovsk, kluichi, norilsk)</li>\n</ul>\n\n<p>Gaussian noise was added to the location features to prevent overfitting to specific locations, whilst allowing xgboost to create its own regions.</p>\n\n<p><strong>All Text:</strong></p>\n\n<ul>\n<li>Length / difference in length</li>\n<li>nGrams Features (n = 1,2,3) for title and description (Both Words and Characters)\n\n<ul><li>Count of Ngrams (#, Sum, Diff, Max, Min)</li>\n<li>Count of Unique Ngrams</li>\n<li>Ratio of Intersect Ngrams</li>\n<li>Ratio of Unique Intersect Ngrams</li></ul></li>\n<li>Distance Features:\n\n<ul><li>Jaccard , Cosine, Levenshtein and Hamming Distance between the titles and descriptions</li></ul></li>\n<li>Special Character Counting &amp; Ratio Features:\n\n<ul><li>Counting &amp; Ratio features of Capital Letters in title and description</li>\n<li>Counting &amp; Ratio features of Special Letters (digits, punctuations, etc.) in title and description</li></ul></li>\n<li>Similarity between sets of words/characters</li>\n<li>Fuzzywuzzy/jellyfish distances</li>\n<li>Number of overlapping sets of n words (n=1,2,3)</li>\n<li>Matching moving windows of strings</li>\n<li>Cross-matching columns (eg. title1 with description2)</li>\n</ul>\n\n<p><strong>Bag of words:</strong></p>\n\n<p>For each of the text columns, we created a bag of words for both the intersection of words and the difference in words and encoded these in a sparse format resulting in ~80,000 columns each. We then used this to build Naive Bayes, SGD and similar models to be used as features.</p>\n\n<p><strong>Price Features:</strong></p>\n\n<ul>\n<li>Price Ratio</li>\n<li>Is both/one price NaN</li>\n<li>Total Price</li>\n</ul>\n\n<p><strong>JSON Features:</strong></p>\n\n<ul>\n<li>Attribute Counting Features:\n\n<ul><li>Sum, diff, max, min</li></ul></li>\n<li>Count of Common Attributes Names</li>\n<li>Count of Common Attributes Values</li>\n<li>Weights of Evidence model on keys/values\nXGBoost model on sparse encoded attributes</li>\n</ul>\n\n<p><strong>Image Features:</strong></p>\n\n<ul>\n<li># of Images in each Set</li>\n<li>Difference Hashing of images</li>\n<li>Hamming distance between each pair of images</li>\n<li>Pairwise comparison of file size of each image</li>\n<li>Image dimension Features:</li>\n<li>Pairwise comparison of dimension of each image</li>\n<li>BRISK keypoint/descriptor matching</li>\n<li>Image histogram comparisons</li>\n<li>Dominant colour analysis</li>\n<li>Uniqueness of images (how many other items have the same images)</li>\n<li>Difference in number of images</li>\n</ul>\n\n<p>I found a possible image metadata leak (the creation dates of the images were embedded in the zip files) but didn&#8217;t try to use these for features - gotta play fair after all!</p>\n\n<p><strong>Clusters:</strong></p>\n\n<p>We found clusters of rows by grouping rows which contain the same items (eg. if row1 has items 123, 456 and row2 has items 456, 789 they are in the same cluster). We discovered that the size of these clusters was a very good feature (larger clusters were more likely to be non-duplicates), as well as the fact that clusters always the same generationMethod. Adding cluster-size features gave us a 0.003 to 0.004 improvement. It would be interesting to know if anyone else found these features :)</p>\n\n<h3>Feature Graveyard:</h3>\n\n<p>Overfitting was probably the biggest problem throughout the competition, and lots of features which destroyed in validation didn&#8217;t do so well on the leaderboard. This is likely because the very powerful features learn to recognise specific products or sellers that do not appear in the test set. Hence, our feature graveyard:</p>\n\n<p><strong>TF-IDF:</strong> This was something we tried very early into the competition, adapting our code from the Home Depot competition. Unfortunately, it overfitted very strongly, netting us 0.98 val-auc and only 0.89 on LB. We tried adding noise, reducing complexity etc. but in the end we gave up.</p>\n\n<p><strong>Word2vec:</strong> We tried both training a model on our cleaned data and using the pretrained model posted in the forums. We tried using word-mover distance from our model as features, but they were rather weak (0.70AUC) so in the end we decided to drop these for simplicity. Using the pre-trained model did not help, as the authors used MyStem for stemming (which is not open-source) so we could not replicate their data cleaning. After doing some transformations on the pre-trained model to try and make it work with our stemming (we got it down to about 20% missing words), it scored the same as our custom word2vec model.</p>\n\n<p><strong>Advanced cluster features:</strong> We tried to expand the gain from our cluster features in several ways. I found that taking the mean prediction for the cluster as well as cluster_size * (1-cluster_mean) provided excellent features in validation (50% of gain in xgb importance), however these overfits for reasons I still have to investigate. We also tried taking features such as the stdev of locations of items in a cluster, but these also overfitted. I suspect that Avito used slightly different methods for generating the test set.</p>\n\n<p><strong>Grammar features:</strong> We tried building features to &#8216;fingerprint&#8217; different types of sellers, such as usage of capital letters, special characters, newlines, punctuation etc. However while these helped a lot in CV, they overfitted on the leaderboard.</p>\n\n<p><strong>Brand violations:</strong> We built some features based around words that could never appear together in duplicate listings. (For example, if one item wrote &#8216;iPhone 4s&#8217; but the other one wrote &#8216;iPhone 5s&#8217;, they could not be duplicates). While they worked well at finding non-duplicates, there were just too few cases where these violations occurred to make a difference to the score.</p>\n\n<h2>Meta-model:</h2>\n\n<p>Using a meta-model improved our score by roughly 0.0015 versus our best XGBoost. In the end, a single model would have been enough to net us second place, but you can never be too careful!</p>\n\n<p>Below is an overview of what our final meta looked like. Each generation is different set of features, with later generations having more features. Note that all scores are private LB scores:</p>\n\n<p><img src=\"https://files.slack.com/files-pri/T18TSG1C3-F1QTWCZ5G/avito-meta.png?pub_secret=14957b8a3e\" alt=\"Meta-model schematic\" title></p>\n\n<p>We also tried using Extra Trees, Random Forest, Adaboost &amp; approximate kNN as base models, however these overfit to our validation-set and so they couldn&#8217;t be used in the final meta-model.</p>\n\n<hr>\n\n<p>If anyone has any questions about our solution, we would be happy to go into more depth!</p>\n\n<p><strong>Finally, a great big thank you to my teammates Peter, Sonny and Marios for making TheQuants a brilliant success!</strong> We couldn&#8217;t have done it without every one of you.</p>\n\n<p>And thanks to everyone else for making this such an enjoyable competition</p>\n\n<p>- Mikel, <br>\nTheQuants</p>",
      "rawMarkdown": "Hi everyone!\r\n\r\nHere's a description of our team's efforts for this competition. I think this competition was very much driven on feature engineering, with meta-modelling more as a finishing touch.\r\n\r\n**Our Strategy**\r\n\r\n - Understanding the Data, developing features until team merger deadline, metamodelling in the last week.\r\n - Merging early based on standing of the leaderboard at the time\r\n - Each individually building features to try and branch out as much as possible, with constant discussion about where to look next\r\n - Onboarding an experienced w̶i̶z̶a̶r̶d̶ Kaggler for the final stage. \r\n\r\n----------\r\n\r\n**Data Cleaning:**\r\n\r\nIn order to clean the text, we applied stemming using the NLTK Snowball Stemmer, and removed stopwords/punctuation as well as transforming to lowercase.\r\n\r\n**Validation Strategy:**\r\n\r\nInitially, we were using a random validation set before switching to a set of non-overlapping items, where none of the items in the valset appeared in the train set. This performed somewhat better, however we had failed to notice that the training set was ordered based on time! We later noticed this and switched to using last 33% as a valset. This set correlated relatively well with the leaderboard until the last week, when we were doing meta-modelling and it fell apart - at a point where it would be too much work to switch to a better set. This hurt us a lot towards the end of the competition.\r\n\r\n##Features:\r\nIn order to pre-emptively find over-fitting features, we built a script that looks at the changes in the properties (histograms and split purity) of a feature over time, which allowed us to quickly (200ms/feature) identify overfitting features without having to run overnight xgboost jobs. If there is interest, I would be willing to put it on github. \r\n\r\nAfter removing overfitting features, our final feature space had 587 features in it. Here’s a summary:\r\n\r\n**General:**\r\n\r\n - CategoryID, parentCategoryID raw CategoryID, parentCategoryID one-hot\r\n   (except overfitting ones) \r\n - Price difference / mean\r\n - Generation3probability (output from model trained to detect   \r\n   generationmethod=3)\r\n\r\n**Location:**\r\n\r\n - LocationID & RegionID raw\r\n - Total latitude/longtitude\r\n - SameMetro, samelocation, same region etc.\r\n - Distance from city centres  (kalingrad, moscow, petersburg, krasnodar, makhachkala, murmansk, perm, omsk, khabarovsk, kluichi, norilsk)\r\n\r\nGaussian noise was added to the location features to prevent overfitting to specific locations, whilst allowing xgboost to create its own regions.\r\n\r\n**All Text:**\r\n\r\n - Length / difference in length\r\n - nGrams Features (n = 1,2,3) for title and description (Both Words and Characters)\r\n- Count of Ngrams (#, Sum, Diff, Max, Min)\r\n- Count of Unique Ngrams\r\n- Ratio of Intersect Ngrams\r\n- Ratio of Unique Intersect Ngrams\r\n - Distance Features:\r\n- Jaccard , Cosine, Levenshtein and Hamming Distance between the titles and descriptions\r\n - Special Character Counting & Ratio Features:\r\n- Counting & Ratio features of Capital Letters in title and description\r\n- Counting & Ratio features of Special Letters (digits, punctuations, etc.) in title and description\r\n - Similarity between sets of words/characters\r\n - Fuzzywuzzy/jellyfish distances\r\n - Number of overlapping sets of n words (n=1,2,3)\r\n - Matching moving windows of strings\r\n - Cross-matching columns (eg. title1 with description2)\r\n\r\n**Bag of words:**\r\n\r\nFor each of the text columns, we created a bag of words for both the intersection of words and the difference in words and encoded these in a sparse format resulting in ~80,000 columns each. We then used this to build Naive Bayes, SGD and similar models to be used as features.\r\n\r\n**Price Features:**\r\n\r\n - Price Ratio\r\n - Is both/one price NaN\r\n - Total Price\r\n\r\n**JSON Features:**\r\n\r\n - Attribute Counting Features:\r\n- Sum, diff, max, min\r\n - Count of Common Attributes Names\r\n - Count of Common Attributes Values\r\n - Weights of Evidence model on keys/values\r\nXGBoost model on sparse encoded attributes\r\n\r\n**Image Features:**\r\n\r\n - \\# of Images in each Set\r\n - Difference Hashing of images\r\n - Hamming distance between each pair of images\r\n - Pairwise comparison of file size of each image\r\n - Image dimension Features:\r\n - Pairwise comparison of dimension of each image\r\n - BRISK keypoint/descriptor matching\r\n - Image histogram comparisons\r\n - Dominant colour analysis\r\n - Uniqueness of images (how many other items have the same images)\r\n - Difference in number of images\r\n\r\nI found a possible image metadata leak (the creation dates of the images were embedded in the zip files) but didn’t try to use these for features - gotta play fair after all!\r\n\r\n**Clusters:**\r\n\r\nWe found clusters of rows by grouping rows which contain the same items (eg. if row1 has items 123, 456 and row2 has items 456, 789 they are in the same cluster). We discovered that the size of these clusters was a very good feature (larger clusters were more likely to be non-duplicates), as well as the fact that clusters always the same generationMethod. Adding cluster-size features gave us a 0.003 to 0.004 improvement. It would be interesting to know if anyone else found these features :)\r\n\r\n###Feature Graveyard:\r\nOverfitting was probably the biggest problem throughout the competition, and lots of features which destroyed in validation didn’t do so well on the leaderboard. This is likely because the very powerful features learn to recognise specific products or sellers that do not appear in the test set. Hence, our feature graveyard:\r\n\r\n**TF-IDF:** This was something we tried very early into the competition, adapting our code from the Home Depot competition. Unfortunately, it overfitted very strongly, netting us 0.98 val-auc and only 0.89 on LB. We tried adding noise, reducing complexity etc. but in the end we gave up.\r\n\r\n**Word2vec:** We tried both training a model on our cleaned data and using the pretrained model posted in the forums. We tried using word-mover distance from our model as features, but they were rather weak (0.70AUC) so in the end we decided to drop these for simplicity. Using the pre-trained model did not help, as the authors used MyStem for stemming (which is not open-source) so we could not replicate their data cleaning. After doing some transformations on the pre-trained model to try and make it work with our stemming (we got it down to about 20% missing words), it scored the same as our custom word2vec model.\r\n\r\n**Advanced cluster features:** We tried to expand the gain from our cluster features in several ways. I found that taking the mean prediction for the cluster as well as cluster_size * (1-cluster_mean) provided excellent features in validation (50% of gain in xgb importance), however these overfits for reasons I still have to investigate. We also tried taking features such as the stdev of locations of items in a cluster, but these also overfitted. I suspect that Avito used slightly different methods for generating the test set.\r\n\r\n**Grammar features:** We tried building features to ‘fingerprint’ different types of sellers, such as usage of capital letters, special characters, newlines, punctuation etc. However while these helped a lot in CV, they overfitted on the leaderboard.\r\n\r\n**Brand violations:** We built some features based around words that could never appear together in duplicate listings. (For example, if one item wrote ‘iPhone 4s’ but the other one wrote ‘iPhone 5s’, they could not be duplicates). While they worked well at finding non-duplicates, there were just too few cases where these violations occurred to make a difference to the score.\r\n\r\n##Meta-model:\r\nUsing a meta-model improved our score by roughly 0.0015 versus our best XGBoost. In the end, a single model would have been enough to net us second place, but you can never be too careful!\r\n\r\nBelow is an overview of what our final meta looked like. Each generation is different set of features, with later generations having more features. Note that all scores are private LB scores:\r\n\r\n![Meta-model schematic][1]\r\n\r\nWe also tried using Extra Trees, Random Forest, Adaboost & approximate kNN as base models, however these overfit to our validation-set and so they couldn’t be used in the final meta-model.\r\n\r\n----------\r\n\r\nIf anyone has any questions about our solution, we would be happy to go into more depth!\r\n\r\n**Finally, a great big thank you to my teammates Peter, Sonny and Marios for making TheQuants a brilliant success!** We couldn’t have done it without every one of you.\r\n\r\nAnd thanks to everyone else for making this such an enjoyable competition\r\n\r\n \\- Mikel,  \r\nTheQuants\r\n\r\n\r\n  [1]: https://files.slack.com/files-pri/T18TSG1C3-F1QTWCZ5G/avito-meta.png?pub_secret=14957b8a3e",
      "votes": null
    },
    {
      "id": "126871",
      "postDate": "07/12/2016 19:13:58",
      "content": "<p>Thanks for sharing!</p>\n\n<blockquote>\n  <p>In order to pre-emptively find over-fitting features, we built a script that looks at the changes in the properties (histograms and split purity) of a feature over time, which allowed us to quickly (200ms/feature) identify overfitting features without having to run overnight xgboost jobs. If there is interest, I would be willing to put it on github.</p>\n</blockquote>\n\n<p>I would love to have a look at the code, please do it! </p>",
      "rawMarkdown": "Thanks for sharing!\r\n\r\n> In order to pre-emptively find over-fitting features, we built a script that looks at the changes in the properties (histograms and split purity) of a feature over time, which allowed us to quickly (200ms/feature) identify overfitting features without having to run overnight xgboost jobs. If there is interest, I would be willing to put it on github.\r\n\r\nI would love to have a look at the code, please do it!",
      "votes": null
    },
    {
      "id": "126873",
      "postDate": "07/12/2016 19:19:15",
      "content": "<p>+1 here</p>",
      "rawMarkdown": "1 here",
      "votes": null
    },
    {
      "id": "127006",
      "postDate": "07/13/2016 01:48:13",
      "content": "<p>Wait... parentCategID is not a 1:1 map with categoryID?</p>",
      "rawMarkdown": "Wait... parentCategID is not a 1:1 map with categoryID?",
      "votes": null
    },
    {
      "id": "127059",
      "postDate": "07/13/2016 07:10:17",
      "content": "<p>[quote=Snow Dog;127006]</p>\n\n<p>Wait... parentCategID is not a 1:1 map with categoryID?</p>\n\n<p>[/quote]</p>\n\n<p>It is not a 1:1 map in the sense that they are not exactly the same. While parentCat can be constructed using categoryID, in order to model a parentCategoryID, XGBoost would need to make multiple splits (one for each underlying categoryID). Hence we provide it with the one-hot parent categories so that it can model these with only one split - it can be viewed almost as a set of pre-determined trees.</p>",
      "rawMarkdown": "[quote=Snow Dog;127006]\r\n\r\nWait... parentCategID is not a 1:1 map with categoryID?\r\n\r\n[/quote]\r\n\r\nIt is not a 1:1 map in the sense that they are not exactly the same. While parentCat can be constructed using categoryID, in order to model a parentCategoryID, XGBoost would need to make multiple splits (one for each underlying categoryID). Hence we provide it with the one-hot parent categories so that it can model these with only one split - it can be viewed almost as a set of pre-determined trees.",
      "votes": null
    },
    {
      "id": "127072",
      "postDate": "07/13/2016 08:11:22",
      "content": "<p>Hello,</p>\n\n<p>Would it be possible to get the neural network architecture you used and the hyper parameter choice ? </p>\n\n<p>Thank you ! </p>",
      "rawMarkdown": "Hello,\r\n\r\nWould it be possible to get the neural network architecture you used and the hyper parameter choice ? \r\n\r\nThank you !",
      "votes": null
    },
    {
      "id": "127085",
      "postDate": "07/13/2016 09:03:00",
      "content": "<p>[quote=myouness;127072]</p>\n\n<p>Hello,</p>\n\n<p>Would it be possible to get the neural network architecture you used and the hyper parameter choice ? </p>\n\n<p>Thank you ! </p>\n\n<p>[/quote]</p>\n\n<p>Hi myouness,</p>\n\n<p>Here is the 3 layer architecture we used (with slight variation):</p>\n\n<pre><code>models = Sequential()\nmodels.add(Dense(800, input_dim=input_dim, init='uniform'))\nmodels.add(Activation('relu'))\nmodels.add(Dropout(0.6))\nmodels.add(BatchNormalization())\nmodels.add(Dense(800, init='uniform'))\nmodels.add(Activation('relu'))\nmodels.add(Dropout(0.6))\nmodels.add(BatchNormalization())\nmodels.add(Dense(400, init='uniform'))\nmodels.add(Activation('relu'))\nmodels.add(Dropout(0.6))\nmodels.add(Dense(output_dim, init='uniform'))\nmodels.add(Activation('softmax'))\nopt = optimizers.Adagrad(lr=0.01)\nmodels.compile(loss='binary_crossentropy', optimizer=opt)\n</code></pre>\n\n<p>and here is the one-layer architecture:</p>\n\n<pre><code>models = Sequential()\nmodels.add(Dense(2000, input_dim=input_dim, init='uniform', W_regularizer=l2(0.00001)))\nmodels.add(PReLU())\nmodels.add(BatchNormalization())\nmodels.add(Dropout(0.6))\nmodels.add(Dense(output_dim, init='uniform'))\nmodels.add(Activation('softmax'))\nopt = optimizers.Adagrad(lr=0.01)\nmodels.compile(loss='binary_crossentropy', optimizer=opt)\n</code></pre>\n\n<p>All our NNs were 10x bagged with different seeds, and were run for 150 epochs.</p>",
      "rawMarkdown": "[quote=myouness;127072]\r\n\r\nHello,\r\n\r\nWould it be possible to get the neural network architecture you used and the hyper parameter choice ? \r\n\r\nThank you ! \r\n\r\n[/quote]\r\n\r\nHi myouness,\r\n\r\nHere is the 3 layer architecture we used (with slight variation):\r\n\r\n    models = Sequential()\r\n    models.add(Dense(800, input_dim=input_dim, init='uniform'))\r\n    models.add(Activation('relu'))\r\n    models.add(Dropout(0.6))\r\n    models.add(BatchNormalization())\r\n    models.add(Dense(800, init='uniform'))\r\n    models.add(Activation('relu'))\r\n    models.add(Dropout(0.6))\r\n    models.add(BatchNormalization())\r\n    models.add(Dense(400, init='uniform'))\r\n    models.add(Activation('relu'))\r\n    models.add(Dropout(0.6))\r\n    models.add(Dense(output_dim, init='uniform'))\r\n    models.add(Activation('softmax'))\r\n    opt = optimizers.Adagrad(lr=0.01)\r\n    models.compile(loss='binary_crossentropy', optimizer=opt)\r\n\r\nand here is the one-layer architecture:\r\n\r\n    models = Sequential()\r\n    models.add(Dense(2000, input_dim=input_dim, init='uniform', W_regularizer=l2(0.00001)))\r\n    models.add(PReLU())\r\n    models.add(BatchNormalization())\r\n    models.add(Dropout(0.6))\r\n    models.add(Dense(output_dim, init='uniform'))\r\n    models.add(Activation('softmax'))\r\n    opt = optimizers.Adagrad(lr=0.01)\r\n    models.compile(loss='binary_crossentropy', optimizer=opt)\r\n\r\nAll our NNs were 10x bagged with different seeds, and were run for 150 epochs.",
      "votes": null
    },
    {
      "id": "127088",
      "postDate": "07/13/2016 09:22:01",
      "content": "<p>Thank you very much !</p>\n\n<p>Does it make a big difference to choose more hidden neurons in the first hidden layers than input variables ? as opposed to an under-complete but deeper NN. </p>",
      "rawMarkdown": "Thank you very much !\r\n\r\nDoes it make a big difference to choose more hidden neurons in the first hidden layers than input variables ? as opposed to an under-complete but deeper NN.",
      "votes": null
    },
    {
      "id": "127140",
      "postDate": "07/13/2016 13:26:42",
      "content": "<p>[quote=anokas;127059]</p>\n\n<p>[quote=Snow Dog;127006]</p>\n\n<p>Wait... parentCategID is not a 1:1 map with categoryID?</p>\n\n<p>[/quote]</p>\n\n<p>It is not a 1:1 map in the sense that they are not exactly the same. While parentCat can be constructed using categoryID, in order to model a parentCategoryID, XGBoost would need to make multiple splits (one for each underlying categoryID). Hence we provide it with the one-hot parent categories so that it can model these with only one split - it can be viewed almost as a set of pre-determined trees.</p>\n\n<p>[/quote]</p>\n\n<p>I just checked the data, it's not a 1:1 mapping. Different categs can belong to the same parentCateg. I remember dropping parentCateg early on thiking it was the same as categ. It's a top 5 feature in a simplified model (no image features) =/</p>",
      "rawMarkdown": "[quote=anokas;127059]\r\n\r\n[quote=Snow Dog;127006]\r\n\r\nWait... parentCategID is not a 1:1 map with categoryID?\r\n\r\n[/quote]\r\n\r\nIt is not a 1:1 map in the sense that they are not exactly the same. While parentCat can be constructed using categoryID, in order to model a parentCategoryID, XGBoost would need to make multiple splits (one for each underlying categoryID). Hence we provide it with the one-hot parent categories so that it can model these with only one split - it can be viewed almost as a set of pre-determined trees.\r\n\r\n[/quote]\r\n\r\nI just checked the data, it's not a 1:1 mapping. Different categs can belong to the same parentCateg. I remember dropping parentCateg early on thiking it was the same as categ. It's a top 5 feature in a simplified model (no image features) =/",
      "votes": null
    },
    {
      "id": "127149",
      "postDate": "07/13/2016 14:04:30",
      "content": "<p>Excellent feature engineering and model ensemble! Thank you for your sharing anokas.</p>\n\n<p>Could you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!</p>",
      "rawMarkdown": "Excellent feature engineering and model ensemble! Thank you for your sharing anokas.\r\n\r\nCould you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!",
      "votes": null
    },
    {
      "id": "127175",
      "postDate": "07/13/2016 15:27:50",
      "content": "<p>[quote=jzbdbeb;127149]</p>\n\n<p>Excellent feature engineering and model ensemble! Thank you for your sharing anokas.</p>\n\n<p>Could you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!</p>\n\n<p>[/quote]</p>\n\n<p>I will have to find an importance file later on, but I know that some of the most important features were:</p>\n\n<ul>\n<li>The cluster features</li>\n<li>BRISK</li>\n<li>dHash</li>\n<li>Bag of words</li>\n</ul>",
      "rawMarkdown": "[quote=jzbdbeb;127149]\r\n\r\nExcellent feature engineering and model ensemble! Thank you for your sharing anokas.\r\n\r\nCould you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!\r\n\r\n[/quote]\r\n\r\nI will have to find an importance file later on, but I know that some of the most important features were:\r\n\r\n - The cluster features\r\n - BRISK\r\n - dHash\r\n - Bag of words",
      "votes": null
    },
    {
      "id": "127212",
      "postDate": "07/13/2016 17:21:27",
      "content": "<p>[quote=anokas;127175]</p>\n\n<p>[quote=jzbdbeb;127149]</p>\n\n<p>Excellent feature engineering and model ensemble! Thank you for your sharing anokas.</p>\n\n<p>Could you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!</p>\n\n<p>[/quote]</p>\n\n<p>I will have to find an importance file later on, but I know that some of the most important features were:</p>\n\n<ul>\n<li>The cluster features</li>\n<li>BRISK</li>\n<li>dHash</li>\n<li>Bag of words</li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>How much RAM  your best single model peaked at?</p>",
      "rawMarkdown": "[quote=anokas;127175]\r\n\r\n[quote=jzbdbeb;127149]\r\n\r\nExcellent feature engineering and model ensemble! Thank you for your sharing anokas.\r\n\r\nCould you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!\r\n\r\n[/quote]\r\n\r\nI will have to find an importance file later on, but I know that some of the most important features were:\r\n\r\n - The cluster features\r\n - BRISK\r\n - dHash\r\n - Bag of words\r\n\r\n[/quote]\r\n\r\nHow much RAM  your best single model peaked at?",
      "votes": null
    },
    {
      "id": "127463",
      "postDate": "07/14/2016 09:31:01",
      "content": "<p>[quote=Snow Dog;127212]</p>\n\n<p>How much RAM  your best single model peaked at?</p>\n\n<p>[/quote]</p>\n\n<p>It was peaking around 70GB during xgb.DMatrix creation.\nWe would load <strong>test</strong> data only after model was built and all train data deleted.</p>",
      "rawMarkdown": "[quote=Snow Dog;127212]\r\n\r\nHow much RAM  your best single model peaked at?\r\n\r\n[/quote]\r\n\r\nIt was peaking around 70GB during xgb.DMatrix creation.\r\nWe would load **test** data only after model was built and all train data deleted.",
      "votes": null
    },
    {
      "id": "127511",
      "postDate": "07/14/2016 12:45:52",
      "content": "<p>@ anokas / The Quants:</p>\n\n<p>Can you please explain quickly how you use BRICK keypoint matching ? </p>\n\n<p>What is something like: for every image in itemID_1, find number of keypoint matchers in every image in itemID_2 and use the calculated number as a feature ?</p>",
      "rawMarkdown": "anokas / The Quants:\r\n\r\nCan you please explain quickly how you use BRICK keypoint matching ? \r\n\r\nWhat is something like: for every image in itemID_1, find number of keypoint matchers in every image in itemID_2 and use the calculated number as a feature ?",
      "votes": null
    },
    {
      "id": "127537",
      "postDate": "07/14/2016 13:50:53",
      "content": "<p>[quote=eagle4;127511]</p>\n\n<p>@ anokas / The Quants:</p>\n\n<p>Can you please explain quickly how you use BRICK keypoint matching ? </p>\n\n<p>What is something like: for every image in itemID_1, find number of keypoint matchers in every image in itemID_2 and use the calculated number as a feature ?</p>\n\n<p>[/quote]</p>\n\n<p>First, the keypoints/descriptors were extracted from every image. And then for every pair of images (x, y):</p>\n\n<ol>\n<li>For every keypoint in x, find closest keypoint in y (brute-force)</li>\n<li>Measure hamming distances between every selected keypoint pair</li>\n<li>Return mean &amp; median hamming distance of keypoint pairs to get values for those pairs of images</li>\n</ol>\n\n<p>Then for each row:</p>\n\n<ol>\n<li>Perform the above operation on every pair of images to get an array of medians and means</li>\n<li>Return minimum mean/median</li>\n<li>Return proportion of pairs where mean/median where distance was under N bits (N=30,80,100)</li>\n<li>Return (<code>min(len(images_array_x), len(images_array_y))/2</code>)th element in array of means/medians (essentially a median adjusted to the smaller set of images)</li>\n</ol>\n\n<p>Hopefully that makes a little bit of sense - there are means and medians everywhere!</p>",
      "rawMarkdown": "[quote=eagle4;127511]\r\n\r\n@ anokas / The Quants:\r\n\r\nCan you please explain quickly how you use BRICK keypoint matching ? \r\n\r\nWhat is something like: for every image in itemID_1, find number of keypoint matchers in every image in itemID_2 and use the calculated number as a feature ?\r\n\r\n[/quote]\r\n\r\nFirst, the keypoints/descriptors were extracted from every image. And then for every pair of images (x, y):\r\n\r\n 1. For every keypoint in x, find closest keypoint in y (brute-force)\r\n 2. Measure hamming distances between every selected keypoint pair\r\n 3. Return mean & median hamming distance of keypoint pairs to get values for those pairs of images\r\n\r\nThen for each row:\r\n\r\n1. Perform the above operation on every pair of images to get an array of medians and means\r\n2. Return minimum mean/median\r\n3. Return proportion of pairs where mean/median where distance was under N bits (N=30,80,100)\r\n4. Return (`min(len(images_array_x), len(images_array_y))/2`)th element in array of means/medians (essentially a median adjusted to the smaller set of images)\r\n\r\nHopefully that makes a little bit of sense - there are means and medians everywhere!",
      "votes": null
    },
    {
      "id": "127604",
      "postDate": "07/14/2016 17:44:39",
      "content": "<p>@anokas <br>\nCongratulations on great result!</p>\n\n<p>Could you please provide more details about discarding overfitting features? This is one technique which is generally useful and it would be beneficial to many of us.</p>",
      "rawMarkdown": "anokas  \r\nCongratulations on great result!\r\n\r\nCould you please provide more details about discarding overfitting features? This is one technique which is generally useful and it would be beneficial to many of us.",
      "votes": null
    },
    {
      "id": "127607",
      "postDate": "07/14/2016 17:53:04",
      "content": "<p>[quote=Ed53;127604]</p>\n\n<p>@anokas <br>\nCongratulations on great result!</p>\n\n<p>Could you please provide more details about discarding overfitting features? This is one technique which is generally useful and it would be beneficial to many of us.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for reminding me. I forgot to post it earlier as it's stuck on another machine - I'll look for it now!</p>",
      "rawMarkdown": "[quote=Ed53;127604]\r\n\r\n@anokas  \r\nCongratulations on great result!\r\n\r\nCould you please provide more details about discarding overfitting features? This is one technique which is generally useful and it would be beneficial to many of us.\r\n\r\n[/quote]\r\n\r\nThanks for reminding me. I forgot to post it earlier as it's stuck on another machine - I'll look for it now!",
      "votes": null
    },
    {
      "id": "127618",
      "postDate": "07/14/2016 19:05:56",
      "content": "<p>@Ed53 @ololo @Gerard</p>\n\n<p>I put the code and a short explanation up in a GitHub repository:</p>\n\n<p><a href=\"https://github.com/mxbi/ftim\">https://github.com/mxbi/ftim</a></p>\n\n<p>- Mikel</p>",
      "rawMarkdown": "Ed53 @ololo @Gerard\r\n\r\nI put the code and a short explanation up in a GitHub repository:\r\n\r\nhttps://github.com/mxbi/ftim\r\n\r\n\\- Mikel",
      "votes": null
    },
    {
      "id": "127780",
      "postDate": "07/15/2016 10:06:46",
      "content": "<p>Really nice work Mikel. </p>\n\n<p>If time_res does not divide the feature list into the same size chunks, the cross-checking of chunks falls over</p>",
      "rawMarkdown": "Really nice work Mikel. \r\n\r\nIf time_res does not divide the feature list into the same size chunks, the cross-checking of chunks falls over",
      "votes": null
    },
    {
      "id": "127815",
      "postDate": "07/15/2016 14:12:46",
      "content": "<p>@anokas \nThank you!</p>",
      "rawMarkdown": "anokas \r\nThank you!",
      "votes": null
    },
    {
      "id": "127903",
      "postDate": "07/16/2016 04:40:32",
      "content": "<p>Congratulations on great result!</p>\n\n<p>Could you please provide more details about &quot;CategoryID, parentCategoryID raw CategoryID, parentCategoryID one-hot (except overfitting ones)&quot;?</p>\n\n<p>I'm a little confused.</p>",
      "rawMarkdown": "Congratulations on great result!\r\n\r\nCould you please provide more details about \"CategoryID, parentCategoryID raw CategoryID, parentCategoryID one-hot (except overfitting ones)\"?\r\n\r\nI'm a little confused.",
      "votes": null
    },
    {
      "id": "128175",
      "postDate": "07/18/2016 10:49:16",
      "content": "<p>Another question,</p>\n\n<p>Could you explain the details the BOW features and models? I use similar BOW features and linear models (LR), as well as FM. But the gap between local CV and LB is huge. How did you ensemble these linear models with XGB?</p>\n\n<p>Thanks a lot!</p>",
      "rawMarkdown": "Another question,\r\n\r\nCould you explain the details the BOW features and models? I use similar BOW features and linear models (LR), as well as FM. But the gap between local CV and LB is huge. How did you ensemble these linear models with XGB?\r\n\r\nThanks a lot!",
      "votes": null
    },
    {
      "id": "142449",
      "postDate": "11/02/2016 10:40:05",
      "content": "<p>@anokas\nCongratulations on the result! I am trying to implement your solution for learning purposes. I looked at your code for detecting overfitting features, <a href=\"https://github.com/mxbi/ftim\">https://github.com/mxbi/ftim</a> and understood what it is doing. But I am unable to understand how we can decide whether a feature is overfitting just by loking at histogram intersection. I am trying to understand the logic behind this. It would be great if you could provide some explanation/intuition in understanding this. Thanks!</p>",
      "rawMarkdown": "anokas\r\nCongratulations on the result! I am trying to implement your solution for learning purposes. I looked at your code for detecting overfitting features, https://github.com/mxbi/ftim and understood what it is doing. But I am unable to understand how we can decide whether a feature is overfitting just by loking at histogram intersection. I am trying to understand the logic behind this. It would be great if you could provide some explanation/intuition in understanding this. Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 126871,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/12/2016 19:13:58",
      "content": "<p>Thanks for sharing!</p>\n\n<blockquote>\n  <p>In order to pre-emptively find over-fitting features, we built a script that looks at the changes in the properties (histograms and split purity) of a feature over time, which allowed us to quickly (200ms/feature) identify overfitting features without having to run overnight xgboost jobs. If there is interest, I would be willing to put it on github.</p>\n</blockquote>\n\n<p>I would love to have a look at the code, please do it! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126873,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/12/2016 19:19:15",
      "content": "<p>+1 here</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127006,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "07/13/2016 01:48:13",
      "content": "<p>Wait... parentCategID is not a 1:1 map with categoryID?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127059,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/13/2016 07:10:17",
      "content": "<p>[quote=Snow Dog;127006]</p>\n\n<p>Wait... parentCategID is not a 1:1 map with categoryID?</p>\n\n<p>[/quote]</p>\n\n<p>It is not a 1:1 map in the sense that they are not exactly the same. While parentCat can be constructed using categoryID, in order to model a parentCategoryID, XGBoost would need to make multiple splits (one for each underlying categoryID). Hence we provide it with the one-hot parent categories so that it can model these with only one split - it can be viewed almost as a set of pre-determined trees.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127072,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "07/13/2016 08:11:22",
      "content": "<p>Hello,</p>\n\n<p>Would it be possible to get the neural network architecture you used and the hyper parameter choice ? </p>\n\n<p>Thank you ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127085,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/13/2016 09:03:00",
      "content": "<p>[quote=myouness;127072]</p>\n\n<p>Hello,</p>\n\n<p>Would it be possible to get the neural network architecture you used and the hyper parameter choice ? </p>\n\n<p>Thank you ! </p>\n\n<p>[/quote]</p>\n\n<p>Hi myouness,</p>\n\n<p>Here is the 3 layer architecture we used (with slight variation):</p>\n\n<pre><code>models = Sequential()\nmodels.add(Dense(800, input_dim=input_dim, init='uniform'))\nmodels.add(Activation('relu'))\nmodels.add(Dropout(0.6))\nmodels.add(BatchNormalization())\nmodels.add(Dense(800, init='uniform'))\nmodels.add(Activation('relu'))\nmodels.add(Dropout(0.6))\nmodels.add(BatchNormalization())\nmodels.add(Dense(400, init='uniform'))\nmodels.add(Activation('relu'))\nmodels.add(Dropout(0.6))\nmodels.add(Dense(output_dim, init='uniform'))\nmodels.add(Activation('softmax'))\nopt = optimizers.Adagrad(lr=0.01)\nmodels.compile(loss='binary_crossentropy', optimizer=opt)\n</code></pre>\n\n<p>and here is the one-layer architecture:</p>\n\n<pre><code>models = Sequential()\nmodels.add(Dense(2000, input_dim=input_dim, init='uniform', W_regularizer=l2(0.00001)))\nmodels.add(PReLU())\nmodels.add(BatchNormalization())\nmodels.add(Dropout(0.6))\nmodels.add(Dense(output_dim, init='uniform'))\nmodels.add(Activation('softmax'))\nopt = optimizers.Adagrad(lr=0.01)\nmodels.compile(loss='binary_crossentropy', optimizer=opt)\n</code></pre>\n\n<p>All our NNs were 10x bagged with different seeds, and were run for 150 epochs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127088,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "07/13/2016 09:22:01",
      "content": "<p>Thank you very much !</p>\n\n<p>Does it make a big difference to choose more hidden neurons in the first hidden layers than input variables ? as opposed to an under-complete but deeper NN. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127140,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "07/13/2016 13:26:42",
      "content": "<p>[quote=anokas;127059]</p>\n\n<p>[quote=Snow Dog;127006]</p>\n\n<p>Wait... parentCategID is not a 1:1 map with categoryID?</p>\n\n<p>[/quote]</p>\n\n<p>It is not a 1:1 map in the sense that they are not exactly the same. While parentCat can be constructed using categoryID, in order to model a parentCategoryID, XGBoost would need to make multiple splits (one for each underlying categoryID). Hence we provide it with the one-hot parent categories so that it can model these with only one split - it can be viewed almost as a set of pre-determined trees.</p>\n\n<p>[/quote]</p>\n\n<p>I just checked the data, it's not a 1:1 mapping. Different categs can belong to the same parentCateg. I remember dropping parentCateg early on thiking it was the same as categ. It's a top 5 feature in a simplified model (no image features) =/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127149,
      "author_name": "jzbjyb",
      "author_url": "",
      "post_date": "07/13/2016 14:04:30",
      "content": "<p>Excellent feature engineering and model ensemble! Thank you for your sharing anokas.</p>\n\n<p>Could you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127175,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/13/2016 15:27:50",
      "content": "<p>[quote=jzbdbeb;127149]</p>\n\n<p>Excellent feature engineering and model ensemble! Thank you for your sharing anokas.</p>\n\n<p>Could you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!</p>\n\n<p>[/quote]</p>\n\n<p>I will have to find an importance file later on, but I know that some of the most important features were:</p>\n\n<ul>\n<li>The cluster features</li>\n<li>BRISK</li>\n<li>dHash</li>\n<li>Bag of words</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127212,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "07/13/2016 17:21:27",
      "content": "<p>[quote=anokas;127175]</p>\n\n<p>[quote=jzbdbeb;127149]</p>\n\n<p>Excellent feature engineering and model ensemble! Thank you for your sharing anokas.</p>\n\n<p>Could you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!</p>\n\n<p>[/quote]</p>\n\n<p>I will have to find an importance file later on, but I know that some of the most important features were:</p>\n\n<ul>\n<li>The cluster features</li>\n<li>BRISK</li>\n<li>dHash</li>\n<li>Bag of words</li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>How much RAM  your best single model peaked at?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127463,
      "author_name": "sonnylaskar",
      "author_url": "",
      "post_date": "07/14/2016 09:31:01",
      "content": "<p>[quote=Snow Dog;127212]</p>\n\n<p>How much RAM  your best single model peaked at?</p>\n\n<p>[/quote]</p>\n\n<p>It was peaking around 70GB during xgb.DMatrix creation.\nWe would load <strong>test</strong> data only after model was built and all train data deleted.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127511,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "07/14/2016 12:45:52",
      "content": "<p>@ anokas / The Quants:</p>\n\n<p>Can you please explain quickly how you use BRICK keypoint matching ? </p>\n\n<p>What is something like: for every image in itemID_1, find number of keypoint matchers in every image in itemID_2 and use the calculated number as a feature ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127537,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/14/2016 13:50:53",
      "content": "<p>[quote=eagle4;127511]</p>\n\n<p>@ anokas / The Quants:</p>\n\n<p>Can you please explain quickly how you use BRICK keypoint matching ? </p>\n\n<p>What is something like: for every image in itemID_1, find number of keypoint matchers in every image in itemID_2 and use the calculated number as a feature ?</p>\n\n<p>[/quote]</p>\n\n<p>First, the keypoints/descriptors were extracted from every image. And then for every pair of images (x, y):</p>\n\n<ol>\n<li>For every keypoint in x, find closest keypoint in y (brute-force)</li>\n<li>Measure hamming distances between every selected keypoint pair</li>\n<li>Return mean &amp; median hamming distance of keypoint pairs to get values for those pairs of images</li>\n</ol>\n\n<p>Then for each row:</p>\n\n<ol>\n<li>Perform the above operation on every pair of images to get an array of medians and means</li>\n<li>Return minimum mean/median</li>\n<li>Return proportion of pairs where mean/median where distance was under N bits (N=30,80,100)</li>\n<li>Return (<code>min(len(images_array_x), len(images_array_y))/2</code>)th element in array of means/medians (essentially a median adjusted to the smaller set of images)</li>\n</ol>\n\n<p>Hopefully that makes a little bit of sense - there are means and medians everywhere!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127604,
      "author_name": "hauserquaid",
      "author_url": "",
      "post_date": "07/14/2016 17:44:39",
      "content": "<p>@anokas <br>\nCongratulations on great result!</p>\n\n<p>Could you please provide more details about discarding overfitting features? This is one technique which is generally useful and it would be beneficial to many of us.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127607,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/14/2016 17:53:04",
      "content": "<p>[quote=Ed53;127604]</p>\n\n<p>@anokas <br>\nCongratulations on great result!</p>\n\n<p>Could you please provide more details about discarding overfitting features? This is one technique which is generally useful and it would be beneficial to many of us.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for reminding me. I forgot to post it earlier as it's stuck on another machine - I'll look for it now!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127618,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/14/2016 19:05:56",
      "content": "<p>@Ed53 @ololo @Gerard</p>\n\n<p>I put the code and a short explanation up in a GitHub repository:</p>\n\n<p><a href=\"https://github.com/mxbi/ftim\">https://github.com/mxbi/ftim</a></p>\n\n<p>- Mikel</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127780,
      "author_name": "tracknut",
      "author_url": "",
      "post_date": "07/15/2016 10:06:46",
      "content": "<p>Really nice work Mikel. </p>\n\n<p>If time_res does not divide the feature list into the same size chunks, the cross-checking of chunks falls over</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127815,
      "author_name": "hauserquaid",
      "author_url": "",
      "post_date": "07/15/2016 14:12:46",
      "content": "<p>@anokas \nThank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127903,
      "author_name": "",
      "author_url": "",
      "post_date": "07/16/2016 04:40:32",
      "content": "<p>Congratulations on great result!</p>\n\n<p>Could you please provide more details about &quot;CategoryID, parentCategoryID raw CategoryID, parentCategoryID one-hot (except overfitting ones)&quot;?</p>\n\n<p>I'm a little confused.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 128175,
      "author_name": "jzbjyb",
      "author_url": "",
      "post_date": "07/18/2016 10:49:16",
      "content": "<p>Another question,</p>\n\n<p>Could you explain the details the BOW features and models? I use similar BOW features and linear models (LR), as well as FM. But the gap between local CV and LB is huge. How did you ensemble these linear models with XGB?</p>\n\n<p>Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 142449,
      "author_name": "rmr949",
      "author_url": "",
      "post_date": "11/02/2016 10:40:05",
      "content": "<p>@anokas\nCongratulations on the result! I am trying to implement your solution for learning purposes. I looked at your code for detecting overfitting features, <a href=\"https://github.com/mxbi/ftim\">https://github.com/mxbi/ftim</a> and understood what it is doing. But I am unable to understand how we can decide whether a feature is overfitting just by loking at histogram intersection. I am trying to understand the logic behind this. It would be great if you could provide some explanation/intuition in understanding this. Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "126869": "Hi everyone!\r\n\r\nHere's a description of our team's efforts for this competition. I think this competition was very much driven on feature engineering, with meta-modelling more as a finishing touch.\r\n\r\n**Our Strategy**\r\n\r\n - Understanding the Data, developing features until team merger deadline, metamodelling in the last week.\r\n - Merging early based on standing of the leaderboard at the time\r\n - Each individually building features to try and branch out as much as possible, with constant discussion about where to look next\r\n - Onboarding an experienced w̶i̶z̶a̶r̶d̶ Kaggler for the final stage. \r\n\r\n----------\r\n\r\n**Data Cleaning:**\r\n\r\nIn order to clean the text, we applied stemming using the NLTK Snowball Stemmer, and removed stopwords/punctuation as well as transforming to lowercase.\r\n\r\n**Validation Strategy:**\r\n\r\nInitially, we were using a random validation set before switching to a set of non-overlapping items, where none of the items in the valset appeared in the train set. This performed somewhat better, however we had failed to notice that the training set was ordered based on time! We later noticed this and switched to using last 33% as a valset. This set correlated relatively well with the leaderboard until the last week, when we were doing meta-modelling and it fell apart - at a point where it would be too much work to switch to a better set. This hurt us a lot towards the end of the competition.\r\n\r\n##Features:\r\nIn order to pre-emptively find over-fitting features, we built a script that looks at the changes in the properties (histograms and split purity) of a feature over time, which allowed us to quickly (200ms/feature) identify overfitting features without having to run overnight xgboost jobs. If there is interest, I would be willing to put it on github. \r\n\r\nAfter removing overfitting features, our final feature space had 587 features in it. Here’s a summary:\r\n\r\n**General:**\r\n\r\n - CategoryID, parentCategoryID raw CategoryID, parentCategoryID one-hot\r\n   (except overfitting ones) \r\n - Price difference / mean\r\n - Generation3probability (output from model trained to detect   \r\n   generationmethod=3)\r\n\r\n**Location:**\r\n\r\n - LocationID & RegionID raw\r\n - Total latitude/longtitude\r\n - SameMetro, samelocation, same region etc.\r\n - Distance from city centres  (kalingrad, moscow, petersburg, krasnodar, makhachkala, murmansk, perm, omsk, khabarovsk, kluichi, norilsk)\r\n\r\nGaussian noise was added to the location features to prevent overfitting to specific locations, whilst allowing xgboost to create its own regions.\r\n\r\n**All Text:**\r\n\r\n - Length / difference in length\r\n - nGrams Features (n = 1,2,3) for title and description (Both Words and Characters)\r\n- Count of Ngrams (#, Sum, Diff, Max, Min)\r\n- Count of Unique Ngrams\r\n- Ratio of Intersect Ngrams\r\n- Ratio of Unique Intersect Ngrams\r\n - Distance Features:\r\n- Jaccard , Cosine, Levenshtein and Hamming Distance between the titles and descriptions\r\n - Special Character Counting & Ratio Features:\r\n- Counting & Ratio features of Capital Letters in title and description\r\n- Counting & Ratio features of Special Letters (digits, punctuations, etc.) in title and description\r\n - Similarity between sets of words/characters\r\n - Fuzzywuzzy/jellyfish distances\r\n - Number of overlapping sets of n words (n=1,2,3)\r\n - Matching moving windows of strings\r\n - Cross-matching columns (eg. title1 with description2)\r\n\r\n**Bag of words:**\r\n\r\nFor each of the text columns, we created a bag of words for both the intersection of words and the difference in words and encoded these in a sparse format resulting in ~80,000 columns each. We then used this to build Naive Bayes, SGD and similar models to be used as features.\r\n\r\n**Price Features:**\r\n\r\n - Price Ratio\r\n - Is both/one price NaN\r\n - Total Price\r\n\r\n**JSON Features:**\r\n\r\n - Attribute Counting Features:\r\n- Sum, diff, max, min\r\n - Count of Common Attributes Names\r\n - Count of Common Attributes Values\r\n - Weights of Evidence model on keys/values\r\nXGBoost model on sparse encoded attributes\r\n\r\n**Image Features:**\r\n\r\n - \\# of Images in each Set\r\n - Difference Hashing of images\r\n - Hamming distance between each pair of images\r\n - Pairwise comparison of file size of each image\r\n - Image dimension Features:\r\n - Pairwise comparison of dimension of each image\r\n - BRISK keypoint/descriptor matching\r\n - Image histogram comparisons\r\n - Dominant colour analysis\r\n - Uniqueness of images (how many other items have the same images)\r\n - Difference in number of images\r\n\r\nI found a possible image metadata leak (the creation dates of the images were embedded in the zip files) but didn’t try to use these for features - gotta play fair after all!\r\n\r\n**Clusters:**\r\n\r\nWe found clusters of rows by grouping rows which contain the same items (eg. if row1 has items 123, 456 and row2 has items 456, 789 they are in the same cluster). We discovered that the size of these clusters was a very good feature (larger clusters were more likely to be non-duplicates), as well as the fact that clusters always the same generationMethod. Adding cluster-size features gave us a 0.003 to 0.004 improvement. It would be interesting to know if anyone else found these features :)\r\n\r\n###Feature Graveyard:\r\nOverfitting was probably the biggest problem throughout the competition, and lots of features which destroyed in validation didn’t do so well on the leaderboard. This is likely because the very powerful features learn to recognise specific products or sellers that do not appear in the test set. Hence, our feature graveyard:\r\n\r\n**TF-IDF:** This was something we tried very early into the competition, adapting our code from the Home Depot competition. Unfortunately, it overfitted very strongly, netting us 0.98 val-auc and only 0.89 on LB. We tried adding noise, reducing complexity etc. but in the end we gave up.\r\n\r\n**Word2vec:** We tried both training a model on our cleaned data and using the pretrained model posted in the forums. We tried using word-mover distance from our model as features, but they were rather weak (0.70AUC) so in the end we decided to drop these for simplicity. Using the pre-trained model did not help, as the authors used MyStem for stemming (which is not open-source) so we could not replicate their data cleaning. After doing some transformations on the pre-trained model to try and make it work with our stemming (we got it down to about 20% missing words), it scored the same as our custom word2vec model.\r\n\r\n**Advanced cluster features:** We tried to expand the gain from our cluster features in several ways. I found that taking the mean prediction for the cluster as well as cluster_size * (1-cluster_mean) provided excellent features in validation (50% of gain in xgb importance), however these overfits for reasons I still have to investigate. We also tried taking features such as the stdev of locations of items in a cluster, but these also overfitted. I suspect that Avito used slightly different methods for generating the test set.\r\n\r\n**Grammar features:** We tried building features to ‘fingerprint’ different types of sellers, such as usage of capital letters, special characters, newlines, punctuation etc. However while these helped a lot in CV, they overfitted on the leaderboard.\r\n\r\n**Brand violations:** We built some features based around words that could never appear together in duplicate listings. (For example, if one item wrote ‘iPhone 4s’ but the other one wrote ‘iPhone 5s’, they could not be duplicates). While they worked well at finding non-duplicates, there were just too few cases where these violations occurred to make a difference to the score.\r\n\r\n##Meta-model:\r\nUsing a meta-model improved our score by roughly 0.0015 versus our best XGBoost. In the end, a single model would have been enough to net us second place, but you can never be too careful!\r\n\r\nBelow is an overview of what our final meta looked like. Each generation is different set of features, with later generations having more features. Note that all scores are private LB scores:\r\n\r\n![Meta-model schematic][1]\r\n\r\nWe also tried using Extra Trees, Random Forest, Adaboost & approximate kNN as base models, however these overfit to our validation-set and so they couldn’t be used in the final meta-model.\r\n\r\n----------\r\n\r\nIf anyone has any questions about our solution, we would be happy to go into more depth!\r\n\r\n**Finally, a great big thank you to my teammates Peter, Sonny and Marios for making TheQuants a brilliant success!** We couldn’t have done it without every one of you.\r\n\r\nAnd thanks to everyone else for making this such an enjoyable competition\r\n\r\n \\- Mikel,  \r\nTheQuants\r\n\r\n\r\n  [1]: https://files.slack.com/files-pri/T18TSG1C3-F1QTWCZ5G/avito-meta.png?pub_secret=14957b8a3e",
    "126871": "Thanks for sharing!\r\n\r\n> In order to pre-emptively find over-fitting features, we built a script that looks at the changes in the properties (histograms and split purity) of a feature over time, which allowed us to quickly (200ms/feature) identify overfitting features without having to run overnight xgboost jobs. If there is interest, I would be willing to put it on github.\r\n\r\nI would love to have a look at the code, please do it!",
    "126873": "1 here",
    "127006": "Wait... parentCategID is not a 1:1 map with categoryID?",
    "127059": "[quote=Snow Dog;127006]\r\n\r\nWait... parentCategID is not a 1:1 map with categoryID?\r\n\r\n[/quote]\r\n\r\nIt is not a 1:1 map in the sense that they are not exactly the same. While parentCat can be constructed using categoryID, in order to model a parentCategoryID, XGBoost would need to make multiple splits (one for each underlying categoryID). Hence we provide it with the one-hot parent categories so that it can model these with only one split - it can be viewed almost as a set of pre-determined trees.",
    "127072": "Hello,\r\n\r\nWould it be possible to get the neural network architecture you used and the hyper parameter choice ? \r\n\r\nThank you !",
    "127085": "[quote=myouness;127072]\r\n\r\nHello,\r\n\r\nWould it be possible to get the neural network architecture you used and the hyper parameter choice ? \r\n\r\nThank you ! \r\n\r\n[/quote]\r\n\r\nHi myouness,\r\n\r\nHere is the 3 layer architecture we used (with slight variation):\r\n\r\n    models = Sequential()\r\n    models.add(Dense(800, input_dim=input_dim, init='uniform'))\r\n    models.add(Activation('relu'))\r\n    models.add(Dropout(0.6))\r\n    models.add(BatchNormalization())\r\n    models.add(Dense(800, init='uniform'))\r\n    models.add(Activation('relu'))\r\n    models.add(Dropout(0.6))\r\n    models.add(BatchNormalization())\r\n    models.add(Dense(400, init='uniform'))\r\n    models.add(Activation('relu'))\r\n    models.add(Dropout(0.6))\r\n    models.add(Dense(output_dim, init='uniform'))\r\n    models.add(Activation('softmax'))\r\n    opt = optimizers.Adagrad(lr=0.01)\r\n    models.compile(loss='binary_crossentropy', optimizer=opt)\r\n\r\nand here is the one-layer architecture:\r\n\r\n    models = Sequential()\r\n    models.add(Dense(2000, input_dim=input_dim, init='uniform', W_regularizer=l2(0.00001)))\r\n    models.add(PReLU())\r\n    models.add(BatchNormalization())\r\n    models.add(Dropout(0.6))\r\n    models.add(Dense(output_dim, init='uniform'))\r\n    models.add(Activation('softmax'))\r\n    opt = optimizers.Adagrad(lr=0.01)\r\n    models.compile(loss='binary_crossentropy', optimizer=opt)\r\n\r\nAll our NNs were 10x bagged with different seeds, and were run for 150 epochs.",
    "127088": "Thank you very much !\r\n\r\nDoes it make a big difference to choose more hidden neurons in the first hidden layers than input variables ? as opposed to an under-complete but deeper NN.",
    "127140": "[quote=anokas;127059]\r\n\r\n[quote=Snow Dog;127006]\r\n\r\nWait... parentCategID is not a 1:1 map with categoryID?\r\n\r\n[/quote]\r\n\r\nIt is not a 1:1 map in the sense that they are not exactly the same. While parentCat can be constructed using categoryID, in order to model a parentCategoryID, XGBoost would need to make multiple splits (one for each underlying categoryID). Hence we provide it with the one-hot parent categories so that it can model these with only one split - it can be viewed almost as a set of pre-determined trees.\r\n\r\n[/quote]\r\n\r\nI just checked the data, it's not a 1:1 mapping. Different categs can belong to the same parentCateg. I remember dropping parentCateg early on thiking it was the same as categ. It's a top 5 feature in a simplified model (no image features) =/",
    "127149": "Excellent feature engineering and model ensemble! Thank you for your sharing anokas.\r\n\r\nCould you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!",
    "127175": "[quote=jzbdbeb;127149]\r\n\r\nExcellent feature engineering and model ensemble! Thank you for your sharing anokas.\r\n\r\nCould you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!\r\n\r\n[/quote]\r\n\r\nI will have to find an importance file later on, but I know that some of the most important features were:\r\n\r\n - The cluster features\r\n - BRISK\r\n - dHash\r\n - Bag of words",
    "127212": "[quote=anokas;127175]\r\n\r\n[quote=jzbdbeb;127149]\r\n\r\nExcellent feature engineering and model ensemble! Thank you for your sharing anokas.\r\n\r\nCould you give us an introduction of the importance of these features? Which features make your submission improve a lot? Maybe the importance output by XGB. Thanks a lot!\r\n\r\n[/quote]\r\n\r\nI will have to find an importance file later on, but I know that some of the most important features were:\r\n\r\n - The cluster features\r\n - BRISK\r\n - dHash\r\n - Bag of words\r\n\r\n[/quote]\r\n\r\nHow much RAM  your best single model peaked at?",
    "127463": "[quote=Snow Dog;127212]\r\n\r\nHow much RAM  your best single model peaked at?\r\n\r\n[/quote]\r\n\r\nIt was peaking around 70GB during xgb.DMatrix creation.\r\nWe would load **test** data only after model was built and all train data deleted.",
    "127511": "anokas / The Quants:\r\n\r\nCan you please explain quickly how you use BRICK keypoint matching ? \r\n\r\nWhat is something like: for every image in itemID_1, find number of keypoint matchers in every image in itemID_2 and use the calculated number as a feature ?",
    "127537": "[quote=eagle4;127511]\r\n\r\n@ anokas / The Quants:\r\n\r\nCan you please explain quickly how you use BRICK keypoint matching ? \r\n\r\nWhat is something like: for every image in itemID_1, find number of keypoint matchers in every image in itemID_2 and use the calculated number as a feature ?\r\n\r\n[/quote]\r\n\r\nFirst, the keypoints/descriptors were extracted from every image. And then for every pair of images (x, y):\r\n\r\n 1. For every keypoint in x, find closest keypoint in y (brute-force)\r\n 2. Measure hamming distances between every selected keypoint pair\r\n 3. Return mean & median hamming distance of keypoint pairs to get values for those pairs of images\r\n\r\nThen for each row:\r\n\r\n1. Perform the above operation on every pair of images to get an array of medians and means\r\n2. Return minimum mean/median\r\n3. Return proportion of pairs where mean/median where distance was under N bits (N=30,80,100)\r\n4. Return (`min(len(images_array_x), len(images_array_y))/2`)th element in array of means/medians (essentially a median adjusted to the smaller set of images)\r\n\r\nHopefully that makes a little bit of sense - there are means and medians everywhere!",
    "127604": "anokas  \r\nCongratulations on great result!\r\n\r\nCould you please provide more details about discarding overfitting features? This is one technique which is generally useful and it would be beneficial to many of us.",
    "127607": "[quote=Ed53;127604]\r\n\r\n@anokas  \r\nCongratulations on great result!\r\n\r\nCould you please provide more details about discarding overfitting features? This is one technique which is generally useful and it would be beneficial to many of us.\r\n\r\n[/quote]\r\n\r\nThanks for reminding me. I forgot to post it earlier as it's stuck on another machine - I'll look for it now!",
    "127618": "Ed53 @ololo @Gerard\r\n\r\nI put the code and a short explanation up in a GitHub repository:\r\n\r\nhttps://github.com/mxbi/ftim\r\n\r\n\\- Mikel",
    "127780": "Really nice work Mikel. \r\n\r\nIf time_res does not divide the feature list into the same size chunks, the cross-checking of chunks falls over",
    "127815": "anokas \r\nThank you!",
    "127903": "Congratulations on great result!\r\n\r\nCould you please provide more details about \"CategoryID, parentCategoryID raw CategoryID, parentCategoryID one-hot (except overfitting ones)\"?\r\n\r\nI'm a little confused.",
    "128175": "Another question,\r\n\r\nCould you explain the details the BOW features and models? I use similar BOW features and linear models (LR), as well as FM. But the gap between local CV and LB is huge. How did you ensemble these linear models with XGB?\r\n\r\nThanks a lot!",
    "142449": "anokas\r\nCongratulations on the result! I am trying to implement your solution for learning purposes. I looked at your code for detecting overfitting features, https://github.com/mxbi/ftim and understood what it is doing. But I am unable to understand how we can decide whether a feature is overfitting just by loking at histogram intersection. I am trying to understand the logic behind this. It would be great if you could provide some explanation/intuition in understanding this. Thanks!"
  },
  "source": "meta"
}