{
  "id": 22190,
  "title": "Solution sharing ",
  "url": "/competitions/avito-duplicate-ads-detection/writeups/ololobhi-solution-sharing",
  "author_name": "",
  "post_date": "2016-07-12T07:45:40.810Z",
  "votes": 5,
  "comment_count": 36,
  "views": 3830,
  "content": "<p>So now when the competition is over I'm very curious what others were doing. So let's discuss the solutions! </p>",
  "messages": [
    {
      "id": "126770",
      "postDate": "07/12/2016 07:45:40",
      "content": "<p>So now when the competition is over I'm very curious what others were doing. So let's discuss the solutions! </p>",
      "rawMarkdown": "So now when the competition is over I'm very curious what others were doing. So let's discuss the solutions!",
      "votes": null
    },
    {
      "id": "126772",
      "postDate": "07/12/2016 08:41:25",
      "content": "<p>So here is my solution:</p>\n\n<h3>Simple features</h3>\n\n<ul>\n<li>category id</li>\n<li>number of images </li>\n<li>price difference</li>\n</ul>\n\n<h3>Simple text features</h3>\n\n<ul>\n<li>number of Russian and English characters in the title and description and number of digits and non alphanumeric characters; their ratio to all the text</li>\n<li>len of title and description</li>\n<li>number of unique chars </li>\n<li>cosine and jaccard of 2-3-4 ngrams on the char level</li>\n<li>fuzzy string distances calculated with FuzzyWuzzy (<a href=\"https://github.com/seatgeek/fuzzywuzzy\">https://github.com/seatgeek/fuzzywuzzy</a>)</li>\n</ul>\n\n<h3>Simple image features</h3>\n\n<ul>\n<li>stats (min, mean, max, std, skew, kurtosis) of each channel (R, G, B) as well as the average of all 3 channels </li>\n<li>file size</li>\n<li>number of geometry matches </li>\n<li>number of exact matches (calculated by md5)</li>\n</ul>\n\n<h3>Simple Geo features</h3>\n\n<ul>\n<li>metro id, location id</li>\n<li>distance between two locations</li>\n</ul>\n\n<p>With these features I was able to beat the avito benchmark </p>\n\n<p>Other features that I included afterwards: </p>\n\n<h3>Attributes</h3>\n\n<ul>\n<li>number of key matches, number of value matches, number of key-value pair matches</li>\n<li>number of fields that both ads didn't fill</li>\n<li>similarity of pairs in the tf-idf space, also svd of this space</li>\n</ul>\n\n<h3>Text features</h3>\n\n<ul>\n<li>jaccard and cosine only on digits and English tokens </li>\n<li>if any of the ads have english chars in a russian word (some of the characters look the same, but have different codes)</li>\n<li>tf, tf-idf and bm25 of title, description and all text</li>\n<li>svd of the above </li>\n<li>tf only on words that both ads have in common (in title, desc, all text), tf on words that only one of the ad has, svd of them</li>\n<li>distances and similarities in word2vec and glove spaces </li>\n<li>word2vec and glove similarity only on nouns</li>\n<li>some variation of word's mover distance for both w2v and glove </li>\n<li>how similar misspellings are in the ads </li>\n</ul>\n\n<p>I also tried to extract contact information (phones, emails, etc) but it didn't help much</p>\n\n<h3>Image features</h3>\n\n<ul>\n<li>image hashes from the imagehash library and from the forums </li>\n<li>phash from imagemagick computed on all individual channels </li>\n<li>diffirent similarities and distances of image histograms </li>\n<li>centroids and image moment invariants computed with imagemagick</li>\n<li>structural similarity of images</li>\n<li>SIFT and keypoint matching </li>\n</ul>\n\n<h3>Geo features</h3>\n\n<ul>\n<li>city, region and zip code extracted from geolocation (same sity, region, etc)</li>\n<li>dbscan and kmeans clusters of geolocations</li>\n</ul>\n\n<h3>Meta features</h3>\n\n<ul>\n<li>I run PCA and SVM on all the feature groups and used them as meta features</li>\n</ul>\n\n<h3>Ensembling</h3>\n\n<p>I computed too many features and it was not possible to fit them all into RAM (I have a 32gb/8cores machine) so I started ensembling pretty early. I randomly picked up 100-150 features and run XGBs or ETs on them. </p>\n\n<p>I mostly trained ETs because I could do ~10 of them per day, while training XGB took 2-3 days. </p>\n\n<p>My best model scored ~0.939 on the public LB (xgb with 2.5k trees). The best submission before merging with Abhishek was 0.94437 public / 0.94394 private which would place me at the 10th position. With Abhishek we reached the 5th. </p>\n\n<h3>Code</h3>\n\n<p>The code is available at <a href=\"https://github.com/alexeygrigorev/avito-duplicates-kaggle\">https://github.com/alexeygrigorev/avito-duplicates-kaggle</a> </p>\n\n<p>Thanks everybody, it was very fun! Looking forward to seeing you in the next competitions. </p>",
      "rawMarkdown": "So here is my solution:\r\n\r\n\r\n### Simple features \r\n\r\n- category id\r\n- number of images \r\n- price difference\r\n\r\n\r\n### Simple text features \r\n\r\n- number of Russian and English characters in the title and description and number of digits and non alphanumeric characters; their ratio to all the text\r\n- len of title and description\r\n- number of unique chars \r\n- cosine and jaccard of 2-3-4 ngrams on the char level\r\n- fuzzy string distances calculated with FuzzyWuzzy (https://github.com/seatgeek/fuzzywuzzy)\r\n\r\n### Simple image features \r\n\r\n- stats (min, mean, max, std, skew, kurtosis) of each channel (R, G, B) as well as the average of all 3 channels \r\n- file size\r\n- number of geometry matches \r\n- number of exact matches (calculated by md5)\r\n\r\n### Simple Geo features\r\n\r\n- metro id, location id\r\n- distance between two locations\r\n\r\nWith these features I was able to beat the avito benchmark \r\n\r\n\r\nOther features that I included afterwards: \r\n\r\n\r\n### Attributes \r\n\r\n- number of key matches, number of value matches, number of key-value pair matches\r\n- number of fields that both ads didn't fill\r\n- similarity of pairs in the tf-idf space, also svd of this space\r\n\r\n\r\n### Text features \r\n\r\n- jaccard and cosine only on digits and English tokens \r\n- if any of the ads have english chars in a russian word (some of the characters look the same, but have different codes)\r\n- tf, tf-idf and bm25 of title, description and all text\r\n- svd of the above \r\n- tf only on words that both ads have in common (in title, desc, all text), tf on words that only one of the ad has, svd of them\r\n- distances and similarities in word2vec and glove spaces \r\n- word2vec and glove similarity only on nouns\r\n- some variation of word's mover distance for both w2v and glove \r\n- how similar misspellings are in the ads \r\n\r\nI also tried to extract contact information (phones, emails, etc) but it didn't help much\r\n\r\n\r\n### Image features \r\n\r\n- image hashes from the imagehash library and from the forums \r\n- phash from imagemagick computed on all individual channels \r\n- diffirent similarities and distances of image histograms \r\n- centroids and image moment invariants computed with imagemagick\r\n- structural similarity of images\r\n- SIFT and keypoint matching \r\n\r\n\r\n### Geo features \r\n\r\n- city, region and zip code extracted from geolocation (same sity, region, etc)\r\n- dbscan and kmeans clusters of geolocations\r\n\r\n\r\n### Meta features\r\n\r\n- I run PCA and SVM on all the feature groups and used them as meta features\r\n\r\n\r\n### Ensembling \r\n\r\nI computed too many features and it was not possible to fit them all into RAM (I have a 32gb/8cores machine) so I started ensembling pretty early. I randomly picked up 100-150 features and run XGBs or ETs on them. \r\n\r\nI mostly trained ETs because I could do ~10 of them per day, while training XGB took 2-3 days. \r\n\r\nMy best model scored ~0.939 on the public LB (xgb with 2.5k trees). The best submission before merging with Abhishek was 0.94437 public / 0.94394 private which would place me at the 10th position. With Abhishek we reached the 5th. \r\n\r\n\r\n### Code\r\n\r\nThe code is available at https://github.com/alexeygrigorev/avito-duplicates-kaggle \r\n\r\nThanks everybody, it was very fun! Looking forward to seeing you in the next competitions.",
      "votes": null
    },
    {
      "id": "126780",
      "postDate": "07/12/2016 09:30:38",
      "content": "<p>Great post Ololo. Thanks for sharing. Our features were similar - obviously a subset. Couple of queries</p>\n\n<p>1) how similar misspellings are in the ads. What did you use to get this information ?</p>\n\n<p>2) So I gather you used geolocation to get City and Zip code and Region ? Did you use a license for google maps api ? \nIf you already have location id and metro id as features - did this additional information help ? </p>\n\n<p>Fantastic feature engineering work.</p>",
      "rawMarkdown": "Great post Ololo. Thanks for sharing. Our features were similar - obviously a subset. Couple of queries\r\n\r\n1) how similar misspellings are in the ads. What did you use to get this information ?\r\n\r\n2) So I gather you used geolocation to get City and Zip code and Region ? Did you use a license for google maps api ? \r\nIf you already have location id and metro id as features - did this additional information help ? \r\n\r\nFantastic feature engineering work.",
      "votes": null
    },
    {
      "id": "126781",
      "postDate": "07/12/2016 09:36:10",
      "content": "<p>[quote=Run2;126780]</p>\n\n<p>Great post Ololo. Thanks for sharing. Our features were similar - obviously a subset. Couple of queries</p>\n\n<p>1) how similar misspellings are in the ads. What did you use to get this information ?</p>\n\n<p>2) So I gather you used geolocation to get City and Zip code and Region ? Did you use a license for google maps api ? \nIf you already have location id and metro id as features - did this additional information help ? </p>\n\n<p>Fantastic feature engineering work.</p>\n\n<p>[/quote]</p>\n\n<p>The answer to both of the questions is here: <a href=\"https://www.kaggle.com/c/avito-duplicate-ads-detection/forums/t/20759/external-data-thread/124803#post124803\">https://www.kaggle.com/c/avito-duplicate-ads-detection/forums/t/20759/external-data-thread/124803#post124803</a></p>\n\n<p>I set up a nominatim instance with a copy of open street map of Russia on my laptop (the server was busy doing other stuff). Then I run all (lat, lon) pairs through it, it took several days. But later I noticed that there are not so many distinct pairs, so it could be done in less than a day. </p>",
      "rawMarkdown": "[quote=Run2;126780]\r\n\r\nGreat post Ololo. Thanks for sharing. Our features were similar - obviously a subset. Couple of queries\r\n\r\n1) how similar misspellings are in the ads. What did you use to get this information ?\r\n\r\n2) So I gather you used geolocation to get City and Zip code and Region ? Did you use a license for google maps api ? \r\nIf you already have location id and metro id as features - did this additional information help ? \r\n\r\nFantastic feature engineering work.\r\n\r\n[/quote]\r\n\r\nThe answer to both of the questions is here: https://www.kaggle.com/c/avito-duplicate-ads-detection/forums/t/20759/external-data-thread/124803#post124803\r\n\r\nI set up a nominatim instance with a copy of open street map of Russia on my laptop (the server was busy doing other stuff). Then I run all (lat, lon) pairs through it, it took several days. But later I noticed that there are not so many distinct pairs, so it could be done in less than a day.",
      "votes": null
    },
    {
      "id": "126782",
      "postDate": "07/12/2016 09:39:44",
      "content": "<p>Hello,</p>\n\n<p>Thanks a lot for sharing ! </p>\n\n<p>What is ET ? <br>\nDid you try anything to reduce the performance score difference between local CV and Public LB ?</p>\n\n<p>How did you combine your different models ?</p>\n\n<p>Thanks again :D </p>",
      "rawMarkdown": "Hello,\r\n\r\nThanks a lot for sharing ! \r\n\r\nWhat is ET ?  \r\nDid you try anything to reduce the performance score difference between local CV and Public LB ?\r\n\r\nHow did you combine your different models ?\r\n\r\nThanks again :D",
      "votes": null
    },
    {
      "id": "126788",
      "postDate": "07/12/2016 10:10:08",
      "content": "<p>Great work Ololo!</p>\n\n<ol>\n<li>How did you get number of geometry matches?</li>\n<li>Does anybody successfuly use the generationMethod and how?</li>\n</ol>\n\n<p>As far as I concerned I used a lot of histograms distances, also SSIM and MSE. I wanted to calculate skimage's BRIEF and ORB matches but unfortunatelly this process was very slow.</p>\n\n<p>@myouness: ETs is Extra Trees classifier I believe (<a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html\">http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html</a>)</p>\n\n<p>P.S. I've also published my code: <a href=\"https://github.com/XAMeLeOH/KAGGLE_AVITO_2016\">https://github.com/XAMeLeOH/KAGGLE_AVITO_2016</a>. It's a quite dirty, I will sanitize it a little bit later.</p>",
      "rawMarkdown": "Great work Ololo!\r\n\r\n 1. How did you get number of geometry matches?\r\n 2. Does anybody successfuly use the generationMethod and how?\r\n\r\nAs far as I concerned I used a lot of histograms distances, also SSIM and MSE. I wanted to calculate skimage's BRIEF and ORB matches but unfortunatelly this process was very slow.\r\n\r\n@myouness: ETs is Extra Trees classifier I believe (http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html)\r\n\r\nP.S. I've also published my code: https://github.com/XAMeLeOH/KAGGLE_AVITO_2016. It's a quite dirty, I will sanitize it a little bit later.",
      "votes": null
    },
    {
      "id": "126791",
      "postDate": "07/12/2016 10:17:45",
      "content": "<p>Thanks :) I did not understand the abbreviation he used </p>",
      "rawMarkdown": "Thanks :) I did not understand the abbreviation he used",
      "votes": null
    },
    {
      "id": "126792",
      "postDate": "07/12/2016 10:19:02",
      "content": "<p>[quote=frist;126788]\n 1. How did you get number of geometry matches?\n 2. Does anybody successfuly use the generationMethod and how?\n[/quote]</p>\n\n<p>By geometry I meant image size (width x height). I used imagemagick to extract that. Then I represented each ad as a &quot;bag of geometries&quot; and computed something similar to jaccard on them. </p>\n\n<p>We tried predicting generationMethod, but in our case it overfit, so we didn't include it. </p>\n\n<p>[quote=myouness;126782]\nWhat is ET ? <br>\nDid you try anything to reduce the performance score difference between local CV and Public LB ?</p>\n\n<p>How did you combine your different models ?\n[/quote]\nET = ExtraTrees from sklearn</p>\n\n<p>Nope, we used 3-fold, and it was more or less good for the 2nd decimal, but somewhat random on the 3rd. So we mostly relied on LB and hoped that with this amount of data it shouldn't shake a lot (and it didn't) </p>\n\n<p>To combine the models we used stacking</p>",
      "rawMarkdown": "[quote=frist;126788]\r\n 1. How did you get number of geometry matches?\r\n 2. Does anybody successfuly use the generationMethod and how?\r\n[/quote]\r\n\r\nBy geometry I meant image size (width x height). I used imagemagick to extract that. Then I represented each ad as a \"bag of geometries\" and computed something similar to jaccard on them. \r\n\r\nWe tried predicting generationMethod, but in our case it overfit, so we didn't include it. \r\n\r\n[quote=myouness;126782]\r\nWhat is ET ?  \r\nDid you try anything to reduce the performance score difference between local CV and Public LB ?\r\n\r\nHow did you combine your different models ?\r\n[/quote]\r\nET = ExtraTrees from sklearn\r\n\r\nNope, we used 3-fold, and it was more or less good for the 2nd decimal, but somewhat random on the 3rd. So we mostly relied on LB and hoped that with this amount of data it shouldn't shake a lot (and it didn't) \r\n\r\nTo combine the models we used stacking",
      "votes": null
    },
    {
      "id": "126794",
      "postDate": "07/12/2016 10:26:19",
      "content": "<p>I was a part of team 8 + 9 = 11. Our close to best solution (0.94700 public LB) was ensemble of my single best model and model from Alex (0.94453 + 0.94132). Best solution (0.94732 public LB) was more complicated ensemble of several solutions. But as you can see it didn't improve much. As I understand in this competition it was more about big set of different features than ensembles. My best single model consists of 505 features and other ones I used for ensembles from 559 features. Half of features was mine and half I got from Alex. I used XGBoost to train model. I found out that train set was ordered by time, so I don't use KFold for final models. I just use as validation set last 2%-5% of train:</p>\n\n<pre><code>split = round((1-test_size)*len(train.index))\nX_train = train[0:split]\nX_valid = train[split:]\n</code></pre>\n\n<p>XGBoost parameters for 0.94453: eta = 0.05, max_depth = 8, subsample = 0.7, colsample_bytree = 0.7. It was around 0.97 local score during training. Increasing of local score leads to increase of leaderboard score. One model required around 2-3 days to finish. It uses ~50 GB of RAM on peak (during DMatrix creation). It was ok to run on my 32 GB system with SSD swap. I made code to store XGBoost model each 1000 iterations. The next XGBoost versions (available on master already, but not available on pip) will have callbacks which simplify this task.</p>\n\n<p>The set of features I used:</p>\n\n<p><strong>Simple features:</strong>\nNumber of images, length/char difference of titles, descriptions, json strings. Same state for all other features, price difference etc\nI also add one-hot encoding for category (not sure if it really needed).</p>\n\n<p><strong>Image features:</strong>\nhave_same - number of same pictures for 2 items (MD5 is used)\nmin_diff_hash1, min_diff_hash2, min_diff_hash3 - minimum difference of image hashes for (1) and (2) ahash, phash, dhash</p>\n\n<p><strong>Distance features:</strong>\nEucleadian distance + Haversine distance</p>\n\n<p><strong>ID based features:</strong>\nI used them because I believe it related to dates. So the more difference between ID, the more time was between Ads.\n'item_id_diff' - maximum from (itemID1)/(itemID2) and (itemID2)/(itemID1)\n'item_id_sub' - absolute diff between itemID1 and itemID2</p>\n\n<p><strong>JSON features:</strong>\nI created the big set of JSON features for every JSON-category. I used &quot;same&quot; 0 or 1 for JSON-categories which is common for train and test. There are some categories which differ like &quot;Address&quot; I compared them using &quot;Damerau Levinshtein&quot; and for address - cousine sim. There are more than 100 features related to JSON in my model.</p>\n\n<p><strong>Text features:</strong>\nFor titles and descriptions I used:</p>\n\n<ol>\n<li>TFIDF cousine sim with Russian NLTK stemmer, punctuation map and Russian stopwords</li>\n<li>Levenshtein distance</li>\n<li>Damerau Levinshtein distance</li>\n<li>Jaro Winkler distance</li>\n</ol>\n\n<p>I also used set of features from Alex related to Inception CNN model for all images as well as his text features (gensim etc).</p>\n\n<p><strong>Special notes:</strong></p>\n\n<p><strong>1.</strong> I find out that there are many IDs which used many times for comparison. I add number of usages of IDs as feature as well.</p>\n\n<p><strong>2.</strong> There are many &quot;logically incorrect&quot; relations in train: </p>\n\n<p>ID1 - ID2 - 1</p>\n\n<p>ID2 - ID3 - 1</p>\n\n<p>ID3 - ID1 - 0</p>\n\n<p>I tried to remove them for training. But it leaded to lower score. I even made an &quot;submission improver&quot; which allows to increase score. It slightly move prediction in 0.5 direction for such triples. It works good on low level models and on validation, but stops give improvement after 0.94.</p>\n\n<p><strong>3.</strong> We increased the train set adding new obvious pairs:</p>\n\n<p>ID1 - ID2 - 1</p>\n\n<p>ID2 - ID3 - 1</p>\n\n<p>then in case ID1 - ID3 not exists we can add it as ID1 - ID3 - 1\nThe same for:</p>\n\n<p>ID1 - ID2 - 1</p>\n\n<p>ID2 - ID3 - 0</p>\n\n<p>then in case ID1 - ID3 not exists we can add it as ID1 - ID3 - 0</p>\n\n<p>Alex trained on enlarged set. I trained on standard set.</p>\n\n<p>......</p>\n\n<p>I published code which I used (as is): <a href=\"https://github.com/ZFTurbo/KAGGLE_AVITO_2016\">https://github.com/ZFTurbo/KAGGLE_AVITO_2016</a></p>\n\n<p>Importance for my features in attachment.</p>",
      "rawMarkdown": "I was a part of team 8 + 9 = 11. Our close to best solution (0.94700 public LB) was ensemble of my single best model and model from Alex (0.94453 + 0.94132). Best solution (0.94732 public LB) was more complicated ensemble of several solutions. But as you can see it didn't improve much. As I understand in this competition it was more about big set of different features than ensembles. My best single model consists of 505 features and other ones I used for ensembles from 559 features. Half of features was mine and half I got from Alex. I used XGBoost to train model. I found out that train set was ordered by time, so I don't use KFold for final models. I just use as validation set last 2%-5% of train:\r\n\r\n\tsplit = round((1-test_size)*len(train.index))\r\n\tX_train = train[0:split]\r\n\tX_valid = train[split:]\r\n\t\r\nXGBoost parameters for 0.94453: eta = 0.05, max_depth = 8, subsample = 0.7, colsample_bytree = 0.7. It was around 0.97 local score during training. Increasing of local score leads to increase of leaderboard score. One model required around 2-3 days to finish. It uses ~50 GB of RAM on peak (during DMatrix creation). It was ok to run on my 32 GB system with SSD swap. I made code to store XGBoost model each 1000 iterations. The next XGBoost versions (available on master already, but not available on pip) will have callbacks which simplify this task.\r\n\r\nThe set of features I used:\r\n\r\n**Simple features:**\r\nNumber of images, length/char difference of titles, descriptions, json strings. Same state for all other features, price difference etc\r\nI also add one-hot encoding for category (not sure if it really needed).\r\n\r\n**Image features:**\r\nhave_same - number of same pictures for 2 items (MD5 is used)\r\nmin_diff_hash1, min_diff_hash2, min_diff_hash3 - minimum difference of image hashes for (1) and (2) ahash, phash, dhash\r\n\r\n**Distance features:**\r\nEucleadian distance + Haversine distance\r\n\r\n**ID based features:**\r\nI used them because I believe it related to dates. So the more difference between ID, the more time was between Ads.\r\n'item_id_diff' - maximum from (itemID1)/(itemID2) and (itemID2)/(itemID1)\r\n'item_id_sub' - absolute diff between itemID1 and itemID2\r\n\r\n**JSON features:**\r\nI created the big set of JSON features for every JSON-category. I used \"same\" 0 or 1 for JSON-categories which is common for train and test. There are some categories which differ like \"Address\" I compared them using \"Damerau Levinshtein\" and for address - cousine sim. There are more than 100 features related to JSON in my model.\r\n\r\n**Text features:**\r\nFor titles and descriptions I used:\r\n\r\n  1. TFIDF cousine sim with Russian NLTK stemmer, punctuation map and Russian stopwords\r\n  2. Levenshtein distance\r\n  3. Damerau Levinshtein distance\r\n  4. Jaro Winkler distance\r\n\r\nI also used set of features from Alex related to Inception CNN model for all images as well as his text features (gensim etc).\r\n\r\n**Special notes:**\r\n\r\n **1.** I find out that there are many IDs which used many times for comparison. I add number of usages of IDs as feature as well.\r\n\r\n **2.** There are many \"logically incorrect\" relations in train: \r\n\r\nID1 - ID2 - 1\r\n\r\nID2 - ID3 - 1\r\n\r\nID3 - ID1 - 0\r\n\r\nI tried to remove them for training. But it leaded to lower score. I even made an \"submission improver\" which allows to increase score. It slightly move prediction in 0.5 direction for such triples. It works good on low level models and on validation, but stops give improvement after 0.94.\r\n  \r\n **3.** We increased the train set adding new obvious pairs:\r\n\r\nID1 - ID2 - 1\r\n\r\nID2 - ID3 - 1\r\n\r\nthen in case ID1 - ID3 not exists we can add it as ID1 - ID3 - 1\r\nThe same for:\r\n\r\nID1 - ID2 - 1\r\n\r\nID2 - ID3 - 0\r\n\r\nthen in case ID1 - ID3 not exists we can add it as ID1 - ID3 - 0\r\n\r\nAlex trained on enlarged set. I trained on standard set.\r\n\r\n......\r\n\r\nI published code which I used (as is): https://github.com/ZFTurbo/KAGGLE_AVITO_2016\r\n\r\nImportance for my features in attachment.",
      "votes": null
    },
    {
      "id": "126795",
      "postDate": "07/12/2016 10:38:54",
      "content": "<p>Good job ZFTurbo! Thanks for sharing the code! I plan to do the same a bit later. </p>",
      "rawMarkdown": "Good job ZFTurbo! Thanks for sharing the code! I plan to do the same a bit later.",
      "votes": null
    },
    {
      "id": "126799",
      "postDate": "07/12/2016 11:17:29",
      "content": "<p>[quote=frist;126788]\nAs far as I concerned I used a lot of histograms distances, also SSIM and MSE. I wanted to calculate skimage's BRIEF and ORB matches but unfortunatelly this process was very slow.</p>\n\n<p>[/quote]</p>\n\n<p>Hi, thanks again for sharing. Can you tell a little more on histogram distances ? And MSE of what ? SSIM of images ?  </p>\n\n<p>Regds</p>",
      "rawMarkdown": "[quote=frist;126788]\r\nAs far as I concerned I used a lot of histograms distances, also SSIM and MSE. I wanted to calculate skimage's BRIEF and ORB matches but unfortunatelly this process was very slow.\r\n\r\n[/quote]\r\n\r\nHi, thanks again for sharing. Can you tell a little more on histogram distances ? And MSE of what ? SSIM of images ?  \r\n\r\nRegds",
      "votes": null
    },
    {
      "id": "126810",
      "postDate": "07/12/2016 12:20:01",
      "content": "<p>[quote=Run2;126799]\n Can you tell a little more on histogram distances ? And MSE of what ? SSIM of images ? <br>\n[/quote]</p>\n\n<p>Image histogram (<a href=\"https://en.wikipedia.org/wiki/Image_histogram\">https://en.wikipedia.org/wiki/Image_histogram</a>) is a way of representing an image as a vector, where dimensionality of this vector is the number of bins of the histogram. Once histograms are calculated, it's possible to use usual distances and similarities to see how similar two histograms are (e.g. euclidean worked fine for me)</p>\n\n<p>SSIM is structural similarity (<a href=\"https://en.wikipedia.org/wiki/Structural_similarity\">https://en.wikipedia.org/wiki/Structural_similarity</a>) - another way of measuring the similarity between two images </p>\n\n<p>As of MSE, maybe @frist meant this: <a href=\"https://en.wikipedia.org/wiki/Peak_signal-to-noise_ratio\">https://en.wikipedia.org/wiki/Peak_signal-to-noise_ratio</a>?</p>",
      "rawMarkdown": "[quote=Run2;126799]\r\n Can you tell a little more on histogram distances ? And MSE of what ? SSIM of images ?  \r\n[/quote]\r\n\r\nImage histogram (https://en.wikipedia.org/wiki/Image_histogram) is a way of representing an image as a vector, where dimensionality of this vector is the number of bins of the histogram. Once histograms are calculated, it's possible to use usual distances and similarities to see how similar two histograms are (e.g. euclidean worked fine for me)\r\n\r\nSSIM is structural similarity (https://en.wikipedia.org/wiki/Structural_similarity) - another way of measuring the similarity between two images \r\n\r\nAs of MSE, maybe @frist meant this: https://en.wikipedia.org/wiki/Peak_signal-to-noise_ratio?",
      "votes": null
    },
    {
      "id": "126812",
      "postDate": "07/12/2016 12:26:55",
      "content": "<p>When we merged with ZFTurbo we discovered that our feature sets where quite different. He'd created a couple hundreds &quot;same/not same/similarity&quot; kinda features on every json attribute and other fields. While I was mostly generating features based on LSI/LDA/word2vec/imagehash/inception-v3. This gave us a good boost right after merging, and our models trained after exchanging features added another 0.01 to the LB score. Unfortunately we've spent all our last week's efforts on building ever more monstrous model ensembles, which turned out to be a complete waste of computing resources. We'd probably improve more if we brainstormed new features instead.</p>\n\n<p>ZFTurbo came up with an idea to add new pairs to the training set.\nIf there were two pairs like that:</p>\n\n<pre><code>id1, id2, 1\nid2, id3, 1\n</code></pre>\n\n<p>We added:</p>\n\n<pre><code>id1, id3, 1\n</code></pre>\n\n<p>These pairs added about 4% to the training set. Correlation between our features and these pairs was not great, about half of what we were seing with the proper training pairs, but I decided to use them in training as semi-noise. To make sure that our models are different, only I was training models on them.</p>\n\n<p>ZFTurbo also created some features based on itemIDs, such as difference between ids (as a proxy to their distance in time, if you assume that IDs where growing in time) or frequency of each id in the dataset. I was not very happy about using those features, but they definetely worked. Anyway, to keep our models different we decided that I won't be using them.</p>\n\n<p>Judging by xgboost gain importance (attached) most usefull features where based on imagehash (phash and dhash being more usefull than ahash).\nI tried different similarity measures. Most usefull were the fraction of images having less that 4 bits difference.</p>\n\n<p>With imagehash, inception-v3 and pre-trained word2vec models I tried different similarity metrics, such as cosine and Earth Mover's Distance, but EMD didn't add much value.</p>\n\n<p>When using inception-v3 I extracted both softmax and pool3 layers. pool3 seemed to have added more gain, but impact was much weaker than simple imagehash. At the same time computational resources required to engineer features using inception-v3 were truly huge.</p>\n\n<p>In text features I found features calculated using n-grams very usefull.</p>\n\n<p>In retrospect, I wish I'd spend more effort coding a better tokenizer. Text fields were very noisy and did call for a lot of spell-checking, carefull work with different number formats etc.</p>\n\n<p>Team 8 + 9 = 11</p>",
      "rawMarkdown": "When we merged with ZFTurbo we discovered that our feature sets where quite different. He'd created a couple hundreds \"same/not same/similarity\" kinda features on every json attribute and other fields. While I was mostly generating features based on LSI/LDA/word2vec/imagehash/inception-v3. This gave us a good boost right after merging, and our models trained after exchanging features added another 0.01 to the LB score. Unfortunately we've spent all our last week's efforts on building ever more monstrous model ensembles, which turned out to be a complete waste of computing resources. We'd probably improve more if we brainstormed new features instead.\r\n\r\nZFTurbo came up with an idea to add new pairs to the training set.\r\nIf there were two pairs like that:\r\n\r\n    id1, id2, 1\r\n    id2, id3, 1\r\n\r\nWe added:\r\n\r\n    id1, id3, 1\r\n\r\nThese pairs added about 4% to the training set. Correlation between our features and these pairs was not great, about half of what we were seing with the proper training pairs, but I decided to use them in training as semi-noise. To make sure that our models are different, only I was training models on them.\r\n\r\nZFTurbo also created some features based on itemIDs, such as difference between ids (as a proxy to their distance in time, if you assume that IDs where growing in time) or frequency of each id in the dataset. I was not very happy about using those features, but they definetely worked. Anyway, to keep our models different we decided that I won't be using them.\r\n\r\nJudging by xgboost gain importance (attached) most usefull features where based on imagehash (phash and dhash being more usefull than ahash).\r\nI tried different similarity measures. Most usefull were the fraction of images having less that 4 bits difference.\r\n\r\nWith imagehash, inception-v3 and pre-trained word2vec models I tried different similarity metrics, such as cosine and Earth Mover's Distance, but EMD didn't add much value.\r\n\r\nWhen using inception-v3 I extracted both softmax and pool3 layers. pool3 seemed to have added more gain, but impact was much weaker than simple imagehash. At the same time computational resources required to engineer features using inception-v3 were truly huge.\r\n\r\nIn text features I found features calculated using n-grams very usefull.\r\n\r\nIn retrospect, I wish I'd spend more effort coding a better tokenizer. Text fields were very noisy and did call for a lot of spell-checking, carefull work with different number formats etc.\r\n\r\nTeam 8 + 9 = 11",
      "votes": null
    },
    {
      "id": "126818",
      "postDate": "07/12/2016 12:53:08",
      "content": "<p>Congrats to the new masters and thanks for sharing your solutions!</p>\n\n<p>How did you handle the description field in particular (all our FE performed quite poorly) ?</p>\n\n<p>note: for ensembling, we used laurae thread to boost our extremely correlated submissions from .9315 area to .9341 using probability^4.  </p>",
      "rawMarkdown": "Congrats to the new masters and thanks for sharing your solutions!\r\n\r\nHow did you handle the description field in particular (all our FE performed quite poorly) ?\r\n\r\n\r\nnote: for ensembling, we used laurae thread to boost our extremely correlated submissions from .9315 area to .9341 using probability^4.",
      "votes": null
    },
    {
      "id": "126821",
      "postDate": "07/12/2016 13:19:42",
      "content": "<p>[quote=Run2;126799]\nCan you tell a little more on histogram distances ? And MSE of what ? SSIM of images ? <br>\n[/quote]</p>\n\n<p>Ololo is right. MSE (mean-squared error) of images not a histograms. Of corse I've got exceptions when the images geometry didn't match.</p>\n\n<p>Measures I've used to compare the images:</p>\n\n<ul>\n<li><a href=\"http://scikit-image.org/docs/stable/api/skimage.measure.html#compare-mse\">skimage.measure.compare_mse</a></li>\n<li><a href=\"http://scikit-image.org/docs/stable/api/skimage.measure.html#compare-ssim\">skimage.measure.compare_ssim</a></li>\n</ul>\n\n<p>You can read about theese measures here: <a href=\"http://www.pyimagesearch.com/2014/09/15/python-compare-two-images/\">http://www.pyimagesearch.com/2014/09/15/python-compare-two-images/</a></p>\n\n<p>For histograms I used OpenCV and scipy distances functions:</p>\n\n<ul>\n<li>OpenCV's cv2.compareHist method with different measures</li>\n<li>scipy.distance methods</li>\n</ul>\n\n<p>I've found some examples comparing histogram distances here: <a href=\"http://www.pyimagesearch.com/2014/07/14/3-ways-compare-histograms-using-opencv-python/\">http://www.pyimagesearch.com/2014/07/14/3-ways-compare-histograms-using-opencv-python/</a></p>\n\n<p>You can find my solution by link a few posts above.</p>",
      "rawMarkdown": "[quote=Run2;126799]\r\nCan you tell a little more on histogram distances ? And MSE of what ? SSIM of images ?  \r\n[/quote]\r\n\r\nOlolo is right. MSE (mean-squared error) of images not a histograms. Of corse I've got exceptions when the images geometry didn't match.\r\n\r\nMeasures I've used to compare the images:\r\n\r\n - [skimage.measure.compare_mse](http://scikit-image.org/docs/stable/api/skimage.measure.html#compare-mse)\r\n - [skimage.measure.compare_ssim](http://scikit-image.org/docs/stable/api/skimage.measure.html#compare-ssim)\r\n\r\nYou can read about theese measures here: http://www.pyimagesearch.com/2014/09/15/python-compare-two-images/\r\n\r\nFor histograms I used OpenCV and scipy distances functions:\r\n\r\n- OpenCV's cv2.compareHist method with different measures\r\n- scipy.distance methods\r\n\r\nI've found some examples comparing histogram distances here: http://www.pyimagesearch.com/2014/07/14/3-ways-compare-histograms-using-opencv-python/\r\n\r\nYou can find my solution by link a few posts above.",
      "votes": null
    },
    {
      "id": "126826",
      "postDate": "07/12/2016 14:16:38",
      "content": "<p>While I didn't do that well, I think in terms of time put in (about one weekend), the results are quite robust. Thanks to the hashes of Yilisg and Dmitry, along with the strong baseline in Kaggle Scripts, you can get 0.89 basically for free by just adding the image intersection count. Getting to about 0.92 was simply adding letter count, categorical histograms and bagging the untuned XGB as provided on Kaggle Scripts. So you can get a 0.92+ single model in about an hour runtime with very minor adjustments over what is known on the forums.</p>",
      "rawMarkdown": "While I didn't do that well, I think in terms of time put in (about one weekend), the results are quite robust. Thanks to the hashes of Yilisg and Dmitry, along with the strong baseline in Kaggle Scripts, you can get 0.89 basically for free by just adding the image intersection count. Getting to about 0.92 was simply adding letter count, categorical histograms and bagging the untuned XGB as provided on Kaggle Scripts. So you can get a 0.92+ single model in about an hour runtime with very minor adjustments over what is known on the forums.",
      "votes": null
    },
    {
      "id": "126827",
      "postDate": "07/12/2016 14:21:29",
      "content": "<p>[quote=ZFTurbo;126794]\n <strong>1.</strong> I find out that there are many IDs which used many times for comparison. I add number of usages of IDs as feature as well.\n[/quote]</p>\n\n<p>[quote=Alexander Vikulin;126812]</p>\n\n<p>ZFTurbo came up with an idea to add new pairs to the training set.\nIf there were two pairs like that:</p>\n\n<pre><code>id1, id2, 1\nid2, id3, 1\n</code></pre>\n\n<p>We added:</p>\n\n<pre><code>id1, id3, 1\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>Alexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.</p>\n\n<p>I will go into more detail in our solution post later today</p>",
      "rawMarkdown": "[quote=ZFTurbo;126794]\r\n **1.** I find out that there are many IDs which used many times for comparison. I add number of usages of IDs as feature as well.\r\n[/quote]\r\n\r\n[quote=Alexander Vikulin;126812]\r\n\r\nZFTurbo came up with an idea to add new pairs to the training set.\r\nIf there were two pairs like that:\r\n\r\n    id1, id2, 1\r\n    id2, id3, 1\r\n\r\nWe added:\r\n\r\n    id1, id3, 1\r\n\r\n[/quote]\r\n\r\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\r\n\r\nI will go into more detail in our solution post later today",
      "votes": null
    },
    {
      "id": "126832",
      "postDate": "07/12/2016 15:03:47",
      "content": "<p>Hi , thanks all for sharing.\nhere are few things that were not yet mentionned \nmodel that score 0.934 was xgb (eta = 0.1, max_depth = 5, subsample = 0.7, colsample_bytree = 0.7), test_size = 0.25 </p>\n\n<p>removing all rows with generationMethod==2 in the training matrix improve this model by '0.005 '</p>\n\n<p>hash comparison:  i used an algo inspired from this (<a href=\"https://7webpages.com/blog/image-duplicates-detection-python/\">https://7webpages.com/blog/image-duplicates-detection-python/</a>) in order to detect potential flip, rotation or change of size/angle in the picture.</p>\n\n<p>json: number of value update, string similarity between value update and OHE of keys for which the values were updated.</p>",
      "rawMarkdown": "Hi , thanks all for sharing.\r\nhere are few things that were not yet mentionned \r\nmodel that score 0.934 was xgb (eta = 0.1, max_depth = 5, subsample = 0.7, colsample_bytree = 0.7), test_size = 0.25 \r\n\r\nremoving all rows with generationMethod==2 in the training matrix improve this model by '0.005 '\r\n\r\nhash comparison:  i used an algo inspired from this (https://7webpages.com/blog/image-duplicates-detection-python/) in order to detect potential flip, rotation or change of size/angle in the picture.\r\n\r\njson: number of value update, string similarity between value update and OHE of keys for which the values were updated.",
      "votes": null
    },
    {
      "id": "126833",
      "postDate": "07/12/2016 15:08:04",
      "content": "<p>[quote=jayjay;126832]</p>\n\n<p>removing all rows with generationMethod==2 in the training matrix improve this model by 0.05 </p>\n\n<p>[/quote]</p>\n\n<p>Surely you mean 0.005?</p>",
      "rawMarkdown": "[quote=jayjay;126832]\r\n\r\nremoving all rows with generationMethod==2 in the training matrix improve this model by 0.05 \r\n\r\n[/quote]\r\n\r\nSurely you mean 0.005?",
      "votes": null
    },
    {
      "id": "126835",
      "postDate": "07/12/2016 15:11:53",
      "content": "<p>yes for sure 0.005</p>",
      "rawMarkdown": "yes for sure 0.005",
      "votes": null
    },
    {
      "id": "126843",
      "postDate": "07/12/2016 16:27:36",
      "content": "<p>[quote=anokas;126827]\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\n[/quote]</p>\n\n<p>Yeah, I understand what you mean and any team willing to win would certainly have to use those, but I strongly feel that it also means the task was poorly designed and Avito will recieve a model, which will be pretty useless. Well, at least with regards to about 0.01 improvement, which ID-derived features were giving. They gave 0.005+ to us. They gave additional 0.004+ to you, because you chose to dig even further, but this improvement will be useless to Avito.</p>\n\n<p>I personally felt guilt using those ID-based features.</p>",
      "rawMarkdown": "[quote=anokas;126827]\r\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\r\n[/quote]\r\n\r\nYeah, I understand what you mean and any team willing to win would certainly have to use those, but I strongly feel that it also means the task was poorly designed and Avito will recieve a model, which will be pretty useless. Well, at least with regards to about 0.01 improvement, which ID-derived features were giving. They gave 0.005+ to us. They gave additional 0.004+ to you, because you chose to dig even further, but this improvement will be useless to Avito.\r\n\r\nI personally felt guilt using those ID-based features.",
      "votes": null
    },
    {
      "id": "126850",
      "postDate": "07/12/2016 17:14:10",
      "content": "<p>[quote=Alexander Vikulin;126843]</p>\n\n<p>[quote=anokas;126827]\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\n[/quote]</p>\n\n<p>Yeah, I understand what you mean and any team willing to win would certainly have to use those, but I strongly feel that it also means the task was poorly designed and Avito will recieve a model, which will be pretty useless. Well, at least with regards to about 0.01 improvement, which ID-derived features were giving. They gave 0.005+ to us. They gave additional 0.004+ to you, because you chose to dig even further, but this improvement will be useless to Avito.</p>\n\n<p>I personally felt guilt using those ID-based features.</p>\n\n<p>[/quote]</p>\n\n<p>We did not use the itemID features at all, but with a 0.005 improvement we could have won!</p>",
      "rawMarkdown": "[quote=Alexander Vikulin;126843]\r\n\r\n[quote=anokas;126827]\r\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\r\n[/quote]\r\n\r\nYeah, I understand what you mean and any team willing to win would certainly have to use those, but I strongly feel that it also means the task was poorly designed and Avito will recieve a model, which will be pretty useless. Well, at least with regards to about 0.01 improvement, which ID-derived features were giving. They gave 0.005+ to us. They gave additional 0.004+ to you, because you chose to dig even further, but this improvement will be useless to Avito.\r\n\r\nI personally felt guilt using those ID-based features.\r\n\r\n[/quote]\r\n\r\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!",
      "votes": null
    },
    {
      "id": "126851",
      "postDate": "07/12/2016 17:18:16",
      "content": "<p>[quote=anokas;126850]\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!\n[/quote]</p>\n\n<p>Previously you said &quot;Using this, you could have added 0.004+ to your score.&quot;, but now you say you haven't use those features. I don't get it.</p>",
      "rawMarkdown": "[quote=anokas;126850]\r\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!\r\n[/quote]\r\n\r\nPreviously you said \"Using this, you could have added 0.004+ to your score.\", but now you say you haven't use those features. I don't get it.",
      "votes": null
    },
    {
      "id": "126855",
      "postDate": "07/12/2016 17:47:31",
      "content": "<p>[quote=Alexander Vikulin;126851]</p>\n\n<p>[quote=anokas;126850]\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!\n[/quote]</p>\n\n<p>Previously you said &quot;Using this, you could have added 0.004+ to your score.&quot;, but now you say you haven't use those features. I don't get it.</p>\n\n<p>[/quote]</p>\n\n<p>We used the size of itemID clusters as a feature, but we did not look at the distance between itemIDs for rows (we did not notice that they are based on time - we expected Avito to randomise this!)</p>",
      "rawMarkdown": "[quote=Alexander Vikulin;126851]\r\n\r\n[quote=anokas;126850]\r\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!\r\n[/quote]\r\n\r\nPreviously you said \"Using this, you could have added 0.004+ to your score.\", but now you say you haven't use those features. I don't get it.\r\n\r\n[/quote]\r\n\r\nWe used the size of itemID clusters as a feature, but we did not look at the distance between itemIDs for rows (we did not notice that they are based on time - we expected Avito to randomise this!)",
      "votes": null
    },
    {
      "id": "126856",
      "postDate": "07/12/2016 18:03:11",
      "content": "<p>[quote=anokas;126855]\nWe used the size of itemID clusters as a feature, but we did not look at the distance between itemIDs for rows (we did not notice that they are based on time - we expected Avito to randomise this!)\n[/quote]</p>\n\n<p>Oh, come on, it's the same thing! If the way items (itemIDs!) are clustered (linked in chaines) has such a strong influence on the score - it's a poorly designed task. It means that the pairs have not been selected randomly. There is some pattern to the way pairs where made and this pattern will unlikely be the same in production.</p>\n\n<p>There is a fairly strong negative correlation (-0.38 AFAIR) between how many times an itemID appears in the dataset and the probability of it being a duplicate. Same with clusters. I cannot believe that the same is true in production.</p>",
      "rawMarkdown": "[quote=anokas;126855]\r\nWe used the size of itemID clusters as a feature, but we did not look at the distance between itemIDs for rows (we did not notice that they are based on time - we expected Avito to randomise this!)\r\n[/quote]\r\n\r\nOh, come on, it's the same thing! If the way items (itemIDs!) are clustered (linked in chaines) has such a strong influence on the score - it's a poorly designed task. It means that the pairs have not been selected randomly. There is some pattern to the way pairs where made and this pattern will unlikely be the same in production.\r\n\r\nThere is a fairly strong negative correlation (-0.38 AFAIR) between how many times an itemID appears in the dataset and the probability of it being a duplicate. Same with clusters. I cannot believe that the same is true in production.",
      "votes": null
    },
    {
      "id": "126858",
      "postDate": "07/12/2016 18:05:24",
      "content": "<p>I do not think that the &quot;time features&quot; won't add value to Avito. It is quite natural that two cars being advertised with a time gap of several weeks or months are NOT the same. </p>\n\n<p>Thus using the &quot;time gap&quot; is really best practice. (We have been blind at this point.) </p>\n\n<p>The cluster features we used are another thing derived by statistical reasoning. If only two items pass the filter for duplicate testing the &quot;isDuplicate&quot; probability is much higher than in the case where 10 or 100 items pass the filter. (The filter may be a simple keyword search.)</p>\n\n<p>Although there is some correlation between cluster size and time gap we should not confuse them.</p>",
      "rawMarkdown": "I do not think that the \"time features\" won't add value to Avito. It is quite natural that two cars being advertised with a time gap of several weeks or months are NOT the same. \r\n\r\nThus using the \"time gap\" is really best practice. (We have been blind at this point.) \r\n\r\nThe cluster features we used are another thing derived by statistical reasoning. If only two items pass the filter for duplicate testing the \"isDuplicate\" probability is much higher than in the case where 10 or 100 items pass the filter. (The filter may be a simple keyword search.)\r\n\r\nAlthough there is some correlation between cluster size and time gap we should not confuse them.",
      "votes": null
    },
    {
      "id": "126860",
      "postDate": "07/12/2016 18:23:18",
      "content": "<p>[quote=Peter Borrmann;126858]\nThe cluster features we used are another thing derived by statistical reasoning. If only two items pass the filter for duplicate testing the &quot;isDuplicate&quot; probability is much higher than in the case where 10 or 100 items pass the filter. (The filter may be a simple keyword search.)\n[/quote]</p>\n\n<p>That is precisely the problem with the dataset. They paired large numbers of non-duplicates with eachother, but duplicates appear in smaller clusters. This adds a hint to how pairs where made, which has nothing to do with the task at hand in production. The clusters should have been similarly sized and contain roughly the same fraction of duplicates and non-duplicates.</p>",
      "rawMarkdown": "[quote=Peter Borrmann;126858]\r\nThe cluster features we used are another thing derived by statistical reasoning. If only two items pass the filter for duplicate testing the \"isDuplicate\" probability is much higher than in the case where 10 or 100 items pass the filter. (The filter may be a simple keyword search.)\r\n[/quote]\r\n\r\nThat is precisely the problem with the dataset. They paired large numbers of non-duplicates with eachother, but duplicates appear in smaller clusters. This adds a hint to how pairs where made, which has nothing to do with the task at hand in production. The clusters should have been similarly sized and contain roughly the same fraction of duplicates and non-duplicates.",
      "votes": null
    },
    {
      "id": "126862",
      "postDate": "07/12/2016 18:32:00",
      "content": "<p>If it makes people feel better, I attempted to look at ad id's, but couldn't place it. I settled with &quot;if itemID_1 appears in itemID_2 then 1 else 0&quot; as a kind of feature (and vice versa), but nothing as strong as a 0.004 gain on the LB. I joined data in my train set that identified a kind of pattern (0.38% 1's in isDuplicate on the low end vs. 0.62% 1's in isDuplicate on the high end), but we were unsure of how to apply this to the test set. I tried to do that with a feature, but it provided me with -0.0002 performance on LB, so we decided not to pursue this further.</p>",
      "rawMarkdown": "If it makes people feel better, I attempted to look at ad id's, but couldn't place it. I settled with \"if itemID_1 appears in itemID_2 then 1 else 0\" as a kind of feature (and vice versa), but nothing as strong as a 0.004 gain on the LB. I joined data in my train set that identified a kind of pattern (0.38% 1's in isDuplicate on the low end vs. 0.62% 1's in isDuplicate on the high end), but we were unsure of how to apply this to the test set. I tried to do that with a feature, but it provided me with -0.0002 performance on LB, so we decided not to pursue this further.",
      "votes": null
    },
    {
      "id": "126880",
      "postDate": "07/12/2016 19:53:26",
      "content": "<p>Another big question that bothers me is whether it was OK to train models (such as word2vec/LSI/LDA) on ALL textual corpora, including test data, which, properly speaking, we shouldn't have used during training.</p>\n\n<p>I had a global LDA model trained on all text data, feeding combined title + description + json for each item as a single document. Then I used it to get similarity metrics for pairs of titles, pairs of descriptions etc.</p>\n\n<p>But I felt uncomfortable using test dataset texts like that. I was comforting myself that LDA can be trained/updated online as more data is available and that in production I could have used that, but still in my opinion this was not a fair use of data and should be prohibited by the rules. I looked for clues in the rules, but couldn't find anything against that.</p>",
      "rawMarkdown": "Another big question that bothers me is whether it was OK to train models (such as word2vec/LSI/LDA) on ALL textual corpora, including test data, which, properly speaking, we shouldn't have used during training.\r\n\r\nI had a global LDA model trained on all text data, feeding combined title + description + json for each item as a single document. Then I used it to get similarity metrics for pairs of titles, pairs of descriptions etc.\r\n\r\nBut I felt uncomfortable using test dataset texts like that. I was comforting myself that LDA can be trained/updated online as more data is available and that in production I could have used that, but still in my opinion this was not a fair use of data and should be prohibited by the rules. I looked for clues in the rules, but couldn't find anything against that.",
      "votes": null
    },
    {
      "id": "126884",
      "postDate": "07/12/2016 20:31:26",
      "content": "<p>Some itemIDs have the same image links (imageID). Perhaps this is the same add but at different points in time. Adding the feature - the number of adds in the cluster did not improve my LB . Someone could use this information?</p>",
      "rawMarkdown": "Some itemIDs have the same image links (imageID). Perhaps this is the same add but at different points in time. Adding the feature - the number of adds in the cluster did not improve my LB . Someone could use this information?",
      "votes": null
    },
    {
      "id": "126886",
      "postDate": "07/12/2016 20:38:52",
      "content": "<p>I updated the solution description to include the link to a github repository with the code. </p>",
      "rawMarkdown": "I updated the solution description to include the link to a github repository with the code.",
      "votes": null
    },
    {
      "id": "126905",
      "postDate": "07/12/2016 21:18:01",
      "content": "<p>Hi, we (theFuture team) also used cluster feature, it is very usefull. Number of corners on the pics is also usefull and so funny.\nDid someone try to find first ad in chain of ad changes? For example same ads, but one description is longer than another. Or user add extra photo.\nIs any usefull information in generationmetod?</p>",
      "rawMarkdown": "Hi, we (theFuture team) also used cluster feature, it is very usefull. Number of corners on the pics is also usefull and so funny.\r\nDid someone try to find first ad in chain of ad changes? For example same ads, but one description is longer than another. Or user add extra photo.\r\nIs any usefull information in generationmetod?",
      "votes": null
    },
    {
      "id": "126923",
      "postDate": "07/12/2016 21:45:23",
      "content": "<p>Thanks all for sharing! I used only 37 features, a subset of the features already mentioned. I could not find any useful information in the generationMethod.</p>\n\n<p>New things:\n - I trained different models for each parentCategory. This added 0.007 to my score. The ratio of duplicates per parentCategory varied, so I had to balance that.\n - I averaged 8 XGBoost models, each with slightly different random hyperparameters</p>\n\n<p>I did not have time to implement ideas to the make use of differences between the test and training data: location and attrsJSON had different distributions.</p>",
      "rawMarkdown": "Thanks all for sharing! I used only 37 features, a subset of the features already mentioned. I could not find any useful information in the generationMethod.\r\n\r\nNew things:\r\n - I trained different models for each parentCategory. This added 0.007 to my score. The ratio of duplicates per parentCategory varied, so I had to balance that.\r\n - I averaged 8 XGBoost models, each with slightly different random hyperparameters\r\n\r\nI did not have time to implement ideas to the make use of differences between the test and training data: location and attrsJSON had different distributions.",
      "votes": null
    },
    {
      "id": "127037",
      "postDate": "07/13/2016 05:21:33",
      "content": "<p>Alexander Vikulin, you're stressing too much about it. It's just a competition :-)</p>",
      "rawMarkdown": "Alexander Vikulin, you're stressing too much about it. It's just a competition :-)",
      "votes": null
    },
    {
      "id": "127224",
      "postDate": "07/13/2016 17:39:35",
      "content": "<p>@ololo Thanks for sharing your solution and code so generously!</p>",
      "rawMarkdown": "ololo Thanks for sharing your solution and code so generously!",
      "votes": null
    },
    {
      "id": "127558",
      "postDate": "07/14/2016 15:04:29",
      "content": "<p>Thank you for sharing ololo and ZFTurbo...</p>\n\n<p>In case, someone cares... :)</p>\n\n<p>My team ended in #89.</p>\n\n<p>It was me and a master's student. The idea was for him to learn some data mining.</p>\n\n<p>We submitted only two models. We were happy to beat Avito's internal benchmark, and didn't have more time to improve it.</p>\n\n<p>The model we used was XGBoost because it seemed very popular. The first submission was with RandomForest from sklearn.</p>\n\n<p>Our features:</p>\n\n<p><strong>CATEGORICALS</strong></p>\n\n<ul>\n<li>absolute differences between: price and metro-id (big number when nan)</li>\n<li>category and parent-category of both items (we extracted parent-category by analyzing the website)</li>\n</ul>\n\n<p><strong>TEXT</strong></p>\n\n<ul>\n<li>TF using only English and Russian colors (between titles and also descriptions) -- the idea was to uncover brands and colors;</li>\n<li>same but only for Russian words (detected as any word that used non-latin alphabet);</li>\n<li>whether both shared the same modal character for each Unicode <a href=\"https://en.wikipedia.org/wiki/Unicode_character_property#General_Category\">general category</a>;</li>\n<li>whether the modal two characters with which both articles started their paragraphs was the same;</li>\n<li>Jaccardian distance for JSON using only keys for which values were numeric, and if there was more than 4 keys;</li>\n<li>absolute difference in several characters like !, and so on.</li>\n</ul>\n\n<p><strong>IMAGES</strong></p>\n\n<ul>\n<li>image hash difference using dhash from ImageHash and Hamming distance (with hash-size=8);</li>\n<li>difference in number of images and boolean saying whether both items have image.</li>\n</ul>\n\n<p>Our code: <a href=\"https://github.com/rpmcruz/avito/\">https://github.com/rpmcruz/avito/</a></p>",
      "rawMarkdown": "Thank you for sharing ololo and ZFTurbo...\r\n\r\nIn case, someone cares... :)\r\n\r\nMy team ended in #89.\r\n\r\nIt was me and a master's student. The idea was for him to learn some data mining.\r\n\r\nWe submitted only two models. We were happy to beat Avito's internal benchmark, and didn't have more time to improve it.\r\n\r\nThe model we used was XGBoost because it seemed very popular. The first submission was with RandomForest from sklearn.\r\n\r\nOur features:\r\n\r\n**CATEGORICALS**\r\n\r\n- absolute differences between: price and metro-id (big number when nan)\r\n- category and parent-category of both items (we extracted parent-category by analyzing the website)\r\n\r\n**TEXT**\r\n\r\n- TF using only English and Russian colors (between titles and also descriptions) -- the idea was to uncover brands and colors;\r\n- same but only for Russian words (detected as any word that used non-latin alphabet);\r\n- whether both shared the same modal character for each Unicode [general category](https://en.wikipedia.org/wiki/Unicode_character_property#General_Category);\r\n- whether the modal two characters with which both articles started their paragraphs was the same;\r\n- Jaccardian distance for JSON using only keys for which values were numeric, and if there was more than 4 keys;\r\n- absolute difference in several characters like !, and so on.\r\n\r\n**IMAGES**\r\n\r\n- image hash difference using dhash from ImageHash and Hamming distance (with hash-size=8);\r\n- difference in number of images and boolean saying whether both items have image.\r\n\r\nOur code: https://github.com/rpmcruz/avito/",
      "votes": null
    },
    {
      "id": "127741",
      "postDate": "07/15/2016 06:25:49",
      "content": "<p>Thanks everyone for sharing.</p>\n\n<p>This our <a href=\"https://github.com/netease-hzdm/avito-duplicate-ads-detection/blob/master/solution.md\">solution</a> which ended in 14th.</p>\n\n<p>It's a single xgboost model with bunch of features(290).</p>\n\n<p>We tries a lot of effort on text(110features) and it doesn't turn out that good.  we definitely should merge.</p>",
      "rawMarkdown": "Thanks everyone for sharing.\r\n\r\nThis our [solution](https://github.com/netease-hzdm/avito-duplicate-ads-detection/blob/master/solution.md) which ended in 14th.\r\n\r\nIt's a single xgboost model with bunch of features(290).\r\n\r\nWe tries a lot of effort on text(110features) and it doesn't turn out that good.  we definitely should merge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 126772,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/12/2016 08:41:25",
      "content": "<p>So here is my solution:</p>\n\n<h3>Simple features</h3>\n\n<ul>\n<li>category id</li>\n<li>number of images </li>\n<li>price difference</li>\n</ul>\n\n<h3>Simple text features</h3>\n\n<ul>\n<li>number of Russian and English characters in the title and description and number of digits and non alphanumeric characters; their ratio to all the text</li>\n<li>len of title and description</li>\n<li>number of unique chars </li>\n<li>cosine and jaccard of 2-3-4 ngrams on the char level</li>\n<li>fuzzy string distances calculated with FuzzyWuzzy (<a href=\"https://github.com/seatgeek/fuzzywuzzy\">https://github.com/seatgeek/fuzzywuzzy</a>)</li>\n</ul>\n\n<h3>Simple image features</h3>\n\n<ul>\n<li>stats (min, mean, max, std, skew, kurtosis) of each channel (R, G, B) as well as the average of all 3 channels </li>\n<li>file size</li>\n<li>number of geometry matches </li>\n<li>number of exact matches (calculated by md5)</li>\n</ul>\n\n<h3>Simple Geo features</h3>\n\n<ul>\n<li>metro id, location id</li>\n<li>distance between two locations</li>\n</ul>\n\n<p>With these features I was able to beat the avito benchmark </p>\n\n<p>Other features that I included afterwards: </p>\n\n<h3>Attributes</h3>\n\n<ul>\n<li>number of key matches, number of value matches, number of key-value pair matches</li>\n<li>number of fields that both ads didn't fill</li>\n<li>similarity of pairs in the tf-idf space, also svd of this space</li>\n</ul>\n\n<h3>Text features</h3>\n\n<ul>\n<li>jaccard and cosine only on digits and English tokens </li>\n<li>if any of the ads have english chars in a russian word (some of the characters look the same, but have different codes)</li>\n<li>tf, tf-idf and bm25 of title, description and all text</li>\n<li>svd of the above </li>\n<li>tf only on words that both ads have in common (in title, desc, all text), tf on words that only one of the ad has, svd of them</li>\n<li>distances and similarities in word2vec and glove spaces </li>\n<li>word2vec and glove similarity only on nouns</li>\n<li>some variation of word's mover distance for both w2v and glove </li>\n<li>how similar misspellings are in the ads </li>\n</ul>\n\n<p>I also tried to extract contact information (phones, emails, etc) but it didn't help much</p>\n\n<h3>Image features</h3>\n\n<ul>\n<li>image hashes from the imagehash library and from the forums </li>\n<li>phash from imagemagick computed on all individual channels </li>\n<li>diffirent similarities and distances of image histograms </li>\n<li>centroids and image moment invariants computed with imagemagick</li>\n<li>structural similarity of images</li>\n<li>SIFT and keypoint matching </li>\n</ul>\n\n<h3>Geo features</h3>\n\n<ul>\n<li>city, region and zip code extracted from geolocation (same sity, region, etc)</li>\n<li>dbscan and kmeans clusters of geolocations</li>\n</ul>\n\n<h3>Meta features</h3>\n\n<ul>\n<li>I run PCA and SVM on all the feature groups and used them as meta features</li>\n</ul>\n\n<h3>Ensembling</h3>\n\n<p>I computed too many features and it was not possible to fit them all into RAM (I have a 32gb/8cores machine) so I started ensembling pretty early. I randomly picked up 100-150 features and run XGBs or ETs on them. </p>\n\n<p>I mostly trained ETs because I could do ~10 of them per day, while training XGB took 2-3 days. </p>\n\n<p>My best model scored ~0.939 on the public LB (xgb with 2.5k trees). The best submission before merging with Abhishek was 0.94437 public / 0.94394 private which would place me at the 10th position. With Abhishek we reached the 5th. </p>\n\n<h3>Code</h3>\n\n<p>The code is available at <a href=\"https://github.com/alexeygrigorev/avito-duplicates-kaggle\">https://github.com/alexeygrigorev/avito-duplicates-kaggle</a> </p>\n\n<p>Thanks everybody, it was very fun! Looking forward to seeing you in the next competitions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126780,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "07/12/2016 09:30:38",
      "content": "<p>Great post Ololo. Thanks for sharing. Our features were similar - obviously a subset. Couple of queries</p>\n\n<p>1) how similar misspellings are in the ads. What did you use to get this information ?</p>\n\n<p>2) So I gather you used geolocation to get City and Zip code and Region ? Did you use a license for google maps api ? \nIf you already have location id and metro id as features - did this additional information help ? </p>\n\n<p>Fantastic feature engineering work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126781,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/12/2016 09:36:10",
      "content": "<p>[quote=Run2;126780]</p>\n\n<p>Great post Ololo. Thanks for sharing. Our features were similar - obviously a subset. Couple of queries</p>\n\n<p>1) how similar misspellings are in the ads. What did you use to get this information ?</p>\n\n<p>2) So I gather you used geolocation to get City and Zip code and Region ? Did you use a license for google maps api ? \nIf you already have location id and metro id as features - did this additional information help ? </p>\n\n<p>Fantastic feature engineering work.</p>\n\n<p>[/quote]</p>\n\n<p>The answer to both of the questions is here: <a href=\"https://www.kaggle.com/c/avito-duplicate-ads-detection/forums/t/20759/external-data-thread/124803#post124803\">https://www.kaggle.com/c/avito-duplicate-ads-detection/forums/t/20759/external-data-thread/124803#post124803</a></p>\n\n<p>I set up a nominatim instance with a copy of open street map of Russia on my laptop (the server was busy doing other stuff). Then I run all (lat, lon) pairs through it, it took several days. But later I noticed that there are not so many distinct pairs, so it could be done in less than a day. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126782,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "07/12/2016 09:39:44",
      "content": "<p>Hello,</p>\n\n<p>Thanks a lot for sharing ! </p>\n\n<p>What is ET ? <br>\nDid you try anything to reduce the performance score difference between local CV and Public LB ?</p>\n\n<p>How did you combine your different models ?</p>\n\n<p>Thanks again :D </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126788,
      "author_name": "xameleoh",
      "author_url": "",
      "post_date": "07/12/2016 10:10:08",
      "content": "<p>Great work Ololo!</p>\n\n<ol>\n<li>How did you get number of geometry matches?</li>\n<li>Does anybody successfuly use the generationMethod and how?</li>\n</ol>\n\n<p>As far as I concerned I used a lot of histograms distances, also SSIM and MSE. I wanted to calculate skimage's BRIEF and ORB matches but unfortunatelly this process was very slow.</p>\n\n<p>@myouness: ETs is Extra Trees classifier I believe (<a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html\">http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html</a>)</p>\n\n<p>P.S. I've also published my code: <a href=\"https://github.com/XAMeLeOH/KAGGLE_AVITO_2016\">https://github.com/XAMeLeOH/KAGGLE_AVITO_2016</a>. It's a quite dirty, I will sanitize it a little bit later.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126791,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "07/12/2016 10:17:45",
      "content": "<p>Thanks :) I did not understand the abbreviation he used </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126792,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/12/2016 10:19:02",
      "content": "<p>[quote=frist;126788]\n 1. How did you get number of geometry matches?\n 2. Does anybody successfuly use the generationMethod and how?\n[/quote]</p>\n\n<p>By geometry I meant image size (width x height). I used imagemagick to extract that. Then I represented each ad as a &quot;bag of geometries&quot; and computed something similar to jaccard on them. </p>\n\n<p>We tried predicting generationMethod, but in our case it overfit, so we didn't include it. </p>\n\n<p>[quote=myouness;126782]\nWhat is ET ? <br>\nDid you try anything to reduce the performance score difference between local CV and Public LB ?</p>\n\n<p>How did you combine your different models ?\n[/quote]\nET = ExtraTrees from sklearn</p>\n\n<p>Nope, we used 3-fold, and it was more or less good for the 2nd decimal, but somewhat random on the 3rd. So we mostly relied on LB and hoped that with this amount of data it shouldn't shake a lot (and it didn't) </p>\n\n<p>To combine the models we used stacking</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126794,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "07/12/2016 10:26:19",
      "content": "<p>I was a part of team 8 + 9 = 11. Our close to best solution (0.94700 public LB) was ensemble of my single best model and model from Alex (0.94453 + 0.94132). Best solution (0.94732 public LB) was more complicated ensemble of several solutions. But as you can see it didn't improve much. As I understand in this competition it was more about big set of different features than ensembles. My best single model consists of 505 features and other ones I used for ensembles from 559 features. Half of features was mine and half I got from Alex. I used XGBoost to train model. I found out that train set was ordered by time, so I don't use KFold for final models. I just use as validation set last 2%-5% of train:</p>\n\n<pre><code>split = round((1-test_size)*len(train.index))\nX_train = train[0:split]\nX_valid = train[split:]\n</code></pre>\n\n<p>XGBoost parameters for 0.94453: eta = 0.05, max_depth = 8, subsample = 0.7, colsample_bytree = 0.7. It was around 0.97 local score during training. Increasing of local score leads to increase of leaderboard score. One model required around 2-3 days to finish. It uses ~50 GB of RAM on peak (during DMatrix creation). It was ok to run on my 32 GB system with SSD swap. I made code to store XGBoost model each 1000 iterations. The next XGBoost versions (available on master already, but not available on pip) will have callbacks which simplify this task.</p>\n\n<p>The set of features I used:</p>\n\n<p><strong>Simple features:</strong>\nNumber of images, length/char difference of titles, descriptions, json strings. Same state for all other features, price difference etc\nI also add one-hot encoding for category (not sure if it really needed).</p>\n\n<p><strong>Image features:</strong>\nhave_same - number of same pictures for 2 items (MD5 is used)\nmin_diff_hash1, min_diff_hash2, min_diff_hash3 - minimum difference of image hashes for (1) and (2) ahash, phash, dhash</p>\n\n<p><strong>Distance features:</strong>\nEucleadian distance + Haversine distance</p>\n\n<p><strong>ID based features:</strong>\nI used them because I believe it related to dates. So the more difference between ID, the more time was between Ads.\n'item_id_diff' - maximum from (itemID1)/(itemID2) and (itemID2)/(itemID1)\n'item_id_sub' - absolute diff between itemID1 and itemID2</p>\n\n<p><strong>JSON features:</strong>\nI created the big set of JSON features for every JSON-category. I used &quot;same&quot; 0 or 1 for JSON-categories which is common for train and test. There are some categories which differ like &quot;Address&quot; I compared them using &quot;Damerau Levinshtein&quot; and for address - cousine sim. There are more than 100 features related to JSON in my model.</p>\n\n<p><strong>Text features:</strong>\nFor titles and descriptions I used:</p>\n\n<ol>\n<li>TFIDF cousine sim with Russian NLTK stemmer, punctuation map and Russian stopwords</li>\n<li>Levenshtein distance</li>\n<li>Damerau Levinshtein distance</li>\n<li>Jaro Winkler distance</li>\n</ol>\n\n<p>I also used set of features from Alex related to Inception CNN model for all images as well as his text features (gensim etc).</p>\n\n<p><strong>Special notes:</strong></p>\n\n<p><strong>1.</strong> I find out that there are many IDs which used many times for comparison. I add number of usages of IDs as feature as well.</p>\n\n<p><strong>2.</strong> There are many &quot;logically incorrect&quot; relations in train: </p>\n\n<p>ID1 - ID2 - 1</p>\n\n<p>ID2 - ID3 - 1</p>\n\n<p>ID3 - ID1 - 0</p>\n\n<p>I tried to remove them for training. But it leaded to lower score. I even made an &quot;submission improver&quot; which allows to increase score. It slightly move prediction in 0.5 direction for such triples. It works good on low level models and on validation, but stops give improvement after 0.94.</p>\n\n<p><strong>3.</strong> We increased the train set adding new obvious pairs:</p>\n\n<p>ID1 - ID2 - 1</p>\n\n<p>ID2 - ID3 - 1</p>\n\n<p>then in case ID1 - ID3 not exists we can add it as ID1 - ID3 - 1\nThe same for:</p>\n\n<p>ID1 - ID2 - 1</p>\n\n<p>ID2 - ID3 - 0</p>\n\n<p>then in case ID1 - ID3 not exists we can add it as ID1 - ID3 - 0</p>\n\n<p>Alex trained on enlarged set. I trained on standard set.</p>\n\n<p>......</p>\n\n<p>I published code which I used (as is): <a href=\"https://github.com/ZFTurbo/KAGGLE_AVITO_2016\">https://github.com/ZFTurbo/KAGGLE_AVITO_2016</a></p>\n\n<p>Importance for my features in attachment.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126795,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/12/2016 10:38:54",
      "content": "<p>Good job ZFTurbo! Thanks for sharing the code! I plan to do the same a bit later. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126799,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "07/12/2016 11:17:29",
      "content": "<p>[quote=frist;126788]\nAs far as I concerned I used a lot of histograms distances, also SSIM and MSE. I wanted to calculate skimage's BRIEF and ORB matches but unfortunatelly this process was very slow.</p>\n\n<p>[/quote]</p>\n\n<p>Hi, thanks again for sharing. Can you tell a little more on histogram distances ? And MSE of what ? SSIM of images ?  </p>\n\n<p>Regds</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126810,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/12/2016 12:20:01",
      "content": "<p>[quote=Run2;126799]\n Can you tell a little more on histogram distances ? And MSE of what ? SSIM of images ? <br>\n[/quote]</p>\n\n<p>Image histogram (<a href=\"https://en.wikipedia.org/wiki/Image_histogram\">https://en.wikipedia.org/wiki/Image_histogram</a>) is a way of representing an image as a vector, where dimensionality of this vector is the number of bins of the histogram. Once histograms are calculated, it's possible to use usual distances and similarities to see how similar two histograms are (e.g. euclidean worked fine for me)</p>\n\n<p>SSIM is structural similarity (<a href=\"https://en.wikipedia.org/wiki/Structural_similarity\">https://en.wikipedia.org/wiki/Structural_similarity</a>) - another way of measuring the similarity between two images </p>\n\n<p>As of MSE, maybe @frist meant this: <a href=\"https://en.wikipedia.org/wiki/Peak_signal-to-noise_ratio\">https://en.wikipedia.org/wiki/Peak_signal-to-noise_ratio</a>?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126812,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "07/12/2016 12:26:55",
      "content": "<p>When we merged with ZFTurbo we discovered that our feature sets where quite different. He'd created a couple hundreds &quot;same/not same/similarity&quot; kinda features on every json attribute and other fields. While I was mostly generating features based on LSI/LDA/word2vec/imagehash/inception-v3. This gave us a good boost right after merging, and our models trained after exchanging features added another 0.01 to the LB score. Unfortunately we've spent all our last week's efforts on building ever more monstrous model ensembles, which turned out to be a complete waste of computing resources. We'd probably improve more if we brainstormed new features instead.</p>\n\n<p>ZFTurbo came up with an idea to add new pairs to the training set.\nIf there were two pairs like that:</p>\n\n<pre><code>id1, id2, 1\nid2, id3, 1\n</code></pre>\n\n<p>We added:</p>\n\n<pre><code>id1, id3, 1\n</code></pre>\n\n<p>These pairs added about 4% to the training set. Correlation between our features and these pairs was not great, about half of what we were seing with the proper training pairs, but I decided to use them in training as semi-noise. To make sure that our models are different, only I was training models on them.</p>\n\n<p>ZFTurbo also created some features based on itemIDs, such as difference between ids (as a proxy to their distance in time, if you assume that IDs where growing in time) or frequency of each id in the dataset. I was not very happy about using those features, but they definetely worked. Anyway, to keep our models different we decided that I won't be using them.</p>\n\n<p>Judging by xgboost gain importance (attached) most usefull features where based on imagehash (phash and dhash being more usefull than ahash).\nI tried different similarity measures. Most usefull were the fraction of images having less that 4 bits difference.</p>\n\n<p>With imagehash, inception-v3 and pre-trained word2vec models I tried different similarity metrics, such as cosine and Earth Mover's Distance, but EMD didn't add much value.</p>\n\n<p>When using inception-v3 I extracted both softmax and pool3 layers. pool3 seemed to have added more gain, but impact was much weaker than simple imagehash. At the same time computational resources required to engineer features using inception-v3 were truly huge.</p>\n\n<p>In text features I found features calculated using n-grams very usefull.</p>\n\n<p>In retrospect, I wish I'd spend more effort coding a better tokenizer. Text fields were very noisy and did call for a lot of spell-checking, carefull work with different number formats etc.</p>\n\n<p>Team 8 + 9 = 11</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126818,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "07/12/2016 12:53:08",
      "content": "<p>Congrats to the new masters and thanks for sharing your solutions!</p>\n\n<p>How did you handle the description field in particular (all our FE performed quite poorly) ?</p>\n\n<p>note: for ensembling, we used laurae thread to boost our extremely correlated submissions from .9315 area to .9341 using probability^4.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126821,
      "author_name": "xameleoh",
      "author_url": "",
      "post_date": "07/12/2016 13:19:42",
      "content": "<p>[quote=Run2;126799]\nCan you tell a little more on histogram distances ? And MSE of what ? SSIM of images ? <br>\n[/quote]</p>\n\n<p>Ololo is right. MSE (mean-squared error) of images not a histograms. Of corse I've got exceptions when the images geometry didn't match.</p>\n\n<p>Measures I've used to compare the images:</p>\n\n<ul>\n<li><a href=\"http://scikit-image.org/docs/stable/api/skimage.measure.html#compare-mse\">skimage.measure.compare_mse</a></li>\n<li><a href=\"http://scikit-image.org/docs/stable/api/skimage.measure.html#compare-ssim\">skimage.measure.compare_ssim</a></li>\n</ul>\n\n<p>You can read about theese measures here: <a href=\"http://www.pyimagesearch.com/2014/09/15/python-compare-two-images/\">http://www.pyimagesearch.com/2014/09/15/python-compare-two-images/</a></p>\n\n<p>For histograms I used OpenCV and scipy distances functions:</p>\n\n<ul>\n<li>OpenCV's cv2.compareHist method with different measures</li>\n<li>scipy.distance methods</li>\n</ul>\n\n<p>I've found some examples comparing histogram distances here: <a href=\"http://www.pyimagesearch.com/2014/07/14/3-ways-compare-histograms-using-opencv-python/\">http://www.pyimagesearch.com/2014/07/14/3-ways-compare-histograms-using-opencv-python/</a></p>\n\n<p>You can find my solution by link a few posts above.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126826,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "07/12/2016 14:16:38",
      "content": "<p>While I didn't do that well, I think in terms of time put in (about one weekend), the results are quite robust. Thanks to the hashes of Yilisg and Dmitry, along with the strong baseline in Kaggle Scripts, you can get 0.89 basically for free by just adding the image intersection count. Getting to about 0.92 was simply adding letter count, categorical histograms and bagging the untuned XGB as provided on Kaggle Scripts. So you can get a 0.92+ single model in about an hour runtime with very minor adjustments over what is known on the forums.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126827,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/12/2016 14:21:29",
      "content": "<p>[quote=ZFTurbo;126794]\n <strong>1.</strong> I find out that there are many IDs which used many times for comparison. I add number of usages of IDs as feature as well.\n[/quote]</p>\n\n<p>[quote=Alexander Vikulin;126812]</p>\n\n<p>ZFTurbo came up with an idea to add new pairs to the training set.\nIf there were two pairs like that:</p>\n\n<pre><code>id1, id2, 1\nid2, id3, 1\n</code></pre>\n\n<p>We added:</p>\n\n<pre><code>id1, id3, 1\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>Alexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.</p>\n\n<p>I will go into more detail in our solution post later today</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126832,
      "author_name": "jayjay75",
      "author_url": "",
      "post_date": "07/12/2016 15:03:47",
      "content": "<p>Hi , thanks all for sharing.\nhere are few things that were not yet mentionned \nmodel that score 0.934 was xgb (eta = 0.1, max_depth = 5, subsample = 0.7, colsample_bytree = 0.7), test_size = 0.25 </p>\n\n<p>removing all rows with generationMethod==2 in the training matrix improve this model by '0.005 '</p>\n\n<p>hash comparison:  i used an algo inspired from this (<a href=\"https://7webpages.com/blog/image-duplicates-detection-python/\">https://7webpages.com/blog/image-duplicates-detection-python/</a>) in order to detect potential flip, rotation or change of size/angle in the picture.</p>\n\n<p>json: number of value update, string similarity between value update and OHE of keys for which the values were updated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126833,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/12/2016 15:08:04",
      "content": "<p>[quote=jayjay;126832]</p>\n\n<p>removing all rows with generationMethod==2 in the training matrix improve this model by 0.05 </p>\n\n<p>[/quote]</p>\n\n<p>Surely you mean 0.005?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126835,
      "author_name": "jayjay75",
      "author_url": "",
      "post_date": "07/12/2016 15:11:53",
      "content": "<p>yes for sure 0.005</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126843,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "07/12/2016 16:27:36",
      "content": "<p>[quote=anokas;126827]\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\n[/quote]</p>\n\n<p>Yeah, I understand what you mean and any team willing to win would certainly have to use those, but I strongly feel that it also means the task was poorly designed and Avito will recieve a model, which will be pretty useless. Well, at least with regards to about 0.01 improvement, which ID-derived features were giving. They gave 0.005+ to us. They gave additional 0.004+ to you, because you chose to dig even further, but this improvement will be useless to Avito.</p>\n\n<p>I personally felt guilt using those ID-based features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126850,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/12/2016 17:14:10",
      "content": "<p>[quote=Alexander Vikulin;126843]</p>\n\n<p>[quote=anokas;126827]\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\n[/quote]</p>\n\n<p>Yeah, I understand what you mean and any team willing to win would certainly have to use those, but I strongly feel that it also means the task was poorly designed and Avito will recieve a model, which will be pretty useless. Well, at least with regards to about 0.01 improvement, which ID-derived features were giving. They gave 0.005+ to us. They gave additional 0.004+ to you, because you chose to dig even further, but this improvement will be useless to Avito.</p>\n\n<p>I personally felt guilt using those ID-based features.</p>\n\n<p>[/quote]</p>\n\n<p>We did not use the itemID features at all, but with a 0.005 improvement we could have won!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126851,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "07/12/2016 17:18:16",
      "content": "<p>[quote=anokas;126850]\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!\n[/quote]</p>\n\n<p>Previously you said &quot;Using this, you could have added 0.004+ to your score.&quot;, but now you say you haven't use those features. I don't get it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126855,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/12/2016 17:47:31",
      "content": "<p>[quote=Alexander Vikulin;126851]</p>\n\n<p>[quote=anokas;126850]\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!\n[/quote]</p>\n\n<p>Previously you said &quot;Using this, you could have added 0.004+ to your score.&quot;, but now you say you haven't use those features. I don't get it.</p>\n\n<p>[/quote]</p>\n\n<p>We used the size of itemID clusters as a feature, but we did not look at the distance between itemIDs for rows (we did not notice that they are based on time - we expected Avito to randomise this!)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126856,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "07/12/2016 18:03:11",
      "content": "<p>[quote=anokas;126855]\nWe used the size of itemID clusters as a feature, but we did not look at the distance between itemIDs for rows (we did not notice that they are based on time - we expected Avito to randomise this!)\n[/quote]</p>\n\n<p>Oh, come on, it's the same thing! If the way items (itemIDs!) are clustered (linked in chaines) has such a strong influence on the score - it's a poorly designed task. It means that the pairs have not been selected randomly. There is some pattern to the way pairs where made and this pattern will unlikely be the same in production.</p>\n\n<p>There is a fairly strong negative correlation (-0.38 AFAIR) between how many times an itemID appears in the dataset and the probability of it being a duplicate. Same with clusters. I cannot believe that the same is true in production.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126858,
      "author_name": "thequants",
      "author_url": "",
      "post_date": "07/12/2016 18:05:24",
      "content": "<p>I do not think that the &quot;time features&quot; won't add value to Avito. It is quite natural that two cars being advertised with a time gap of several weeks or months are NOT the same. </p>\n\n<p>Thus using the &quot;time gap&quot; is really best practice. (We have been blind at this point.) </p>\n\n<p>The cluster features we used are another thing derived by statistical reasoning. If only two items pass the filter for duplicate testing the &quot;isDuplicate&quot; probability is much higher than in the case where 10 or 100 items pass the filter. (The filter may be a simple keyword search.)</p>\n\n<p>Although there is some correlation between cluster size and time gap we should not confuse them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126860,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "07/12/2016 18:23:18",
      "content": "<p>[quote=Peter Borrmann;126858]\nThe cluster features we used are another thing derived by statistical reasoning. If only two items pass the filter for duplicate testing the &quot;isDuplicate&quot; probability is much higher than in the case where 10 or 100 items pass the filter. (The filter may be a simple keyword search.)\n[/quote]</p>\n\n<p>That is precisely the problem with the dataset. They paired large numbers of non-duplicates with eachother, but duplicates appear in smaller clusters. This adds a hint to how pairs where made, which has nothing to do with the task at hand in production. The clusters should have been similarly sized and contain roughly the same fraction of duplicates and non-duplicates.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126862,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/12/2016 18:32:00",
      "content": "<p>If it makes people feel better, I attempted to look at ad id's, but couldn't place it. I settled with &quot;if itemID_1 appears in itemID_2 then 1 else 0&quot; as a kind of feature (and vice versa), but nothing as strong as a 0.004 gain on the LB. I joined data in my train set that identified a kind of pattern (0.38% 1's in isDuplicate on the low end vs. 0.62% 1's in isDuplicate on the high end), but we were unsure of how to apply this to the test set. I tried to do that with a feature, but it provided me with -0.0002 performance on LB, so we decided not to pursue this further.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126880,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "07/12/2016 19:53:26",
      "content": "<p>Another big question that bothers me is whether it was OK to train models (such as word2vec/LSI/LDA) on ALL textual corpora, including test data, which, properly speaking, we shouldn't have used during training.</p>\n\n<p>I had a global LDA model trained on all text data, feeding combined title + description + json for each item as a single document. Then I used it to get similarity metrics for pairs of titles, pairs of descriptions etc.</p>\n\n<p>But I felt uncomfortable using test dataset texts like that. I was comforting myself that LDA can be trained/updated online as more data is available and that in production I could have used that, but still in my opinion this was not a fair use of data and should be prohibited by the rules. I looked for clues in the rules, but couldn't find anything against that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126884,
      "author_name": "dshulchevdkii",
      "author_url": "",
      "post_date": "07/12/2016 20:31:26",
      "content": "<p>Some itemIDs have the same image links (imageID). Perhaps this is the same add but at different points in time. Adding the feature - the number of adds in the cluster did not improve my LB . Someone could use this information?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126886,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/12/2016 20:38:52",
      "content": "<p>I updated the solution description to include the link to a github repository with the code. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126905,
      "author_name": "dansheb",
      "author_url": "",
      "post_date": "07/12/2016 21:18:01",
      "content": "<p>Hi, we (theFuture team) also used cluster feature, it is very usefull. Number of corners on the pics is also usefull and so funny.\nDid someone try to find first ad in chain of ad changes? For example same ads, but one description is longer than another. Or user add extra photo.\nIs any usefull information in generationmetod?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126923,
      "author_name": "joostgp",
      "author_url": "",
      "post_date": "07/12/2016 21:45:23",
      "content": "<p>Thanks all for sharing! I used only 37 features, a subset of the features already mentioned. I could not find any useful information in the generationMethod.</p>\n\n<p>New things:\n - I trained different models for each parentCategory. This added 0.007 to my score. The ratio of duplicates per parentCategory varied, so I had to balance that.\n - I averaged 8 XGBoost models, each with slightly different random hyperparameters</p>\n\n<p>I did not have time to implement ideas to the make use of differences between the test and training data: location and attrsJSON had different distributions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127037,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "07/13/2016 05:21:33",
      "content": "<p>Alexander Vikulin, you're stressing too much about it. It's just a competition :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127224,
      "author_name": "yilihome",
      "author_url": "",
      "post_date": "07/13/2016 17:39:35",
      "content": "<p>@ololo Thanks for sharing your solution and code so generously!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127558,
      "author_name": "rpmcruz",
      "author_url": "",
      "post_date": "07/14/2016 15:04:29",
      "content": "<p>Thank you for sharing ololo and ZFTurbo...</p>\n\n<p>In case, someone cares... :)</p>\n\n<p>My team ended in #89.</p>\n\n<p>It was me and a master's student. The idea was for him to learn some data mining.</p>\n\n<p>We submitted only two models. We were happy to beat Avito's internal benchmark, and didn't have more time to improve it.</p>\n\n<p>The model we used was XGBoost because it seemed very popular. The first submission was with RandomForest from sklearn.</p>\n\n<p>Our features:</p>\n\n<p><strong>CATEGORICALS</strong></p>\n\n<ul>\n<li>absolute differences between: price and metro-id (big number when nan)</li>\n<li>category and parent-category of both items (we extracted parent-category by analyzing the website)</li>\n</ul>\n\n<p><strong>TEXT</strong></p>\n\n<ul>\n<li>TF using only English and Russian colors (between titles and also descriptions) -- the idea was to uncover brands and colors;</li>\n<li>same but only for Russian words (detected as any word that used non-latin alphabet);</li>\n<li>whether both shared the same modal character for each Unicode <a href=\"https://en.wikipedia.org/wiki/Unicode_character_property#General_Category\">general category</a>;</li>\n<li>whether the modal two characters with which both articles started their paragraphs was the same;</li>\n<li>Jaccardian distance for JSON using only keys for which values were numeric, and if there was more than 4 keys;</li>\n<li>absolute difference in several characters like !, and so on.</li>\n</ul>\n\n<p><strong>IMAGES</strong></p>\n\n<ul>\n<li>image hash difference using dhash from ImageHash and Hamming distance (with hash-size=8);</li>\n<li>difference in number of images and boolean saying whether both items have image.</li>\n</ul>\n\n<p>Our code: <a href=\"https://github.com/rpmcruz/avito/\">https://github.com/rpmcruz/avito/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 127741,
      "author_name": "valueq",
      "author_url": "",
      "post_date": "07/15/2016 06:25:49",
      "content": "<p>Thanks everyone for sharing.</p>\n\n<p>This our <a href=\"https://github.com/netease-hzdm/avito-duplicate-ads-detection/blob/master/solution.md\">solution</a> which ended in 14th.</p>\n\n<p>It's a single xgboost model with bunch of features(290).</p>\n\n<p>We tries a lot of effort on text(110features) and it doesn't turn out that good.  we definitely should merge.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "126770": "So now when the competition is over I'm very curious what others were doing. So let's discuss the solutions!",
    "126772": "So here is my solution:\r\n\r\n\r\n### Simple features \r\n\r\n- category id\r\n- number of images \r\n- price difference\r\n\r\n\r\n### Simple text features \r\n\r\n- number of Russian and English characters in the title and description and number of digits and non alphanumeric characters; their ratio to all the text\r\n- len of title and description\r\n- number of unique chars \r\n- cosine and jaccard of 2-3-4 ngrams on the char level\r\n- fuzzy string distances calculated with FuzzyWuzzy (https://github.com/seatgeek/fuzzywuzzy)\r\n\r\n### Simple image features \r\n\r\n- stats (min, mean, max, std, skew, kurtosis) of each channel (R, G, B) as well as the average of all 3 channels \r\n- file size\r\n- number of geometry matches \r\n- number of exact matches (calculated by md5)\r\n\r\n### Simple Geo features\r\n\r\n- metro id, location id\r\n- distance between two locations\r\n\r\nWith these features I was able to beat the avito benchmark \r\n\r\n\r\nOther features that I included afterwards: \r\n\r\n\r\n### Attributes \r\n\r\n- number of key matches, number of value matches, number of key-value pair matches\r\n- number of fields that both ads didn't fill\r\n- similarity of pairs in the tf-idf space, also svd of this space\r\n\r\n\r\n### Text features \r\n\r\n- jaccard and cosine only on digits and English tokens \r\n- if any of the ads have english chars in a russian word (some of the characters look the same, but have different codes)\r\n- tf, tf-idf and bm25 of title, description and all text\r\n- svd of the above \r\n- tf only on words that both ads have in common (in title, desc, all text), tf on words that only one of the ad has, svd of them\r\n- distances and similarities in word2vec and glove spaces \r\n- word2vec and glove similarity only on nouns\r\n- some variation of word's mover distance for both w2v and glove \r\n- how similar misspellings are in the ads \r\n\r\nI also tried to extract contact information (phones, emails, etc) but it didn't help much\r\n\r\n\r\n### Image features \r\n\r\n- image hashes from the imagehash library and from the forums \r\n- phash from imagemagick computed on all individual channels \r\n- diffirent similarities and distances of image histograms \r\n- centroids and image moment invariants computed with imagemagick\r\n- structural similarity of images\r\n- SIFT and keypoint matching \r\n\r\n\r\n### Geo features \r\n\r\n- city, region and zip code extracted from geolocation (same sity, region, etc)\r\n- dbscan and kmeans clusters of geolocations\r\n\r\n\r\n### Meta features\r\n\r\n- I run PCA and SVM on all the feature groups and used them as meta features\r\n\r\n\r\n### Ensembling \r\n\r\nI computed too many features and it was not possible to fit them all into RAM (I have a 32gb/8cores machine) so I started ensembling pretty early. I randomly picked up 100-150 features and run XGBs or ETs on them. \r\n\r\nI mostly trained ETs because I could do ~10 of them per day, while training XGB took 2-3 days. \r\n\r\nMy best model scored ~0.939 on the public LB (xgb with 2.5k trees). The best submission before merging with Abhishek was 0.94437 public / 0.94394 private which would place me at the 10th position. With Abhishek we reached the 5th. \r\n\r\n\r\n### Code\r\n\r\nThe code is available at https://github.com/alexeygrigorev/avito-duplicates-kaggle \r\n\r\nThanks everybody, it was very fun! Looking forward to seeing you in the next competitions.",
    "126780": "Great post Ololo. Thanks for sharing. Our features were similar - obviously a subset. Couple of queries\r\n\r\n1) how similar misspellings are in the ads. What did you use to get this information ?\r\n\r\n2) So I gather you used geolocation to get City and Zip code and Region ? Did you use a license for google maps api ? \r\nIf you already have location id and metro id as features - did this additional information help ? \r\n\r\nFantastic feature engineering work.",
    "126781": "[quote=Run2;126780]\r\n\r\nGreat post Ololo. Thanks for sharing. Our features were similar - obviously a subset. Couple of queries\r\n\r\n1) how similar misspellings are in the ads. What did you use to get this information ?\r\n\r\n2) So I gather you used geolocation to get City and Zip code and Region ? Did you use a license for google maps api ? \r\nIf you already have location id and metro id as features - did this additional information help ? \r\n\r\nFantastic feature engineering work.\r\n\r\n[/quote]\r\n\r\nThe answer to both of the questions is here: https://www.kaggle.com/c/avito-duplicate-ads-detection/forums/t/20759/external-data-thread/124803#post124803\r\n\r\nI set up a nominatim instance with a copy of open street map of Russia on my laptop (the server was busy doing other stuff). Then I run all (lat, lon) pairs through it, it took several days. But later I noticed that there are not so many distinct pairs, so it could be done in less than a day.",
    "126782": "Hello,\r\n\r\nThanks a lot for sharing ! \r\n\r\nWhat is ET ?  \r\nDid you try anything to reduce the performance score difference between local CV and Public LB ?\r\n\r\nHow did you combine your different models ?\r\n\r\nThanks again :D",
    "126788": "Great work Ololo!\r\n\r\n 1. How did you get number of geometry matches?\r\n 2. Does anybody successfuly use the generationMethod and how?\r\n\r\nAs far as I concerned I used a lot of histograms distances, also SSIM and MSE. I wanted to calculate skimage's BRIEF and ORB matches but unfortunatelly this process was very slow.\r\n\r\n@myouness: ETs is Extra Trees classifier I believe (http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html)\r\n\r\nP.S. I've also published my code: https://github.com/XAMeLeOH/KAGGLE_AVITO_2016. It's a quite dirty, I will sanitize it a little bit later.",
    "126791": "Thanks :) I did not understand the abbreviation he used",
    "126792": "[quote=frist;126788]\r\n 1. How did you get number of geometry matches?\r\n 2. Does anybody successfuly use the generationMethod and how?\r\n[/quote]\r\n\r\nBy geometry I meant image size (width x height). I used imagemagick to extract that. Then I represented each ad as a \"bag of geometries\" and computed something similar to jaccard on them. \r\n\r\nWe tried predicting generationMethod, but in our case it overfit, so we didn't include it. \r\n\r\n[quote=myouness;126782]\r\nWhat is ET ?  \r\nDid you try anything to reduce the performance score difference between local CV and Public LB ?\r\n\r\nHow did you combine your different models ?\r\n[/quote]\r\nET = ExtraTrees from sklearn\r\n\r\nNope, we used 3-fold, and it was more or less good for the 2nd decimal, but somewhat random on the 3rd. So we mostly relied on LB and hoped that with this amount of data it shouldn't shake a lot (and it didn't) \r\n\r\nTo combine the models we used stacking",
    "126794": "I was a part of team 8 + 9 = 11. Our close to best solution (0.94700 public LB) was ensemble of my single best model and model from Alex (0.94453 + 0.94132). Best solution (0.94732 public LB) was more complicated ensemble of several solutions. But as you can see it didn't improve much. As I understand in this competition it was more about big set of different features than ensembles. My best single model consists of 505 features and other ones I used for ensembles from 559 features. Half of features was mine and half I got from Alex. I used XGBoost to train model. I found out that train set was ordered by time, so I don't use KFold for final models. I just use as validation set last 2%-5% of train:\r\n\r\n\tsplit = round((1-test_size)*len(train.index))\r\n\tX_train = train[0:split]\r\n\tX_valid = train[split:]\r\n\t\r\nXGBoost parameters for 0.94453: eta = 0.05, max_depth = 8, subsample = 0.7, colsample_bytree = 0.7. It was around 0.97 local score during training. Increasing of local score leads to increase of leaderboard score. One model required around 2-3 days to finish. It uses ~50 GB of RAM on peak (during DMatrix creation). It was ok to run on my 32 GB system with SSD swap. I made code to store XGBoost model each 1000 iterations. The next XGBoost versions (available on master already, but not available on pip) will have callbacks which simplify this task.\r\n\r\nThe set of features I used:\r\n\r\n**Simple features:**\r\nNumber of images, length/char difference of titles, descriptions, json strings. Same state for all other features, price difference etc\r\nI also add one-hot encoding for category (not sure if it really needed).\r\n\r\n**Image features:**\r\nhave_same - number of same pictures for 2 items (MD5 is used)\r\nmin_diff_hash1, min_diff_hash2, min_diff_hash3 - minimum difference of image hashes for (1) and (2) ahash, phash, dhash\r\n\r\n**Distance features:**\r\nEucleadian distance + Haversine distance\r\n\r\n**ID based features:**\r\nI used them because I believe it related to dates. So the more difference between ID, the more time was between Ads.\r\n'item_id_diff' - maximum from (itemID1)/(itemID2) and (itemID2)/(itemID1)\r\n'item_id_sub' - absolute diff between itemID1 and itemID2\r\n\r\n**JSON features:**\r\nI created the big set of JSON features for every JSON-category. I used \"same\" 0 or 1 for JSON-categories which is common for train and test. There are some categories which differ like \"Address\" I compared them using \"Damerau Levinshtein\" and for address - cousine sim. There are more than 100 features related to JSON in my model.\r\n\r\n**Text features:**\r\nFor titles and descriptions I used:\r\n\r\n  1. TFIDF cousine sim with Russian NLTK stemmer, punctuation map and Russian stopwords\r\n  2. Levenshtein distance\r\n  3. Damerau Levinshtein distance\r\n  4. Jaro Winkler distance\r\n\r\nI also used set of features from Alex related to Inception CNN model for all images as well as his text features (gensim etc).\r\n\r\n**Special notes:**\r\n\r\n **1.** I find out that there are many IDs which used many times for comparison. I add number of usages of IDs as feature as well.\r\n\r\n **2.** There are many \"logically incorrect\" relations in train: \r\n\r\nID1 - ID2 - 1\r\n\r\nID2 - ID3 - 1\r\n\r\nID3 - ID1 - 0\r\n\r\nI tried to remove them for training. But it leaded to lower score. I even made an \"submission improver\" which allows to increase score. It slightly move prediction in 0.5 direction for such triples. It works good on low level models and on validation, but stops give improvement after 0.94.\r\n  \r\n **3.** We increased the train set adding new obvious pairs:\r\n\r\nID1 - ID2 - 1\r\n\r\nID2 - ID3 - 1\r\n\r\nthen in case ID1 - ID3 not exists we can add it as ID1 - ID3 - 1\r\nThe same for:\r\n\r\nID1 - ID2 - 1\r\n\r\nID2 - ID3 - 0\r\n\r\nthen in case ID1 - ID3 not exists we can add it as ID1 - ID3 - 0\r\n\r\nAlex trained on enlarged set. I trained on standard set.\r\n\r\n......\r\n\r\nI published code which I used (as is): https://github.com/ZFTurbo/KAGGLE_AVITO_2016\r\n\r\nImportance for my features in attachment.",
    "126795": "Good job ZFTurbo! Thanks for sharing the code! I plan to do the same a bit later.",
    "126799": "[quote=frist;126788]\r\nAs far as I concerned I used a lot of histograms distances, also SSIM and MSE. I wanted to calculate skimage's BRIEF and ORB matches but unfortunatelly this process was very slow.\r\n\r\n[/quote]\r\n\r\nHi, thanks again for sharing. Can you tell a little more on histogram distances ? And MSE of what ? SSIM of images ?  \r\n\r\nRegds",
    "126810": "[quote=Run2;126799]\r\n Can you tell a little more on histogram distances ? And MSE of what ? SSIM of images ?  \r\n[/quote]\r\n\r\nImage histogram (https://en.wikipedia.org/wiki/Image_histogram) is a way of representing an image as a vector, where dimensionality of this vector is the number of bins of the histogram. Once histograms are calculated, it's possible to use usual distances and similarities to see how similar two histograms are (e.g. euclidean worked fine for me)\r\n\r\nSSIM is structural similarity (https://en.wikipedia.org/wiki/Structural_similarity) - another way of measuring the similarity between two images \r\n\r\nAs of MSE, maybe @frist meant this: https://en.wikipedia.org/wiki/Peak_signal-to-noise_ratio?",
    "126812": "When we merged with ZFTurbo we discovered that our feature sets where quite different. He'd created a couple hundreds \"same/not same/similarity\" kinda features on every json attribute and other fields. While I was mostly generating features based on LSI/LDA/word2vec/imagehash/inception-v3. This gave us a good boost right after merging, and our models trained after exchanging features added another 0.01 to the LB score. Unfortunately we've spent all our last week's efforts on building ever more monstrous model ensembles, which turned out to be a complete waste of computing resources. We'd probably improve more if we brainstormed new features instead.\r\n\r\nZFTurbo came up with an idea to add new pairs to the training set.\r\nIf there were two pairs like that:\r\n\r\n    id1, id2, 1\r\n    id2, id3, 1\r\n\r\nWe added:\r\n\r\n    id1, id3, 1\r\n\r\nThese pairs added about 4% to the training set. Correlation between our features and these pairs was not great, about half of what we were seing with the proper training pairs, but I decided to use them in training as semi-noise. To make sure that our models are different, only I was training models on them.\r\n\r\nZFTurbo also created some features based on itemIDs, such as difference between ids (as a proxy to their distance in time, if you assume that IDs where growing in time) or frequency of each id in the dataset. I was not very happy about using those features, but they definetely worked. Anyway, to keep our models different we decided that I won't be using them.\r\n\r\nJudging by xgboost gain importance (attached) most usefull features where based on imagehash (phash and dhash being more usefull than ahash).\r\nI tried different similarity measures. Most usefull were the fraction of images having less that 4 bits difference.\r\n\r\nWith imagehash, inception-v3 and pre-trained word2vec models I tried different similarity metrics, such as cosine and Earth Mover's Distance, but EMD didn't add much value.\r\n\r\nWhen using inception-v3 I extracted both softmax and pool3 layers. pool3 seemed to have added more gain, but impact was much weaker than simple imagehash. At the same time computational resources required to engineer features using inception-v3 were truly huge.\r\n\r\nIn text features I found features calculated using n-grams very usefull.\r\n\r\nIn retrospect, I wish I'd spend more effort coding a better tokenizer. Text fields were very noisy and did call for a lot of spell-checking, carefull work with different number formats etc.\r\n\r\nTeam 8 + 9 = 11",
    "126818": "Congrats to the new masters and thanks for sharing your solutions!\r\n\r\nHow did you handle the description field in particular (all our FE performed quite poorly) ?\r\n\r\n\r\nnote: for ensembling, we used laurae thread to boost our extremely correlated submissions from .9315 area to .9341 using probability^4.",
    "126821": "[quote=Run2;126799]\r\nCan you tell a little more on histogram distances ? And MSE of what ? SSIM of images ?  \r\n[/quote]\r\n\r\nOlolo is right. MSE (mean-squared error) of images not a histograms. Of corse I've got exceptions when the images geometry didn't match.\r\n\r\nMeasures I've used to compare the images:\r\n\r\n - [skimage.measure.compare_mse](http://scikit-image.org/docs/stable/api/skimage.measure.html#compare-mse)\r\n - [skimage.measure.compare_ssim](http://scikit-image.org/docs/stable/api/skimage.measure.html#compare-ssim)\r\n\r\nYou can read about theese measures here: http://www.pyimagesearch.com/2014/09/15/python-compare-two-images/\r\n\r\nFor histograms I used OpenCV and scipy distances functions:\r\n\r\n- OpenCV's cv2.compareHist method with different measures\r\n- scipy.distance methods\r\n\r\nI've found some examples comparing histogram distances here: http://www.pyimagesearch.com/2014/07/14/3-ways-compare-histograms-using-opencv-python/\r\n\r\nYou can find my solution by link a few posts above.",
    "126826": "While I didn't do that well, I think in terms of time put in (about one weekend), the results are quite robust. Thanks to the hashes of Yilisg and Dmitry, along with the strong baseline in Kaggle Scripts, you can get 0.89 basically for free by just adding the image intersection count. Getting to about 0.92 was simply adding letter count, categorical histograms and bagging the untuned XGB as provided on Kaggle Scripts. So you can get a 0.92+ single model in about an hour runtime with very minor adjustments over what is known on the forums.",
    "126827": "[quote=ZFTurbo;126794]\r\n **1.** I find out that there are many IDs which used many times for comparison. I add number of usages of IDs as feature as well.\r\n[/quote]\r\n\r\n[quote=Alexander Vikulin;126812]\r\n\r\nZFTurbo came up with an idea to add new pairs to the training set.\r\nIf there were two pairs like that:\r\n\r\n    id1, id2, 1\r\n    id2, id3, 1\r\n\r\nWe added:\r\n\r\n    id1, id3, 1\r\n\r\n[/quote]\r\n\r\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\r\n\r\nI will go into more detail in our solution post later today",
    "126832": "Hi , thanks all for sharing.\r\nhere are few things that were not yet mentionned \r\nmodel that score 0.934 was xgb (eta = 0.1, max_depth = 5, subsample = 0.7, colsample_bytree = 0.7), test_size = 0.25 \r\n\r\nremoving all rows with generationMethod==2 in the training matrix improve this model by '0.005 '\r\n\r\nhash comparison:  i used an algo inspired from this (https://7webpages.com/blog/image-duplicates-detection-python/) in order to detect potential flip, rotation or change of size/angle in the picture.\r\n\r\njson: number of value update, string similarity between value update and OHE of keys for which the values were updated.",
    "126833": "[quote=jayjay;126832]\r\n\r\nremoving all rows with generationMethod==2 in the training matrix improve this model by 0.05 \r\n\r\n[/quote]\r\n\r\nSurely you mean 0.005?",
    "126835": "yes for sure 0.005",
    "126843": "[quote=anokas;126827]\r\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\r\n[/quote]\r\n\r\nYeah, I understand what you mean and any team willing to win would certainly have to use those, but I strongly feel that it also means the task was poorly designed and Avito will recieve a model, which will be pretty useless. Well, at least with regards to about 0.01 improvement, which ID-derived features were giving. They gave 0.005+ to us. They gave additional 0.004+ to you, because you chose to dig even further, but this improvement will be useless to Avito.\r\n\r\nI personally felt guilt using those ID-based features.",
    "126850": "[quote=Alexander Vikulin;126843]\r\n\r\n[quote=anokas;126827]\r\nAlexander and Roman, you actually missed something very good here! By matching rows with overlapping itemIDs (as you have done), you can build up 'clusters' of rows. Rows in large clusters are much more likely to be non-duplicates. Using this, you could have added 0.004+ to your score.\r\n[/quote]\r\n\r\nYeah, I understand what you mean and any team willing to win would certainly have to use those, but I strongly feel that it also means the task was poorly designed and Avito will recieve a model, which will be pretty useless. Well, at least with regards to about 0.01 improvement, which ID-derived features were giving. They gave 0.005+ to us. They gave additional 0.004+ to you, because you chose to dig even further, but this improvement will be useless to Avito.\r\n\r\nI personally felt guilt using those ID-based features.\r\n\r\n[/quote]\r\n\r\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!",
    "126851": "[quote=anokas;126850]\r\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!\r\n[/quote]\r\n\r\nPreviously you said \"Using this, you could have added 0.004+ to your score.\", but now you say you haven't use those features. I don't get it.",
    "126855": "[quote=Alexander Vikulin;126851]\r\n\r\n[quote=anokas;126850]\r\nWe did not use the itemID features at all, but with a 0.005 improvement we could have won!\r\n[/quote]\r\n\r\nPreviously you said \"Using this, you could have added 0.004+ to your score.\", but now you say you haven't use those features. I don't get it.\r\n\r\n[/quote]\r\n\r\nWe used the size of itemID clusters as a feature, but we did not look at the distance between itemIDs for rows (we did not notice that they are based on time - we expected Avito to randomise this!)",
    "126856": "[quote=anokas;126855]\r\nWe used the size of itemID clusters as a feature, but we did not look at the distance between itemIDs for rows (we did not notice that they are based on time - we expected Avito to randomise this!)\r\n[/quote]\r\n\r\nOh, come on, it's the same thing! If the way items (itemIDs!) are clustered (linked in chaines) has such a strong influence on the score - it's a poorly designed task. It means that the pairs have not been selected randomly. There is some pattern to the way pairs where made and this pattern will unlikely be the same in production.\r\n\r\nThere is a fairly strong negative correlation (-0.38 AFAIR) between how many times an itemID appears in the dataset and the probability of it being a duplicate. Same with clusters. I cannot believe that the same is true in production.",
    "126858": "I do not think that the \"time features\" won't add value to Avito. It is quite natural that two cars being advertised with a time gap of several weeks or months are NOT the same. \r\n\r\nThus using the \"time gap\" is really best practice. (We have been blind at this point.) \r\n\r\nThe cluster features we used are another thing derived by statistical reasoning. If only two items pass the filter for duplicate testing the \"isDuplicate\" probability is much higher than in the case where 10 or 100 items pass the filter. (The filter may be a simple keyword search.)\r\n\r\nAlthough there is some correlation between cluster size and time gap we should not confuse them.",
    "126860": "[quote=Peter Borrmann;126858]\r\nThe cluster features we used are another thing derived by statistical reasoning. If only two items pass the filter for duplicate testing the \"isDuplicate\" probability is much higher than in the case where 10 or 100 items pass the filter. (The filter may be a simple keyword search.)\r\n[/quote]\r\n\r\nThat is precisely the problem with the dataset. They paired large numbers of non-duplicates with eachother, but duplicates appear in smaller clusters. This adds a hint to how pairs where made, which has nothing to do with the task at hand in production. The clusters should have been similarly sized and contain roughly the same fraction of duplicates and non-duplicates.",
    "126862": "If it makes people feel better, I attempted to look at ad id's, but couldn't place it. I settled with \"if itemID_1 appears in itemID_2 then 1 else 0\" as a kind of feature (and vice versa), but nothing as strong as a 0.004 gain on the LB. I joined data in my train set that identified a kind of pattern (0.38% 1's in isDuplicate on the low end vs. 0.62% 1's in isDuplicate on the high end), but we were unsure of how to apply this to the test set. I tried to do that with a feature, but it provided me with -0.0002 performance on LB, so we decided not to pursue this further.",
    "126880": "Another big question that bothers me is whether it was OK to train models (such as word2vec/LSI/LDA) on ALL textual corpora, including test data, which, properly speaking, we shouldn't have used during training.\r\n\r\nI had a global LDA model trained on all text data, feeding combined title + description + json for each item as a single document. Then I used it to get similarity metrics for pairs of titles, pairs of descriptions etc.\r\n\r\nBut I felt uncomfortable using test dataset texts like that. I was comforting myself that LDA can be trained/updated online as more data is available and that in production I could have used that, but still in my opinion this was not a fair use of data and should be prohibited by the rules. I looked for clues in the rules, but couldn't find anything against that.",
    "126884": "Some itemIDs have the same image links (imageID). Perhaps this is the same add but at different points in time. Adding the feature - the number of adds in the cluster did not improve my LB . Someone could use this information?",
    "126886": "I updated the solution description to include the link to a github repository with the code.",
    "126905": "Hi, we (theFuture team) also used cluster feature, it is very usefull. Number of corners on the pics is also usefull and so funny.\r\nDid someone try to find first ad in chain of ad changes? For example same ads, but one description is longer than another. Or user add extra photo.\r\nIs any usefull information in generationmetod?",
    "126923": "Thanks all for sharing! I used only 37 features, a subset of the features already mentioned. I could not find any useful information in the generationMethod.\r\n\r\nNew things:\r\n - I trained different models for each parentCategory. This added 0.007 to my score. The ratio of duplicates per parentCategory varied, so I had to balance that.\r\n - I averaged 8 XGBoost models, each with slightly different random hyperparameters\r\n\r\nI did not have time to implement ideas to the make use of differences between the test and training data: location and attrsJSON had different distributions.",
    "127037": "Alexander Vikulin, you're stressing too much about it. It's just a competition :-)",
    "127224": "ololo Thanks for sharing your solution and code so generously!",
    "127558": "Thank you for sharing ololo and ZFTurbo...\r\n\r\nIn case, someone cares... :)\r\n\r\nMy team ended in #89.\r\n\r\nIt was me and a master's student. The idea was for him to learn some data mining.\r\n\r\nWe submitted only two models. We were happy to beat Avito's internal benchmark, and didn't have more time to improve it.\r\n\r\nThe model we used was XGBoost because it seemed very popular. The first submission was with RandomForest from sklearn.\r\n\r\nOur features:\r\n\r\n**CATEGORICALS**\r\n\r\n- absolute differences between: price and metro-id (big number when nan)\r\n- category and parent-category of both items (we extracted parent-category by analyzing the website)\r\n\r\n**TEXT**\r\n\r\n- TF using only English and Russian colors (between titles and also descriptions) -- the idea was to uncover brands and colors;\r\n- same but only for Russian words (detected as any word that used non-latin alphabet);\r\n- whether both shared the same modal character for each Unicode [general category](https://en.wikipedia.org/wiki/Unicode_character_property#General_Category);\r\n- whether the modal two characters with which both articles started their paragraphs was the same;\r\n- Jaccardian distance for JSON using only keys for which values were numeric, and if there was more than 4 keys;\r\n- absolute difference in several characters like !, and so on.\r\n\r\n**IMAGES**\r\n\r\n- image hash difference using dhash from ImageHash and Hamming distance (with hash-size=8);\r\n- difference in number of images and boolean saying whether both items have image.\r\n\r\nOur code: https://github.com/rpmcruz/avito/",
    "127741": "Thanks everyone for sharing.\r\n\r\nThis our [solution](https://github.com/netease-hzdm/avito-duplicate-ads-detection/blob/master/solution.md) which ended in 14th.\r\n\r\nIt's a single xgboost model with bunch of features(290).\r\n\r\nWe tries a lot of effort on text(110features) and it doesn't turn out that good.  we definitely should merge."
  },
  "source": "meta"
}