{
  "id": 20745,
  "title": "Welcome to Duplicate Ads Detection competition",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/20745",
  "author_name": "",
  "post_date": "2016-05-05T21:19:49.367Z",
  "votes": 3,
  "comment_count": 44,
  "views": 6711,
  "content": "<p>Dear Kagglers,</p>\n\n<p>Welcome to to Duplicate Ads Detection competition where you are asked to predict if two ads are a duplicate i.e. if they are semantically the same (describe same object) or not.\nThis a unique chance for you to create models that uses both natural language and image information on a large scale.</p>\n\n<p>Don't be afraid of the data size. Although there are 10mln+ images to analyse the trick is to extract some signal from them and add it to your full set of engineered features to train your model upon. You can either use simple tricks as <a href=\"http://www.hackerfactor.com/blog/?/archives/529-Kind-of-Like-That.html\">distortion tolerant hashes</a> or more advanced techniques like deep embeddings.</p>\n\n<p>Also don't be afraid of the Russian language in text. Its just a set of words and their relationships. All our previous competitions were won by people who did not know a single word in Russian =) You can check on our <a href=\"https://www.kaggle.com/c/avito-prohibited-content/data\">previous competitions</a> and introductory code on how to read Russian in Python.</p>\n\n<p>If you have any data-related questions we will be glad to help.</p>",
  "messages": [
    {
      "id": "118876",
      "postDate": "05/05/2016 21:19:49",
      "content": "<p>Dear Kagglers,</p>\n\n<p>Welcome to to Duplicate Ads Detection competition where you are asked to predict if two ads are a duplicate i.e. if they are semantically the same (describe same object) or not.\nThis a unique chance for you to create models that uses both natural language and image information on a large scale.</p>\n\n<p>Don't be afraid of the data size. Although there are 10mln+ images to analyse the trick is to extract some signal from them and add it to your full set of engineered features to train your model upon. You can either use simple tricks as <a href=\"http://www.hackerfactor.com/blog/?/archives/529-Kind-of-Like-That.html\">distortion tolerant hashes</a> or more advanced techniques like deep embeddings.</p>\n\n<p>Also don't be afraid of the Russian language in text. Its just a set of words and their relationships. All our previous competitions were won by people who did not know a single word in Russian =) You can check on our <a href=\"https://www.kaggle.com/c/avito-prohibited-content/data\">previous competitions</a> and introductory code on how to read Russian in Python.</p>\n\n<p>If you have any data-related questions we will be glad to help.</p>",
      "rawMarkdown": "Dear Kagglers,\r\n\r\nWelcome to to Duplicate Ads Detection competition where you are asked to predict if two ads are a duplicate i.e. if they are semantically the same (describe same object) or not.\r\nThis a unique chance for you to create models that uses both natural language and image information on a large scale.\r\n\r\nDon't be afraid of the data size. Although there are 10mln+ images to analyse the trick is to extract some signal from them and add it to your full set of engineered features to train your model upon. You can either use simple tricks as [distortion tolerant hashes][1] or more advanced techniques like deep embeddings.\r\n\r\nAlso don't be afraid of the Russian language in text. Its just a set of words and their relationships. All our previous competitions were won by people who did not know a single word in Russian =) You can check on our [previous competitions][2] and introductory code on how to read Russian in Python.\r\n\r\nIf you have any data-related questions we will be glad to help.\r\n\r\n  [1]: http://www.hackerfactor.com/blog/?/archives/529-Kind-of-Like-That.html\r\n  [2]: https://www.kaggle.com/c/avito-prohibited-content/data",
      "votes": null
    },
    {
      "id": "118905",
      "postDate": "05/06/2016 01:49:51",
      "content": "<p>[quote=Ivan Guz;118876]\nIf you have any data-related questions we will be glad to help.\n[/quote]</p>\n\n<p>The competition got restarted a few hours ago, is the data exactly the same? if not, what changed? and....how do I know what to re-download?</p>",
      "rawMarkdown": "[quote=Ivan Guz;118876]\r\nIf you have any data-related questions we will be glad to help.\r\n[/quote]\r\n\r\nThe competition got restarted a few hours ago, is the data exactly the same? if not, what changed? and....how do I know what to re-download?",
      "votes": null
    },
    {
      "id": "118906",
      "postDate": "05/06/2016 01:54:54",
      "content": "<p>No need to re-download. All the files are the same. </p>",
      "rawMarkdown": "No need to re-download. All the files are the same.",
      "votes": null
    },
    {
      "id": "118931",
      "postDate": "05/06/2016 07:27:10",
      "content": "<p>[quote=Ivan Guz;118876]\nIf you have any data-related questions we will be glad to help.\n[/quote]</p>\n\n<p>Would you be able to tell if the Avito benchmark is with or without features extracted from images data.</p>",
      "rawMarkdown": "[quote=Ivan Guz;118876]\r\nIf you have any data-related questions we will be glad to help.\r\n[/quote]\r\n\r\nWould you be able to tell if the Avito benchmark is with or without features extracted from images data.",
      "votes": null
    },
    {
      "id": "119010",
      "postDate": "05/06/2016 17:02:17",
      "content": "<p>[quote=DataGeek;118931]</p>\n\n<p>[quote=Ivan Guz;118876]\nIf you have any data-related questions we will be glad to help.\n[/quote]</p>\n\n<p>Would you be able to tell if the Avito benchmark is with or without features extracted from images data.</p>\n\n<p>[/quote]</p>\n\n<p>Images are used in this benchmark</p>",
      "rawMarkdown": "[quote=DataGeek;118931]\r\n\r\n[quote=Ivan Guz;118876]\r\nIf you have any data-related questions we will be glad to help.\r\n[/quote]\r\n\r\nWould you be able to tell if the Avito benchmark is with or without features extracted from images data.\r\n\r\n[/quote]\r\n\r\nImages are used in this benchmark",
      "votes": null
    },
    {
      "id": "119015",
      "postDate": "05/06/2016 18:22:25",
      "content": "<p>The Evaluation page states that submission files must have a header row that reads &quot;id,probability&quot; but the sample submission has a header of &quot;id,isDuplicate&quot;.</p>",
      "rawMarkdown": "The Evaluation page states that submission files must have a header row that reads \"id,probability\" but the sample submission has a header of \"id,isDuplicate\".",
      "votes": null
    },
    {
      "id": "119020",
      "postDate": "05/06/2016 18:37:55",
      "content": "<p>@David, </p>\n\n<p>I just downloaded the Random_submission.csv file and it's showing the correct header of &quot;id,probability&quot;. Can you double check?</p>",
      "rawMarkdown": "David, \r\n\r\nI just downloaded the Random_submission.csv file and it's showing the correct header of \"id,probability\". Can you double check?",
      "votes": null
    },
    {
      "id": "119022",
      "postDate": "05/06/2016 18:39:22",
      "content": "<p>I can confirm that its id,probability </p>",
      "rawMarkdown": "I can confirm that its id,probability",
      "votes": null
    },
    {
      "id": "119024",
      "postDate": "05/06/2016 18:44:33",
      "content": "<p>I downloaded Random_submission.csv before the competition was temporarily removed yesterday, so that file seems to have changed. Thanks for the confirmation that &quot;id,probability&quot; is correct.</p>",
      "rawMarkdown": "I downloaded Random_submission.csv before the competition was temporarily removed yesterday, so that file seems to have changed. Thanks for the confirmation that \"id,probability\" is correct.",
      "votes": null
    },
    {
      "id": "119041",
      "postDate": "05/06/2016 21:59:56",
      "content": "<p>[quote=Ivan Guz;119010]</p>\n\n<p>[quote=DataGeek;118931]</p>\n\n<p>[quote=Ivan Guz;118876]\nIf you have any data-related questions we will be glad to help.\n[/quote]</p>\n\n<p>Would you be able to tell if the Avito benchmark is with or without features extracted from images data.</p>\n\n<p>[/quote]</p>\n\n<p>Images are used in this benchmark</p>\n\n<p>[/quote]</p>\n\n<p>The first thing I would try in this competition is to directly calculate the probability based on the similarity of two images. The images are relative clean and not that noisy . Again, pretrained model is nice to try first, but not sure whether it is allowed.</p>",
      "rawMarkdown": "[quote=Ivan Guz;119010]\r\n\r\n[quote=DataGeek;118931]\r\n\r\n[quote=Ivan Guz;118876]\r\nIf you have any data-related questions we will be glad to help.\r\n[/quote]\r\n\r\nWould you be able to tell if the Avito benchmark is with or without features extracted from images data.\r\n\r\n[/quote]\r\n\r\nImages are used in this benchmark\r\n\r\n[/quote]\r\n\r\nThe first thing I would try in this competition is to directly calculate the probability based on the similarity of two images. The images are relative clean and not that noisy . Again, pretrained model is nice to try first, but not sure whether it is allowed.",
      "votes": null
    },
    {
      "id": "119048",
      "postDate": "05/06/2016 22:18:38",
      "content": "<p>[quote=SecondPlan;119041]\n Again, pretrained model is nice to try first, but not sure whether it is allowed.</p>\n\n<p>[/quote]</p>\n\n<p>From the rules page:</p>\n\n<p><em>External data (for example, language models and computer vision pre-trained models) that is publicly available and free to use is allowed.</em></p>",
      "rawMarkdown": "[quote=SecondPlan;119041]\r\n Again, pretrained model is nice to try first, but not sure whether it is allowed.\r\n\r\n[/quote]\r\n\r\nFrom the rules page:\r\n\r\n_External data (for example, language models and computer vision pre-trained models) that is publicly available and free to use is allowed._",
      "votes": null
    },
    {
      "id": "119084",
      "postDate": "05/07/2016 04:47:23",
      "content": "<p>Is there any other way than http to download images? Something like ftp would be nice ;)</p>",
      "rawMarkdown": "Is there any other way than http to download images? Something like ftp would be nice ;)",
      "votes": null
    },
    {
      "id": "119161",
      "postDate": "05/07/2016 17:51:31",
      "content": "<p>I could have sworn I read something in the rules about only being allowed to use pairwise comparisons between ads and not being allowed to make transitive inferences... but I can't find it anymore. Help?</p>",
      "rawMarkdown": "I could have sworn I read something in the rules about only being allowed to use pairwise comparisons between ads and not being allowed to make transitive inferences... but I can't find it anymore. Help?",
      "votes": null
    },
    {
      "id": "119162",
      "postDate": "05/07/2016 17:54:48",
      "content": "<p>@David, you're correct. That is what we removed after we took down the competition for a few hours. </p>",
      "rawMarkdown": "David, you're correct. That is what we removed after we took down the competition for a few hours.",
      "votes": null
    },
    {
      "id": "119164",
      "postDate": "05/07/2016 17:58:14",
      "content": "<p>Ah, good. Thank you, Wendy.</p>",
      "rawMarkdown": "Ah, good. Thank you, Wendy.",
      "votes": null
    },
    {
      "id": "119171",
      "postDate": "05/07/2016 19:26:24",
      "content": "<p>Does XGBoost have a transitive inference module?</p>",
      "rawMarkdown": "Does XGBoost have a transitive inference module?",
      "votes": null
    },
    {
      "id": "119210",
      "postDate": "05/08/2016 04:00:15",
      "content": "<p>What exactly does <em>generationMethod</em> means? Is this the procedure used to get the labels? </p>",
      "rawMarkdown": "What exactly does *generationMethod* means? Is this the procedure used to get the labels?",
      "votes": null
    },
    {
      "id": "119215",
      "postDate": "05/08/2016 05:36:28",
      "content": "<p>It's a procedure to both select pairs for analyses and label them.</p>",
      "rawMarkdown": "It's a procedure to both select pairs for analyses and label them.",
      "votes": null
    },
    {
      "id": "119263",
      "postDate": "05/08/2016 16:55:59",
      "content": "<p>So the labels are not 100% correct, and each generation method has it's own noise level?</p>",
      "rawMarkdown": "So the labels are not 100% correct, and each generation method has it's own noise level?",
      "votes": null
    },
    {
      "id": "119265",
      "postDate": "05/08/2016 17:02:19",
      "content": "<p>Labeling is done by humans and they make errors. Even algorithms make errors. It is written in data description that there is noise in the data.</p>",
      "rawMarkdown": "Labeling is done by humans and they make errors. Even algorithms make errors. It is written in data description that there is noise in the data.",
      "votes": null
    },
    {
      "id": "119276",
      "postDate": "05/08/2016 18:55:01",
      "content": "<p>I know, but 'noise in the data' and 'noise in the target variable' are different things. Maybe you guys should edit the description to make it more clear. Just my opinion, thanks.</p>",
      "rawMarkdown": "I know, but 'noise in the data' and 'noise in the target variable' are different things. Maybe you guys should edit the description to make it more clear. Just my opinion, thanks.",
      "votes": null
    },
    {
      "id": "119285",
      "postDate": "05/08/2016 20:32:23",
      "content": "<p>generationMethod is not present in the &quot;test&quot; dataset. Did I skip any info on that?</p>",
      "rawMarkdown": "generationMethod is not present in the \"test\" dataset. Did I skip any info on that?",
      "votes": null
    },
    {
      "id": "119314",
      "postDate": "05/09/2016 05:47:37",
      "content": "<p>It's not present, because we do not want to you to use it at scoring phase. You are given a pair of ads and you need to provide a prediction. However it might be the case that learning only on part of the training data or give each part different weights might lead to a better results at test.</p>",
      "rawMarkdown": "It's not present, because we do not want to you to use it at scoring phase. You are given a pair of ads and you need to provide a prediction. However it might be the case that learning only on part of the training data or give each part different weights might lead to a better results at test.",
      "votes": null
    },
    {
      "id": "119316",
      "postDate": "05/09/2016 06:32:24",
      "content": "<p>[quote=Ivan Guz;119314]</p>\n\n<p>It's not present, because we do not want to you to use it at scoring phase. You are given a pair of ads and you need to provide a prediction. However it might be the case that learning only on part of the training data or give each part different weights might lead to a better results at test.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks Ivan for responding back.\nSo would it be correct to assume that the actual scores for the TEST set are also been generated by all of these methods. If yes, then I am not sure how the weights would be properly used.</p>",
      "rawMarkdown": "[quote=Ivan Guz;119314]\r\n\r\nIt's not present, because we do not want to you to use it at scoring phase. You are given a pair of ads and you need to provide a prediction. However it might be the case that learning only on part of the training data or give each part different weights might lead to a better results at test.\r\n\r\n[/quote]\r\n\r\nThanks Ivan for responding back.\r\nSo would it be correct to assume that the actual scores for the TEST set are also been generated by all of these methods. If yes, then I am not sure how the weights would be properly used.",
      "votes": null
    },
    {
      "id": "119321",
      "postDate": "05/09/2016 07:54:11",
      "content": "<p>Yes. All the same 3 methods are used for generating test data.</p>",
      "rawMarkdown": "Yes. All the same 3 methods are used for generating test data.",
      "votes": null
    },
    {
      "id": "119322",
      "postDate": "05/09/2016 07:55:56",
      "content": "<p>I think a way of using this information could be to weight different samples according to the generation method in the loss function (during learning).. but I haven't tried yet.</p>",
      "rawMarkdown": "I think a way of using this information could be to weight different samples according to the generation method in the loss function (during learning).. but I haven't tried yet.",
      "votes": null
    },
    {
      "id": "119455",
      "postDate": "05/10/2016 12:25:37",
      "content": "<p>@Ivan could you also confirm that the test score is been generated only by those three methods, i.e. there's no additional 'source of truth'. </p>",
      "rawMarkdown": "Ivan could you also confirm that the test score is been generated only by those three methods, i.e. there's no additional 'source of truth'.",
      "votes": null
    },
    {
      "id": "119473",
      "postDate": "05/10/2016 15:07:52",
      "content": "<p>[quote=sh1ng;119455]</p>\n\n<p>@Ivan could you also confirm that the test score is been generated only by those three methods, i.e. there's no additional 'source of truth'. </p>\n\n<p>[/quote]</p>\n\n<p>I confirm.</p>",
      "rawMarkdown": "[quote=sh1ng;119455]\r\n\r\n@Ivan could you also confirm that the test score is been generated only by those three methods, i.e. there's no additional 'source of truth'. \r\n\r\n[/quote]\r\n\r\nI confirm.",
      "votes": null
    },
    {
      "id": "119475",
      "postDate": "05/10/2016 15:15:39",
      "content": "<p>@Ivan one more thought, coefficients of all thee methods should be non-negative. If so better score in local CV for all of them will result in better target score. Is it also true? </p>",
      "rawMarkdown": "Ivan one more thought, coefficients of all thee methods should be non-negative. If so better score in local CV for all of them will result in better target score. Is it also true?",
      "votes": null
    },
    {
      "id": "119489",
      "postDate": "05/10/2016 19:13:03",
      "content": "<p>You will have to figure this out yourself. Baseline is produced with equal weights.</p>",
      "rawMarkdown": "You will have to figure this out yourself. Baseline is produced with equal weights.",
      "votes": null
    },
    {
      "id": "119583",
      "postDate": "05/11/2016 15:11:40",
      "content": "<p>@Ivan why don't we have owners of ads ids? or I missed something?</p>",
      "rawMarkdown": "Ivan why don't we have owners of ads ids? or I missed something?",
      "votes": null
    },
    {
      "id": "119606",
      "postDate": "05/11/2016 19:56:14",
      "content": "<p>@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\nBut Open SIFT (<a href=\"https://robwhess.github.io/opensift/\">https://robwhess.github.io/opensift/</a>) should be allowed.</p>\n\n<p>Can you confirm please?</p>",
      "rawMarkdown": "Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\r\nBut Open SIFT (https://robwhess.github.io/opensift/) should be allowed.\r\n\r\nCan you confirm please?",
      "votes": null
    },
    {
      "id": "119621",
      "postDate": "05/11/2016 22:38:18",
      "content": "<p>[quote=Andrey Kudryavets;119583]</p>\n\n<p>@Ivan why don't we have owners of ads ids? or I missed something?</p>\n\n<p>[/quote]</p>\n\n<p>Because we don't want you to use it :-)</p>",
      "rawMarkdown": "[quote=Andrey Kudryavets;119583]\r\n\r\n@Ivan why don't we have owners of ads ids? or I missed something?\r\n\r\n[/quote]\r\n\r\nBecause we don't want you to use it :-)",
      "votes": null
    },
    {
      "id": "119622",
      "postDate": "05/11/2016 22:44:13",
      "content": "<p>[quote=Sonny Laskar;119606]</p>\n\n<p>@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\nBut Open SIFT (<a href=\"https://robwhess.github.io/opensift/\">https://robwhess.github.io/opensift/</a>) should be allowed.</p>\n\n<p>Can you confirm please?</p>\n\n<p>[/quote]</p>\n\n<p>Looks to me that OpenSIFT should be OK. </p>",
      "rawMarkdown": "[quote=Sonny Laskar;119606]\r\n\r\n@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\r\nBut Open SIFT (https://robwhess.github.io/opensift/) should be allowed.\r\n\r\nCan you confirm please?\r\n\r\n[/quote]\r\n\r\nLooks to me that OpenSIFT should be OK.",
      "votes": null
    },
    {
      "id": "119782",
      "postDate": "05/12/2016 19:08:05",
      "content": "<p>How were the pairs of ads selected out of the universe of all possible ad-ad pairs on your site?  Did you oversample to get more duplicate ads in the training and test sets? </p>",
      "rawMarkdown": "How were the pairs of ads selected out of the universe of all possible ad-ad pairs on your site?  Did you oversample to get more duplicate ads in the training and test sets?",
      "votes": null
    },
    {
      "id": "119816",
      "postDate": "05/12/2016 21:17:31",
      "content": "<p>Question 1: How did you generate geographic latitude and longitude for each ad? I doubt the users entered in those fields manually. Did they enter in a city name, which was converted to lat/long? In densely populated areas, distance between ad locations might be less significant than in less crowded areas. Does the geographic distribution of ads in the sample dataset reflect the distribution in either the population as a whole or in the test set?</p>\n\n<p>Question 2: How were the 208x156 .jpg image files generated by Avito? My concerns here would be whether two identical image files (identical down to filesize/checksum) that were submitted by ad authors could possibly be processed and encoded on different machines with different graphics and CPU specifications and result in small JPEG files that would not have identical filesize and checksums, but nonetheless would look identical to a human or fairly smart computer. Even more complications come in when the ad creator used the same image with different resolutions cropped the image in two different aspect ratios. One thing that might help Avito detect duplicate images is by keeping the original source image when the size is small enough to make that possible.  </p>\n\n<p>Question 3: Are ad authors permitted to add telephone numbers or real email addresses to the ad text? I know Craigslist tries to mask those out if the author includes them. </p>",
      "rawMarkdown": "Question 1: How did you generate geographic latitude and longitude for each ad? I doubt the users entered in those fields manually. Did they enter in a city name, which was converted to lat/long? In densely populated areas, distance between ad locations might be less significant than in less crowded areas. Does the geographic distribution of ads in the sample dataset reflect the distribution in either the population as a whole or in the test set?\r\n\r\nQuestion 2: How were the 208x156 .jpg image files generated by Avito? My concerns here would be whether two identical image files (identical down to filesize/checksum) that were submitted by ad authors could possibly be processed and encoded on different machines with different graphics and CPU specifications and result in small JPEG files that would not have identical filesize and checksums, but nonetheless would look identical to a human or fairly smart computer. Even more complications come in when the ad creator used the same image with different resolutions cropped the image in two different aspect ratios. One thing that might help Avito detect duplicate images is by keeping the original source image when the size is small enough to make that possible.  \r\n\r\nQuestion 3: Are ad authors permitted to add telephone numbers or real email addresses to the ad text? I know Craigslist tries to mask those out if the author includes them.",
      "votes": null
    },
    {
      "id": "120683",
      "postDate": "05/19/2016 22:40:35",
      "content": "<p>[quote=Ivan Guz;119622]</p>\n\n<p>[quote=Sonny Laskar;119606]</p>\n\n<p>@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\nBut Open SIFT (<a href=\"https://robwhess.github.io/opensift/\">https://robwhess.github.io/opensift/</a>) should be allowed.</p>\n\n<p>Can you confirm please?</p>\n\n<p>[/quote]</p>\n\n<p>Looks to me that OpenSIFT should be OK. </p>\n\n<p>[/quote]</p>\n\n<p>As per the OpenSift website it makes notice/warning that it uses patented components  So why would this be ok?</p>",
      "rawMarkdown": "[quote=Ivan Guz;119622]\r\n\r\n[quote=Sonny Laskar;119606]\r\n\r\n@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\r\nBut Open SIFT (https://robwhess.github.io/opensift/) should be allowed.\r\n\r\nCan you confirm please?\r\n\r\n[/quote]\r\n\r\nLooks to me that OpenSIFT should be OK. \r\n\r\n[/quote]\r\n\r\nAs per the OpenSift website it makes notice/warning that it uses patented components  So why would this be ok?",
      "votes": null
    },
    {
      "id": "121264",
      "postDate": "05/25/2016 08:04:57",
      "content": "<p>Admins</p>\n\n<p>So generationMethod =2 has every row with isDuplicate as 1. Can you please explain this ?</p>\n\n<p>Regds</p>",
      "rawMarkdown": "Admins\r\n\r\nSo generationMethod =2 has every row with isDuplicate as 1. Can you please explain this ?\r\n\r\nRegds",
      "votes": null
    },
    {
      "id": "124711",
      "postDate": "06/21/2016 14:34:47",
      "content": "<p>[quote=Ivan Guz;119622]</p>\n\n<p>[quote=Sonny Laskar;119606]</p>\n\n<p>@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\nBut Open SIFT (<a href=\"https://robwhess.github.io/opensift/\">https://robwhess.github.io/opensift/</a>) should be allowed.</p>\n\n<p>Can you confirm please?</p>\n\n<p>[/quote]</p>\n\n<p>Looks to me that OpenSIFT should be OK. </p>\n\n<p>[/quote]</p>\n\n<p>@Ivan Guz, can you please confirm that we can use SIFT as implemented in openIMAJ (<a href=\"http://www.openimaj.org/openimaj-image/image-local-features/apidocs/org/openimaj/image/feature/local/engine/DoGSIFTEngine.html\">here</a>)</p>",
      "rawMarkdown": "[quote=Ivan Guz;119622]\r\n\r\n[quote=Sonny Laskar;119606]\r\n\r\n@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\r\nBut Open SIFT (https://robwhess.github.io/opensift/) should be allowed.\r\n\r\nCan you confirm please?\r\n\r\n[/quote]\r\n\r\nLooks to me that OpenSIFT should be OK. \r\n\r\n[/quote]\r\n\r\n@Ivan Guz, can you please confirm that we can use SIFT as implemented in openIMAJ ([here][1])\r\n\r\n\r\n  [1]: http://www.openimaj.org/openimaj-image/image-local-features/apidocs/org/openimaj/image/feature/local/engine/DoGSIFTEngine.html",
      "votes": null
    },
    {
      "id": "125575",
      "postDate": "06/30/2016 13:56:44",
      "content": "<p>Please use openIMAJ</p>",
      "rawMarkdown": "Please use openIMAJ",
      "votes": null
    },
    {
      "id": "125792",
      "postDate": "07/02/2016 18:23:38",
      "content": "<p>Hi Ivan,</p>\n\n<p>Could you tell us whether the private/public leaderboard split is random, or whether it is again based on time?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi Ivan,\r\n\r\nCould you tell us whether the private/public leaderboard split is random, or whether it is again based on time?\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "125796",
      "postDate": "07/02/2016 19:20:47",
      "content": "<p>Hi Anokas, </p>\n\n<p>It is based on time.</p>",
      "rawMarkdown": "Hi Anokas, \r\n\r\nIt is based on time.",
      "votes": null
    },
    {
      "id": "126061",
      "postDate": "07/05/2016 23:13:50",
      "content": "<p>Hi Competition admins,</p>\n\n<p>I just realized I missed the competition initial submission deadline by a day. Is there any means that I can still participate in the competition? Am I allowed to join a team that is already participating/submitted? Can I post a request to join a team in the forum?</p>\n\n<p>Kiran</p>",
      "rawMarkdown": "Hi Competition admins,\r\n\r\nI just realized I missed the competition initial submission deadline by a day. Is there any means that I can still participate in the competition? Am I allowed to join a team that is already participating/submitted? Can I post a request to join a team in the forum?\r\n\r\nKiran",
      "votes": null
    },
    {
      "id": "126065",
      "postDate": "07/05/2016 23:39:04",
      "content": "<p>@skyrus - As far as I know, there has never been an exception to the initial submission / team merger deadline rule.</p>",
      "rawMarkdown": "skyrus - As far as I know, there has never been an exception to the initial submission / team merger deadline rule.",
      "votes": null
    },
    {
      "id": "126071",
      "postDate": "07/06/2016 00:06:49",
      "content": "<p>In the same spot as Skyrus. </p>\n\n<p>This was my first Kaggle competition attempt and this initial deadline was not very obvious to me. All I saw everyday was the ticker which mentioned x days for competition to end! </p>\n\n<p>Have been working on this project quite hard for some time. We actually have  a pretty good model and would just like an opportunity to compete. Don't care much about the prize. </p>\n\n<p>Even 1 days extension of the initial submission deadline would be a really huge thing for us.  </p>\n\n<p>Thanks and Regards</p>",
      "rawMarkdown": "In the same spot as Skyrus. \r\n\r\nThis was my first Kaggle competition attempt and this initial deadline was not very obvious to me. All I saw everyday was the ticker which mentioned x days for competition to end! \r\n\r\nHave been working on this project quite hard for some time. We actually have  a pretty good model and would just like an opportunity to compete. Don't care much about the prize. \r\n\r\nEven 1 days extension of the initial submission deadline would be a really huge thing for us.  \r\n\r\nThanks and Regards",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118905,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "05/06/2016 01:49:51",
      "content": "<p>[quote=Ivan Guz;118876]\nIf you have any data-related questions we will be glad to help.\n[/quote]</p>\n\n<p>The competition got restarted a few hours ago, is the data exactly the same? if not, what changed? and....how do I know what to re-download?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118906,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "05/06/2016 01:54:54",
      "content": "<p>No need to re-download. All the files are the same. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118931,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "05/06/2016 07:27:10",
      "content": "<p>[quote=Ivan Guz;118876]\nIf you have any data-related questions we will be glad to help.\n[/quote]</p>\n\n<p>Would you be able to tell if the Avito benchmark is with or without features extracted from images data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119010,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/06/2016 17:02:17",
      "content": "<p>[quote=DataGeek;118931]</p>\n\n<p>[quote=Ivan Guz;118876]\nIf you have any data-related questions we will be glad to help.\n[/quote]</p>\n\n<p>Would you be able to tell if the Avito benchmark is with or without features extracted from images data.</p>\n\n<p>[/quote]</p>\n\n<p>Images are used in this benchmark</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119015,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "05/06/2016 18:22:25",
      "content": "<p>The Evaluation page states that submission files must have a header row that reads &quot;id,probability&quot; but the sample submission has a header of &quot;id,isDuplicate&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119020,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "05/06/2016 18:37:55",
      "content": "<p>@David, </p>\n\n<p>I just downloaded the Random_submission.csv file and it's showing the correct header of &quot;id,probability&quot;. Can you double check?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119022,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "05/06/2016 18:39:22",
      "content": "<p>I can confirm that its id,probability </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119024,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "05/06/2016 18:44:33",
      "content": "<p>I downloaded Random_submission.csv before the competition was temporarily removed yesterday, so that file seems to have changed. Thanks for the confirmation that &quot;id,probability&quot; is correct.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119041,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "05/06/2016 21:59:56",
      "content": "<p>[quote=Ivan Guz;119010]</p>\n\n<p>[quote=DataGeek;118931]</p>\n\n<p>[quote=Ivan Guz;118876]\nIf you have any data-related questions we will be glad to help.\n[/quote]</p>\n\n<p>Would you be able to tell if the Avito benchmark is with or without features extracted from images data.</p>\n\n<p>[/quote]</p>\n\n<p>Images are used in this benchmark</p>\n\n<p>[/quote]</p>\n\n<p>The first thing I would try in this competition is to directly calculate the probability based on the similarity of two images. The images are relative clean and not that noisy . Again, pretrained model is nice to try first, but not sure whether it is allowed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119048,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/06/2016 22:18:38",
      "content": "<p>[quote=SecondPlan;119041]\n Again, pretrained model is nice to try first, but not sure whether it is allowed.</p>\n\n<p>[/quote]</p>\n\n<p>From the rules page:</p>\n\n<p><em>External data (for example, language models and computer vision pre-trained models) that is publicly available and free to use is allowed.</em></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119084,
      "author_name": "alexlzzz",
      "author_url": "",
      "post_date": "05/07/2016 04:47:23",
      "content": "<p>Is there any other way than http to download images? Something like ftp would be nice ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119161,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "05/07/2016 17:51:31",
      "content": "<p>I could have sworn I read something in the rules about only being allowed to use pairwise comparisons between ads and not being allowed to make transitive inferences... but I can't find it anymore. Help?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119162,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "05/07/2016 17:54:48",
      "content": "<p>@David, you're correct. That is what we removed after we took down the competition for a few hours. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119164,
      "author_name": "davidtran",
      "author_url": "",
      "post_date": "05/07/2016 17:58:14",
      "content": "<p>Ah, good. Thank you, Wendy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119171,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/07/2016 19:26:24",
      "content": "<p>Does XGBoost have a transitive inference module?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119210,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "05/08/2016 04:00:15",
      "content": "<p>What exactly does <em>generationMethod</em> means? Is this the procedure used to get the labels? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119215,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/08/2016 05:36:28",
      "content": "<p>It's a procedure to both select pairs for analyses and label them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119263,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "05/08/2016 16:55:59",
      "content": "<p>So the labels are not 100% correct, and each generation method has it's own noise level?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119265,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/08/2016 17:02:19",
      "content": "<p>Labeling is done by humans and they make errors. Even algorithms make errors. It is written in data description that there is noise in the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119276,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "05/08/2016 18:55:01",
      "content": "<p>I know, but 'noise in the data' and 'noise in the target variable' are different things. Maybe you guys should edit the description to make it more clear. Just my opinion, thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119285,
      "author_name": "sonnylaskar",
      "author_url": "",
      "post_date": "05/08/2016 20:32:23",
      "content": "<p>generationMethod is not present in the &quot;test&quot; dataset. Did I skip any info on that?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119314,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/09/2016 05:47:37",
      "content": "<p>It's not present, because we do not want to you to use it at scoring phase. You are given a pair of ads and you need to provide a prediction. However it might be the case that learning only on part of the training data or give each part different weights might lead to a better results at test.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119316,
      "author_name": "sonnylaskar",
      "author_url": "",
      "post_date": "05/09/2016 06:32:24",
      "content": "<p>[quote=Ivan Guz;119314]</p>\n\n<p>It's not present, because we do not want to you to use it at scoring phase. You are given a pair of ads and you need to provide a prediction. However it might be the case that learning only on part of the training data or give each part different weights might lead to a better results at test.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks Ivan for responding back.\nSo would it be correct to assume that the actual scores for the TEST set are also been generated by all of these methods. If yes, then I am not sure how the weights would be properly used.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119321,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/09/2016 07:54:11",
      "content": "<p>Yes. All the same 3 methods are used for generating test data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119322,
      "author_name": "lfiaschi",
      "author_url": "",
      "post_date": "05/09/2016 07:55:56",
      "content": "<p>I think a way of using this information could be to weight different samples according to the generation method in the loss function (during learning).. but I haven't tried yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119455,
      "author_name": "sh1ngg",
      "author_url": "",
      "post_date": "05/10/2016 12:25:37",
      "content": "<p>@Ivan could you also confirm that the test score is been generated only by those three methods, i.e. there's no additional 'source of truth'. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119473,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/10/2016 15:07:52",
      "content": "<p>[quote=sh1ng;119455]</p>\n\n<p>@Ivan could you also confirm that the test score is been generated only by those three methods, i.e. there's no additional 'source of truth'. </p>\n\n<p>[/quote]</p>\n\n<p>I confirm.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119475,
      "author_name": "sh1ngg",
      "author_url": "",
      "post_date": "05/10/2016 15:15:39",
      "content": "<p>@Ivan one more thought, coefficients of all thee methods should be non-negative. If so better score in local CV for all of them will result in better target score. Is it also true? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119489,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/10/2016 19:13:03",
      "content": "<p>You will have to figure this out yourself. Baseline is produced with equal weights.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119583,
      "author_name": "kudryavets",
      "author_url": "",
      "post_date": "05/11/2016 15:11:40",
      "content": "<p>@Ivan why don't we have owners of ads ids? or I missed something?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119606,
      "author_name": "sonnylaskar",
      "author_url": "",
      "post_date": "05/11/2016 19:56:14",
      "content": "<p>@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\nBut Open SIFT (<a href=\"https://robwhess.github.io/opensift/\">https://robwhess.github.io/opensift/</a>) should be allowed.</p>\n\n<p>Can you confirm please?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119621,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/11/2016 22:38:18",
      "content": "<p>[quote=Andrey Kudryavets;119583]</p>\n\n<p>@Ivan why don't we have owners of ads ids? or I missed something?</p>\n\n<p>[/quote]</p>\n\n<p>Because we don't want you to use it :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119622,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/11/2016 22:44:13",
      "content": "<p>[quote=Sonny Laskar;119606]</p>\n\n<p>@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\nBut Open SIFT (<a href=\"https://robwhess.github.io/opensift/\">https://robwhess.github.io/opensift/</a>) should be allowed.</p>\n\n<p>Can you confirm please?</p>\n\n<p>[/quote]</p>\n\n<p>Looks to me that OpenSIFT should be OK. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119782,
      "author_name": "alicegif",
      "author_url": "",
      "post_date": "05/12/2016 19:08:05",
      "content": "<p>How were the pairs of ads selected out of the universe of all possible ad-ad pairs on your site?  Did you oversample to get more duplicate ads in the training and test sets? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119816,
      "author_name": "alicegif",
      "author_url": "",
      "post_date": "05/12/2016 21:17:31",
      "content": "<p>Question 1: How did you generate geographic latitude and longitude for each ad? I doubt the users entered in those fields manually. Did they enter in a city name, which was converted to lat/long? In densely populated areas, distance between ad locations might be less significant than in less crowded areas. Does the geographic distribution of ads in the sample dataset reflect the distribution in either the population as a whole or in the test set?</p>\n\n<p>Question 2: How were the 208x156 .jpg image files generated by Avito? My concerns here would be whether two identical image files (identical down to filesize/checksum) that were submitted by ad authors could possibly be processed and encoded on different machines with different graphics and CPU specifications and result in small JPEG files that would not have identical filesize and checksums, but nonetheless would look identical to a human or fairly smart computer. Even more complications come in when the ad creator used the same image with different resolutions cropped the image in two different aspect ratios. One thing that might help Avito detect duplicate images is by keeping the original source image when the size is small enough to make that possible.  </p>\n\n<p>Question 3: Are ad authors permitted to add telephone numbers or real email addresses to the ad text? I know Craigslist tries to mask those out if the author includes them. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120683,
      "author_name": "bigdatahabits",
      "author_url": "",
      "post_date": "05/19/2016 22:40:35",
      "content": "<p>[quote=Ivan Guz;119622]</p>\n\n<p>[quote=Sonny Laskar;119606]</p>\n\n<p>@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\nBut Open SIFT (<a href=\"https://robwhess.github.io/opensift/\">https://robwhess.github.io/opensift/</a>) should be allowed.</p>\n\n<p>Can you confirm please?</p>\n\n<p>[/quote]</p>\n\n<p>Looks to me that OpenSIFT should be OK. </p>\n\n<p>[/quote]</p>\n\n<p>As per the OpenSift website it makes notice/warning that it uses patented components  So why would this be ok?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121264,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "05/25/2016 08:04:57",
      "content": "<p>Admins</p>\n\n<p>So generationMethod =2 has every row with isDuplicate as 1. Can you please explain this ?</p>\n\n<p>Regds</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124711,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "06/21/2016 14:34:47",
      "content": "<p>[quote=Ivan Guz;119622]</p>\n\n<p>[quote=Sonny Laskar;119606]</p>\n\n<p>@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\nBut Open SIFT (<a href=\"https://robwhess.github.io/opensift/\">https://robwhess.github.io/opensift/</a>) should be allowed.</p>\n\n<p>Can you confirm please?</p>\n\n<p>[/quote]</p>\n\n<p>Looks to me that OpenSIFT should be OK. </p>\n\n<p>[/quote]</p>\n\n<p>@Ivan Guz, can you please confirm that we can use SIFT as implemented in openIMAJ (<a href=\"http://www.openimaj.org/openimaj-image/image-local-features/apidocs/org/openimaj/image/feature/local/engine/DoGSIFTEngine.html\">here</a>)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125575,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "06/30/2016 13:56:44",
      "content": "<p>Please use openIMAJ</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125792,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/02/2016 18:23:38",
      "content": "<p>Hi Ivan,</p>\n\n<p>Could you tell us whether the private/public leaderboard split is random, or whether it is again based on time?</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125796,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "07/02/2016 19:20:47",
      "content": "<p>Hi Anokas, </p>\n\n<p>It is based on time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126061,
      "author_name": "skyrus",
      "author_url": "",
      "post_date": "07/05/2016 23:13:50",
      "content": "<p>Hi Competition admins,</p>\n\n<p>I just realized I missed the competition initial submission deadline by a day. Is there any means that I can still participate in the competition? Am I allowed to join a team that is already participating/submitted? Can I post a request to join a team in the forum?</p>\n\n<p>Kiran</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126065,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "07/05/2016 23:39:04",
      "content": "<p>@skyrus - As far as I know, there has never been an exception to the initial submission / team merger deadline rule.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126071,
      "author_name": "explorr",
      "author_url": "",
      "post_date": "07/06/2016 00:06:49",
      "content": "<p>In the same spot as Skyrus. </p>\n\n<p>This was my first Kaggle competition attempt and this initial deadline was not very obvious to me. All I saw everyday was the ticker which mentioned x days for competition to end! </p>\n\n<p>Have been working on this project quite hard for some time. We actually have  a pretty good model and would just like an opportunity to compete. Don't care much about the prize. </p>\n\n<p>Even 1 days extension of the initial submission deadline would be a really huge thing for us.  </p>\n\n<p>Thanks and Regards</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118876": "Dear Kagglers,\r\n\r\nWelcome to to Duplicate Ads Detection competition where you are asked to predict if two ads are a duplicate i.e. if they are semantically the same (describe same object) or not.\r\nThis a unique chance for you to create models that uses both natural language and image information on a large scale.\r\n\r\nDon't be afraid of the data size. Although there are 10mln+ images to analyse the trick is to extract some signal from them and add it to your full set of engineered features to train your model upon. You can either use simple tricks as [distortion tolerant hashes][1] or more advanced techniques like deep embeddings.\r\n\r\nAlso don't be afraid of the Russian language in text. Its just a set of words and their relationships. All our previous competitions were won by people who did not know a single word in Russian =) You can check on our [previous competitions][2] and introductory code on how to read Russian in Python.\r\n\r\nIf you have any data-related questions we will be glad to help.\r\n\r\n  [1]: http://www.hackerfactor.com/blog/?/archives/529-Kind-of-Like-That.html\r\n  [2]: https://www.kaggle.com/c/avito-prohibited-content/data",
    "118905": "[quote=Ivan Guz;118876]\r\nIf you have any data-related questions we will be glad to help.\r\n[/quote]\r\n\r\nThe competition got restarted a few hours ago, is the data exactly the same? if not, what changed? and....how do I know what to re-download?",
    "118906": "No need to re-download. All the files are the same.",
    "118931": "[quote=Ivan Guz;118876]\r\nIf you have any data-related questions we will be glad to help.\r\n[/quote]\r\n\r\nWould you be able to tell if the Avito benchmark is with or without features extracted from images data.",
    "119010": "[quote=DataGeek;118931]\r\n\r\n[quote=Ivan Guz;118876]\r\nIf you have any data-related questions we will be glad to help.\r\n[/quote]\r\n\r\nWould you be able to tell if the Avito benchmark is with or without features extracted from images data.\r\n\r\n[/quote]\r\n\r\nImages are used in this benchmark",
    "119015": "The Evaluation page states that submission files must have a header row that reads \"id,probability\" but the sample submission has a header of \"id,isDuplicate\".",
    "119020": "David, \r\n\r\nI just downloaded the Random_submission.csv file and it's showing the correct header of \"id,probability\". Can you double check?",
    "119022": "I can confirm that its id,probability",
    "119024": "I downloaded Random_submission.csv before the competition was temporarily removed yesterday, so that file seems to have changed. Thanks for the confirmation that \"id,probability\" is correct.",
    "119041": "[quote=Ivan Guz;119010]\r\n\r\n[quote=DataGeek;118931]\r\n\r\n[quote=Ivan Guz;118876]\r\nIf you have any data-related questions we will be glad to help.\r\n[/quote]\r\n\r\nWould you be able to tell if the Avito benchmark is with or without features extracted from images data.\r\n\r\n[/quote]\r\n\r\nImages are used in this benchmark\r\n\r\n[/quote]\r\n\r\nThe first thing I would try in this competition is to directly calculate the probability based on the similarity of two images. The images are relative clean and not that noisy . Again, pretrained model is nice to try first, but not sure whether it is allowed.",
    "119048": "[quote=SecondPlan;119041]\r\n Again, pretrained model is nice to try first, but not sure whether it is allowed.\r\n\r\n[/quote]\r\n\r\nFrom the rules page:\r\n\r\n_External data (for example, language models and computer vision pre-trained models) that is publicly available and free to use is allowed._",
    "119084": "Is there any other way than http to download images? Something like ftp would be nice ;)",
    "119161": "I could have sworn I read something in the rules about only being allowed to use pairwise comparisons between ads and not being allowed to make transitive inferences... but I can't find it anymore. Help?",
    "119162": "David, you're correct. That is what we removed after we took down the competition for a few hours.",
    "119164": "Ah, good. Thank you, Wendy.",
    "119171": "Does XGBoost have a transitive inference module?",
    "119210": "What exactly does *generationMethod* means? Is this the procedure used to get the labels?",
    "119215": "It's a procedure to both select pairs for analyses and label them.",
    "119263": "So the labels are not 100% correct, and each generation method has it's own noise level?",
    "119265": "Labeling is done by humans and they make errors. Even algorithms make errors. It is written in data description that there is noise in the data.",
    "119276": "I know, but 'noise in the data' and 'noise in the target variable' are different things. Maybe you guys should edit the description to make it more clear. Just my opinion, thanks.",
    "119285": "generationMethod is not present in the \"test\" dataset. Did I skip any info on that?",
    "119314": "It's not present, because we do not want to you to use it at scoring phase. You are given a pair of ads and you need to provide a prediction. However it might be the case that learning only on part of the training data or give each part different weights might lead to a better results at test.",
    "119316": "[quote=Ivan Guz;119314]\r\n\r\nIt's not present, because we do not want to you to use it at scoring phase. You are given a pair of ads and you need to provide a prediction. However it might be the case that learning only on part of the training data or give each part different weights might lead to a better results at test.\r\n\r\n[/quote]\r\n\r\nThanks Ivan for responding back.\r\nSo would it be correct to assume that the actual scores for the TEST set are also been generated by all of these methods. If yes, then I am not sure how the weights would be properly used.",
    "119321": "Yes. All the same 3 methods are used for generating test data.",
    "119322": "I think a way of using this information could be to weight different samples according to the generation method in the loss function (during learning).. but I haven't tried yet.",
    "119455": "Ivan could you also confirm that the test score is been generated only by those three methods, i.e. there's no additional 'source of truth'.",
    "119473": "[quote=sh1ng;119455]\r\n\r\n@Ivan could you also confirm that the test score is been generated only by those three methods, i.e. there's no additional 'source of truth'. \r\n\r\n[/quote]\r\n\r\nI confirm.",
    "119475": "Ivan one more thought, coefficients of all thee methods should be non-negative. If so better score in local CV for all of them will result in better target score. Is it also true?",
    "119489": "You will have to figure this out yourself. Baseline is produced with equal weights.",
    "119583": "Ivan why don't we have owners of ads ids? or I missed something?",
    "119606": "Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\r\nBut Open SIFT (https://robwhess.github.io/opensift/) should be allowed.\r\n\r\nCan you confirm please?",
    "119621": "[quote=Andrey Kudryavets;119583]\r\n\r\n@Ivan why don't we have owners of ads ids? or I missed something?\r\n\r\n[/quote]\r\n\r\nBecause we don't want you to use it :-)",
    "119622": "[quote=Sonny Laskar;119606]\r\n\r\n@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\r\nBut Open SIFT (https://robwhess.github.io/opensift/) should be allowed.\r\n\r\nCan you confirm please?\r\n\r\n[/quote]\r\n\r\nLooks to me that OpenSIFT should be OK.",
    "119782": "How were the pairs of ads selected out of the universe of all possible ad-ad pairs on your site?  Did you oversample to get more duplicate ads in the training and test sets?",
    "119816": "Question 1: How did you generate geographic latitude and longitude for each ad? I doubt the users entered in those fields manually. Did they enter in a city name, which was converted to lat/long? In densely populated areas, distance between ad locations might be less significant than in less crowded areas. Does the geographic distribution of ads in the sample dataset reflect the distribution in either the population as a whole or in the test set?\r\n\r\nQuestion 2: How were the 208x156 .jpg image files generated by Avito? My concerns here would be whether two identical image files (identical down to filesize/checksum) that were submitted by ad authors could possibly be processed and encoded on different machines with different graphics and CPU specifications and result in small JPEG files that would not have identical filesize and checksums, but nonetheless would look identical to a human or fairly smart computer. Even more complications come in when the ad creator used the same image with different resolutions cropped the image in two different aspect ratios. One thing that might help Avito detect duplicate images is by keeping the original source image when the size is small enough to make that possible.  \r\n\r\nQuestion 3: Are ad authors permitted to add telephone numbers or real email addresses to the ad text? I know Craigslist tries to mask those out if the author includes them.",
    "120683": "[quote=Ivan Guz;119622]\r\n\r\n[quote=Sonny Laskar;119606]\r\n\r\n@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\r\nBut Open SIFT (https://robwhess.github.io/opensift/) should be allowed.\r\n\r\nCan you confirm please?\r\n\r\n[/quote]\r\n\r\nLooks to me that OpenSIFT should be OK. \r\n\r\n[/quote]\r\n\r\nAs per the OpenSift website it makes notice/warning that it uses patented components  So why would this be ok?",
    "121264": "Admins\r\n\r\nSo generationMethod =2 has every row with isDuplicate as 1. Can you please explain this ?\r\n\r\nRegds",
    "124711": "[quote=Ivan Guz;119622]\r\n\r\n[quote=Sonny Laskar;119606]\r\n\r\n@Ivan, I think we can't use SIFT from OpenCV because of its non-commercial license nature.\r\nBut Open SIFT (https://robwhess.github.io/opensift/) should be allowed.\r\n\r\nCan you confirm please?\r\n\r\n[/quote]\r\n\r\nLooks to me that OpenSIFT should be OK. \r\n\r\n[/quote]\r\n\r\n@Ivan Guz, can you please confirm that we can use SIFT as implemented in openIMAJ ([here][1])\r\n\r\n\r\n  [1]: http://www.openimaj.org/openimaj-image/image-local-features/apidocs/org/openimaj/image/feature/local/engine/DoGSIFTEngine.html",
    "125575": "Please use openIMAJ",
    "125792": "Hi Ivan,\r\n\r\nCould you tell us whether the private/public leaderboard split is random, or whether it is again based on time?\r\n\r\nThanks",
    "125796": "Hi Anokas, \r\n\r\nIt is based on time.",
    "126061": "Hi Competition admins,\r\n\r\nI just realized I missed the competition initial submission deadline by a day. Is there any means that I can still participate in the competition? Am I allowed to join a team that is already participating/submitted? Can I post a request to join a team in the forum?\r\n\r\nKiran",
    "126065": "skyrus - As far as I know, there has never been an exception to the initial submission / team merger deadline rule.",
    "126071": "In the same spot as Skyrus. \r\n\r\nThis was my first Kaggle competition attempt and this initial deadline was not very obvious to me. All I saw everyday was the ticker which mentioned x days for competition to end! \r\n\r\nHave been working on this project quite hard for some time. We actually have  a pretty good model and would just like an opportunity to compete. Don't care much about the prize. \r\n\r\nEven 1 days extension of the initial submission deadline would be a really huge thing for us.  \r\n\r\nThanks and Regards"
  },
  "source": "meta"
}