{
  "id": 20794,
  "title": "How to use the image data",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/20794",
  "author_name": "",
  "post_date": "2016-05-07T20:30:05.087Z",
  "votes": 1,
  "comment_count": 15,
  "views": 2830,
  "content": "<p>Does anybody started using the actual images of the Ads? Any ideas how to turn this into features? </p>",
  "messages": [
    {
      "id": "119179",
      "postDate": "05/07/2016 20:30:05",
      "content": "<p>Does anybody started using the actual images of the Ads? Any ideas how to turn this into features? </p>",
      "rawMarkdown": "Does anybody started using the actual images of the Ads? Any ideas how to turn this into features?",
      "votes": null
    },
    {
      "id": "119182",
      "postDate": "05/07/2016 20:50:44",
      "content": "<p>Probably none of the submissions used the images yet, since they are still far below the benchmark.. That's understandable due to the size of the dataset. You may want to look into past competitions that involved computer vision. The usual approach is to use a conv neural net to extract features from images in the form of a vector, and then feed the vector into a classification model of your choice. That may require a GPU, but can be done with CPU only, though it takes a lot longer.</p>",
      "rawMarkdown": "Probably none of the submissions used the images yet, since they are still far below the benchmark.. That's understandable due to the size of the dataset. You may want to look into past competitions that involved computer vision. The usual approach is to use a conv neural net to extract features from images in the form of a vector, and then feed the vector into a classification model of your choice. That may require a GPU, but can be done with CPU only, though it takes a lot longer.",
      "votes": null
    },
    {
      "id": "119190",
      "postDate": "05/07/2016 21:52:50",
      "content": "<p>[quote=FernandoProcy;119182]</p>\n\n<p>Probably none of the submissions used the images yet, since they are still far below the benchmark.. That's understandable due to the size of the dataset. You may want to look into past competitions that involved computer vision. The usual approach is to use a conv neural net to extract features from images in the form of a vector, and then feed the vector into a classification model of your choice. That may require a GPU, but can be done with CPU only, though it takes a lot longer.</p>\n\n<p>[/quote]</p>\n\n<p>47.3 Gb of images is absolutely ridiculous. I doubt many people have the hardware to process that many images.  Very few people might beat the Avito Benchmark.</p>",
      "rawMarkdown": "[quote=FernandoProcy;119182]\r\n\r\nProbably none of the submissions used the images yet, since they are still far below the benchmark.. That's understandable due to the size of the dataset. You may want to look into past competitions that involved computer vision. The usual approach is to use a conv neural net to extract features from images in the form of a vector, and then feed the vector into a classification model of your choice. That may require a GPU, but can be done with CPU only, though it takes a lot longer.\r\n\r\n[/quote]\r\n\r\n47.3 Gb of images is absolutely ridiculous. I doubt many people have the hardware to process that many images.  Very few people might beat the Avito Benchmark.",
      "votes": null
    },
    {
      "id": "119192",
      "postDate": "05/07/2016 22:12:55",
      "content": "<p>The welcome message has a hint on how it could be used. As is seen with the statefarm competition, neural nets on images can be very expensive to train. Because this is about determining how x and y are the same (or different), it's helpful if you think about the distance between x and y, thus also the distance between image x and y.</p>",
      "rawMarkdown": "The welcome message has a hint on how it could be used. As is seen with the statefarm competition, neural nets on images can be very expensive to train. Because this is about determining how x and y are the same (or different), it's helpful if you think about the distance between x and y, thus also the distance between image x and y.",
      "votes": null
    },
    {
      "id": "119268",
      "postDate": "05/08/2016 17:48:51",
      "content": "<p>I started to convert the image data. It is not really a big deal. On my  Macbook Pro (2012) unzipping for one image file takes about 1 hour 20 minutes and analyzing the images less than an hour. It is an overnight job.  Unzipped images take about 80 GB disk space.</p>",
      "rawMarkdown": "I started to convert the image data. It is not really a big deal. On my  Macbook Pro (2012) unzipping for one image file takes about 1 hour 20 minutes and analyzing the images less than an hour. It is an overnight job.  Unzipped images take about 80 GB disk space.",
      "votes": null
    },
    {
      "id": "119291",
      "postDate": "05/08/2016 22:11:21",
      "content": "<p>[quote=Peter Borrmann;119268]</p>\n\n<p>analyzing the images less than an hour</p>\n\n<p>[/quote]</p>\n\n<p>Really? I'm guessing this is fairly simple feature extraction, right?</p>",
      "rawMarkdown": "[quote=Peter Borrmann;119268]\r\n\r\nanalyzing the images less than an hour\r\n\r\n[/quote]\r\n\r\nReally? I'm guessing this is fairly simple feature extraction, right?",
      "votes": null
    },
    {
      "id": "119295",
      "postDate": "05/08/2016 22:52:05",
      "content": "<p>[quote=inversion;119291]</p>\n\n<p>[quote=Peter Borrmann;119268]</p>\n\n<p>analyzing the images less than an hour</p>\n\n<p>[/quote]</p>\n\n<p>Really? I'm guessing this is fairly simple feature extraction, right?</p>\n\n<p>[/quote]</p>\n\n<p>Yes, the simple hashing method mentioned  in the welcome message takes about 1 hour for one of the image files.</p>",
      "rawMarkdown": "[quote=inversion;119291]\r\n\r\n[quote=Peter Borrmann;119268]\r\n\r\nanalyzing the images less than an hour\r\n\r\n[/quote]\r\n\r\nReally? I'm guessing this is fairly simple feature extraction, right?\r\n\r\n[/quote]\r\n\r\nYes, the simple hashing method mentioned  in the welcome message takes about 1 hour for one of the image files.",
      "votes": null
    },
    {
      "id": "119363",
      "postDate": "05/09/2016 14:47:37",
      "content": "<p>Found a lot of stuff on the internet</p>\n\n<p><a href=\"http://blog.iconfinder.com/detecting-duplicate-images-using-python/\">http://blog.iconfinder.com/detecting-duplicate-images-using-python/</a></p>\n\n<p><a href=\"https://pypi.python.org/pypi/PhotoHash\">https://pypi.python.org/pypi/PhotoHash</a></p>\n\n<p><a href=\"https://pypi.python.org/pypi/ImageHash\">https://pypi.python.org/pypi/ImageHash</a></p>\n\n<p><a href=\"http://www.hackerfactor.com/blog/?/archives/432-Looks-Like-It.html\">http://www.hackerfactor.com/blog/?/archives/432-Looks-Like-It.html</a></p>\n\n<p>I was also learning about pHash <a href=\"http://www.phash.org/licensing/\">http://www.phash.org/licensing/</a></p>\n\n<p>I understand that cannot be used - right ? - or can it ?</p>",
      "rawMarkdown": "Found a lot of stuff on the internet\r\n\r\nhttp://blog.iconfinder.com/detecting-duplicate-images-using-python/\r\n\r\nhttps://pypi.python.org/pypi/PhotoHash\r\n\r\nhttps://pypi.python.org/pypi/ImageHash\r\n\r\nhttp://www.hackerfactor.com/blog/?/archives/432-Looks-Like-It.html\r\n\r\n\r\nI was also learning about pHash http://www.phash.org/licensing/\r\n\r\nI understand that cannot be used - right ? - or can it ?",
      "votes": null
    },
    {
      "id": "119466",
      "postDate": "05/10/2016 13:05:16",
      "content": "<p>Any answer on pHash ? I suppose we cannot use it - right ?</p>",
      "rawMarkdown": "Any answer on pHash ? I suppose we cannot use it - right ?",
      "votes": null
    },
    {
      "id": "119474",
      "postDate": "05/10/2016 15:14:30",
      "content": "<p>You cannot use implementation from phash.org. That is correct since their licence restrict it.\nHowever <a href=\"https://pypi.python.org/pypi/ImageHash\">https://pypi.python.org/pypi/ImageHash</a> license does not contain these restrictions as far as I see. And it does have pHash implementation. I guess you can check it yourself in the source code.</p>",
      "rawMarkdown": "You cannot use implementation from phash.org. That is correct since their licence restrict it.\r\nHowever https://pypi.python.org/pypi/ImageHash license does not contain these restrictions as far as I see. And it does have pHash implementation. I guess you can check it yourself in the source code.",
      "votes": null
    },
    {
      "id": "119507",
      "postDate": "05/11/2016 01:48:00",
      "content": "<p>[quote=Ivan Guz;119474]</p>\n\n<p>You cannot use implementation from phash.org. That is correct since their licence restrict it.</p>\n\n<p>[/quote]</p>\n\n<p>I would like to know what exactly is restricted so that we cannot use it. I don't know much about licenses and thank you in advance.</p>",
      "rawMarkdown": "[quote=Ivan Guz;119474]\r\n\r\nYou cannot use implementation from phash.org. That is correct since their licence restrict it.\r\n\r\n[/quote]\r\n\r\nI would like to know what exactly is restricted so that we cannot use it. I don't know much about licenses and thank you in advance.",
      "votes": null
    },
    {
      "id": "119955",
      "postDate": "05/14/2016 01:48:50",
      "content": "<p>I don't even seem to have the hardware to successfully download and extract the image files. :-(  The almost new computer with the hardwired ethernet connection but a HDD can download the zip files in about an hour apiece, but it appears to freeze up when I try to unzip the files. The laptop with a 1 TB SSD takes 3 hours per image zip file download over wifi but has been unzipping for over 12 hours. </p>\n\n<p>Has anyone uploaded the image files to AWS where it would be possible for other people to access them? </p>",
      "rawMarkdown": "I don't even seem to have the hardware to successfully download and extract the image files. :-(  The almost new computer with the hardwired ethernet connection but a HDD can download the zip files in about an hour apiece, but it appears to freeze up when I try to unzip the files. The laptop with a 1 TB SSD takes 3 hours per image zip file download over wifi but has been unzipping for over 12 hours. \r\n\r\nHas anyone uploaded the image files to AWS where it would be possible for other people to access them?",
      "votes": null
    },
    {
      "id": "120779",
      "postDate": "05/20/2016 14:25:43",
      "content": "<p>Hi,</p>\n\n<p>I might be looking wrongly at the images but, when I open two images supposed to be duplicates, it's a picture of a bear and the other a picture of pants Oo... I tried with 10 paies so far (from the itemPairs_train) and none of them are actual duplicates. What did I miss ?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi,\r\n\r\nI might be looking wrongly at the images but, when I open two images supposed to be duplicates, it's a picture of a bear and the other a picture of pants Oo... I tried with 10 paies so far (from the itemPairs_train) and none of them are actual duplicates. What did I miss ?\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "121044",
      "postDate": "05/23/2016 07:52:35",
      "content": "<p>Any image processing solution in r? </p>",
      "rawMarkdown": "Any image processing solution in r?",
      "votes": null
    },
    {
      "id": "121082",
      "postDate": "05/23/2016 17:15:37",
      "content": "<p>[quote=Floran Gmehlin;120779]</p>\n\n<p>Hi,</p>\n\n<p>I might be looking wrongly at the images but, when I open two images supposed to be duplicates, it's a picture of a bear and the other a picture of pants Oo... I tried with 10 paies so far (from the itemPairs_train) and none of them are actual duplicates. What did I miss ?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>The images aren't necessarily supposed to be duplicates; the ads are.  They could have identical textual content but different images, possibly.  Also - maybe they have several images apiece, and only a few are duplicate images while the rest are bears and pants.</p>",
      "rawMarkdown": "[quote=Floran Gmehlin;120779]\r\n\r\nHi,\r\n\r\nI might be looking wrongly at the images but, when I open two images supposed to be duplicates, it's a picture of a bear and the other a picture of pants Oo... I tried with 10 paies so far (from the itemPairs_train) and none of them are actual duplicates. What did I miss ?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nThe images aren't necessarily supposed to be duplicates; the ads are.  They could have identical textual content but different images, possibly.  Also - maybe they have several images apiece, and only a few are duplicate images while the rest are bears and pants.",
      "votes": null
    },
    {
      "id": "125090",
      "postDate": "06/25/2016 16:04:33",
      "content": "<p>Hi, I takes insanely long to download these images on my computer and also, I am runnnig out of disk space. If anybody is willing to share some precoumputed image hashes / feature it would be REALLY great</p>",
      "rawMarkdown": "Hi, I takes insanely long to download these images on my computer and also, I am runnnig out of disk space. If anybody is willing to share some precoumputed image hashes / feature it would be REALLY great",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119182,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "05/07/2016 20:50:44",
      "content": "<p>Probably none of the submissions used the images yet, since they are still far below the benchmark.. That's understandable due to the size of the dataset. You may want to look into past competitions that involved computer vision. The usual approach is to use a conv neural net to extract features from images in the form of a vector, and then feed the vector into a classification model of your choice. That may require a GPU, but can be done with CPU only, though it takes a lot longer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119190,
      "author_name": "notaapple",
      "author_url": "",
      "post_date": "05/07/2016 21:52:50",
      "content": "<p>[quote=FernandoProcy;119182]</p>\n\n<p>Probably none of the submissions used the images yet, since they are still far below the benchmark.. That's understandable due to the size of the dataset. You may want to look into past competitions that involved computer vision. The usual approach is to use a conv neural net to extract features from images in the form of a vector, and then feed the vector into a classification model of your choice. That may require a GPU, but can be done with CPU only, though it takes a lot longer.</p>\n\n<p>[/quote]</p>\n\n<p>47.3 Gb of images is absolutely ridiculous. I doubt many people have the hardware to process that many images.  Very few people might beat the Avito Benchmark.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119192,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "05/07/2016 22:12:55",
      "content": "<p>The welcome message has a hint on how it could be used. As is seen with the statefarm competition, neural nets on images can be very expensive to train. Because this is about determining how x and y are the same (or different), it's helpful if you think about the distance between x and y, thus also the distance between image x and y.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119268,
      "author_name": "thequants",
      "author_url": "",
      "post_date": "05/08/2016 17:48:51",
      "content": "<p>I started to convert the image data. It is not really a big deal. On my  Macbook Pro (2012) unzipping for one image file takes about 1 hour 20 minutes and analyzing the images less than an hour. It is an overnight job.  Unzipped images take about 80 GB disk space.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119291,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "05/08/2016 22:11:21",
      "content": "<p>[quote=Peter Borrmann;119268]</p>\n\n<p>analyzing the images less than an hour</p>\n\n<p>[/quote]</p>\n\n<p>Really? I'm guessing this is fairly simple feature extraction, right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119295,
      "author_name": "thequants",
      "author_url": "",
      "post_date": "05/08/2016 22:52:05",
      "content": "<p>[quote=inversion;119291]</p>\n\n<p>[quote=Peter Borrmann;119268]</p>\n\n<p>analyzing the images less than an hour</p>\n\n<p>[/quote]</p>\n\n<p>Really? I'm guessing this is fairly simple feature extraction, right?</p>\n\n<p>[/quote]</p>\n\n<p>Yes, the simple hashing method mentioned  in the welcome message takes about 1 hour for one of the image files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119363,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "05/09/2016 14:47:37",
      "content": "<p>Found a lot of stuff on the internet</p>\n\n<p><a href=\"http://blog.iconfinder.com/detecting-duplicate-images-using-python/\">http://blog.iconfinder.com/detecting-duplicate-images-using-python/</a></p>\n\n<p><a href=\"https://pypi.python.org/pypi/PhotoHash\">https://pypi.python.org/pypi/PhotoHash</a></p>\n\n<p><a href=\"https://pypi.python.org/pypi/ImageHash\">https://pypi.python.org/pypi/ImageHash</a></p>\n\n<p><a href=\"http://www.hackerfactor.com/blog/?/archives/432-Looks-Like-It.html\">http://www.hackerfactor.com/blog/?/archives/432-Looks-Like-It.html</a></p>\n\n<p>I was also learning about pHash <a href=\"http://www.phash.org/licensing/\">http://www.phash.org/licensing/</a></p>\n\n<p>I understand that cannot be used - right ? - or can it ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119466,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "05/10/2016 13:05:16",
      "content": "<p>Any answer on pHash ? I suppose we cannot use it - right ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119474,
      "author_name": "ivanguz",
      "author_url": "",
      "post_date": "05/10/2016 15:14:30",
      "content": "<p>You cannot use implementation from phash.org. That is correct since their licence restrict it.\nHowever <a href=\"https://pypi.python.org/pypi/ImageHash\">https://pypi.python.org/pypi/ImageHash</a> license does not contain these restrictions as far as I see. And it does have pHash implementation. I guess you can check it yourself in the source code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119507,
      "author_name": "contemplator",
      "author_url": "",
      "post_date": "05/11/2016 01:48:00",
      "content": "<p>[quote=Ivan Guz;119474]</p>\n\n<p>You cannot use implementation from phash.org. That is correct since their licence restrict it.</p>\n\n<p>[/quote]</p>\n\n<p>I would like to know what exactly is restricted so that we cannot use it. I don't know much about licenses and thank you in advance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119955,
      "author_name": "alicegif",
      "author_url": "",
      "post_date": "05/14/2016 01:48:50",
      "content": "<p>I don't even seem to have the hardware to successfully download and extract the image files. :-(  The almost new computer with the hardwired ethernet connection but a HDD can download the zip files in about an hour apiece, but it appears to freeze up when I try to unzip the files. The laptop with a 1 TB SSD takes 3 hours per image zip file download over wifi but has been unzipping for over 12 hours. </p>\n\n<p>Has anyone uploaded the image files to AWS where it would be possible for other people to access them? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120779,
      "author_name": "",
      "author_url": "",
      "post_date": "05/20/2016 14:25:43",
      "content": "<p>Hi,</p>\n\n<p>I might be looking wrongly at the images but, when I open two images supposed to be duplicates, it's a picture of a bear and the other a picture of pants Oo... I tried with 10 paies so far (from the itemPairs_train) and none of them are actual duplicates. What did I miss ?</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121044,
      "author_name": "nitsourish",
      "author_url": "",
      "post_date": "05/23/2016 07:52:35",
      "content": "<p>Any image processing solution in r? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121082,
      "author_name": "phillipadkins",
      "author_url": "",
      "post_date": "05/23/2016 17:15:37",
      "content": "<p>[quote=Floran Gmehlin;120779]</p>\n\n<p>Hi,</p>\n\n<p>I might be looking wrongly at the images but, when I open two images supposed to be duplicates, it's a picture of a bear and the other a picture of pants Oo... I tried with 10 paies so far (from the itemPairs_train) and none of them are actual duplicates. What did I miss ?</p>\n\n<p>Thanks</p>\n\n<p>[/quote]</p>\n\n<p>The images aren't necessarily supposed to be duplicates; the ads are.  They could have identical textual content but different images, possibly.  Also - maybe they have several images apiece, and only a few are duplicate images while the rest are bears and pants.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125090,
      "author_name": "lfiaschi",
      "author_url": "",
      "post_date": "06/25/2016 16:04:33",
      "content": "<p>Hi, I takes insanely long to download these images on my computer and also, I am runnnig out of disk space. If anybody is willing to share some precoumputed image hashes / feature it would be REALLY great</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119179": "Does anybody started using the actual images of the Ads? Any ideas how to turn this into features?",
    "119182": "Probably none of the submissions used the images yet, since they are still far below the benchmark.. That's understandable due to the size of the dataset. You may want to look into past competitions that involved computer vision. The usual approach is to use a conv neural net to extract features from images in the form of a vector, and then feed the vector into a classification model of your choice. That may require a GPU, but can be done with CPU only, though it takes a lot longer.",
    "119190": "[quote=FernandoProcy;119182]\r\n\r\nProbably none of the submissions used the images yet, since they are still far below the benchmark.. That's understandable due to the size of the dataset. You may want to look into past competitions that involved computer vision. The usual approach is to use a conv neural net to extract features from images in the form of a vector, and then feed the vector into a classification model of your choice. That may require a GPU, but can be done with CPU only, though it takes a lot longer.\r\n\r\n[/quote]\r\n\r\n47.3 Gb of images is absolutely ridiculous. I doubt many people have the hardware to process that many images.  Very few people might beat the Avito Benchmark.",
    "119192": "The welcome message has a hint on how it could be used. As is seen with the statefarm competition, neural nets on images can be very expensive to train. Because this is about determining how x and y are the same (or different), it's helpful if you think about the distance between x and y, thus also the distance between image x and y.",
    "119268": "I started to convert the image data. It is not really a big deal. On my  Macbook Pro (2012) unzipping for one image file takes about 1 hour 20 minutes and analyzing the images less than an hour. It is an overnight job.  Unzipped images take about 80 GB disk space.",
    "119291": "[quote=Peter Borrmann;119268]\r\n\r\nanalyzing the images less than an hour\r\n\r\n[/quote]\r\n\r\nReally? I'm guessing this is fairly simple feature extraction, right?",
    "119295": "[quote=inversion;119291]\r\n\r\n[quote=Peter Borrmann;119268]\r\n\r\nanalyzing the images less than an hour\r\n\r\n[/quote]\r\n\r\nReally? I'm guessing this is fairly simple feature extraction, right?\r\n\r\n[/quote]\r\n\r\nYes, the simple hashing method mentioned  in the welcome message takes about 1 hour for one of the image files.",
    "119363": "Found a lot of stuff on the internet\r\n\r\nhttp://blog.iconfinder.com/detecting-duplicate-images-using-python/\r\n\r\nhttps://pypi.python.org/pypi/PhotoHash\r\n\r\nhttps://pypi.python.org/pypi/ImageHash\r\n\r\nhttp://www.hackerfactor.com/blog/?/archives/432-Looks-Like-It.html\r\n\r\n\r\nI was also learning about pHash http://www.phash.org/licensing/\r\n\r\nI understand that cannot be used - right ? - or can it ?",
    "119466": "Any answer on pHash ? I suppose we cannot use it - right ?",
    "119474": "You cannot use implementation from phash.org. That is correct since their licence restrict it.\r\nHowever https://pypi.python.org/pypi/ImageHash license does not contain these restrictions as far as I see. And it does have pHash implementation. I guess you can check it yourself in the source code.",
    "119507": "[quote=Ivan Guz;119474]\r\n\r\nYou cannot use implementation from phash.org. That is correct since their licence restrict it.\r\n\r\n[/quote]\r\n\r\nI would like to know what exactly is restricted so that we cannot use it. I don't know much about licenses and thank you in advance.",
    "119955": "I don't even seem to have the hardware to successfully download and extract the image files. :-(  The almost new computer with the hardwired ethernet connection but a HDD can download the zip files in about an hour apiece, but it appears to freeze up when I try to unzip the files. The laptop with a 1 TB SSD takes 3 hours per image zip file download over wifi but has been unzipping for over 12 hours. \r\n\r\nHas anyone uploaded the image files to AWS where it would be possible for other people to access them?",
    "120779": "Hi,\r\n\r\nI might be looking wrongly at the images but, when I open two images supposed to be duplicates, it's a picture of a bear and the other a picture of pants Oo... I tried with 10 paies so far (from the itemPairs_train) and none of them are actual duplicates. What did I miss ?\r\n\r\nThanks",
    "121044": "Any image processing solution in r?",
    "121082": "[quote=Floran Gmehlin;120779]\r\n\r\nHi,\r\n\r\nI might be looking wrongly at the images but, when I open two images supposed to be duplicates, it's a picture of a bear and the other a picture of pants Oo... I tried with 10 paies so far (from the itemPairs_train) and none of them are actual duplicates. What did I miss ?\r\n\r\nThanks\r\n\r\n[/quote]\r\n\r\nThe images aren't necessarily supposed to be duplicates; the ads are.  They could have identical textual content but different images, possibly.  Also - maybe they have several images apiece, and only a few are duplicate images while the rest are bears and pants.",
    "125090": "Hi, I takes insanely long to download these images on my computer and also, I am runnnig out of disk space. If anybody is willing to share some precoumputed image hashes / feature it would be REALLY great"
  },
  "source": "meta"
}