{
  "id": 18675,
  "title": "Duplicate photo ids in test data test_photo_to_biz.csv",
  "url": "/competitions/yelp-restaurant-photo-classification/discussion/18675",
  "author_name": "",
  "post_date": "2016-01-31T22:28:00.427Z",
  "votes": null,
  "comment_count": 4,
  "views": 1122,
  "content": "<p>E.g. photo id 1 is repeated several times for different businesses. How should I understand that? </p>\n\n<p>E.g. for businesses 3tv9h and 4udt2</p>",
  "messages": [
    {
      "id": "106453",
      "postDate": "01/31/2016 22:28:00",
      "content": "<p>E.g. photo id 1 is repeated several times for different businesses. How should I understand that? </p>\n\n<p>E.g. for businesses 3tv9h and 4udt2</p>",
      "rawMarkdown": "E.g. photo id 1 is repeated several times for different businesses. How should I understand that? \r\n\r\nE.g. for businesses 3tv9h and 4udt2",
      "votes": null
    },
    {
      "id": "106638",
      "postDate": "02/02/2016 19:07:33",
      "content": "<p>In the test set there are some businesses which have been added for noise.  As mentioned in the data page: 'To deter hand labeling, Kaggle has supplemented the test set with additional &quot;ignored&quot; businesses. These are not counted in the scoring.'  See this thread as well: <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18335/business-ids-in-test\">https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18335/business-ids-in-test</a>.</p>",
      "rawMarkdown": "In the test set there are some businesses which have been added for noise.  As mentioned in the data page: 'To deter hand labeling, Kaggle has supplemented the test set with additional \"ignored\" businesses. These are not counted in the scoring.'  See this thread as well: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18335/business-ids-in-test.",
      "votes": null
    },
    {
      "id": "109066",
      "postDate": "02/22/2016 23:48:33",
      "content": "<p>Apparently they did this a <strong>lot</strong>. IMHO it was unnecessary. </p>\n\n<p>We already have a ridiculous number of images and businesses to process in a competition with a $0 prize... no need to artificially jack up the size of the data and make it more convoluted with fake businesses that break logical algorithms based on the idea that each business has a unique set of photos. </p>",
      "rawMarkdown": "Apparently they did this a **lot**. IMHO it was unnecessary. \r\n\r\nWe already have a ridiculous number of images and businesses to process in a competition with a $0 prize... no need to artificially jack up the size of the data and make it more convoluted with fake businesses that break logical algorithms based on the idea that each business has a unique set of photos.",
      "votes": null
    },
    {
      "id": "109151",
      "postDate": "02/23/2016 19:19:40",
      "content": "<p>There is indeed quite some processing overhead if you process those images multiple times. </p>\n\n<p>To avoid that, I  do two stages: First scoring of each test image individually and storing the results in a dataframe, and then joining the scores to the businesses in the second stage. This way I don't need to process the images multiple times. Maybe that can work for your approach as well.</p>",
      "rawMarkdown": "There is indeed quite some processing overhead if you process those images multiple times. \r\n\r\nTo avoid that, I  do two stages: First scoring of each test image individually and storing the results in a dataframe, and then joining the scores to the businesses in the second stage. This way I don't need to process the images multiple times. Maybe that can work for your approach as well.",
      "votes": null
    },
    {
      "id": "112917",
      "postDate": "03/25/2016 02:27:03",
      "content": "<p>Yeah, this is pervasive :( .. </p>\n\n<p>There are 1,190,225 rows in test_photo_to_biz.csv</p>\n\n<p>There are only 237,153 files in test_photos</p>\n\n<p>So among the 237,152 photos, they're reused on average 5x</p>\n\n<p>That probably means that the real business only account for 1/5 of the actual entries and they just sampled from the other businesses randomly to fill in the gaps. </p>\n\n<p>Attached some screenshots of a table and a graph showing the dupes.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/112917/3990/09 PM.png?sv=2012-02-12&se=2016-03-28T02:27:50Z&sr=b&sp=r&sig=2T6%2BROJpats1b2oYcPq3A4XGGoytjrlVYvJdgcYQ6Ng%3D\" alt=\"graph\" title> <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/112917/3991/52 PM.png?sv=2012-02-12&se=2016-03-28T02:28:01Z&sr=b&sp=r&sig=2rT%2B0eYPNvKMMYVlK%2FVK%2FwILswX7z2wCKDgXLPBAH0A%3D\" alt=\"table\" title></p>",
      "rawMarkdown": "Yeah, this is pervasive :( .. \r\n\r\nThere are 1,190,225 rows in test_photo_to_biz.csv\r\n\r\nThere are only 237,153 files in test_photos\r\n\r\nSo among the 237,152 photos, they're reused on average 5x\r\n\r\nThat probably means that the real business only account for 1/5 of the actual entries and they just sampled from the other businesses randomly to fill in the gaps. \r\n\r\nAttached some screenshots of a table and a graph showing the dupes.\r\n\r\n![graph][1] ![table][2]\r\n\r\n\r\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/112917/3990/09%20PM.png?sv=2012-02-12&se=2016-03-28T02%3A27%3A50Z&sr=b&sp=r&sig=2T6%2BROJpats1b2oYcPq3A4XGGoytjrlVYvJdgcYQ6Ng%3D\r\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/112917/3991/52%20PM.png?sv=2012-02-12&se=2016-03-28T02%3A28%3A01Z&sr=b&sp=r&sig=2rT%2B0eYPNvKMMYVlK%2FVK%2FwILswX7z2wCKDgXLPBAH0A%3D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 106638,
      "author_name": "ody30357",
      "author_url": "",
      "post_date": "02/02/2016 19:07:33",
      "content": "<p>In the test set there are some businesses which have been added for noise.  As mentioned in the data page: 'To deter hand labeling, Kaggle has supplemented the test set with additional &quot;ignored&quot; businesses. These are not counted in the scoring.'  See this thread as well: <a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18335/business-ids-in-test\">https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18335/business-ids-in-test</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109066,
      "author_name": "millerintllc",
      "author_url": "",
      "post_date": "02/22/2016 23:48:33",
      "content": "<p>Apparently they did this a <strong>lot</strong>. IMHO it was unnecessary. </p>\n\n<p>We already have a ridiculous number of images and businesses to process in a competition with a $0 prize... no need to artificially jack up the size of the data and make it more convoluted with fake businesses that break logical algorithms based on the idea that each business has a unique set of photos. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109151,
      "author_name": "allexius",
      "author_url": "",
      "post_date": "02/23/2016 19:19:40",
      "content": "<p>There is indeed quite some processing overhead if you process those images multiple times. </p>\n\n<p>To avoid that, I  do two stages: First scoring of each test image individually and storing the results in a dataframe, and then joining the scores to the businesses in the second stage. This way I don't need to process the images multiple times. Maybe that can work for your approach as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 112917,
      "author_name": "frankcarey",
      "author_url": "",
      "post_date": "03/25/2016 02:27:03",
      "content": "<p>Yeah, this is pervasive :( .. </p>\n\n<p>There are 1,190,225 rows in test_photo_to_biz.csv</p>\n\n<p>There are only 237,153 files in test_photos</p>\n\n<p>So among the 237,152 photos, they're reused on average 5x</p>\n\n<p>That probably means that the real business only account for 1/5 of the actual entries and they just sampled from the other businesses randomly to fill in the gaps. </p>\n\n<p>Attached some screenshots of a table and a graph showing the dupes.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/112917/3990/09 PM.png?sv=2012-02-12&se=2016-03-28T02:27:50Z&sr=b&sp=r&sig=2T6%2BROJpats1b2oYcPq3A4XGGoytjrlVYvJdgcYQ6Ng%3D\" alt=\"graph\" title> <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/112917/3991/52 PM.png?sv=2012-02-12&se=2016-03-28T02:28:01Z&sr=b&sp=r&sig=2rT%2B0eYPNvKMMYVlK%2FVK%2FwILswX7z2wCKDgXLPBAH0A%3D\" alt=\"table\" title></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "106453": "E.g. photo id 1 is repeated several times for different businesses. How should I understand that? \r\n\r\nE.g. for businesses 3tv9h and 4udt2",
    "106638": "In the test set there are some businesses which have been added for noise.  As mentioned in the data page: 'To deter hand labeling, Kaggle has supplemented the test set with additional \"ignored\" businesses. These are not counted in the scoring.'  See this thread as well: https://www.kaggle.com/c/yelp-restaurant-photo-classification/forums/t/18335/business-ids-in-test.",
    "109066": "Apparently they did this a **lot**. IMHO it was unnecessary. \r\n\r\nWe already have a ridiculous number of images and businesses to process in a competition with a $0 prize... no need to artificially jack up the size of the data and make it more convoluted with fake businesses that break logical algorithms based on the idea that each business has a unique set of photos.",
    "109151": "There is indeed quite some processing overhead if you process those images multiple times. \r\n\r\nTo avoid that, I  do two stages: First scoring of each test image individually and storing the results in a dataframe, and then joining the scores to the businesses in the second stage. This way I don't need to process the images multiple times. Maybe that can work for your approach as well.",
    "112917": "Yeah, this is pervasive :( .. \r\n\r\nThere are 1,190,225 rows in test_photo_to_biz.csv\r\n\r\nThere are only 237,153 files in test_photos\r\n\r\nSo among the 237,152 photos, they're reused on average 5x\r\n\r\nThat probably means that the real business only account for 1/5 of the actual entries and they just sampled from the other businesses randomly to fill in the gaps. \r\n\r\nAttached some screenshots of a table and a graph showing the dupes.\r\n\r\n![graph][1] ![table][2]\r\n\r\n\r\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/112917/3990/09%20PM.png?sv=2012-02-12&se=2016-03-28T02%3A27%3A50Z&sr=b&sp=r&sig=2T6%2BROJpats1b2oYcPq3A4XGGoytjrlVYvJdgcYQ6Ng%3D\r\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/112917/3991/52%20PM.png?sv=2012-02-12&se=2016-03-28T02%3A28%3A01Z&sr=b&sp=r&sig=2rT%2B0eYPNvKMMYVlK%2FVK%2FwILswX7z2wCKDgXLPBAH0A%3D"
  },
  "source": "meta"
}