{
  "id": 199539,
  "title": "Can we use old competition data? (probably)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199539",
  "author_name": "Kerem Turgutlu",
  "post_date": "2020-11-26T05:37:59.508000",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>We know that there is an old competition with exactly same targets and with a lot of unlabelled data to be potentially used for semi-supervised or unsupervised methods. But there is no explanation whether those are duplicates of what we already have here. </p>\n<p>I looked at image hashes to compare new vs old data and only found <code>3721/22025</code> (using average hash) and <code>4902/ 22025</code> (using perceptual hash) duplicates , but there may be more depending hash function.</p>\n<p>For more details about image duplicates notebook:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/keremt/01-cassava-eda\" target=\"_blank\">https://www.kaggle.com/keremt/01-cassava-eda</a></li>\n</ul>\n<p>Data is here made public:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/keremt/cassavaold\" target=\"_blank\">https://www.kaggle.com/keremt/cassavaold</a> (this is already zipped opposed to original competition data which you need to unzip in kernel before you use)</li>\n</ul>\n<p>But it would be great if organizers can just tell us if all are duplicated. <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> </p>\n<p>Greater the hashsize, more inform can be encoded hence similarity becomes more sensitive. Best is to try different hash sizes; default is 8.</p>\n<p>Phash, hashsize=4 3493/22025<br>\nPhash, hashsize=8 4902/22025<br>\nPhash, hashsize=16 3524/22025</p>",
  "messages": [
    {
      "id": 1091573,
      "postDate": "2020-11-26T05:37:59.510Z",
      "content": "<p>We know that there is an old competition with exactly same targets and with a lot of unlabelled data to be potentially used for semi-supervised or unsupervised methods. But there is no explanation whether those are duplicates of what we already have here. </p>\n<p>I looked at image hashes to compare new vs old data and only found <code>3721/22025</code> (using average hash) and <code>4902/ 22025</code> (using perceptual hash) duplicates , but there may be more depending hash function.</p>\n<p>For more details about image duplicates notebook:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/keremt/01-cassava-eda\" target=\"_blank\">https://www.kaggle.com/keremt/01-cassava-eda</a></li>\n</ul>\n<p>Data is here made public:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/keremt/cassavaold\" target=\"_blank\">https://www.kaggle.com/keremt/cassavaold</a> (this is already zipped opposed to original competition data which you need to unzip in kernel before you use)</li>\n</ul>\n<p>But it would be great if organizers can just tell us if all are duplicated. <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> </p>\n<p>Greater the hashsize, more inform can be encoded hence similarity becomes more sensitive. Best is to try different hash sizes; default is 8.</p>\n<p>Phash, hashsize=4 3493/22025<br>\nPhash, hashsize=8 4902/22025<br>\nPhash, hashsize=16 3524/22025</p>",
      "rawMarkdown": "We know that there is an old competition with exactly same targets and with a lot of unlabelled data to be potentially used for semi-supervised or unsupervised methods. But there is no explanation whether those are duplicates of what we already have here. \n\nI looked at image hashes to compare new vs old data and only found `3721/22025` (using average hash) and `4902/ 22025` (using perceptual hash) duplicates , but there may be more depending hash function.\n\nFor more details about image duplicates notebook:\n-  https://www.kaggle.com/keremt/01-cassava-eda\n\nData is here made public:\n-  https://www.kaggle.com/keremt/cassavaold (this is already zipped opposed to original competition data which you need to unzip in kernel before you use)\n\nBut it would be great if organizers can just tell us if all are duplicated. @juliaelliott \n\n\nGreater the hashsize, more inform can be encoded hence similarity becomes more sensitive. Best is to try different hash sizes; default is 8.\n\nPhash, hashsize=4 3493/22025\nPhash, hashsize=8 4902/22025\nPhash, hashsize=16 3524/22025",
      "votes": 8
    },
    {
      "id": 1185700,
      "postDate": "2021-02-04T09:50:06.510Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1185700,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-04T09:50:06.510000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1091573": "We know that there is an old competition with exactly same targets and with a lot of unlabelled data to be potentially used for semi-supervised or unsupervised methods. But there is no explanation whether those are duplicates of what we already have here. \n\nI looked at image hashes to compare new vs old data and only found `3721/22025` (using average hash) and `4902/ 22025` (using perceptual hash) duplicates , but there may be more depending hash function.\n\nFor more details about image duplicates notebook:\n-  https://www.kaggle.com/keremt/01-cassava-eda\n\nData is here made public:\n-  https://www.kaggle.com/keremt/cassavaold (this is already zipped opposed to original competition data which you need to unzip in kernel before you use)\n\nBut it would be great if organizers can just tell us if all are duplicated. @juliaelliott \n\n\nGreater the hashsize, more inform can be encoded hence similarity becomes more sensitive. Best is to try different hash sizes; default is 8.\n\nPhash, hashsize=4 3493/22025\nPhash, hashsize=8 4902/22025\nPhash, hashsize=16 3524/22025",
    "1185700": ""
  }
}