{
  "id": 73928,
  "title": "Other uses for the HPA data",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/73928",
  "author_name": "",
  "post_date": "2018-12-06T19:09:40.844539600Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Now that the leak score boost is available to all, there is still a question of how to use the external data. I would like to try using it simply to augment the challenge data. I have a question about how to do this.</p>\n\n<p>I could simply resize appropriately and just throw it in the training data but I think that this approach will not work. The problem is that if there are external duplicates either with the challenge train data or within the external data, the train splitting will be putting the same image in train and validate sets. This creates a leak between train and validate and will make the validation step useless in detecting overfitting. A couple of possibilities:</p>\n\n<p>1 rigorously find and remove all such duplicates.\n2 make a rotation or dihedral transform on ALL validation images after splitting\n3 throw out duplicates based on their labels only\n4 ignore the problem and dont use validation in the workflow\n5 dont use the external data for this purpose</p>\n\n<p>The first option means I will have to have a robust understanding of how to find duplicates in large sets of large images. As Phil mentioned in the leak thread, this is not a trivial task. There are some lists out but I imagine that this is a field with lots of opinions. So this option might take quite a bit of time but seems the most certain to succeed.\nThe second option might fail. For example, if you have a training image and flip it l-r, I think the result will leak to its training partner. Any augmentation used may not be enough (although, for example, the dihedral will produce an image that will likely fail any duplicate test).\nThird might be done in by noisy labels\nFourth sounds dicey\nFifth is a major possibility</p>\n\n<p>Any suggestions? This might be a non-issue - if so, I hope someone will tell me.</p>",
  "messages": [
    {
      "id": "434678",
      "postDate": "12/06/2018 19:09:40",
      "content": "<p>Now that the leak score boost is available to all, there is still a question of how to use the external data. I would like to try using it simply to augment the challenge data. I have a question about how to do this.</p>\n\n<p>I could simply resize appropriately and just throw it in the training data but I think that this approach will not work. The problem is that if there are external duplicates either with the challenge train data or within the external data, the train splitting will be putting the same image in train and validate sets. This creates a leak between train and validate and will make the validation step useless in detecting overfitting. A couple of possibilities:</p>\n\n<p>1 rigorously find and remove all such duplicates.\n2 make a rotation or dihedral transform on ALL validation images after splitting\n3 throw out duplicates based on their labels only\n4 ignore the problem and dont use validation in the workflow\n5 dont use the external data for this purpose</p>\n\n<p>The first option means I will have to have a robust understanding of how to find duplicates in large sets of large images. As Phil mentioned in the leak thread, this is not a trivial task. There are some lists out but I imagine that this is a field with lots of opinions. So this option might take quite a bit of time but seems the most certain to succeed.\nThe second option might fail. For example, if you have a training image and flip it l-r, I think the result will leak to its training partner. Any augmentation used may not be enough (although, for example, the dihedral will produce an image that will likely fail any duplicate test).\nThird might be done in by noisy labels\nFourth sounds dicey\nFifth is a major possibility</p>\n\n<p>Any suggestions? This might be a non-issue - if so, I hope someone will tell me.</p>",
      "rawMarkdown": "Now that the leak score boost is available to all, there is still a question of how to use the external data. I would like to try using it simply to augment the challenge data. I have a question about how to do this.\n\nI could simply resize appropriately and just throw it in the training data but I think that this approach will not work. The problem is that if there are external duplicates either with the challenge train data or within the external data, the train splitting will be putting the same image in train and validate sets. This creates a leak between train and validate and will make the validation step useless in detecting overfitting. A couple of possibilities:\n\n1 rigorously find and remove all such duplicates.\n2 make a rotation or dihedral transform on ALL validation images after splitting\n3 throw out duplicates based on their labels only\n4 ignore the problem and dont use validation in the workflow\n5 dont use the external data for this purpose\n\nThe first option means I will have to have a robust understanding of how to find duplicates in large sets of large images. As Phil mentioned in the leak thread, this is not a trivial task. There are some lists out but I imagine that this is a field with lots of opinions. So this option might take quite a bit of time but seems the most certain to succeed.\nThe second option might fail. For example, if you have a training image and flip it l-r, I think the result will leak to its training partner. Any augmentation used may not be enough (although, for example, the dihedral will produce an image that will likely fail any duplicate test).\nThird might be done in by noisy labels\nFourth sounds dicey\nFifth is a major possibility\n\nAny suggestions? This might be a non-issue - if so, I hope someone will tell me.",
      "votes": null
    },
    {
      "id": "435475",
      "postDate": "12/08/2018 04:35:57",
      "content": "<p>I don't use external data.</p>",
      "rawMarkdown": "I don't use external data.",
      "votes": null
    },
    {
      "id": "439460",
      "postDate": "12/15/2018 14:13:48",
      "content": "<p>Hello, Could you tell me what thresholds are you choice?</p>",
      "rawMarkdown": "Hello, Could you tell me what thresholds are you choice?",
      "votes": null
    },
    {
      "id": "442467",
      "postDate": "12/20/2018 02:06:10",
      "content": "<p>0.2 or 0.3</p>",
      "rawMarkdown": "0.2 or 0.3",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 435475,
      "author_name": "crilinux",
      "author_url": "",
      "post_date": "12/08/2018 04:35:57",
      "content": "<p>I don't use external data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 439460,
          "author_name": "lianandrew",
          "author_url": "",
          "post_date": "12/15/2018 14:13:48",
          "content": "<p>Hello, Could you tell me what thresholds are you choice?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442467,
          "author_name": "crilinux",
          "author_url": "",
          "post_date": "12/20/2018 02:06:10",
          "content": "<p>0.2 or 0.3</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "434678": "Now that the leak score boost is available to all, there is still a question of how to use the external data. I would like to try using it simply to augment the challenge data. I have a question about how to do this.\n\nI could simply resize appropriately and just throw it in the training data but I think that this approach will not work. The problem is that if there are external duplicates either with the challenge train data or within the external data, the train splitting will be putting the same image in train and validate sets. This creates a leak between train and validate and will make the validation step useless in detecting overfitting. A couple of possibilities:\n\n1 rigorously find and remove all such duplicates.\n2 make a rotation or dihedral transform on ALL validation images after splitting\n3 throw out duplicates based on their labels only\n4 ignore the problem and dont use validation in the workflow\n5 dont use the external data for this purpose\n\nThe first option means I will have to have a robust understanding of how to find duplicates in large sets of large images. As Phil mentioned in the leak thread, this is not a trivial task. There are some lists out but I imagine that this is a field with lots of opinions. So this option might take quite a bit of time but seems the most certain to succeed.\nThe second option might fail. For example, if you have a training image and flip it l-r, I think the result will leak to its training partner. Any augmentation used may not be enough (although, for example, the dihedral will produce an image that will likely fail any duplicate test).\nThird might be done in by noisy labels\nFourth sounds dicey\nFifth is a major possibility\n\nAny suggestions? This might be a non-issue - if so, I hope someone will tell me.",
    "435475": "I don't use external data.",
    "439460": "Hello, Could you tell me what thresholds are you choice?",
    "442467": "0.2 or 0.3"
  },
  "source": "meta"
}