{
  "id": 342505,
  "title": "Generalization from HPA to HuBMAP",
  "url": "/competitions/hubmap-organ-segmentation/discussion/342505",
  "author_name": "",
  "post_date": "2022-08-07T17:52:59.024641100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In the train dataset, the data_source field is set to HPA only. How is it possible to generalize to HuBMAP data? What ideas might there be?</p>",
  "messages": [
    {
      "id": "1888644",
      "postDate": "08/07/2022 17:52:59",
      "content": "<p>In the train dataset, the data_source field is set to HPA only. How is it possible to generalize to HuBMAP data? What ideas might there be?</p>",
      "rawMarkdown": "In the train dataset, the data_source field is set to HPA only. How is it possible to generalize to HuBMAP data? What ideas might there be?",
      "votes": null
    },
    {
      "id": "1888696",
      "postDate": "08/07/2022 18:15:16",
      "content": "<p>Use base model without image augmentation, and then when you get a submission score for your base model, you slowly add 1 augmentation type to the training image, train your model with that, and if your submission score increases you have good augmentation that generalizes your model to HuBMAP data. And you iterate over the augmentations that gives you best results.</p>",
      "rawMarkdown": "Use base model without image augmentation, and then when you get a submission score for your base model, you slowly add 1 augmentation type to the training image, train your model with that, and if your submission score increases you have good augmentation that generalizes your model to HuBMAP data. And you iterate over the augmentations that gives you best results.",
      "votes": null
    },
    {
      "id": "1888711",
      "postDate": "08/07/2022 18:30:58",
      "content": "<p>Thank you! This solution is the worst thing you can think of that works :) We do not control overfitting :( What else can you think of? That's the question!  :)</p>",
      "rawMarkdown": "Thank you! This solution is the worst thing you can think of that works :) We do not control overfitting :( What else can you think of? That's the question!  :)",
      "votes": null
    },
    {
      "id": "1890735",
      "postDate": "08/09/2022 03:06:44",
      "content": "<p>\"you slowly add 1 augmentation type to the training image, train your model with that, and if your submission score increases \"</p>\n<p>a better way is to \"create\" your own H&amp;E dataset and slowly add images so that the score of this images is the same as public lb-humap score</p>\n<p>\"create\" can mean:</p>\n<ul>\n<li>augment from HPA (using stain tools, changing color, simulate thickness/resolution, artifacts)</li>\n<li>GAN generated</li>\n<li>download from internet</li>\n</ul>\n<hr>\n<p>if you are good are modeling, you need to reach a loss point where the score will not change when you image is perturbed.<br>\nThis can be done by:</p>\n<ul>\n<li>flat valley optimizer, e.g. Shaperness aware, SWA</li>\n<li>adversarial perturbation training</li>\n</ul>\n<p>this method is the same as above. but now the \"augmented\" samples are \"virtual\" samples in the feature space (instead of image space)</p>",
      "rawMarkdown": "\"you slowly add 1 augmentation type to the training image, train your model with that, and if your submission score increases \"\n\na better way is to \"create\" your own H&E dataset and slowly add images so that the score of this images is the same as public lb-humap score\n\n\"create\" can mean:\n- augment from HPA (using stain tools, changing color, simulate thickness/resolution, artifacts)\n- GAN generated\n- download from internet\n\n---\n\nif you are good are modeling, you need to reach a loss point where the score will not change when you image is perturbed.\nThis can be done by:\n- flat valley optimizer, e.g. Shaperness aware, SWA\n- adversarial perturbation training\n\nthis method is the same as above. but now the \"augmented\" samples are \"virtual\" samples in the feature space (instead of image space)",
      "votes": null
    },
    {
      "id": "1892269",
      "postDate": "08/10/2022 02:08:54",
      "content": "<p>My approach is basically to not allow the model to learn color based patterns, since the only difference between HPA (the train dataset) and the HuBMAP (part of the hidden test dataset), besides their image size, is that they have different stain agent used (different colors). In my case, I try generalizing it by \"randomly\" switching RGB channels of the train data.</p>",
      "rawMarkdown": "My approach is basically to not allow the model to learn color based patterns, since the only difference between HPA (the train dataset) and the HuBMAP (part of the hidden test dataset), besides their image size, is that they have different stain agent used (different colors). In my case, I try generalizing it by \"randomly\" switching RGB channels of the train data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1888696,
      "author_name": "urosjarc",
      "author_url": "",
      "post_date": "08/07/2022 18:15:16",
      "content": "<p>Use base model without image augmentation, and then when you get a submission score for your base model, you slowly add 1 augmentation type to the training image, train your model with that, and if your submission score increases you have good augmentation that generalizes your model to HuBMAP data. And you iterate over the augmentations that gives you best results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1890735,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/09/2022 03:06:44",
          "content": "<p>\"you slowly add 1 augmentation type to the training image, train your model with that, and if your submission score increases \"</p>\n<p>a better way is to \"create\" your own H&amp;E dataset and slowly add images so that the score of this images is the same as public lb-humap score</p>\n<p>\"create\" can mean:</p>\n<ul>\n<li>augment from HPA (using stain tools, changing color, simulate thickness/resolution, artifacts)</li>\n<li>GAN generated</li>\n<li>download from internet</li>\n</ul>\n<hr>\n<p>if you are good are modeling, you need to reach a loss point where the score will not change when you image is perturbed.<br>\nThis can be done by:</p>\n<ul>\n<li>flat valley optimizer, e.g. Shaperness aware, SWA</li>\n<li>adversarial perturbation training</li>\n</ul>\n<p>this method is the same as above. but now the \"augmented\" samples are \"virtual\" samples in the feature space (instead of image space)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1888711,
      "author_name": "alien308",
      "author_url": "",
      "post_date": "08/07/2022 18:30:58",
      "content": "<p>Thank you! This solution is the worst thing you can think of that works :) We do not control overfitting :( What else can you think of? That's the question!  :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1892269,
      "author_name": "seraphwedd18",
      "author_url": "",
      "post_date": "08/10/2022 02:08:54",
      "content": "<p>My approach is basically to not allow the model to learn color based patterns, since the only difference between HPA (the train dataset) and the HuBMAP (part of the hidden test dataset), besides their image size, is that they have different stain agent used (different colors). In my case, I try generalizing it by \"randomly\" switching RGB channels of the train data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1888644": "In the train dataset, the data_source field is set to HPA only. How is it possible to generalize to HuBMAP data? What ideas might there be?",
    "1888696": "Use base model without image augmentation, and then when you get a submission score for your base model, you slowly add 1 augmentation type to the training image, train your model with that, and if your submission score increases you have good augmentation that generalizes your model to HuBMAP data. And you iterate over the augmentations that gives you best results.",
    "1888711": "Thank you! This solution is the worst thing you can think of that works :) We do not control overfitting :( What else can you think of? That's the question!  :)",
    "1890735": "\"you slowly add 1 augmentation type to the training image, train your model with that, and if your submission score increases \"\n\na better way is to \"create\" your own H&E dataset and slowly add images so that the score of this images is the same as public lb-humap score\n\n\"create\" can mean:\n- augment from HPA (using stain tools, changing color, simulate thickness/resolution, artifacts)\n- GAN generated\n- download from internet\n\n---\n\nif you are good are modeling, you need to reach a loss point where the score will not change when you image is perturbed.\nThis can be done by:\n- flat valley optimizer, e.g. Shaperness aware, SWA\n- adversarial perturbation training\n\nthis method is the same as above. but now the \"augmented\" samples are \"virtual\" samples in the feature space (instead of image space)",
    "1892269": "My approach is basically to not allow the model to learn color based patterns, since the only difference between HPA (the train dataset) and the HuBMAP (part of the hidden test dataset), besides their image size, is that they have different stain agent used (different colors). In my case, I try generalizing it by \"randomly\" switching RGB channels of the train data."
  },
  "source": "meta"
}