{
  "id": 87069,
  "title": "Better way to normalize slides?: WSI Normalization",
  "url": "/competitions/histopathologic-cancer-detection/discussion/87069",
  "author_name": "",
  "post_date": "2019-03-28T14:59:17.287205200Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p><img src=\"https://i.postimg.cc/c12VKGZr/WSI-normilization.png\" alt=\"\">\nYou can train a classification network using WSI ids as labels. The slides with the same stain will be classified together. And then you can normalize each image according to its predicted WSI ids. Notice that the classification scores may not be high, but that should not be a problem since one kind of stain can have multiple WSI ids.\nI think it is better than stain normalization. What do you think?</p>",
  "messages": [
    {
      "id": "502417",
      "postDate": "03/28/2019 14:59:17",
      "content": "<p><img src=\"https://i.postimg.cc/c12VKGZr/WSI-normilization.png\" alt=\"\">\nYou can train a classification network using WSI ids as labels. The slides with the same stain will be classified together. And then you can normalize each image according to its predicted WSI ids. Notice that the classification scores may not be high, but that should not be a problem since one kind of stain can have multiple WSI ids.\nI think it is better than stain normalization. What do you think?</p>",
      "rawMarkdown": "![](https://i.postimg.cc/c12VKGZr/WSI-normilization.png)\nYou can train a classification network using WSI ids as labels. The slides with the same stain will be classified together. And then you can normalize each image according to its predicted WSI ids. Notice that the classification scores may not be high, but that should not be a problem since one kind of stain can have multiple WSI ids.\nI think it is better than stain normalization. What do you think?",
      "votes": null
    },
    {
      "id": "502513",
      "postDate": "03/28/2019 17:09:27",
      "content": "<p>This highlights something that I’ve been thinking about for a while, which is the medical utility of this patch-based benchmark in the first place. </p>\n\n<p>I do not believe real diagnoses will be made from such disconnected image patches as are given in the pcam set. However, the tractability of the patches vs the full slide dataset which is many hundreds of GB could be a very nice benefit if our models are able to perform well on the larger images after being trained on patches. I have not tested this yet but I will once this is over and I have a good idea of how well my models actually did. Many people here have already achieved an AUC score above the published benchmark (actually the leak will let us check our models against the full public and private set, I doubt the difference between them is great but with Kaggle you never know), so it will be useful to see the generalization. </p>\n\n<p>Overall I feel using information that is not available from the patches alone defeats the purpose of a patch-based approach but I’m not convinced that patches are the most useful thing in the first place, though personally I did not and will not use any non-patch info for this task. </p>",
      "rawMarkdown": "This highlights something that I’ve been thinking about for a while, which is the medical utility of this patch-based benchmark in the first place. \n\nI do not believe real diagnoses will be made from such disconnected image patches as are given in the pcam set. However, the tractability of the patches vs the full slide dataset which is many hundreds of GB could be a very nice benefit if our models are able to perform well on the larger images after being trained on patches. I have not tested this yet but I will once this is over and I have a good idea of how well my models actually did. Many people here have already achieved an AUC score above the published benchmark (actually the leak will let us check our models against the full public and private set, I doubt the difference between them is great but with Kaggle you never know), so it will be useful to see the generalization. \n\nOverall I feel using information that is not available from the patches alone defeats the purpose of a patch-based approach but I’m not convinced that patches are the most useful thing in the first place, though personally I did not and will not use any non-patch info for this task.",
      "votes": null
    },
    {
      "id": "502536",
      "postDate": "03/28/2019 17:44:15",
      "content": "<p><code>I have not tested this yet but I will once this is over and I have a good idea of how well my models actually did</code> You probably should not evaluate your model based on LB (both public and private) since they only have about 1/5 data compared to your training set. Did you calculate the expected shakeup?</p>",
      "rawMarkdown": "`I have not tested this yet but I will once this is over and I have a good idea of how well my models actually did` You probably should not evaluate your model based on LB (both public and private) since they only have about 1/5 data compared to your training set. Did you calculate the expected shakeup?",
      "votes": null
    },
    {
      "id": "502566",
      "postDate": "03/28/2019 18:41:11",
      "content": "<p>Good point, and I did not really bother about the shakeup. It does pretty well on holdouts from training set, I’ll try varying the size a lot more before I try it on the big set. </p>",
      "rawMarkdown": "Good point, and I did not really bother about the shakeup. It does pretty well on holdouts from training set, I’ll try varying the size a lot more before I try it on the big set.",
      "votes": null
    },
    {
      "id": "503773",
      "postDate": "03/30/2019 14:34:53",
      "content": "<p>In my view the WSI thing is really about trying to eliminate dataset bias. By separating out a validation set that uses different slides than the training set we are improving our model while not necessarily improving our scores on our validation set. I moved to this competition late because I was previously focused on hmnist (the skin cancer dataset) but quit working on that when it became clear that dataset bias was really the biggest problem and that the test dataset likely suffered from the same bias as the training set. By focusing on heat maps it became clear that the model was learning the idiosyncracies of the dataset as a significant predictive factor rather than just features of the lesions. For example images *not * obtained through a dermoscope had a vastly higher probability of being benign nevi rather than  melanoma. The dermatologist classified the lesion as nevi on its characteristics and never went to the next step of imaging it through a dermoscope. This created a large bias in the dataset so the model was learning to look for artifacts of a dermoscope and use this to classify as nevi. I spent most of my effort debiasing the data. This of course leads to a model with lower scores even though the model would likely be better in the wild. So I determined that the competition was really a lot about recognizing the dataset rather than recognizing skin cancer. When I moved over to the histopathology dataset it seemed likely that the models could be learning to detect the slide (via stain density) as a factor in classification in addition to the polluted train/validation problem. So my goal was to focus on de-biasing the data again. The discussion about detecting WSI led me to look a bit deeper, I realized the filenames were a hash and it didn't take much effort detecting the full WSI list. Unfortunately in making that public I seriously damaged the competition. Personally I believe there is utility in a good patch based model. A lot of what we are looking for is detection of small metastases so detecting them in small patches seems a reasonable step. I do however believe that using things like WSI that can also aid in stain normalization and can make for less biased models trained on the same dataset. </p>",
      "rawMarkdown": "In my view the WSI thing is really about trying to eliminate dataset bias. By separating out a validation set that uses different slides than the training set we are improving our model while not necessarily improving our scores on our validation set. I moved to this competition late because I was previously focused on hmnist (the skin cancer dataset) but quit working on that when it became clear that dataset bias was really the biggest problem and that the test dataset likely suffered from the same bias as the training set. By focusing on heat maps it became clear that the model was learning the idiosyncracies of the dataset as a significant predictive factor rather than just features of the lesions. For example images *not * obtained through a dermoscope had a vastly higher probability of being benign nevi rather than  melanoma. The dermatologist classified the lesion as nevi on its characteristics and never went to the next step of imaging it through a dermoscope. This created a large bias in the dataset so the model was learning to look for artifacts of a dermoscope and use this to classify as nevi. I spent most of my effort debiasing the data. This of course leads to a model with lower scores even though the model would likely be better in the wild. So I determined that the competition was really a lot about recognizing the dataset rather than recognizing skin cancer. When I moved over to the histopathology dataset it seemed likely that the models could be learning to detect the slide (via stain density) as a factor in classification in addition to the polluted train/validation problem. So my goal was to focus on de-biasing the data again. The discussion about detecting WSI led me to look a bit deeper, I realized the filenames were a hash and it didn't take much effort detecting the full WSI list. Unfortunately in making that public I seriously damaged the competition. Personally I believe there is utility in a good patch based model. A lot of what we are looking for is detection of small metastases so detecting them in small patches seems a reasonable step. I do however believe that using things like WSI that can also aid in stain normalization and can make for less biased models trained on the same dataset.",
      "votes": null
    },
    {
      "id": "503873",
      "postDate": "03/30/2019 16:45:15",
      "content": "<p>Thanks for sharing your experience with this sort of data! Those are great points and things we should all consider more carefully. </p>",
      "rawMarkdown": "Thanks for sharing your experience with this sort of data! Those are great points and things we should all consider more carefully.",
      "votes": null
    },
    {
      "id": "504079",
      "postDate": "03/30/2019 23:59:27",
      "content": "<p>Great insights! Models will work horribly given that the test set (and industrial usages) and the training set are not in the same distribution. If Kaggle gives us a biased test set, the winners' model would have to be biased in the same way. It is up to people whether they want to develop a practical model or models favored by Kaggle's evaluation.</p>",
      "rawMarkdown": "Great insights! Models will work horribly given that the test set (and industrial usages) and the training set are not in the same distribution. If Kaggle gives us a biased test set, the winners' model would have to be biased in the same way. It is up to people whether they want to develop a practical model or models favored by Kaggle's evaluation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 502513,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "03/28/2019 17:09:27",
      "content": "<p>This highlights something that I’ve been thinking about for a while, which is the medical utility of this patch-based benchmark in the first place. </p>\n\n<p>I do not believe real diagnoses will be made from such disconnected image patches as are given in the pcam set. However, the tractability of the patches vs the full slide dataset which is many hundreds of GB could be a very nice benefit if our models are able to perform well on the larger images after being trained on patches. I have not tested this yet but I will once this is over and I have a good idea of how well my models actually did. Many people here have already achieved an AUC score above the published benchmark (actually the leak will let us check our models against the full public and private set, I doubt the difference between them is great but with Kaggle you never know), so it will be useful to see the generalization. </p>\n\n<p>Overall I feel using information that is not available from the patches alone defeats the purpose of a patch-based approach but I’m not convinced that patches are the most useful thing in the first place, though personally I did not and will not use any non-patch info for this task. </p>",
      "votes": null,
      "replies": [
        {
          "id": 502536,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/28/2019 17:44:15",
          "content": "<p><code>I have not tested this yet but I will once this is over and I have a good idea of how well my models actually did</code> You probably should not evaluate your model based on LB (both public and private) since they only have about 1/5 data compared to your training set. Did you calculate the expected shakeup?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 502566,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "03/28/2019 18:41:11",
          "content": "<p>Good point, and I did not really bother about the shakeup. It does pretty well on holdouts from training set, I’ll try varying the size a lot more before I try it on the big set. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 503773,
      "author_name": "markdesimone",
      "author_url": "",
      "post_date": "03/30/2019 14:34:53",
      "content": "<p>In my view the WSI thing is really about trying to eliminate dataset bias. By separating out a validation set that uses different slides than the training set we are improving our model while not necessarily improving our scores on our validation set. I moved to this competition late because I was previously focused on hmnist (the skin cancer dataset) but quit working on that when it became clear that dataset bias was really the biggest problem and that the test dataset likely suffered from the same bias as the training set. By focusing on heat maps it became clear that the model was learning the idiosyncracies of the dataset as a significant predictive factor rather than just features of the lesions. For example images *not * obtained through a dermoscope had a vastly higher probability of being benign nevi rather than  melanoma. The dermatologist classified the lesion as nevi on its characteristics and never went to the next step of imaging it through a dermoscope. This created a large bias in the dataset so the model was learning to look for artifacts of a dermoscope and use this to classify as nevi. I spent most of my effort debiasing the data. This of course leads to a model with lower scores even though the model would likely be better in the wild. So I determined that the competition was really a lot about recognizing the dataset rather than recognizing skin cancer. When I moved over to the histopathology dataset it seemed likely that the models could be learning to detect the slide (via stain density) as a factor in classification in addition to the polluted train/validation problem. So my goal was to focus on de-biasing the data again. The discussion about detecting WSI led me to look a bit deeper, I realized the filenames were a hash and it didn't take much effort detecting the full WSI list. Unfortunately in making that public I seriously damaged the competition. Personally I believe there is utility in a good patch based model. A lot of what we are looking for is detection of small metastases so detecting them in small patches seems a reasonable step. I do however believe that using things like WSI that can also aid in stain normalization and can make for less biased models trained on the same dataset. </p>",
      "votes": null,
      "replies": [
        {
          "id": 503873,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "03/30/2019 16:45:15",
          "content": "<p>Thanks for sharing your experience with this sort of data! Those are great points and things we should all consider more carefully. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 504079,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/30/2019 23:59:27",
          "content": "<p>Great insights! Models will work horribly given that the test set (and industrial usages) and the training set are not in the same distribution. If Kaggle gives us a biased test set, the winners' model would have to be biased in the same way. It is up to people whether they want to develop a practical model or models favored by Kaggle's evaluation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "502417": "![](https://i.postimg.cc/c12VKGZr/WSI-normilization.png)\nYou can train a classification network using WSI ids as labels. The slides with the same stain will be classified together. And then you can normalize each image according to its predicted WSI ids. Notice that the classification scores may not be high, but that should not be a problem since one kind of stain can have multiple WSI ids.\nI think it is better than stain normalization. What do you think?",
    "502513": "This highlights something that I’ve been thinking about for a while, which is the medical utility of this patch-based benchmark in the first place. \n\nI do not believe real diagnoses will be made from such disconnected image patches as are given in the pcam set. However, the tractability of the patches vs the full slide dataset which is many hundreds of GB could be a very nice benefit if our models are able to perform well on the larger images after being trained on patches. I have not tested this yet but I will once this is over and I have a good idea of how well my models actually did. Many people here have already achieved an AUC score above the published benchmark (actually the leak will let us check our models against the full public and private set, I doubt the difference between them is great but with Kaggle you never know), so it will be useful to see the generalization. \n\nOverall I feel using information that is not available from the patches alone defeats the purpose of a patch-based approach but I’m not convinced that patches are the most useful thing in the first place, though personally I did not and will not use any non-patch info for this task.",
    "502536": "`I have not tested this yet but I will once this is over and I have a good idea of how well my models actually did` You probably should not evaluate your model based on LB (both public and private) since they only have about 1/5 data compared to your training set. Did you calculate the expected shakeup?",
    "502566": "Good point, and I did not really bother about the shakeup. It does pretty well on holdouts from training set, I’ll try varying the size a lot more before I try it on the big set.",
    "503773": "In my view the WSI thing is really about trying to eliminate dataset bias. By separating out a validation set that uses different slides than the training set we are improving our model while not necessarily improving our scores on our validation set. I moved to this competition late because I was previously focused on hmnist (the skin cancer dataset) but quit working on that when it became clear that dataset bias was really the biggest problem and that the test dataset likely suffered from the same bias as the training set. By focusing on heat maps it became clear that the model was learning the idiosyncracies of the dataset as a significant predictive factor rather than just features of the lesions. For example images *not * obtained through a dermoscope had a vastly higher probability of being benign nevi rather than  melanoma. The dermatologist classified the lesion as nevi on its characteristics and never went to the next step of imaging it through a dermoscope. This created a large bias in the dataset so the model was learning to look for artifacts of a dermoscope and use this to classify as nevi. I spent most of my effort debiasing the data. This of course leads to a model with lower scores even though the model would likely be better in the wild. So I determined that the competition was really a lot about recognizing the dataset rather than recognizing skin cancer. When I moved over to the histopathology dataset it seemed likely that the models could be learning to detect the slide (via stain density) as a factor in classification in addition to the polluted train/validation problem. So my goal was to focus on de-biasing the data again. The discussion about detecting WSI led me to look a bit deeper, I realized the filenames were a hash and it didn't take much effort detecting the full WSI list. Unfortunately in making that public I seriously damaged the competition. Personally I believe there is utility in a good patch based model. A lot of what we are looking for is detection of small metastases so detecting them in small patches seems a reasonable step. I do however believe that using things like WSI that can also aid in stain normalization and can make for less biased models trained on the same dataset.",
    "503873": "Thanks for sharing your experience with this sort of data! Those are great points and things we should all consider more carefully.",
    "504079": "Great insights! Models will work horribly given that the test set (and industrial usages) and the training set are not in the same distribution. If Kaggle gives us a biased test set, the winners' model would have to be biased in the same way. It is up to people whether they want to develop a practical model or models favored by Kaggle's evaluation."
  },
  "source": "meta"
}