{
  "id": 337992,
  "title": "Make Histopathologic Models Robust To Domain Shift [Author: Heather D. Couture Ph.D.]",
  "url": "/competitions/hubmap-organ-segmentation/discussion/337992",
  "author_name": "",
  "post_date": "2022-07-18T17:00:03.451720500Z",
  "votes": 47,
  "comment_count": 5,
  "views": 0,
  "content": "<p><br></p>\n<h3><strong><em>NOTE: This is essentially just a repost/share of the <a href=\"https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html\" target=\"_blank\">article/information written here</a> with updates to allow for better formatting in markdown.</em></strong></h3>\n<p><strong><em>PLEASE GIVE CREDIT TO THE ORIGINAL AUTHOR --&gt; <a href=\"http://www.hdcouture.com/\" target=\"_blank\">Heather D. Couture Ph.D.</a></em></strong></p>\n<p><em>I'm sharing it as I found it incredibly helpful and instead of posting the link I thought reiterating the content here would help more people</em></p>\n<hr>\n<ul>\n<li><strong>TITLE:</strong><ul>\n<li>5 WAYS TO MAKE HISTOPATHOLOGY IMAGE MODELS MORE ROBUST TO DOMAIN SHIFTS</li></ul></li>\n<li><strong>AUTHOR:</strong><ul>\n<li>Heather D. Couture Ph.D.</li></ul></li>\n<li><strong>DATE PUBLISHED:</strong><ul>\n<li>August 2, 2021</li></ul></li>\n<li><strong>LOCATION PUBLISHED INITIALLY:</strong><ul>\n<li><a href=\"https://towardsdatascience.com/\" target=\"_blank\">Towards Data Science</a></li></ul></li>\n<li><strong>LINK:</strong><ul>\n<li><a href=\"https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html\" target=\"_blank\">https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html</a></li></ul></li>\n</ul>\n<hr>\n<p><br><br></p>\n<h3><strong>5 WAYS TO MAKE HISTOPATHOLOGY IMAGE MODELS MORE ROBUST TO DOMAIN SHIFTS</strong></h3>\n<hr>\n<p>One of the largest challenges in histopathology image analysis is creating models that are robust to the variations across different labs and imaging systems. These variations can be caused by:</p>\n<ul>\n<li>different color responses of slide scanners</li>\n<li>raw materials</li>\n<li>manufacturing techniques</li>\n<li>protocols for staining.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/CWSsS7k/domain-shift1.jpg\"></p>\n<p><strong><em>Figure 1:</em></strong> <em>H&amp;E image variations induced by different scanners <a href=\"https://arxiv.org/abs/2103.16515\" target=\"_blank\">[<strong>ref: Aubreville2021</strong>]</a></em></p>\n<p><br></p>\n<p>Different setups can produce images with different stain intensities or other changes, <strong>creating a domain shift between the source data that the model was trained on and the target data on which a deployed solution would need to operate</strong>. When the domain shift is too large, a model trained on one type of data will fail on another, often in unpredictable ways.</p>\n<p>I’ve [Heather D. Couture Ph.D.] been working with a client who is interested in selecting the best object detection model for their use case. But the images their model will examine once deployed come from a different lab and a variety of scanners.</p>\n<ul>\n<li>The domain shift from their training dataset to the target one is likely to be a larger challenge than getting state-of-the-art results on the training set.</li>\n<li>My [Heather D. Couture Ph.D.] advice to them was to <strong>tackle this domain adaptation challenge early.</strong></li>\n<li>They can always experiment with better object detection models later, once they’ve learned how to handle the domain shift.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/7bJ3yKT/domain-shift2.jpg\" alt=\"domain-shift2\"></p>\n<p><strong><em>Figure 2:</em></strong> <em>Mitosis appearance in two different datasets: Radboudumc (top) and TUPAC (bottom) <a href=\"https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf\" target=\"_blank\">[<strong>ref: Tellez2018</strong>]</a></em></p>\n<p><br></p>\n<p>So how do you handle the domain shift? There are a few different options:</p>\n<ol>\n<li>Standardize the appearance of your images with stain normalization techniques</li>\n<li>Color augmentation during training to take advantage of variations in staining</li>\n<li>Domain adversarial training to learn domain-invariant features</li>\n<li>Adapt the model at test time to handle the new image distribution</li>\n<li>Finetune the model on the target domain</li>\n</ol>\n<p>Some of these approaches have opposing goals. </p>\n<ul>\n<li>For example, color augmentation increases the diversity of images, while stain normalization tries to reduce the variations. * Domain adversarial training tries to learn domain-invariant features, while adapting or finetuning the model transforms the model to be suitable for the target domain only.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/8Kd56k0/domain-shift3.jpg\" alt=\"domain-shift3\"></p>\n<p><strong><em>Figure 3:</em></strong> <em>The different goals of color augmentation, stain normalization, and domain alignment by with a GAN <a href=\"https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1478702&amp;dswid=9029\" target=\"_blank\">[<strong>ref: Stacke2020</strong>]</a></em></p>\n<p><br></p>\n<p><strong>This article [Kaggle Discussion Post] will review each of the five strategies, followed by a summary of studies revealing which method(s) work best.</strong></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>1. STAIN NORMALIZATION</strong></h4>\n<p><br></p>\n<p>Different labs and scanners can produce images with different color profiles for a particular stain. <strong>The goal of stain normalization is to standardize the appearance of these stains.</strong></p>\n<p>Traditionally, methods like color matching [<a href=\"https://ieeexplore.ieee.org/abstract/document/946629\" target=\"_blank\">Reinhard2001</a>] and stain separation [<a href=\"https://ieeexplore.ieee.org/abstract/document/5193250\" target=\"_blank\">Macenko2009</a>, <a href=\"https://ieeexplore.ieee.org/abstract/document/6727397\" target=\"_blank\">Khan2014</a>, <a href=\"https://ieeexplore.ieee.org/abstract/document/7460968\" target=\"_blank\">Vahadane2016</a>] were used. However, these methods rely on selection of a single reference slide. Ren et al. have since shown that using an ensemble with different reference slides is one possible solution [<a href=\"https://www.frontiersin.org/articles/10.3389/fbioe.2019.00102/full\" target=\"_blank\">Ren2019</a>].</p>\n<p><img src=\"https://i.ibb.co/MVKpDX9/domain-shift4.jpg\" alt=\"domain-shift4\"></p>\n<p><strong><em>Figure 4:</em></strong> <em>H&amp;E image (left) and deconvolved hematoxylin (middle) and eosin (right) <a href=\"https://ieeexplore.ieee.org/abstract/document/5193250\" target=\"_blank\">[<strong>ref: Macenko2009</strong>]</a></em></p>\n<p><br></p>\n<p>The larger problem is that these techniques do not consider spatial features, which can lead to tissue structure not being preserved.</p>\n<p><strong>G</strong>enerative <strong>A</strong>dversarial <strong>N</strong>etworks (<strong>GAN</strong>s) are the state-of-the-art in stain normalization today. </p>\n<ul>\n<li>Given an image from domain A, a generator converts it into domain B. </li>\n<li>A discriminator network tries to distinguish real domain B images from fake ones, helping the generator to improve.</li>\n</ul>\n<p>If paired and aligned images from domains A and B are available, this setup performs well. </p>\n<ul>\n<li>However, it typically requires scanning each slide on two different scanners -- or potentially even restaining and rescanning each slide.</li>\n</ul>\n<p><strong>But there is a simpler solution to obtaining paired images:</strong> </p>\n<ul>\n<li>convert a color image to grayscale (domain A) and pair it with the original color image (domain B) <a href=\"https://arxiv.org/abs/2002.00647\" target=\"_blank\">[Salehi2020]</a>. </li>\n<li>The two are perfectly aligned and a Conditional GAN can be trained to reconstruct the color image.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/pjpknGP/domain-shift5.jpg\" alt=\"domain-shift5\"></p>\n<p><strong><em>Figure 5:</em></strong> <em>Stain style transfer from grayscale to H&amp;E with a Conditional GAN <a href=\"https://arxiv.org/abs/1710.08543\" target=\"_blank\">[</a></em><em>ref: Cho2017]</em></p>\n<p><br></p>\n<p>One major advantage of this approach is that <strong>a restaining model trained for one particular domain may work for a variety of different labs and scanners</strong> as there is less variation in the input grayscale images than in color.</p>\n<p>An alternative approach when paired images are not available is a <a href=\"https://openaccess.thecvf.com/content_iccv_2017/html/Zhu_Unpaired_Image-To-Image_Translation_ICCV_2017_paper.html\" target=\"_blank\">CycleGAN [Zhu2017]</a>. </p>\n<ul>\n<li>In this setup there are two generators: one to convert from domain A to B and another to go from domain B to A. </li>\n<li>The goal of these two models is to be able to reconstruct an original image: <ul>\n<li>A -&gt; B -&gt; A or B -&gt; A -&gt; B. </li></ul></li>\n<li>CycleGANs also makes use of discriminators to predict real versus generated images for each domain.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/3psgzCS/domain-shift6.jpg\" alt=\"domain-shift6\"></p>\n<p><strong><em>Figure 6:</em></strong> <em>Stain style transfer with a CycleGAN <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S1568494620307602\" target=\"_blank\">[</a></em><em>ref: Lo2021]</em></p>\n<p><br></p>\n<p>Stain normalization methods using deep learning have become increasingly complex. As a first pass to see if this type of standardization is helpful for your task, I [Heather D. Couture Ph.D.] suggest starting simple. </p>\n<ul>\n<li>StainTools and HistomicsTK both implement some of the color matching and stain separation methods.</li>\n<li>These simpler methods are sufficient in some cases, but not all. </li>\n<li>The figure below demonstrates how different methods perform on five different datasets.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/crDH8mx/domain-shift7.jpg\" alt=\"domain-shift7\"></p>\n<p><strong><em>Figure 7:</em></strong> <em>Comparison of stain normalization results on different datasets <a href=\"https://arxiv.org/abs/2002.00647\" target=\"_blank\">[</a></em><em>ref: Salehi2020]</em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>2. COLOR AUGMENTATION</strong></h4>\n<p><br></p>\n<p>Image augmentation by applying random affine transforms or adding noise is one of the most common regularization techniques for combating overfitting. Similarly, variations in staining can be taken advantage of to increase the diversity of image appearance presented to the model during training.</p>\n<ul>\n<li>While drastic changes in color are not realistic for histology, more subtle ones generated through random additive and multiplicative changes to each color channel have been shown to improve model performance.</li>\n<li>The intensity of color augmentation is an additional hyperparameter that should be experimented with during training and validated on test sets from different labs or scanners.</li>\n</ul>\n<p>Faryna et al. demonstrated the <strong><code>RandAugment</code></strong> technique on histopathology <a href=\"https://openreview.net/forum?id=JrBfXaoxbA2\" target=\"_blank\">[Faryna2021]</a>. </p>\n<ul>\n<li>This approach parameterizes augmentation as the number of random transformations selected and their magnitude.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/Hdk0rrS/domain-shift8.jpg\" alt=\"domain-shift8\"></p>\n<p><strong><em>Figure 8:</em></strong> <em>Image augmentation techniques <a href=\"https://openreview.net/forum?id=JrBfXaoxbA2\" target=\"_blank\">[<strong>ref: Faryna2021</strong>]</a></em></p>\n<p><br></p>\n<p>Tellez et al. studied the effect of different augmentation techniques (individually and combined) on mitosis detection and proposed an H&amp;E-specific transform <a href=\"https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf\" target=\"_blank\">[Tellez2018]</a>. </p>\n<ul>\n<li>They performed color deconvolution (as is used in the stain separation methods mentioned above), then applied the random shifts in hematoxylin and eosin space before converting back to RGB. </li>\n<li>The H&amp;E transform was the best performing individual augmentation method. </li>\n<li>A combination of all augmentation methods was critical in generalizing performance to a new dataset.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/QfHDR7S/domain-shift9.jpg\" alt=\"domain-shift9\"></p>\n<p><strong><em>Figure 9:</em></strong> <em>Color augmentation on hematoxylin and eosin channels <a href=\"https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf\" target=\"_blank\">[<strong>ref: Tellez2018</strong>]</a></em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>3. UNSUPERVISED DOMAIN ADVERSARIAL TRAINING</strong></h4>\n<p><br></p>\n<p>The next technique for domain adaptation is domain adversarial training <a href=\"https://dl.acm.org/doi/abs/10.5555/2946645.2946704\" target=\"_blank\">[Ganin2016]</a>. This approach makes use of unlabeled images from the target domain.</p>\n<p>A domain adversarial module is added to an existing model. </p>\n<ul>\n<li>The goal of this classifier is to predict whether an image belongs to the source or the target domain. </li>\n<li>A gradient reversal layer connects this module to the existing networking so that training optimizes the original task and encourages the network to learn domain-invariant features.</li>\n</ul>\n<p>During training, labeled images from the source domain and unlabeled ones from the target domain are used. </p>\n<ul>\n<li>For labeled source images, both the loss of the original network and the domain loss are applied. </li>\n<li>For unlabeled target images, only the domain loss is used.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/qRmDYY7/domain-shift10.jpg\" alt=\"domain-shift10\"></p>\n<p><strong><em>Figure 10:</em></strong> <em>A domain adversarial module attached to a CNN classifier <a href=\"https://dl.acm.org/doi/abs/10.5555/2946645.2946704\" target=\"_blank\">[<strong>ref: Ganin2016</strong>]</a></em></p>\n<p><br></p>\n<p>This module can be added to a variety of deep learning models. </p>\n<ul>\n<li>For <strong>classification</strong>, it is typically connected to a layer near the output. </li>\n<li>For segmentation, it is usually applied to the bottleneck layer - although it can also be applied to multiple layers. </li>\n<li>For detection, it can be applied to the feature pyramid network. </li>\n<li>For challenging object detectors like mitoses that require an additional classifier network, the domain adversarial network may only be applied to the second stage <a href=\"https://arxiv.org/abs/1911.10873\" target=\"_blank\">[Aubreville2020a]</a>.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/R9VpHhq/domain-shift11.jpg\" alt=\"domain-shift11\"></p>\n<p><strong><em>Figure 11:</em></strong> <em>Improved domain invariance with a domain adapted model <a href=\"https://arxiv.org/abs/1911.10873\" target=\"_blank\">[<strong>ref: Aubreville2020a</strong>]</a></em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>4. ADAPT MODEL AT TEST TIME</strong></h4>\n<p><br></p>\n<p>Instead of accommodating domain shifts during training, the model may be modified at test time. </p>\n<ul>\n<li>Domain shifts are reflected in a change in distribution in feature space -- a covariate shift. </li>\n<li>So for models using batch normalization layers, the <strong>mean and standard deviation may be recalculated for a new test set</strong>.</li>\n</ul>\n<p>These new statistics can be calculated over the whole test set and updated in the model before running inference. </p>\n<ul>\n<li>Or they can be calculated for each new batch of data. </li>\n<li>Nado et al. found that the latter approach, called prediction-time batch normalization, was sufficient <a href=\"https://arxiv.org/abs/2006.10963\" target=\"_blank\">[Nado2020]</a>. </li>\n<li>Further, a single batch of 500 images was sufficient to get a substantial improvement in model accuracy.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/9ZgCPy5/domain-shift12.jpg\" alt=\"domain-shift12\"></p>\n<p><strong><em>Figure 12:</em></strong> <em>Prediction-time batch normalization aligns the shifted activations with the training distribution <a href=\"https://arxiv.org/abs/2006.10963\" target=\"_blank\">[<strong>ref: Nado2020</strong>]</a></em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>5. MODEL FINETUNING</strong></h4>\n<p><br></p>\n<p>Finally, a model may be finetuned on a test set with a domain shift <a href=\"https://www.nature.com/articles/s41597-020-00756-z\" target=\"_blank\">[Aubreville2020b]</a>. </p>\n<ul>\n<li>If enough labeled examples are available in the test set, this approach is likely to produce the best result. </li>\n<li>However, it is the most time-consuming and least generalizable. </li>\n<li>The model may need to be finetuned again on other test sets in the future.</li>\n</ul>\n<p><br></p>\n<hr>\n<p><br><br></p>\n<h4><strong>COMPARISON OF APPROACHES</strong></h4>\n<p>Color augmentation and stain normalization are used extensively in pathology image applications, especially for whole slide H&amp;E images. Domain adversarial training and model adaptation at test time are less studied thus far.</p>\n<p><br></p>\n<h5><strong>STAIN NORMALIZATION v. COLOR AUGMENTATION</strong></h5>\n<p>Khan et al. studied stain normalization and color augmentation when applied individually and together <a href=\"https://www.researchgate.net/publication/338826217_Generalizing_Convolution_Neural_Networks_on_Stain_Color_Heterogeneous_Data_for_Computational_Pathology\" target=\"_blank\">[Khan2020]</a>. </p>\n<ul>\n<li>The best results were obtained when the methods were used together.</li>\n</ul>\n<p>Tellez et al. tested different combinations of image augmentation and normalization strategies for a variety of different histology classification tasks <a href=\"https://arxiv.org/abs/1902.06543\" target=\"_blank\">[Tellez2019]</a>. </p>\n<ul>\n<li>The best configuration turned out to be randomly shifting the color channels and applying no stain normalization. </li>\n<li>Experiments using stain normalization performed only slightly poorer. </li>\n<li>This validates the importance of image augmentation in creating a robust classifier for histology and emphasizes the importance of color transformations. </li>\n<li>While applying stain normalization didn’t really hurt, this extra computation may not be necessary.</li>\n</ul>\n<p><br></p>\n<h5><strong>STAIN NORMALIZATION v. DOMAIN ADVERSARIAL TRAINING</strong></h5>\n<p>Ren et al. compared domain adversarial training with some stain normalization and color augmentation approaches, demonstrating that <strong>the domain adversarial approach was superior for generalizing to new image sets</strong> <a href=\"https://www.frontiersin.org/articles/10.3389/fbioe.2019.00102/full\" target=\"_blank\">[Ren2019]</a>.</p>\n<p><br></p>\n<h5><strong>STAIN NORMALIZATION v. COLOR AUGMENTATION v. DOMAIN ADVERSARIAL TRAINING</strong></h5>\n<p>Larfarge et al. performed a similar study that also included domain adversarial training. </p>\n<p>They compared domain adversarial training with color augmentation and stain normalization on mitosis classification and nuclei segmentation tasks <a href=\"https://www.frontiersin.org/articles/10.3389/fmed.2019.00162/full\" target=\"_blank\">[Lafarge2019]</a>.</p>\n<ul>\n<li>On mitosis <strong>classification</strong>, they found that color augmentation performed best for test images from the same lab on which the model was trained</li>\n<li>However, on images from other labs, domain adversarial training combined with color augmentation was best</li>\n</ul>\n<p>For nuclei segmentation, the results were a bit different. </p>\n<ul>\n<li>Stain normalization was key when tested on images of the same tissue type. </li>\n<li>On different tissue types, domain adversarial training with stain normalization was best.</li>\n</ul>\n<p><strong>Clearly domain adversarial training was beneficial for both types of domain shifts!</strong></p>\n<ul>\n<li>However, the best preprocessing and augmentation strategy varied with the dataset. </li>\n<li>Lafarge speculated that this is due to the amount of domain variability in the training set.</li>\n</ul>\n<p><br></p>\n<h5><strong>SIDE NOTE:</strong></h5>\n<p>The stain normalization methods tested in the above analysis of methods are only the traditional color matching and stain separation techniques. </p>\n<ul>\n<li>Many newer deep learning-based approaches exist that were not included in these benchmarks.</li>\n</ul>\n<p><br></p>\n<hr>\n<p><br><br></p>\n<h4><strong>RECOMMENDATIONS</strong></h4>\n<p>Traditional stain normalization techniques are worth trying as a first pass as they’re easier to implement and often faster to run. </p>\n<ul>\n<li>For some domain shifts, these may even be sufficient, especially when combined with color augmentation or domain adaptation. </li>\n<li>Experiments with a simpler method can also provide insights into whether some amount of stain normalization will improve model generalization performance.</li>\n<li>For more robust stain normalization that preserves tissue structure, evaluate the methods discussed above.</li>\n</ul>\n<p>The available data may also be a deciding factor in selecting appropriate techniques. </p>\n<ul>\n<li>Stain normalization and color augmentation do not require target domain images during training, while the other three approaches do. </li>\n<li>Model adaptation requires unlabeled target data, while finetuning needs it labeled. </li>\n</ul>\n<p>For these reasons, stain normalization and color augmentation are often tried first, with adversarial domain adaptation included when needed. </p>\n<ul>\n<li>These three methods are also the best approaches for training a single generalizable model. </li>\n<li>If a large set of target images is available, then adapting the model (with unlabeled data) or finetuning (with labeled data) will likely be most effective.</li>\n</ul>\n<p>Monitoring a deployed system for unexpected domain shifts is also critical. </p>\n<ul>\n<li>Stacke et al. developed a way to quantify domain shift <a href=\"https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1478702&amp;dswid=9029\" target=\"_blank\">[Stacke2020]</a>.</li>\n<li>Their metric does not require annotated data, so can serve as a simple test to see if new data is likely to be handled well by an existing model.</li>\n</ul>\n<hr>\n<h2><strong>FIN</strong></h2>",
  "messages": [
    {
      "id": "1860944",
      "postDate": "07/18/2022 17:00:03",
      "content": "<p><br></p>\n<h3><strong><em>NOTE: This is essentially just a repost/share of the <a href=\"https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html\" target=\"_blank\">article/information written here</a> with updates to allow for better formatting in markdown.</em></strong></h3>\n<p><strong><em>PLEASE GIVE CREDIT TO THE ORIGINAL AUTHOR --&gt; <a href=\"http://www.hdcouture.com/\" target=\"_blank\">Heather D. Couture Ph.D.</a></em></strong></p>\n<p><em>I'm sharing it as I found it incredibly helpful and instead of posting the link I thought reiterating the content here would help more people</em></p>\n<hr>\n<ul>\n<li><strong>TITLE:</strong><ul>\n<li>5 WAYS TO MAKE HISTOPATHOLOGY IMAGE MODELS MORE ROBUST TO DOMAIN SHIFTS</li></ul></li>\n<li><strong>AUTHOR:</strong><ul>\n<li>Heather D. Couture Ph.D.</li></ul></li>\n<li><strong>DATE PUBLISHED:</strong><ul>\n<li>August 2, 2021</li></ul></li>\n<li><strong>LOCATION PUBLISHED INITIALLY:</strong><ul>\n<li><a href=\"https://towardsdatascience.com/\" target=\"_blank\">Towards Data Science</a></li></ul></li>\n<li><strong>LINK:</strong><ul>\n<li><a href=\"https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html\" target=\"_blank\">https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html</a></li></ul></li>\n</ul>\n<hr>\n<p><br><br></p>\n<h3><strong>5 WAYS TO MAKE HISTOPATHOLOGY IMAGE MODELS MORE ROBUST TO DOMAIN SHIFTS</strong></h3>\n<hr>\n<p>One of the largest challenges in histopathology image analysis is creating models that are robust to the variations across different labs and imaging systems. These variations can be caused by:</p>\n<ul>\n<li>different color responses of slide scanners</li>\n<li>raw materials</li>\n<li>manufacturing techniques</li>\n<li>protocols for staining.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/CWSsS7k/domain-shift1.jpg\"></p>\n<p><strong><em>Figure 1:</em></strong> <em>H&amp;E image variations induced by different scanners <a href=\"https://arxiv.org/abs/2103.16515\" target=\"_blank\">[<strong>ref: Aubreville2021</strong>]</a></em></p>\n<p><br></p>\n<p>Different setups can produce images with different stain intensities or other changes, <strong>creating a domain shift between the source data that the model was trained on and the target data on which a deployed solution would need to operate</strong>. When the domain shift is too large, a model trained on one type of data will fail on another, often in unpredictable ways.</p>\n<p>I’ve [Heather D. Couture Ph.D.] been working with a client who is interested in selecting the best object detection model for their use case. But the images their model will examine once deployed come from a different lab and a variety of scanners.</p>\n<ul>\n<li>The domain shift from their training dataset to the target one is likely to be a larger challenge than getting state-of-the-art results on the training set.</li>\n<li>My [Heather D. Couture Ph.D.] advice to them was to <strong>tackle this domain adaptation challenge early.</strong></li>\n<li>They can always experiment with better object detection models later, once they’ve learned how to handle the domain shift.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/7bJ3yKT/domain-shift2.jpg\" alt=\"domain-shift2\"></p>\n<p><strong><em>Figure 2:</em></strong> <em>Mitosis appearance in two different datasets: Radboudumc (top) and TUPAC (bottom) <a href=\"https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf\" target=\"_blank\">[<strong>ref: Tellez2018</strong>]</a></em></p>\n<p><br></p>\n<p>So how do you handle the domain shift? There are a few different options:</p>\n<ol>\n<li>Standardize the appearance of your images with stain normalization techniques</li>\n<li>Color augmentation during training to take advantage of variations in staining</li>\n<li>Domain adversarial training to learn domain-invariant features</li>\n<li>Adapt the model at test time to handle the new image distribution</li>\n<li>Finetune the model on the target domain</li>\n</ol>\n<p>Some of these approaches have opposing goals. </p>\n<ul>\n<li>For example, color augmentation increases the diversity of images, while stain normalization tries to reduce the variations. * Domain adversarial training tries to learn domain-invariant features, while adapting or finetuning the model transforms the model to be suitable for the target domain only.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/8Kd56k0/domain-shift3.jpg\" alt=\"domain-shift3\"></p>\n<p><strong><em>Figure 3:</em></strong> <em>The different goals of color augmentation, stain normalization, and domain alignment by with a GAN <a href=\"https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1478702&amp;dswid=9029\" target=\"_blank\">[<strong>ref: Stacke2020</strong>]</a></em></p>\n<p><br></p>\n<p><strong>This article [Kaggle Discussion Post] will review each of the five strategies, followed by a summary of studies revealing which method(s) work best.</strong></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>1. STAIN NORMALIZATION</strong></h4>\n<p><br></p>\n<p>Different labs and scanners can produce images with different color profiles for a particular stain. <strong>The goal of stain normalization is to standardize the appearance of these stains.</strong></p>\n<p>Traditionally, methods like color matching [<a href=\"https://ieeexplore.ieee.org/abstract/document/946629\" target=\"_blank\">Reinhard2001</a>] and stain separation [<a href=\"https://ieeexplore.ieee.org/abstract/document/5193250\" target=\"_blank\">Macenko2009</a>, <a href=\"https://ieeexplore.ieee.org/abstract/document/6727397\" target=\"_blank\">Khan2014</a>, <a href=\"https://ieeexplore.ieee.org/abstract/document/7460968\" target=\"_blank\">Vahadane2016</a>] were used. However, these methods rely on selection of a single reference slide. Ren et al. have since shown that using an ensemble with different reference slides is one possible solution [<a href=\"https://www.frontiersin.org/articles/10.3389/fbioe.2019.00102/full\" target=\"_blank\">Ren2019</a>].</p>\n<p><img src=\"https://i.ibb.co/MVKpDX9/domain-shift4.jpg\" alt=\"domain-shift4\"></p>\n<p><strong><em>Figure 4:</em></strong> <em>H&amp;E image (left) and deconvolved hematoxylin (middle) and eosin (right) <a href=\"https://ieeexplore.ieee.org/abstract/document/5193250\" target=\"_blank\">[<strong>ref: Macenko2009</strong>]</a></em></p>\n<p><br></p>\n<p>The larger problem is that these techniques do not consider spatial features, which can lead to tissue structure not being preserved.</p>\n<p><strong>G</strong>enerative <strong>A</strong>dversarial <strong>N</strong>etworks (<strong>GAN</strong>s) are the state-of-the-art in stain normalization today. </p>\n<ul>\n<li>Given an image from domain A, a generator converts it into domain B. </li>\n<li>A discriminator network tries to distinguish real domain B images from fake ones, helping the generator to improve.</li>\n</ul>\n<p>If paired and aligned images from domains A and B are available, this setup performs well. </p>\n<ul>\n<li>However, it typically requires scanning each slide on two different scanners -- or potentially even restaining and rescanning each slide.</li>\n</ul>\n<p><strong>But there is a simpler solution to obtaining paired images:</strong> </p>\n<ul>\n<li>convert a color image to grayscale (domain A) and pair it with the original color image (domain B) <a href=\"https://arxiv.org/abs/2002.00647\" target=\"_blank\">[Salehi2020]</a>. </li>\n<li>The two are perfectly aligned and a Conditional GAN can be trained to reconstruct the color image.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/pjpknGP/domain-shift5.jpg\" alt=\"domain-shift5\"></p>\n<p><strong><em>Figure 5:</em></strong> <em>Stain style transfer from grayscale to H&amp;E with a Conditional GAN <a href=\"https://arxiv.org/abs/1710.08543\" target=\"_blank\">[</a></em><em>ref: Cho2017]</em></p>\n<p><br></p>\n<p>One major advantage of this approach is that <strong>a restaining model trained for one particular domain may work for a variety of different labs and scanners</strong> as there is less variation in the input grayscale images than in color.</p>\n<p>An alternative approach when paired images are not available is a <a href=\"https://openaccess.thecvf.com/content_iccv_2017/html/Zhu_Unpaired_Image-To-Image_Translation_ICCV_2017_paper.html\" target=\"_blank\">CycleGAN [Zhu2017]</a>. </p>\n<ul>\n<li>In this setup there are two generators: one to convert from domain A to B and another to go from domain B to A. </li>\n<li>The goal of these two models is to be able to reconstruct an original image: <ul>\n<li>A -&gt; B -&gt; A or B -&gt; A -&gt; B. </li></ul></li>\n<li>CycleGANs also makes use of discriminators to predict real versus generated images for each domain.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/3psgzCS/domain-shift6.jpg\" alt=\"domain-shift6\"></p>\n<p><strong><em>Figure 6:</em></strong> <em>Stain style transfer with a CycleGAN <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S1568494620307602\" target=\"_blank\">[</a></em><em>ref: Lo2021]</em></p>\n<p><br></p>\n<p>Stain normalization methods using deep learning have become increasingly complex. As a first pass to see if this type of standardization is helpful for your task, I [Heather D. Couture Ph.D.] suggest starting simple. </p>\n<ul>\n<li>StainTools and HistomicsTK both implement some of the color matching and stain separation methods.</li>\n<li>These simpler methods are sufficient in some cases, but not all. </li>\n<li>The figure below demonstrates how different methods perform on five different datasets.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/crDH8mx/domain-shift7.jpg\" alt=\"domain-shift7\"></p>\n<p><strong><em>Figure 7:</em></strong> <em>Comparison of stain normalization results on different datasets <a href=\"https://arxiv.org/abs/2002.00647\" target=\"_blank\">[</a></em><em>ref: Salehi2020]</em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>2. COLOR AUGMENTATION</strong></h4>\n<p><br></p>\n<p>Image augmentation by applying random affine transforms or adding noise is one of the most common regularization techniques for combating overfitting. Similarly, variations in staining can be taken advantage of to increase the diversity of image appearance presented to the model during training.</p>\n<ul>\n<li>While drastic changes in color are not realistic for histology, more subtle ones generated through random additive and multiplicative changes to each color channel have been shown to improve model performance.</li>\n<li>The intensity of color augmentation is an additional hyperparameter that should be experimented with during training and validated on test sets from different labs or scanners.</li>\n</ul>\n<p>Faryna et al. demonstrated the <strong><code>RandAugment</code></strong> technique on histopathology <a href=\"https://openreview.net/forum?id=JrBfXaoxbA2\" target=\"_blank\">[Faryna2021]</a>. </p>\n<ul>\n<li>This approach parameterizes augmentation as the number of random transformations selected and their magnitude.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/Hdk0rrS/domain-shift8.jpg\" alt=\"domain-shift8\"></p>\n<p><strong><em>Figure 8:</em></strong> <em>Image augmentation techniques <a href=\"https://openreview.net/forum?id=JrBfXaoxbA2\" target=\"_blank\">[<strong>ref: Faryna2021</strong>]</a></em></p>\n<p><br></p>\n<p>Tellez et al. studied the effect of different augmentation techniques (individually and combined) on mitosis detection and proposed an H&amp;E-specific transform <a href=\"https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf\" target=\"_blank\">[Tellez2018]</a>. </p>\n<ul>\n<li>They performed color deconvolution (as is used in the stain separation methods mentioned above), then applied the random shifts in hematoxylin and eosin space before converting back to RGB. </li>\n<li>The H&amp;E transform was the best performing individual augmentation method. </li>\n<li>A combination of all augmentation methods was critical in generalizing performance to a new dataset.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/QfHDR7S/domain-shift9.jpg\" alt=\"domain-shift9\"></p>\n<p><strong><em>Figure 9:</em></strong> <em>Color augmentation on hematoxylin and eosin channels <a href=\"https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf\" target=\"_blank\">[<strong>ref: Tellez2018</strong>]</a></em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>3. UNSUPERVISED DOMAIN ADVERSARIAL TRAINING</strong></h4>\n<p><br></p>\n<p>The next technique for domain adaptation is domain adversarial training <a href=\"https://dl.acm.org/doi/abs/10.5555/2946645.2946704\" target=\"_blank\">[Ganin2016]</a>. This approach makes use of unlabeled images from the target domain.</p>\n<p>A domain adversarial module is added to an existing model. </p>\n<ul>\n<li>The goal of this classifier is to predict whether an image belongs to the source or the target domain. </li>\n<li>A gradient reversal layer connects this module to the existing networking so that training optimizes the original task and encourages the network to learn domain-invariant features.</li>\n</ul>\n<p>During training, labeled images from the source domain and unlabeled ones from the target domain are used. </p>\n<ul>\n<li>For labeled source images, both the loss of the original network and the domain loss are applied. </li>\n<li>For unlabeled target images, only the domain loss is used.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/qRmDYY7/domain-shift10.jpg\" alt=\"domain-shift10\"></p>\n<p><strong><em>Figure 10:</em></strong> <em>A domain adversarial module attached to a CNN classifier <a href=\"https://dl.acm.org/doi/abs/10.5555/2946645.2946704\" target=\"_blank\">[<strong>ref: Ganin2016</strong>]</a></em></p>\n<p><br></p>\n<p>This module can be added to a variety of deep learning models. </p>\n<ul>\n<li>For <strong>classification</strong>, it is typically connected to a layer near the output. </li>\n<li>For segmentation, it is usually applied to the bottleneck layer - although it can also be applied to multiple layers. </li>\n<li>For detection, it can be applied to the feature pyramid network. </li>\n<li>For challenging object detectors like mitoses that require an additional classifier network, the domain adversarial network may only be applied to the second stage <a href=\"https://arxiv.org/abs/1911.10873\" target=\"_blank\">[Aubreville2020a]</a>.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/R9VpHhq/domain-shift11.jpg\" alt=\"domain-shift11\"></p>\n<p><strong><em>Figure 11:</em></strong> <em>Improved domain invariance with a domain adapted model <a href=\"https://arxiv.org/abs/1911.10873\" target=\"_blank\">[<strong>ref: Aubreville2020a</strong>]</a></em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>4. ADAPT MODEL AT TEST TIME</strong></h4>\n<p><br></p>\n<p>Instead of accommodating domain shifts during training, the model may be modified at test time. </p>\n<ul>\n<li>Domain shifts are reflected in a change in distribution in feature space -- a covariate shift. </li>\n<li>So for models using batch normalization layers, the <strong>mean and standard deviation may be recalculated for a new test set</strong>.</li>\n</ul>\n<p>These new statistics can be calculated over the whole test set and updated in the model before running inference. </p>\n<ul>\n<li>Or they can be calculated for each new batch of data. </li>\n<li>Nado et al. found that the latter approach, called prediction-time batch normalization, was sufficient <a href=\"https://arxiv.org/abs/2006.10963\" target=\"_blank\">[Nado2020]</a>. </li>\n<li>Further, a single batch of 500 images was sufficient to get a substantial improvement in model accuracy.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/9ZgCPy5/domain-shift12.jpg\" alt=\"domain-shift12\"></p>\n<p><strong><em>Figure 12:</em></strong> <em>Prediction-time batch normalization aligns the shifted activations with the training distribution <a href=\"https://arxiv.org/abs/2006.10963\" target=\"_blank\">[<strong>ref: Nado2020</strong>]</a></em></p>\n<p><br></p>\n<hr>\n<p><br></p>\n<h4><strong>5. MODEL FINETUNING</strong></h4>\n<p><br></p>\n<p>Finally, a model may be finetuned on a test set with a domain shift <a href=\"https://www.nature.com/articles/s41597-020-00756-z\" target=\"_blank\">[Aubreville2020b]</a>. </p>\n<ul>\n<li>If enough labeled examples are available in the test set, this approach is likely to produce the best result. </li>\n<li>However, it is the most time-consuming and least generalizable. </li>\n<li>The model may need to be finetuned again on other test sets in the future.</li>\n</ul>\n<p><br></p>\n<hr>\n<p><br><br></p>\n<h4><strong>COMPARISON OF APPROACHES</strong></h4>\n<p>Color augmentation and stain normalization are used extensively in pathology image applications, especially for whole slide H&amp;E images. Domain adversarial training and model adaptation at test time are less studied thus far.</p>\n<p><br></p>\n<h5><strong>STAIN NORMALIZATION v. COLOR AUGMENTATION</strong></h5>\n<p>Khan et al. studied stain normalization and color augmentation when applied individually and together <a href=\"https://www.researchgate.net/publication/338826217_Generalizing_Convolution_Neural_Networks_on_Stain_Color_Heterogeneous_Data_for_Computational_Pathology\" target=\"_blank\">[Khan2020]</a>. </p>\n<ul>\n<li>The best results were obtained when the methods were used together.</li>\n</ul>\n<p>Tellez et al. tested different combinations of image augmentation and normalization strategies for a variety of different histology classification tasks <a href=\"https://arxiv.org/abs/1902.06543\" target=\"_blank\">[Tellez2019]</a>. </p>\n<ul>\n<li>The best configuration turned out to be randomly shifting the color channels and applying no stain normalization. </li>\n<li>Experiments using stain normalization performed only slightly poorer. </li>\n<li>This validates the importance of image augmentation in creating a robust classifier for histology and emphasizes the importance of color transformations. </li>\n<li>While applying stain normalization didn’t really hurt, this extra computation may not be necessary.</li>\n</ul>\n<p><br></p>\n<h5><strong>STAIN NORMALIZATION v. DOMAIN ADVERSARIAL TRAINING</strong></h5>\n<p>Ren et al. compared domain adversarial training with some stain normalization and color augmentation approaches, demonstrating that <strong>the domain adversarial approach was superior for generalizing to new image sets</strong> <a href=\"https://www.frontiersin.org/articles/10.3389/fbioe.2019.00102/full\" target=\"_blank\">[Ren2019]</a>.</p>\n<p><br></p>\n<h5><strong>STAIN NORMALIZATION v. COLOR AUGMENTATION v. DOMAIN ADVERSARIAL TRAINING</strong></h5>\n<p>Larfarge et al. performed a similar study that also included domain adversarial training. </p>\n<p>They compared domain adversarial training with color augmentation and stain normalization on mitosis classification and nuclei segmentation tasks <a href=\"https://www.frontiersin.org/articles/10.3389/fmed.2019.00162/full\" target=\"_blank\">[Lafarge2019]</a>.</p>\n<ul>\n<li>On mitosis <strong>classification</strong>, they found that color augmentation performed best for test images from the same lab on which the model was trained</li>\n<li>However, on images from other labs, domain adversarial training combined with color augmentation was best</li>\n</ul>\n<p>For nuclei segmentation, the results were a bit different. </p>\n<ul>\n<li>Stain normalization was key when tested on images of the same tissue type. </li>\n<li>On different tissue types, domain adversarial training with stain normalization was best.</li>\n</ul>\n<p><strong>Clearly domain adversarial training was beneficial for both types of domain shifts!</strong></p>\n<ul>\n<li>However, the best preprocessing and augmentation strategy varied with the dataset. </li>\n<li>Lafarge speculated that this is due to the amount of domain variability in the training set.</li>\n</ul>\n<p><br></p>\n<h5><strong>SIDE NOTE:</strong></h5>\n<p>The stain normalization methods tested in the above analysis of methods are only the traditional color matching and stain separation techniques. </p>\n<ul>\n<li>Many newer deep learning-based approaches exist that were not included in these benchmarks.</li>\n</ul>\n<p><br></p>\n<hr>\n<p><br><br></p>\n<h4><strong>RECOMMENDATIONS</strong></h4>\n<p>Traditional stain normalization techniques are worth trying as a first pass as they’re easier to implement and often faster to run. </p>\n<ul>\n<li>For some domain shifts, these may even be sufficient, especially when combined with color augmentation or domain adaptation. </li>\n<li>Experiments with a simpler method can also provide insights into whether some amount of stain normalization will improve model generalization performance.</li>\n<li>For more robust stain normalization that preserves tissue structure, evaluate the methods discussed above.</li>\n</ul>\n<p>The available data may also be a deciding factor in selecting appropriate techniques. </p>\n<ul>\n<li>Stain normalization and color augmentation do not require target domain images during training, while the other three approaches do. </li>\n<li>Model adaptation requires unlabeled target data, while finetuning needs it labeled. </li>\n</ul>\n<p>For these reasons, stain normalization and color augmentation are often tried first, with adversarial domain adaptation included when needed. </p>\n<ul>\n<li>These three methods are also the best approaches for training a single generalizable model. </li>\n<li>If a large set of target images is available, then adapting the model (with unlabeled data) or finetuning (with labeled data) will likely be most effective.</li>\n</ul>\n<p>Monitoring a deployed system for unexpected domain shifts is also critical. </p>\n<ul>\n<li>Stacke et al. developed a way to quantify domain shift <a href=\"https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1478702&amp;dswid=9029\" target=\"_blank\">[Stacke2020]</a>.</li>\n<li>Their metric does not require annotated data, so can serve as a simple test to see if new data is likely to be handled well by an existing model.</li>\n</ul>\n<hr>\n<h2><strong>FIN</strong></h2>",
      "rawMarkdown": "<br>\n\n### ***NOTE: This is essentially just a repost/share of the [article/information written here](https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html) with updates to allow for better formatting in markdown.***\n\n***PLEASE GIVE CREDIT TO THE ORIGINAL AUTHOR --> [Heather D. Couture Ph.D.](http://www.hdcouture.com/)***\n\n*I'm sharing it as I found it incredibly helpful and instead of posting the link I thought reiterating the content here would help more people*\n\n---\n\n* **TITLE:**\n  * 5 WAYS TO MAKE HISTOPATHOLOGY IMAGE MODELS MORE ROBUST TO DOMAIN SHIFTS\n* **AUTHOR:**\n  * Heather D. Couture Ph.D.\n* **DATE PUBLISHED:**\n  * August 2, 2021\n* **LOCATION PUBLISHED INITIALLY:**\n  * [Towards Data Science](https://towardsdatascience.com/)\n* **LINK:**\n  * https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html\n\n---\n\n<br><br>\n\n### **5 WAYS TO MAKE HISTOPATHOLOGY IMAGE MODELS MORE ROBUST TO DOMAIN SHIFTS**\n\n---\n\nOne of the largest challenges in histopathology image analysis is creating models that are robust to the variations across different labs and imaging systems. These variations can be caused by:\n* different color responses of slide scanners\n* raw materials\n* manufacturing techniques\n* protocols for staining.\n\n<center><img src=\"https://i.ibb.co/CWSsS7k/domain-shift1.jpg\" width=100%></center>\n\n***Figure 1:*** *H&E image variations induced by different scanners [[**ref: Aubreville2021**]](https://arxiv.org/abs/2103.16515)*\n\n<br>\n\nDifferent setups can produce images with different stain intensities or other changes, **creating a domain shift between the source data that the model was trained on and the target data on which a deployed solution would need to operate**. When the domain shift is too large, a model trained on one type of data will fail on another, often in unpredictable ways.\n\nI’ve [Heather D. Couture Ph.D.] been working with a client who is interested in selecting the best object detection model for their use case. But the images their model will examine once deployed come from a different lab and a variety of scanners.\n* The domain shift from their training dataset to the target one is likely to be a larger challenge than getting state-of-the-art results on the training set.\n* My [Heather D. Couture Ph.D.] advice to them was to **tackle this domain adaptation challenge early.**\n* They can always experiment with better object detection models later, once they’ve learned how to handle the domain shift.\n\n<center><img src=\"https://i.ibb.co/7bJ3yKT/domain-shift2.jpg\" alt=\"domain-shift2\" width=100%></center>\n\n***Figure 2:*** *Mitosis appearance in two different datasets: Radboudumc (top) and TUPAC (bottom) [[**ref: Tellez2018**]](https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf)*\n\n<br>\n\nSo how do you handle the domain shift? There are a few different options:\n1. Standardize the appearance of your images with stain normalization techniques\n2. Color augmentation during training to take advantage of variations in staining\n3. Domain adversarial training to learn domain-invariant features\n4. Adapt the model at test time to handle the new image distribution\n5. Finetune the model on the target domain\n\nSome of these approaches have opposing goals. \n* For example, color augmentation increases the diversity of images, while stain normalization tries to reduce the variations. * Domain adversarial training tries to learn domain-invariant features, while adapting or finetuning the model transforms the model to be suitable for the target domain only.\n\n<center><img src=\"https://i.ibb.co/8Kd56k0/domain-shift3.jpg\" alt=\"domain-shift3\" width=100%></center>\n\n***Figure 3:*** *The different goals of color augmentation, stain normalization, and domain alignment by with a GAN [[**ref: Stacke2020**]](https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1478702&dswid=9029)*\n\n<br>\n\n**This article [Kaggle Discussion Post] will review each of the five strategies, followed by a summary of studies revealing which method(s) work best.**\n\n<br>\n\n---\n\n<br>\n\n#### **1. STAIN NORMALIZATION**\n\n<br>\n\nDifferent labs and scanners can produce images with different color profiles for a particular stain. **The goal of stain normalization is to standardize the appearance of these stains.**\n\nTraditionally, methods like color matching [<a href=\"https://ieeexplore.ieee.org/abstract/document/946629\" target=\"_blank\">Reinhard2001</a>] and stain separation [<a href=\"https://ieeexplore.ieee.org/abstract/document/5193250\" target=\"_blank\">Macenko2009</a>, <a href=\"https://ieeexplore.ieee.org/abstract/document/6727397\" target=\"_blank\">Khan2014</a>, <a href=\"https://ieeexplore.ieee.org/abstract/document/7460968\" target=\"_blank\">Vahadane2016</a>] were used. However, these methods rely on selection of a single reference slide. Ren et al. have since shown that using an ensemble with different reference slides is one possible solution [<a href=\"https://www.frontiersin.org/articles/10.3389/fbioe.2019.00102/full\" target=\"_blank\">Ren2019</a>].\n\n<center><img src=\"https://i.ibb.co/MVKpDX9/domain-shift4.jpg\" alt=\"domain-shift4\" width=100%></center>\n\n***Figure 4:*** *H&E image (left) and deconvolved hematoxylin (middle) and eosin (right) [[**ref: Macenko2009**]](https://ieeexplore.ieee.org/abstract/document/5193250)*\n\n<br>\n\nThe larger problem is that these techniques do not consider spatial features, which can lead to tissue structure not being preserved.\n\n**G**enerative **A**dversarial **N**etworks (**GAN**s) are the state-of-the-art in stain normalization today. \n* Given an image from domain A, a generator converts it into domain B. \n* A discriminator network tries to distinguish real domain B images from fake ones, helping the generator to improve.\n\nIf paired and aligned images from domains A and B are available, this setup performs well. \n* However, it typically requires scanning each slide on two different scanners -- or potentially even restaining and rescanning each slide.\n\n**But there is a simpler solution to obtaining paired images:** \n* convert a color image to grayscale (domain A) and pair it with the original color image (domain B) [[Salehi2020]](https://arxiv.org/abs/2002.00647). \n* The two are perfectly aligned and a Conditional GAN can be trained to reconstruct the color image.\n\n<center><img src=\"https://i.ibb.co/pjpknGP/domain-shift5.jpg\" alt=\"domain-shift5\" width=100%></center>\n\n***Figure 5:*** *Stain style transfer from grayscale to H&E with a Conditional GAN [[**ref: Cho2017]](https://arxiv.org/abs/1710.08543)*\n\n<br>\n\nOne major advantage of this approach is that **a restaining model trained for one particular domain may work for a variety of different labs and scanners** as there is less variation in the input grayscale images than in color.\n\nAn alternative approach when paired images are not available is a [CycleGAN [Zhu2017]](https://openaccess.thecvf.com/content_iccv_2017/html/Zhu_Unpaired_Image-To-Image_Translation_ICCV_2017_paper.html). \n* In this setup there are two generators: one to convert from domain A to B and another to go from domain B to A. \n* The goal of these two models is to be able to reconstruct an original image: \n  * A -> B -> A or B -> A -> B. \n* CycleGANs also makes use of discriminators to predict real versus generated images for each domain.\n\n<center><img src=\"https://i.ibb.co/3psgzCS/domain-shift6.jpg\" alt=\"domain-shift6\" width=100%></center>\n\n***Figure 6:*** *Stain style transfer with a CycleGAN [[**ref: Lo2021]](https://www.sciencedirect.com/science/article/abs/pii/S1568494620307602)*\n\n<br>\n\nStain normalization methods using deep learning have become increasingly complex. As a first pass to see if this type of standardization is helpful for your task, I [Heather D. Couture Ph.D.] suggest starting simple. \n* StainTools and HistomicsTK both implement some of the color matching and stain separation methods.\n* These simpler methods are sufficient in some cases, but not all. \n* The figure below demonstrates how different methods perform on five different datasets.\n\n<center><img src=\"https://i.ibb.co/crDH8mx/domain-shift7.jpg\" alt=\"domain-shift7\" width=100%></center>\n\n***Figure 7:*** *Comparison of stain normalization results on different datasets [[**ref: Salehi2020]](https://arxiv.org/abs/2002.00647)*\n\n<br>\n\n---\n\n<br>\n\n#### **2. COLOR AUGMENTATION**\n\n<br>\n\nImage augmentation by applying random affine transforms or adding noise is one of the most common regularization techniques for combating overfitting. Similarly, variations in staining can be taken advantage of to increase the diversity of image appearance presented to the model during training.\n* While drastic changes in color are not realistic for histology, more subtle ones generated through random additive and multiplicative changes to each color channel have been shown to improve model performance.\n* The intensity of color augmentation is an additional hyperparameter that should be experimented with during training and validated on test sets from different labs or scanners.\n\nFaryna et al. demonstrated the **`RandAugment`** technique on histopathology [[Faryna2021]](https://openreview.net/forum?id=JrBfXaoxbA2). \n* This approach parameterizes augmentation as the number of random transformations selected and their magnitude.\n\n<center><img src=\"https://i.ibb.co/Hdk0rrS/domain-shift8.jpg\" alt=\"domain-shift8\" width=100%></center>\n\n***Figure 8:*** *Image augmentation techniques [[**ref: Faryna2021**]](https://openreview.net/forum?id=JrBfXaoxbA2)*\n\n<br>\n\nTellez et al. studied the effect of different augmentation techniques (individually and combined) on mitosis detection and proposed an H&E-specific transform [[Tellez2018]](https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf). \n* They performed color deconvolution (as is used in the stain separation methods mentioned above), then applied the random shifts in hematoxylin and eosin space before converting back to RGB. \n* The H&E transform was the best performing individual augmentation method. \n* A combination of all augmentation methods was critical in generalizing performance to a new dataset.\n\n<center><img src=\"https://i.ibb.co/QfHDR7S/domain-shift9.jpg\" alt=\"domain-shift9\" width=100%></center>\n\n***Figure 9:*** *Color augmentation on hematoxylin and eosin channels [[**ref: Tellez2018**]](https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf)*\n\n<br>\n\n---\n\n<br>\n\n#### **3. UNSUPERVISED DOMAIN ADVERSARIAL TRAINING**\n\n<br>\n\nThe next technique for domain adaptation is domain adversarial training [[Ganin2016]](https://dl.acm.org/doi/abs/10.5555/2946645.2946704). This approach makes use of unlabeled images from the target domain.\n\nA domain adversarial module is added to an existing model. \n* The goal of this classifier is to predict whether an image belongs to the source or the target domain. \n* A gradient reversal layer connects this module to the existing networking so that training optimizes the original task and encourages the network to learn domain-invariant features.\n\nDuring training, labeled images from the source domain and unlabeled ones from the target domain are used. \n* For labeled source images, both the loss of the original network and the domain loss are applied. \n* For unlabeled target images, only the domain loss is used.\n\n<center><img src=\"https://i.ibb.co/qRmDYY7/domain-shift10.jpg\" alt=\"domain-shift10\" width=100%></center>\n\n***Figure 10:*** *A domain adversarial module attached to a CNN classifier [[**ref: Ganin2016**]](https://dl.acm.org/doi/abs/10.5555/2946645.2946704)*\n\n<br>\n\nThis module can be added to a variety of deep learning models. \n* For **classification**, it is typically connected to a layer near the output. \n* For segmentation, it is usually applied to the bottleneck layer - although it can also be applied to multiple layers. \n* For detection, it can be applied to the feature pyramid network. \n* For challenging object detectors like mitoses that require an additional classifier network, the domain adversarial network may only be applied to the second stage [[Aubreville2020a]](https://arxiv.org/abs/1911.10873).\n\n<center><img src=\"https://i.ibb.co/R9VpHhq/domain-shift11.jpg\" alt=\"domain-shift11\" width=100%></center>\n\n***Figure 11:*** *Improved domain invariance with a domain adapted model [[**ref: Aubreville2020a**]](https://arxiv.org/abs/1911.10873)*\n\n<br>\n\n---\n\n<br>\n\n#### **4. ADAPT MODEL AT TEST TIME**\n\n<br>\n\nInstead of accommodating domain shifts during training, the model may be modified at test time. \n* Domain shifts are reflected in a change in distribution in feature space -- a covariate shift. \n* So for models using batch normalization layers, the **mean and standard deviation may be recalculated for a new test set**.\n\nThese new statistics can be calculated over the whole test set and updated in the model before running inference. \n* Or they can be calculated for each new batch of data. \n* Nado et al. found that the latter approach, called prediction-time batch normalization, was sufficient [[Nado2020]](https://arxiv.org/abs/2006.10963). \n* Further, a single batch of 500 images was sufficient to get a substantial improvement in model accuracy.\n\n<center><img src=\"https://i.ibb.co/9ZgCPy5/domain-shift12.jpg\" alt=\"domain-shift12\" width=100%></center>\n\n***Figure 12:*** *Prediction-time batch normalization aligns the shifted activations with the training distribution [[**ref: Nado2020**]](https://arxiv.org/abs/2006.10963)*\n\n<br>\n\n---\n\n<br>\n\n#### **5. MODEL FINETUNING**\n\n<br>\n\nFinally, a model may be finetuned on a test set with a domain shift [[Aubreville2020b]](https://www.nature.com/articles/s41597-020-00756-z). \n* If enough labeled examples are available in the test set, this approach is likely to produce the best result. \n* However, it is the most time-consuming and least generalizable. \n* The model may need to be finetuned again on other test sets in the future.\n\n<br>\n\n---\n\n<br><br>\n\n#### **COMPARISON OF APPROACHES**\n\nColor augmentation and stain normalization are used extensively in pathology image applications, especially for whole slide H&E images. Domain adversarial training and model adaptation at test time are less studied thus far.\n\n<br>\n\n##### **STAIN NORMALIZATION v. COLOR AUGMENTATION**\n\nKhan et al. studied stain normalization and color augmentation when applied individually and together [[Khan2020]](https://www.researchgate.net/publication/338826217_Generalizing_Convolution_Neural_Networks_on_Stain_Color_Heterogeneous_Data_for_Computational_Pathology). \n* The best results were obtained when the methods were used together.\n\nTellez et al. tested different combinations of image augmentation and normalization strategies for a variety of different histology classification tasks [[Tellez2019]](https://arxiv.org/abs/1902.06543). \n* The best configuration turned out to be randomly shifting the color channels and applying no stain normalization. \n* Experiments using stain normalization performed only slightly poorer. \n* This validates the importance of image augmentation in creating a robust classifier for histology and emphasizes the importance of color transformations. \n* While applying stain normalization didn’t really hurt, this extra computation may not be necessary.\n\n<br>\n\n##### **STAIN NORMALIZATION v. DOMAIN ADVERSARIAL TRAINING**\n\nRen et al. compared domain adversarial training with some stain normalization and color augmentation approaches, demonstrating that **the domain adversarial approach was superior for generalizing to new image sets** [[Ren2019]](https://www.frontiersin.org/articles/10.3389/fbioe.2019.00102/full).\n\n<br>\n\n##### **STAIN NORMALIZATION v. COLOR AUGMENTATION v. DOMAIN ADVERSARIAL TRAINING**\n\nLarfarge et al. performed a similar study that also included domain adversarial training. \n\nThey compared domain adversarial training with color augmentation and stain normalization on mitosis classification and nuclei segmentation tasks [[Lafarge2019]](https://www.frontiersin.org/articles/10.3389/fmed.2019.00162/full).\n* On mitosis **classification**, they found that color augmentation performed best for test images from the same lab on which the model was trained\n* However, on images from other labs, domain adversarial training combined with color augmentation was best\n\nFor nuclei segmentation, the results were a bit different. \n* Stain normalization was key when tested on images of the same tissue type. \n* On different tissue types, domain adversarial training with stain normalization was best.\n\n**Clearly domain adversarial training was beneficial for both types of domain shifts!**\n* However, the best preprocessing and augmentation strategy varied with the dataset. \n* Lafarge speculated that this is due to the amount of domain variability in the training set.\n\n<br>\n\n##### **SIDE NOTE:**\n\nThe stain normalization methods tested in the above analysis of methods are only the traditional color matching and stain separation techniques. \n* Many newer deep learning-based approaches exist that were not included in these benchmarks.\n\n<br>\n\n---\n\n<br><br>\n\n#### **RECOMMENDATIONS**\n\nTraditional stain normalization techniques are worth trying as a first pass as they’re easier to implement and often faster to run. \n* For some domain shifts, these may even be sufficient, especially when combined with color augmentation or domain adaptation. \n* Experiments with a simpler method can also provide insights into whether some amount of stain normalization will improve model generalization performance.\n* For more robust stain normalization that preserves tissue structure, evaluate the methods discussed above.\n\nThe available data may also be a deciding factor in selecting appropriate techniques. \n* Stain normalization and color augmentation do not require target domain images during training, while the other three approaches do. \n* Model adaptation requires unlabeled target data, while finetuning needs it labeled. \n\nFor these reasons, stain normalization and color augmentation are often tried first, with adversarial domain adaptation included when needed. \n* These three methods are also the best approaches for training a single generalizable model. \n* If a large set of target images is available, then adapting the model (with unlabeled data) or finetuning (with labeled data) will likely be most effective.\n\nMonitoring a deployed system for unexpected domain shifts is also critical. \n* Stacke et al. developed a way to quantify domain shift [[Stacke2020]](https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1478702&dswid=9029).\n* Their metric does not require annotated data, so can serve as a simple test to see if new data is likely to be handled well by an existing model.\n\n---\n\n## **FIN**",
      "votes": null
    },
    {
      "id": "1860952",
      "postDate": "07/18/2022 17:03:23",
      "content": "<p>Follow and/or visit the author <a href=\"https://www.linkedin.com/in/hdcouture/\" target=\"_blank\"><strong>Heather D. Couture</strong></a> Ph.D. on <a href=\"https://www.linkedin.com/in/hdcouture/\" target=\"_blank\"><strong>LinkedIn</strong></a> for other insightful posts and information related to the domain of Artificial Intelligence in Pathology (and much more!).</p>",
      "rawMarkdown": "Follow and/or visit the author [**Heather D. Couture**](https://www.linkedin.com/in/hdcouture/) Ph.D. on [**LinkedIn**](https://www.linkedin.com/in/hdcouture/) for other insightful posts and information related to the domain of Artificial Intelligence in Pathology (and much more!).",
      "votes": null
    },
    {
      "id": "1862642",
      "postDate": "07/19/2022 21:37:08",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": null
    },
    {
      "id": "1862737",
      "postDate": "07/20/2022 00:32:26",
      "content": "<p>Super! Thank you very much for this useful sharing!</p>",
      "rawMarkdown": "Super! Thank you very much for this useful sharing!",
      "votes": null
    },
    {
      "id": "1867486",
      "postDate": "07/23/2022 08:46:25",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!",
      "votes": null
    },
    {
      "id": "1892616",
      "postDate": "08/10/2022 07:40:31",
      "content": "<p>I have implemented <a href=\"https://www.kaggle.com/code/shir0mani/hubmap-stain-transfer-w-pix2pix-pytorch\" target=\"_blank\">stain transfer using Pix2Pix</a> (mentioned in Section 1 above) which uses gray/color image pairs. I used HPA images to train the GAN but external HuBMAP images can be used as well. Training with HuBMAP images will give HuBMAP (H&amp;E) stained images.</p>",
      "rawMarkdown": "I have implemented [stain transfer using Pix2Pix](https://www.kaggle.com/code/shir0mani/hubmap-stain-transfer-w-pix2pix-pytorch) (mentioned in Section 1 above) which uses gray/color image pairs. I used HPA images to train the GAN but external HuBMAP images can be used as well. Training with HuBMAP images will give HuBMAP (H&E) stained images.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1860952,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "07/18/2022 17:03:23",
      "content": "<p>Follow and/or visit the author <a href=\"https://www.linkedin.com/in/hdcouture/\" target=\"_blank\"><strong>Heather D. Couture</strong></a> Ph.D. on <a href=\"https://www.linkedin.com/in/hdcouture/\" target=\"_blank\"><strong>LinkedIn</strong></a> for other insightful posts and information related to the domain of Artificial Intelligence in Pathology (and much more!).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1862642,
      "author_name": "muhammad4hmed",
      "author_url": "",
      "post_date": "07/19/2022 21:37:08",
      "content": "<p>Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1862737,
      "author_name": "mawanda",
      "author_url": "",
      "post_date": "07/20/2022 00:32:26",
      "content": "<p>Super! Thank you very much for this useful sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1867486,
      "author_name": "bridgeoverwater",
      "author_url": "",
      "post_date": "07/23/2022 08:46:25",
      "content": "<p>thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1892616,
      "author_name": "shir0mani",
      "author_url": "",
      "post_date": "08/10/2022 07:40:31",
      "content": "<p>I have implemented <a href=\"https://www.kaggle.com/code/shir0mani/hubmap-stain-transfer-w-pix2pix-pytorch\" target=\"_blank\">stain transfer using Pix2Pix</a> (mentioned in Section 1 above) which uses gray/color image pairs. I used HPA images to train the GAN but external HuBMAP images can be used as well. Training with HuBMAP images will give HuBMAP (H&amp;E) stained images.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1860944": "<br>\n\n### ***NOTE: This is essentially just a repost/share of the [article/information written here](https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html) with updates to allow for better formatting in markdown.***\n\n***PLEASE GIVE CREDIT TO THE ORIGINAL AUTHOR --> [Heather D. Couture Ph.D.](http://www.hdcouture.com/)***\n\n*I'm sharing it as I found it incredibly helpful and instead of posting the link I thought reiterating the content here would help more people*\n\n---\n\n* **TITLE:**\n  * 5 WAYS TO MAKE HISTOPATHOLOGY IMAGE MODELS MORE ROBUST TO DOMAIN SHIFTS\n* **AUTHOR:**\n  * Heather D. Couture Ph.D.\n* **DATE PUBLISHED:**\n  * August 2, 2021\n* **LOCATION PUBLISHED INITIALLY:**\n  * [Towards Data Science](https://towardsdatascience.com/)\n* **LINK:**\n  * https://pixelscientia.com/article-5-ways-to-make-histopathology-image-models-more-robust-to-domain-shifts.html\n\n---\n\n<br><br>\n\n### **5 WAYS TO MAKE HISTOPATHOLOGY IMAGE MODELS MORE ROBUST TO DOMAIN SHIFTS**\n\n---\n\nOne of the largest challenges in histopathology image analysis is creating models that are robust to the variations across different labs and imaging systems. These variations can be caused by:\n* different color responses of slide scanners\n* raw materials\n* manufacturing techniques\n* protocols for staining.\n\n<center><img src=\"https://i.ibb.co/CWSsS7k/domain-shift1.jpg\" width=100%></center>\n\n***Figure 1:*** *H&E image variations induced by different scanners [[**ref: Aubreville2021**]](https://arxiv.org/abs/2103.16515)*\n\n<br>\n\nDifferent setups can produce images with different stain intensities or other changes, **creating a domain shift between the source data that the model was trained on and the target data on which a deployed solution would need to operate**. When the domain shift is too large, a model trained on one type of data will fail on another, often in unpredictable ways.\n\nI’ve [Heather D. Couture Ph.D.] been working with a client who is interested in selecting the best object detection model for their use case. But the images their model will examine once deployed come from a different lab and a variety of scanners.\n* The domain shift from their training dataset to the target one is likely to be a larger challenge than getting state-of-the-art results on the training set.\n* My [Heather D. Couture Ph.D.] advice to them was to **tackle this domain adaptation challenge early.**\n* They can always experiment with better object detection models later, once they’ve learned how to handle the domain shift.\n\n<center><img src=\"https://i.ibb.co/7bJ3yKT/domain-shift2.jpg\" alt=\"domain-shift2\" width=100%></center>\n\n***Figure 2:*** *Mitosis appearance in two different datasets: Radboudumc (top) and TUPAC (bottom) [[**ref: Tellez2018**]](https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf)*\n\n<br>\n\nSo how do you handle the domain shift? There are a few different options:\n1. Standardize the appearance of your images with stain normalization techniques\n2. Color augmentation during training to take advantage of variations in staining\n3. Domain adversarial training to learn domain-invariant features\n4. Adapt the model at test time to handle the new image distribution\n5. Finetune the model on the target domain\n\nSome of these approaches have opposing goals. \n* For example, color augmentation increases the diversity of images, while stain normalization tries to reduce the variations. * Domain adversarial training tries to learn domain-invariant features, while adapting or finetuning the model transforms the model to be suitable for the target domain only.\n\n<center><img src=\"https://i.ibb.co/8Kd56k0/domain-shift3.jpg\" alt=\"domain-shift3\" width=100%></center>\n\n***Figure 3:*** *The different goals of color augmentation, stain normalization, and domain alignment by with a GAN [[**ref: Stacke2020**]](https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1478702&dswid=9029)*\n\n<br>\n\n**This article [Kaggle Discussion Post] will review each of the five strategies, followed by a summary of studies revealing which method(s) work best.**\n\n<br>\n\n---\n\n<br>\n\n#### **1. STAIN NORMALIZATION**\n\n<br>\n\nDifferent labs and scanners can produce images with different color profiles for a particular stain. **The goal of stain normalization is to standardize the appearance of these stains.**\n\nTraditionally, methods like color matching [<a href=\"https://ieeexplore.ieee.org/abstract/document/946629\" target=\"_blank\">Reinhard2001</a>] and stain separation [<a href=\"https://ieeexplore.ieee.org/abstract/document/5193250\" target=\"_blank\">Macenko2009</a>, <a href=\"https://ieeexplore.ieee.org/abstract/document/6727397\" target=\"_blank\">Khan2014</a>, <a href=\"https://ieeexplore.ieee.org/abstract/document/7460968\" target=\"_blank\">Vahadane2016</a>] were used. However, these methods rely on selection of a single reference slide. Ren et al. have since shown that using an ensemble with different reference slides is one possible solution [<a href=\"https://www.frontiersin.org/articles/10.3389/fbioe.2019.00102/full\" target=\"_blank\">Ren2019</a>].\n\n<center><img src=\"https://i.ibb.co/MVKpDX9/domain-shift4.jpg\" alt=\"domain-shift4\" width=100%></center>\n\n***Figure 4:*** *H&E image (left) and deconvolved hematoxylin (middle) and eosin (right) [[**ref: Macenko2009**]](https://ieeexplore.ieee.org/abstract/document/5193250)*\n\n<br>\n\nThe larger problem is that these techniques do not consider spatial features, which can lead to tissue structure not being preserved.\n\n**G**enerative **A**dversarial **N**etworks (**GAN**s) are the state-of-the-art in stain normalization today. \n* Given an image from domain A, a generator converts it into domain B. \n* A discriminator network tries to distinguish real domain B images from fake ones, helping the generator to improve.\n\nIf paired and aligned images from domains A and B are available, this setup performs well. \n* However, it typically requires scanning each slide on two different scanners -- or potentially even restaining and rescanning each slide.\n\n**But there is a simpler solution to obtaining paired images:** \n* convert a color image to grayscale (domain A) and pair it with the original color image (domain B) [[Salehi2020]](https://arxiv.org/abs/2002.00647). \n* The two are perfectly aligned and a Conditional GAN can be trained to reconstruct the color image.\n\n<center><img src=\"https://i.ibb.co/pjpknGP/domain-shift5.jpg\" alt=\"domain-shift5\" width=100%></center>\n\n***Figure 5:*** *Stain style transfer from grayscale to H&E with a Conditional GAN [[**ref: Cho2017]](https://arxiv.org/abs/1710.08543)*\n\n<br>\n\nOne major advantage of this approach is that **a restaining model trained for one particular domain may work for a variety of different labs and scanners** as there is less variation in the input grayscale images than in color.\n\nAn alternative approach when paired images are not available is a [CycleGAN [Zhu2017]](https://openaccess.thecvf.com/content_iccv_2017/html/Zhu_Unpaired_Image-To-Image_Translation_ICCV_2017_paper.html). \n* In this setup there are two generators: one to convert from domain A to B and another to go from domain B to A. \n* The goal of these two models is to be able to reconstruct an original image: \n  * A -> B -> A or B -> A -> B. \n* CycleGANs also makes use of discriminators to predict real versus generated images for each domain.\n\n<center><img src=\"https://i.ibb.co/3psgzCS/domain-shift6.jpg\" alt=\"domain-shift6\" width=100%></center>\n\n***Figure 6:*** *Stain style transfer with a CycleGAN [[**ref: Lo2021]](https://www.sciencedirect.com/science/article/abs/pii/S1568494620307602)*\n\n<br>\n\nStain normalization methods using deep learning have become increasingly complex. As a first pass to see if this type of standardization is helpful for your task, I [Heather D. Couture Ph.D.] suggest starting simple. \n* StainTools and HistomicsTK both implement some of the color matching and stain separation methods.\n* These simpler methods are sufficient in some cases, but not all. \n* The figure below demonstrates how different methods perform on five different datasets.\n\n<center><img src=\"https://i.ibb.co/crDH8mx/domain-shift7.jpg\" alt=\"domain-shift7\" width=100%></center>\n\n***Figure 7:*** *Comparison of stain normalization results on different datasets [[**ref: Salehi2020]](https://arxiv.org/abs/2002.00647)*\n\n<br>\n\n---\n\n<br>\n\n#### **2. COLOR AUGMENTATION**\n\n<br>\n\nImage augmentation by applying random affine transforms or adding noise is one of the most common regularization techniques for combating overfitting. Similarly, variations in staining can be taken advantage of to increase the diversity of image appearance presented to the model during training.\n* While drastic changes in color are not realistic for histology, more subtle ones generated through random additive and multiplicative changes to each color channel have been shown to improve model performance.\n* The intensity of color augmentation is an additional hyperparameter that should be experimented with during training and validated on test sets from different labs or scanners.\n\nFaryna et al. demonstrated the **`RandAugment`** technique on histopathology [[Faryna2021]](https://openreview.net/forum?id=JrBfXaoxbA2). \n* This approach parameterizes augmentation as the number of random transformations selected and their magnitude.\n\n<center><img src=\"https://i.ibb.co/Hdk0rrS/domain-shift8.jpg\" alt=\"domain-shift8\" width=100%></center>\n\n***Figure 8:*** *Image augmentation techniques [[**ref: Faryna2021**]](https://openreview.net/forum?id=JrBfXaoxbA2)*\n\n<br>\n\nTellez et al. studied the effect of different augmentation techniques (individually and combined) on mitosis detection and proposed an H&E-specific transform [[Tellez2018]](https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf). \n* They performed color deconvolution (as is used in the stain separation methods mentioned above), then applied the random shifts in hematoxylin and eosin space before converting back to RGB. \n* The H&E transform was the best performing individual augmentation method. \n* A combination of all augmentation methods was critical in generalizing performance to a new dataset.\n\n<center><img src=\"https://i.ibb.co/QfHDR7S/domain-shift9.jpg\" alt=\"domain-shift9\" width=100%></center>\n\n***Figure 9:*** *Color augmentation on hematoxylin and eosin channels [[**ref: Tellez2018**]](https://geertlitjens.nl/publication/tell-18-a/tell-18-a.pdf)*\n\n<br>\n\n---\n\n<br>\n\n#### **3. UNSUPERVISED DOMAIN ADVERSARIAL TRAINING**\n\n<br>\n\nThe next technique for domain adaptation is domain adversarial training [[Ganin2016]](https://dl.acm.org/doi/abs/10.5555/2946645.2946704). This approach makes use of unlabeled images from the target domain.\n\nA domain adversarial module is added to an existing model. \n* The goal of this classifier is to predict whether an image belongs to the source or the target domain. \n* A gradient reversal layer connects this module to the existing networking so that training optimizes the original task and encourages the network to learn domain-invariant features.\n\nDuring training, labeled images from the source domain and unlabeled ones from the target domain are used. \n* For labeled source images, both the loss of the original network and the domain loss are applied. \n* For unlabeled target images, only the domain loss is used.\n\n<center><img src=\"https://i.ibb.co/qRmDYY7/domain-shift10.jpg\" alt=\"domain-shift10\" width=100%></center>\n\n***Figure 10:*** *A domain adversarial module attached to a CNN classifier [[**ref: Ganin2016**]](https://dl.acm.org/doi/abs/10.5555/2946645.2946704)*\n\n<br>\n\nThis module can be added to a variety of deep learning models. \n* For **classification**, it is typically connected to a layer near the output. \n* For segmentation, it is usually applied to the bottleneck layer - although it can also be applied to multiple layers. \n* For detection, it can be applied to the feature pyramid network. \n* For challenging object detectors like mitoses that require an additional classifier network, the domain adversarial network may only be applied to the second stage [[Aubreville2020a]](https://arxiv.org/abs/1911.10873).\n\n<center><img src=\"https://i.ibb.co/R9VpHhq/domain-shift11.jpg\" alt=\"domain-shift11\" width=100%></center>\n\n***Figure 11:*** *Improved domain invariance with a domain adapted model [[**ref: Aubreville2020a**]](https://arxiv.org/abs/1911.10873)*\n\n<br>\n\n---\n\n<br>\n\n#### **4. ADAPT MODEL AT TEST TIME**\n\n<br>\n\nInstead of accommodating domain shifts during training, the model may be modified at test time. \n* Domain shifts are reflected in a change in distribution in feature space -- a covariate shift. \n* So for models using batch normalization layers, the **mean and standard deviation may be recalculated for a new test set**.\n\nThese new statistics can be calculated over the whole test set and updated in the model before running inference. \n* Or they can be calculated for each new batch of data. \n* Nado et al. found that the latter approach, called prediction-time batch normalization, was sufficient [[Nado2020]](https://arxiv.org/abs/2006.10963). \n* Further, a single batch of 500 images was sufficient to get a substantial improvement in model accuracy.\n\n<center><img src=\"https://i.ibb.co/9ZgCPy5/domain-shift12.jpg\" alt=\"domain-shift12\" width=100%></center>\n\n***Figure 12:*** *Prediction-time batch normalization aligns the shifted activations with the training distribution [[**ref: Nado2020**]](https://arxiv.org/abs/2006.10963)*\n\n<br>\n\n---\n\n<br>\n\n#### **5. MODEL FINETUNING**\n\n<br>\n\nFinally, a model may be finetuned on a test set with a domain shift [[Aubreville2020b]](https://www.nature.com/articles/s41597-020-00756-z). \n* If enough labeled examples are available in the test set, this approach is likely to produce the best result. \n* However, it is the most time-consuming and least generalizable. \n* The model may need to be finetuned again on other test sets in the future.\n\n<br>\n\n---\n\n<br><br>\n\n#### **COMPARISON OF APPROACHES**\n\nColor augmentation and stain normalization are used extensively in pathology image applications, especially for whole slide H&E images. Domain adversarial training and model adaptation at test time are less studied thus far.\n\n<br>\n\n##### **STAIN NORMALIZATION v. COLOR AUGMENTATION**\n\nKhan et al. studied stain normalization and color augmentation when applied individually and together [[Khan2020]](https://www.researchgate.net/publication/338826217_Generalizing_Convolution_Neural_Networks_on_Stain_Color_Heterogeneous_Data_for_Computational_Pathology). \n* The best results were obtained when the methods were used together.\n\nTellez et al. tested different combinations of image augmentation and normalization strategies for a variety of different histology classification tasks [[Tellez2019]](https://arxiv.org/abs/1902.06543). \n* The best configuration turned out to be randomly shifting the color channels and applying no stain normalization. \n* Experiments using stain normalization performed only slightly poorer. \n* This validates the importance of image augmentation in creating a robust classifier for histology and emphasizes the importance of color transformations. \n* While applying stain normalization didn’t really hurt, this extra computation may not be necessary.\n\n<br>\n\n##### **STAIN NORMALIZATION v. DOMAIN ADVERSARIAL TRAINING**\n\nRen et al. compared domain adversarial training with some stain normalization and color augmentation approaches, demonstrating that **the domain adversarial approach was superior for generalizing to new image sets** [[Ren2019]](https://www.frontiersin.org/articles/10.3389/fbioe.2019.00102/full).\n\n<br>\n\n##### **STAIN NORMALIZATION v. COLOR AUGMENTATION v. DOMAIN ADVERSARIAL TRAINING**\n\nLarfarge et al. performed a similar study that also included domain adversarial training. \n\nThey compared domain adversarial training with color augmentation and stain normalization on mitosis classification and nuclei segmentation tasks [[Lafarge2019]](https://www.frontiersin.org/articles/10.3389/fmed.2019.00162/full).\n* On mitosis **classification**, they found that color augmentation performed best for test images from the same lab on which the model was trained\n* However, on images from other labs, domain adversarial training combined with color augmentation was best\n\nFor nuclei segmentation, the results were a bit different. \n* Stain normalization was key when tested on images of the same tissue type. \n* On different tissue types, domain adversarial training with stain normalization was best.\n\n**Clearly domain adversarial training was beneficial for both types of domain shifts!**\n* However, the best preprocessing and augmentation strategy varied with the dataset. \n* Lafarge speculated that this is due to the amount of domain variability in the training set.\n\n<br>\n\n##### **SIDE NOTE:**\n\nThe stain normalization methods tested in the above analysis of methods are only the traditional color matching and stain separation techniques. \n* Many newer deep learning-based approaches exist that were not included in these benchmarks.\n\n<br>\n\n---\n\n<br><br>\n\n#### **RECOMMENDATIONS**\n\nTraditional stain normalization techniques are worth trying as a first pass as they’re easier to implement and often faster to run. \n* For some domain shifts, these may even be sufficient, especially when combined with color augmentation or domain adaptation. \n* Experiments with a simpler method can also provide insights into whether some amount of stain normalization will improve model generalization performance.\n* For more robust stain normalization that preserves tissue structure, evaluate the methods discussed above.\n\nThe available data may also be a deciding factor in selecting appropriate techniques. \n* Stain normalization and color augmentation do not require target domain images during training, while the other three approaches do. \n* Model adaptation requires unlabeled target data, while finetuning needs it labeled. \n\nFor these reasons, stain normalization and color augmentation are often tried first, with adversarial domain adaptation included when needed. \n* These three methods are also the best approaches for training a single generalizable model. \n* If a large set of target images is available, then adapting the model (with unlabeled data) or finetuning (with labeled data) will likely be most effective.\n\nMonitoring a deployed system for unexpected domain shifts is also critical. \n* Stacke et al. developed a way to quantify domain shift [[Stacke2020]](https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1478702&dswid=9029).\n* Their metric does not require annotated data, so can serve as a simple test to see if new data is likely to be handled well by an existing model.\n\n---\n\n## **FIN**",
    "1860952": "Follow and/or visit the author [**Heather D. Couture**](https://www.linkedin.com/in/hdcouture/) Ph.D. on [**LinkedIn**](https://www.linkedin.com/in/hdcouture/) for other insightful posts and information related to the domain of Artificial Intelligence in Pathology (and much more!).",
    "1862642": "Great work!",
    "1862737": "Super! Thank you very much for this useful sharing!",
    "1867486": "thanks for sharing!",
    "1892616": "I have implemented [stain transfer using Pix2Pix](https://www.kaggle.com/code/shir0mani/hubmap-stain-transfer-w-pix2pix-pytorch) (mentioned in Section 1 above) which uses gray/color image pairs. I used HPA images to train the GAN but external HuBMAP images can be used as well. Training with HuBMAP images will give HuBMAP (H&E) stained images."
  },
  "source": "meta"
}