{
  "id": 341235,
  "title": "Data Augmentation: Gaussian-Laplacian Pyramid Blending",
  "url": "/competitions/hubmap-organ-segmentation/discussion/341235",
  "author_name": "",
  "post_date": "2022-08-02T00:29:45.535367Z",
  "votes": 25,
  "comment_count": 4,
  "views": 0,
  "content": "<p>While working on this challenge, I realize some problems that we are facing. Thus, I want to suggest a potential solution for those problems. I also create a notebook: <a href=\"https://www.kaggle.com/code/nghihuynh/data-augmentation-laplacian-pyramid-blending\" target=\"_blank\">Data Augmentation: Gaussian-Laplacian Pyramid Blending</a> to demonstrate the following technique.</p>\n<hr>\n<h3>Motivation:</h3>\n<p>There are <strong>3</strong> main problems:</p>\n<ol>\n<li><p><strong><a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5374099/\" target=\"_blank\">Stain Variations</a></strong>: </p>\n<ul>\n<li><p>Histopathological images (HIs) stained with <strong>DAB + H</strong> based on immunohistochemistry (IHC) technique leads to inter-batch variations. In DAB staining, brown chromogen reveals protein expression, while hematoxylin is for tissue counterstaining.</p></li>\n<li><p>HIs stained with <strong>H&amp;E</strong>: hematoxylin highlights the nuclei with a blueish color, and eosin highlights the cytoplasm and extracellular matrix in pink</p></li>\n<li><p>HIs stained with <strong>PAS</strong>: magenta to red color for PAS positive material, blue color for cell nuclei<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fc9065ce2c07b8f4daaa4dfcee70f1deb%2Fstain_variation_renal.jpeg?generation=1659399459257410&amp;alt=media\" alt=\"\"></p></li>\n<li><p><strong><a href=\"https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&amp;p=PMC3&amp;id=5374099_AHC16025f02.jpg\" target=\"_blank\">Figure 1</a></strong>: A-E (left-right, top-bottom): Results of hematoxylin and eosin (H&amp;E), periodic acid-Schiff (PAS), and immunohistochemical staining (3,3'-diaminobenzidine and hematoxylin, DAB&amp;H) of renal tissue from control rats. <strong>A.</strong> H&amp;E staining. D, distal tubules; G, glomerulus; P, proximal tubules. <strong>B.</strong> PAS staining. Arrows, brush borders; BM, tubular basement membranes; D, distal tubules; P, proximal tubules. <strong>C.</strong> iNOS immunostaining (DAB&amp;H). <strong>D.</strong> BAX immunostaining (DAB&amp;H). <strong>E.</strong> VDR immunostaining (DAB&amp;H).</p></li></ul>\n<p>=&gt; All techniques commonly face differences in intensity, saturation, and hue in the HIs. Color variations may introduce a bias to ML algorithms.</p></li>\n<li><p><strong>Data Imbalance</strong>: DL models require very large datasets for training to avoid overfitting the models. We have some class imbalance for <strong>spleen</strong>, <strong>lung</strong> and <strong>large intestine</strong>. HI datasets are usually small due to expensive labeling. Data imbalance leads to impacts on supervised learning algorithms</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F733fb443451b1261f6acd166c62d1a3b%2Forgan_distribution.png?generation=1659399541136515&amp;alt=media\" alt=\"\"></p></li>\n<li><p><strong>Inter- and Intra-Class Variability</strong>: HIs are not only different between classes, but they are also different within classes</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F23b79729bf500835b4246b1e2f94d8e2%2Finter_intra_variability.png?generation=1659399558402586&amp;alt=media\" alt=\"\"></p></li>\n</ol>\n<p>=&gt; Data augmentation has been actively used to circumvent these problems. So, I found this paper, using the <strong><a href=\"https://arxiv.org/pdf/2002.00072.pdf\" target=\"_blank\">Gaussian-Laplacian pyramid blending</a></strong> technique to generate more data within classes. This technique aims to improve the generalization ability of ML algorithms dealing with HIs.</p>\n<hr>\n<h3>Gaussian-Laplacian Pyramid Blending</h3>\n<p><strong>Objective</strong>: Given 2 histopathological images within the same class, and an image mask, blend the images in a seamless way</p>\n<p><strong>Algorithm Overview</strong>:</p>\n<p><strong>1.</strong> Build <strong>Laplacian</strong> pyramids <em>LA</em> and <em>LB</em> from images A and B</p>\n<p><strong>2.</strong> Build a <strong>Gaussian</strong> pyramid <em>GR</em> from selected region R (mask that says which pixels come from left and which from right)</p>\n<p><strong>3.</strong> Form a <strong>combined</strong> pyramid <em>LS</em> from <em>LA</em> and <em>LB</em> using nodes of <em>GR</em> as weights:</p>\n<p><em>LS(i,j) = GR(i,j)*LA(i,j) + (1-GR(i,j))*LB(i,j)</em></p>\n<p><strong>4.</strong> Collapse the <em>LS</em> pyramid to get the final blended image</p>\n<p>Here is the result from applying this technique:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F0e94b8beab368ee4f93bf63822b8dec8%2Fblended_kidney.png?generation=1659399813375510&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://arxiv.org/pdf/2002.00072.pdf\" target=\"_blank\"><strong>Reference</strong>: Data Augmentation for Histopathological Images Based on Gaussian-Laplacian Pyramid Blending</a></p>\n<hr>\n<p>Let me know if anything is missing or if I need to make any corrections. </p>\n<p>Thank you😊</p>",
  "messages": [
    {
      "id": "1880686",
      "postDate": "08/02/2022 00:29:45",
      "content": "<p>While working on this challenge, I realize some problems that we are facing. Thus, I want to suggest a potential solution for those problems. I also create a notebook: <a href=\"https://www.kaggle.com/code/nghihuynh/data-augmentation-laplacian-pyramid-blending\" target=\"_blank\">Data Augmentation: Gaussian-Laplacian Pyramid Blending</a> to demonstrate the following technique.</p>\n<hr>\n<h3>Motivation:</h3>\n<p>There are <strong>3</strong> main problems:</p>\n<ol>\n<li><p><strong><a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5374099/\" target=\"_blank\">Stain Variations</a></strong>: </p>\n<ul>\n<li><p>Histopathological images (HIs) stained with <strong>DAB + H</strong> based on immunohistochemistry (IHC) technique leads to inter-batch variations. In DAB staining, brown chromogen reveals protein expression, while hematoxylin is for tissue counterstaining.</p></li>\n<li><p>HIs stained with <strong>H&amp;E</strong>: hematoxylin highlights the nuclei with a blueish color, and eosin highlights the cytoplasm and extracellular matrix in pink</p></li>\n<li><p>HIs stained with <strong>PAS</strong>: magenta to red color for PAS positive material, blue color for cell nuclei<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fc9065ce2c07b8f4daaa4dfcee70f1deb%2Fstain_variation_renal.jpeg?generation=1659399459257410&amp;alt=media\" alt=\"\"></p></li>\n<li><p><strong><a href=\"https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&amp;p=PMC3&amp;id=5374099_AHC16025f02.jpg\" target=\"_blank\">Figure 1</a></strong>: A-E (left-right, top-bottom): Results of hematoxylin and eosin (H&amp;E), periodic acid-Schiff (PAS), and immunohistochemical staining (3,3'-diaminobenzidine and hematoxylin, DAB&amp;H) of renal tissue from control rats. <strong>A.</strong> H&amp;E staining. D, distal tubules; G, glomerulus; P, proximal tubules. <strong>B.</strong> PAS staining. Arrows, brush borders; BM, tubular basement membranes; D, distal tubules; P, proximal tubules. <strong>C.</strong> iNOS immunostaining (DAB&amp;H). <strong>D.</strong> BAX immunostaining (DAB&amp;H). <strong>E.</strong> VDR immunostaining (DAB&amp;H).</p></li></ul>\n<p>=&gt; All techniques commonly face differences in intensity, saturation, and hue in the HIs. Color variations may introduce a bias to ML algorithms.</p></li>\n<li><p><strong>Data Imbalance</strong>: DL models require very large datasets for training to avoid overfitting the models. We have some class imbalance for <strong>spleen</strong>, <strong>lung</strong> and <strong>large intestine</strong>. HI datasets are usually small due to expensive labeling. Data imbalance leads to impacts on supervised learning algorithms</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F733fb443451b1261f6acd166c62d1a3b%2Forgan_distribution.png?generation=1659399541136515&amp;alt=media\" alt=\"\"></p></li>\n<li><p><strong>Inter- and Intra-Class Variability</strong>: HIs are not only different between classes, but they are also different within classes</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F23b79729bf500835b4246b1e2f94d8e2%2Finter_intra_variability.png?generation=1659399558402586&amp;alt=media\" alt=\"\"></p></li>\n</ol>\n<p>=&gt; Data augmentation has been actively used to circumvent these problems. So, I found this paper, using the <strong><a href=\"https://arxiv.org/pdf/2002.00072.pdf\" target=\"_blank\">Gaussian-Laplacian pyramid blending</a></strong> technique to generate more data within classes. This technique aims to improve the generalization ability of ML algorithms dealing with HIs.</p>\n<hr>\n<h3>Gaussian-Laplacian Pyramid Blending</h3>\n<p><strong>Objective</strong>: Given 2 histopathological images within the same class, and an image mask, blend the images in a seamless way</p>\n<p><strong>Algorithm Overview</strong>:</p>\n<p><strong>1.</strong> Build <strong>Laplacian</strong> pyramids <em>LA</em> and <em>LB</em> from images A and B</p>\n<p><strong>2.</strong> Build a <strong>Gaussian</strong> pyramid <em>GR</em> from selected region R (mask that says which pixels come from left and which from right)</p>\n<p><strong>3.</strong> Form a <strong>combined</strong> pyramid <em>LS</em> from <em>LA</em> and <em>LB</em> using nodes of <em>GR</em> as weights:</p>\n<p><em>LS(i,j) = GR(i,j)*LA(i,j) + (1-GR(i,j))*LB(i,j)</em></p>\n<p><strong>4.</strong> Collapse the <em>LS</em> pyramid to get the final blended image</p>\n<p>Here is the result from applying this technique:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F0e94b8beab368ee4f93bf63822b8dec8%2Fblended_kidney.png?generation=1659399813375510&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://arxiv.org/pdf/2002.00072.pdf\" target=\"_blank\"><strong>Reference</strong>: Data Augmentation for Histopathological Images Based on Gaussian-Laplacian Pyramid Blending</a></p>\n<hr>\n<p>Let me know if anything is missing or if I need to make any corrections. </p>\n<p>Thank you😊</p>",
      "rawMarkdown": "While working on this challenge, I realize some problems that we are facing. Thus, I want to suggest a potential solution for those problems. I also create a notebook: [Data Augmentation: Gaussian-Laplacian Pyramid Blending](https://www.kaggle.com/code/nghihuynh/data-augmentation-laplacian-pyramid-blending) to demonstrate the following technique.\n\n---\n\n### Motivation:\n\nThere are **3** main problems:\n\n1. **[Stain Variations](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5374099/)**: \n    * Histopathological images (HIs) stained with **DAB + H** based on immunohistochemistry (IHC) technique leads to inter-batch variations. In DAB staining, brown chromogen reveals protein expression, while hematoxylin is for tissue counterstaining.\n\n    * HIs stained with **H&E**: hematoxylin highlights the nuclei with a blueish color, and eosin highlights the cytoplasm and extracellular matrix in pink\n    \n    * HIs stained with **PAS**: magenta to red color for PAS positive material, blue color for cell nuclei\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fc9065ce2c07b8f4daaa4dfcee70f1deb%2Fstain_variation_renal.jpeg?generation=1659399459257410&alt=media)\n    \n    * **[Figure 1](https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&p=PMC3&id=5374099_AHC16025f02.jpg)**: A-E (left-right, top-bottom): Results of hematoxylin and eosin (H&E), periodic acid-Schiff (PAS), and immunohistochemical staining (3,3'-diaminobenzidine and hematoxylin, DAB&H) of renal tissue from control rats. **A.** H&E staining. D, distal tubules; G, glomerulus; P, proximal tubules. **B.** PAS staining. Arrows, brush borders; BM, tubular basement membranes; D, distal tubules; P, proximal tubules. **C.** iNOS immunostaining (DAB&H). **D.** BAX immunostaining (DAB&H). **E.** VDR immunostaining (DAB&H).\n\n    => All techniques commonly face differences in intensity, saturation, and hue in the HIs. Color variations may introduce a bias to ML algorithms.\n\n2. **Data Imbalance**: DL models require very large datasets for training to avoid overfitting the models. We have some class imbalance for **spleen**, **lung** and **large intestine**. HI datasets are usually small due to expensive labeling. Data imbalance leads to impacts on supervised learning algorithms\n\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F733fb443451b1261f6acd166c62d1a3b%2Forgan_distribution.png?generation=1659399541136515&alt=media)\n\n\n3. **Inter- and Intra-Class Variability**: HIs are not only different between classes, but they are also different within classes\n\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F23b79729bf500835b4246b1e2f94d8e2%2Finter_intra_variability.png?generation=1659399558402586&alt=media)\n\n=> Data augmentation has been actively used to circumvent these problems. So, I found this paper, using the **[Gaussian-Laplacian pyramid blending](https://arxiv.org/pdf/2002.00072.pdf)** technique to generate more data within classes. This technique aims to improve the generalization ability of ML algorithms dealing with HIs.\n\n---\n\n### Gaussian-Laplacian Pyramid Blending\n\n**Objective**: Given 2 histopathological images within the same class, and an image mask, blend the images in a seamless way\n\n**Algorithm Overview**:\n\n**1.** Build **Laplacian** pyramids *LA* and *LB* from images A and B\n\n**2.** Build a **Gaussian** pyramid *GR* from selected region R (mask that says which pixels come from left and which from right)\n\n**3.** Form a **combined** pyramid *LS* from *LA* and *LB* using nodes of *GR* as weights:\n\n *LS(i,j) = GR(i,j)\\*LA(i,j) + (1-GR(i,j))\\*LB(i,j)*\n    \n**4.** Collapse the *LS* pyramid to get the final blended image\n\nHere is the result from applying this technique:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F0e94b8beab368ee4f93bf63822b8dec8%2Fblended_kidney.png?generation=1659399813375510&alt=media)\n\n[**Reference**: Data Augmentation for Histopathological Images Based on Gaussian-Laplacian Pyramid Blending](https://arxiv.org/pdf/2002.00072.pdf)\n\n---\n\nLet me know if anything is missing or if I need to make any corrections. \n\nThank you😊",
      "votes": null
    },
    {
      "id": "1894795",
      "postDate": "08/11/2022 17:58:36",
      "content": "<p>Our team just updated the mask blending in the notebook: <a href=\"https://www.kaggle.com/code/nghihuynh/data-augmentation-laplacian-pyramid-blending\" target=\"_blank\">Data Augmentation: Gaussian-Laplacian Pyramid Blending </a>. Here are the results from blending images and masks:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fe6a0e48a7b787201537f75428bb52cd5%2FScreen%20Shot%202022-08-11%20at%201.48.41%20PM.png?generation=1660240637051345&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fe9e33ab2cd9c466658fade318451e65e%2FScreen%20Shot%202022-08-11%20at%201.48.50%20PM.png?generation=1660240653895250&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F0c7ece988d33e79c2d197bc4463c3450%2FScreen%20Shot%202022-08-11%20at%201.48.59%20PM.png?generation=1660240668831984&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fd7777ba07be13fc25c999315041195f0%2FScreen%20Shot%202022-08-11%20at%201.49.07%20PM.png?generation=1660240685565325&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F2bac1c4817528c236466ea3afb6f21d3%2FScreen%20Shot%202022-08-11%20at%201.49.16%20PM.png?generation=1660240697407107&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Our team just updated the mask blending in the notebook: [Data Augmentation: Gaussian-Laplacian Pyramid Blending ](https://www.kaggle.com/code/nghihuynh/data-augmentation-laplacian-pyramid-blending). Here are the results from blending images and masks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fe6a0e48a7b787201537f75428bb52cd5%2FScreen%20Shot%202022-08-11%20at%201.48.41%20PM.png?generation=1660240637051345&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fe9e33ab2cd9c466658fade318451e65e%2FScreen%20Shot%202022-08-11%20at%201.48.50%20PM.png?generation=1660240653895250&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F0c7ece988d33e79c2d197bc4463c3450%2FScreen%20Shot%202022-08-11%20at%201.48.59%20PM.png?generation=1660240668831984&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fd7777ba07be13fc25c999315041195f0%2FScreen%20Shot%202022-08-11%20at%201.49.07%20PM.png?generation=1660240685565325&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F2bac1c4817528c236466ea3afb6f21d3%2FScreen%20Shot%202022-08-11%20at%201.49.16%20PM.png?generation=1660240697407107&alt=media)",
      "votes": null
    },
    {
      "id": "1921377",
      "postDate": "08/31/2022 19:26:08",
      "content": "<p><strong>Update:</strong> <br>\nI tried this method out, but unfortunately, it gave poorer performance on LB. Our approach was to use this blending technique to augment spleen, lung, and large intestine. Then, my teammate fine-tuned the pre-trained models on this dataset and saved the new models.<br>\nResults:</p>\n<ul>\n<li>LB: 0.71: just ensemble different combination of my models (combination A)</li>\n<li>LB: 0.65: combined combination A + models fine-tuned on the augmented dataset</li>\n</ul>",
      "rawMarkdown": "**Update:** \nI tried this method out, but unfortunately, it gave poorer performance on LB. Our approach was to use this blending technique to augment spleen, lung, and large intestine. Then, my teammate fine-tuned the pre-trained models on this dataset and saved the new models.\nResults:\n+ LB: 0.71: just ensemble different combination of my models (combination A)\n+ LB: 0.65: combined combination A + models fine-tuned on the augmented dataset",
      "votes": null
    },
    {
      "id": "1921816",
      "postDate": "09/01/2022 04:40:53",
      "content": "<p>Maybe it overfitting on HPA. I use cutmix-aug and get a lower score than without cutmix</p>",
      "rawMarkdown": "Maybe it overfitting on HPA. I use cutmix-aug and get a lower score than without cutmix",
      "votes": null
    },
    {
      "id": "1922485",
      "postDate": "09/01/2022 14:19:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jiageng\" target=\"_blank\">@jiageng</a> ! Yes, it overfitted on HPA. If heavy augmentation leads to overfitting on HPA, then a way around is to introduce some HuBMAP data for training. Did you use any external HuBMAP images for training?</p>",
      "rawMarkdown": "Hi @jiageng ! Yes, it overfitted on HPA. If heavy augmentation leads to overfitting on HPA, then a way around is to introduce some HuBMAP data for training. Did you use any external HuBMAP images for training?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1894795,
      "author_name": "nghihuynh",
      "author_url": "",
      "post_date": "08/11/2022 17:58:36",
      "content": "<p>Our team just updated the mask blending in the notebook: <a href=\"https://www.kaggle.com/code/nghihuynh/data-augmentation-laplacian-pyramid-blending\" target=\"_blank\">Data Augmentation: Gaussian-Laplacian Pyramid Blending </a>. Here are the results from blending images and masks:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fe6a0e48a7b787201537f75428bb52cd5%2FScreen%20Shot%202022-08-11%20at%201.48.41%20PM.png?generation=1660240637051345&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fe9e33ab2cd9c466658fade318451e65e%2FScreen%20Shot%202022-08-11%20at%201.48.50%20PM.png?generation=1660240653895250&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F0c7ece988d33e79c2d197bc4463c3450%2FScreen%20Shot%202022-08-11%20at%201.48.59%20PM.png?generation=1660240668831984&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fd7777ba07be13fc25c999315041195f0%2FScreen%20Shot%202022-08-11%20at%201.49.07%20PM.png?generation=1660240685565325&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F2bac1c4817528c236466ea3afb6f21d3%2FScreen%20Shot%202022-08-11%20at%201.49.16%20PM.png?generation=1660240697407107&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1921377,
      "author_name": "nghihuynh",
      "author_url": "",
      "post_date": "08/31/2022 19:26:08",
      "content": "<p><strong>Update:</strong> <br>\nI tried this method out, but unfortunately, it gave poorer performance on LB. Our approach was to use this blending technique to augment spleen, lung, and large intestine. Then, my teammate fine-tuned the pre-trained models on this dataset and saved the new models.<br>\nResults:</p>\n<ul>\n<li>LB: 0.71: just ensemble different combination of my models (combination A)</li>\n<li>LB: 0.65: combined combination A + models fine-tuned on the augmented dataset</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1921816,
          "author_name": "jiageng",
          "author_url": "",
          "post_date": "09/01/2022 04:40:53",
          "content": "<p>Maybe it overfitting on HPA. I use cutmix-aug and get a lower score than without cutmix</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1922485,
          "author_name": "nghihuynh",
          "author_url": "",
          "post_date": "09/01/2022 14:19:02",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jiageng\" target=\"_blank\">@jiageng</a> ! Yes, it overfitted on HPA. If heavy augmentation leads to overfitting on HPA, then a way around is to introduce some HuBMAP data for training. Did you use any external HuBMAP images for training?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1880686": "While working on this challenge, I realize some problems that we are facing. Thus, I want to suggest a potential solution for those problems. I also create a notebook: [Data Augmentation: Gaussian-Laplacian Pyramid Blending](https://www.kaggle.com/code/nghihuynh/data-augmentation-laplacian-pyramid-blending) to demonstrate the following technique.\n\n---\n\n### Motivation:\n\nThere are **3** main problems:\n\n1. **[Stain Variations](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5374099/)**: \n    * Histopathological images (HIs) stained with **DAB + H** based on immunohistochemistry (IHC) technique leads to inter-batch variations. In DAB staining, brown chromogen reveals protein expression, while hematoxylin is for tissue counterstaining.\n\n    * HIs stained with **H&E**: hematoxylin highlights the nuclei with a blueish color, and eosin highlights the cytoplasm and extracellular matrix in pink\n    \n    * HIs stained with **PAS**: magenta to red color for PAS positive material, blue color for cell nuclei\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fc9065ce2c07b8f4daaa4dfcee70f1deb%2Fstain_variation_renal.jpeg?generation=1659399459257410&alt=media)\n    \n    * **[Figure 1](https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&p=PMC3&id=5374099_AHC16025f02.jpg)**: A-E (left-right, top-bottom): Results of hematoxylin and eosin (H&E), periodic acid-Schiff (PAS), and immunohistochemical staining (3,3'-diaminobenzidine and hematoxylin, DAB&H) of renal tissue from control rats. **A.** H&E staining. D, distal tubules; G, glomerulus; P, proximal tubules. **B.** PAS staining. Arrows, brush borders; BM, tubular basement membranes; D, distal tubules; P, proximal tubules. **C.** iNOS immunostaining (DAB&H). **D.** BAX immunostaining (DAB&H). **E.** VDR immunostaining (DAB&H).\n\n    => All techniques commonly face differences in intensity, saturation, and hue in the HIs. Color variations may introduce a bias to ML algorithms.\n\n2. **Data Imbalance**: DL models require very large datasets for training to avoid overfitting the models. We have some class imbalance for **spleen**, **lung** and **large intestine**. HI datasets are usually small due to expensive labeling. Data imbalance leads to impacts on supervised learning algorithms\n\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F733fb443451b1261f6acd166c62d1a3b%2Forgan_distribution.png?generation=1659399541136515&alt=media)\n\n\n3. **Inter- and Intra-Class Variability**: HIs are not only different between classes, but they are also different within classes\n\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F23b79729bf500835b4246b1e2f94d8e2%2Finter_intra_variability.png?generation=1659399558402586&alt=media)\n\n=> Data augmentation has been actively used to circumvent these problems. So, I found this paper, using the **[Gaussian-Laplacian pyramid blending](https://arxiv.org/pdf/2002.00072.pdf)** technique to generate more data within classes. This technique aims to improve the generalization ability of ML algorithms dealing with HIs.\n\n---\n\n### Gaussian-Laplacian Pyramid Blending\n\n**Objective**: Given 2 histopathological images within the same class, and an image mask, blend the images in a seamless way\n\n**Algorithm Overview**:\n\n**1.** Build **Laplacian** pyramids *LA* and *LB* from images A and B\n\n**2.** Build a **Gaussian** pyramid *GR* from selected region R (mask that says which pixels come from left and which from right)\n\n**3.** Form a **combined** pyramid *LS* from *LA* and *LB* using nodes of *GR* as weights:\n\n *LS(i,j) = GR(i,j)\\*LA(i,j) + (1-GR(i,j))\\*LB(i,j)*\n    \n**4.** Collapse the *LS* pyramid to get the final blended image\n\nHere is the result from applying this technique:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F0e94b8beab368ee4f93bf63822b8dec8%2Fblended_kidney.png?generation=1659399813375510&alt=media)\n\n[**Reference**: Data Augmentation for Histopathological Images Based on Gaussian-Laplacian Pyramid Blending](https://arxiv.org/pdf/2002.00072.pdf)\n\n---\n\nLet me know if anything is missing or if I need to make any corrections. \n\nThank you😊",
    "1894795": "Our team just updated the mask blending in the notebook: [Data Augmentation: Gaussian-Laplacian Pyramid Blending ](https://www.kaggle.com/code/nghihuynh/data-augmentation-laplacian-pyramid-blending). Here are the results from blending images and masks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fe6a0e48a7b787201537f75428bb52cd5%2FScreen%20Shot%202022-08-11%20at%201.48.41%20PM.png?generation=1660240637051345&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fe9e33ab2cd9c466658fade318451e65e%2FScreen%20Shot%202022-08-11%20at%201.48.50%20PM.png?generation=1660240653895250&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F0c7ece988d33e79c2d197bc4463c3450%2FScreen%20Shot%202022-08-11%20at%201.48.59%20PM.png?generation=1660240668831984&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fd7777ba07be13fc25c999315041195f0%2FScreen%20Shot%202022-08-11%20at%201.49.07%20PM.png?generation=1660240685565325&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2F2bac1c4817528c236466ea3afb6f21d3%2FScreen%20Shot%202022-08-11%20at%201.49.16%20PM.png?generation=1660240697407107&alt=media)",
    "1921377": "**Update:** \nI tried this method out, but unfortunately, it gave poorer performance on LB. Our approach was to use this blending technique to augment spleen, lung, and large intestine. Then, my teammate fine-tuned the pre-trained models on this dataset and saved the new models.\nResults:\n+ LB: 0.71: just ensemble different combination of my models (combination A)\n+ LB: 0.65: combined combination A + models fine-tuned on the augmented dataset",
    "1921816": "Maybe it overfitting on HPA. I use cutmix-aug and get a lower score than without cutmix",
    "1922485": "Hi @jiageng ! Yes, it overfitted on HPA. If heavy augmentation leads to overfitting on HPA, then a way around is to introduce some HuBMAP data for training. Did you use any external HuBMAP images for training?"
  },
  "source": "meta"
}