{
  "id": 219143,
  "title": "HuBMAP visualization of errors:   Why isn't my LB score 1.000 ?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/219143",
  "author_name": "Mark A Lavin",
  "post_date": "2021-02-13T16:15:48.551000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I was curious about what the errors would look like in my solution for identifying glomeruli, so I tried the following:</p>\n<p>First, I trained a model following the approach suggested by Wojtek Rosa <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train</a>.   Then I used the model to predict glomeruli in the test sets, public and private.   My best submission so far got a public leaderboard score of 0.850   My next question was, what was the nature of the errors in this score?</p>\n<p>Since I had the ground truth glom mask images only for the training set, I decided to run predictions using the images in the training set as inputs.   I understand that the resulting scores are not valid, due to leakage from the training test cases, but nevertheless believed that errors I saw between predicted and ground truth would provide some insights as to where my real submissions were deficient.</p>\n<p>I created and ran a public notebook <br>\n<a href=\"https://www.kaggle.com/markalavin/hubmap-validate-against-training-set\" target=\"_blank\">https://www.kaggle.com/markalavin/hubmap-validate-against-training-set</a>.   The notebook performs several comparisons between the predicted and ground truth glom mask images:</p>\n<ul>\n<li>For each of the 8 training images, I calculated the individual image DICE coefficient, then took the average of the DICE coefficient, weighted by the total number of 1's in the predicted and ground truth images.</li>\n<li>I divided each of the images (predicted and ground truth) into 1024 x 1024 windows, and computed the DICE coefficient for each window pair, then plotted histograms of the values.</li>\n</ul>\n<p>I realized that looking only at the DICE coefficients, which normalize by the number of 1's in predicted and ground truth, I was not getting a picture of what the magnitude and \"shape\" of the discrepancies was.  So, I tried a different measure I call the XOR coefficient, which is simply the count of the number of pixels that differ between predicted and ground truth.</p>\n<ul>\n<li>For each predicted/ground truth image pair, I computed the XOR coefficient for each window pair, then thresholded the coefficient so that I only considered window pairs with at least 5000 discrepant pixels out of 1 M.   For each such pair, I plotted the predicted and ground truth window images, along with the two-way differences between each, i.e., predicted and not ground truth, and ground truth and not predicted.   So I end up plotting for each discrepant window four images, which I then examined \"manually\".</li>\n</ul>\n<p>I suggest that you take a look at the XOR images yourself, but here's a brief summary of what I saw.   In the vast majority of cases, the set of gloms resembled each other and the differences were thin \"halos\" around the periphery of each glom.   In of these cases, I believe the halo reflects differences because the ground truth ultimately comes from polygonal approximations.   I also observed that the halos were roughly balanced in the two 1-sided image differences, implying that the threshold on the real-valued prediction was probably close to correct.</p>\n<p>In a very small number of cases, I found that there were gloms apparent in the ground truth images that were completely missing in the predicted image; call these \"false negatives\".</p>\n<p>In a slightly larger number of cases (a few %) I found that there were gloms apparent in the predicted image that had no corresponding glom in the ground truth image; call these \"false positives\".   In some of these cases, the \"hallucinated gloms\" were smaller than the average size for gloms; in a subset of those, the gloms were partial or malformed, not the general size and convex shape of most gloms.</p>\n<p>It's possible we might be able to clean up <em>some</em> of these cases in several ways:</p>\n<ul>\n<li>In the predicted image, perform connectivity analysis to identify \"blobs\" and then threshold these blobs according to area and convexity (perimeter ^ 2 / area)</li>\n<li>In the predicted image, perform a morphological opening using a structuring element that's approximately the size of the smallest acceptable glom, something like a diameter of 200 pixels; this of course depends on a more or less uniform scale across the images.</li>\n</ul>",
  "messages": [
    {
      "id": 1199226,
      "postDate": "2021-02-13T16:15:48.553Z",
      "content": "<p>I was curious about what the errors would look like in my solution for identifying glomeruli, so I tried the following:</p>\n<p>First, I trained a model following the approach suggested by Wojtek Rosa <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train</a>.   Then I used the model to predict glomeruli in the test sets, public and private.   My best submission so far got a public leaderboard score of 0.850   My next question was, what was the nature of the errors in this score?</p>\n<p>Since I had the ground truth glom mask images only for the training set, I decided to run predictions using the images in the training set as inputs.   I understand that the resulting scores are not valid, due to leakage from the training test cases, but nevertheless believed that errors I saw between predicted and ground truth would provide some insights as to where my real submissions were deficient.</p>\n<p>I created and ran a public notebook <br>\n<a href=\"https://www.kaggle.com/markalavin/hubmap-validate-against-training-set\" target=\"_blank\">https://www.kaggle.com/markalavin/hubmap-validate-against-training-set</a>.   The notebook performs several comparisons between the predicted and ground truth glom mask images:</p>\n<ul>\n<li>For each of the 8 training images, I calculated the individual image DICE coefficient, then took the average of the DICE coefficient, weighted by the total number of 1's in the predicted and ground truth images.</li>\n<li>I divided each of the images (predicted and ground truth) into 1024 x 1024 windows, and computed the DICE coefficient for each window pair, then plotted histograms of the values.</li>\n</ul>\n<p>I realized that looking only at the DICE coefficients, which normalize by the number of 1's in predicted and ground truth, I was not getting a picture of what the magnitude and \"shape\" of the discrepancies was.  So, I tried a different measure I call the XOR coefficient, which is simply the count of the number of pixels that differ between predicted and ground truth.</p>\n<ul>\n<li>For each predicted/ground truth image pair, I computed the XOR coefficient for each window pair, then thresholded the coefficient so that I only considered window pairs with at least 5000 discrepant pixels out of 1 M.   For each such pair, I plotted the predicted and ground truth window images, along with the two-way differences between each, i.e., predicted and not ground truth, and ground truth and not predicted.   So I end up plotting for each discrepant window four images, which I then examined \"manually\".</li>\n</ul>\n<p>I suggest that you take a look at the XOR images yourself, but here's a brief summary of what I saw.   In the vast majority of cases, the set of gloms resembled each other and the differences were thin \"halos\" around the periphery of each glom.   In of these cases, I believe the halo reflects differences because the ground truth ultimately comes from polygonal approximations.   I also observed that the halos were roughly balanced in the two 1-sided image differences, implying that the threshold on the real-valued prediction was probably close to correct.</p>\n<p>In a very small number of cases, I found that there were gloms apparent in the ground truth images that were completely missing in the predicted image; call these \"false negatives\".</p>\n<p>In a slightly larger number of cases (a few %) I found that there were gloms apparent in the predicted image that had no corresponding glom in the ground truth image; call these \"false positives\".   In some of these cases, the \"hallucinated gloms\" were smaller than the average size for gloms; in a subset of those, the gloms were partial or malformed, not the general size and convex shape of most gloms.</p>\n<p>It's possible we might be able to clean up <em>some</em> of these cases in several ways:</p>\n<ul>\n<li>In the predicted image, perform connectivity analysis to identify \"blobs\" and then threshold these blobs according to area and convexity (perimeter ^ 2 / area)</li>\n<li>In the predicted image, perform a morphological opening using a structuring element that's approximately the size of the smallest acceptable glom, something like a diameter of 200 pixels; this of course depends on a more or less uniform scale across the images.</li>\n</ul>",
      "rawMarkdown": "I was curious about what the errors would look like in my solution for identifying glomeruli, so I tried the following:\n\nFirst, I trained a model following the approach suggested by Wojtek Rosa https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train.   Then I used the model to predict glomeruli in the test sets, public and private.   My best submission so far got a public leaderboard score of 0.850   My next question was, what was the nature of the errors in this score?\n\nSince I had the ground truth glom mask images only for the training set, I decided to run predictions using the images in the training set as inputs.   I understand that the resulting scores are not valid, due to leakage from the training test cases, but nevertheless believed that errors I saw between predicted and ground truth would provide some insights as to where my real submissions were deficient.\n\nI created and ran a public notebook \nhttps://www.kaggle.com/markalavin/hubmap-validate-against-training-set.   The notebook performs several comparisons between the predicted and ground truth glom mask images:\n\n*  For each of the 8 training images, I calculated the individual image DICE coefficient, then took the average of the DICE coefficient, weighted by the total number of 1's in the predicted and ground truth images.\n*   I divided each of the images (predicted and ground truth) into 1024 x 1024 windows, and computed the DICE coefficient for each window pair, then plotted histograms of the values.\n\nI realized that looking only at the DICE coefficients, which normalize by the number of 1's in predicted and ground truth, I was not getting a picture of what the magnitude and \"shape\" of the discrepancies was.  So, I tried a different measure I call the XOR coefficient, which is simply the count of the number of pixels that differ between predicted and ground truth.\n\n*   For each predicted/ground truth image pair, I computed the XOR coefficient for each window pair, then thresholded the coefficient so that I only considered window pairs with at least 5000 discrepant pixels out of 1 M.   For each such pair, I plotted the predicted and ground truth window images, along with the two-way differences between each, i.e., predicted and not ground truth, and ground truth and not predicted.   So I end up plotting for each discrepant window four images, which I then examined \"manually\".\n\nI suggest that you take a look at the XOR images yourself, but here's a brief summary of what I saw.   In the vast majority of cases, the set of gloms resembled each other and the differences were thin \"halos\" around the periphery of each glom.   In of these cases, I believe the halo reflects differences because the ground truth ultimately comes from polygonal approximations.   I also observed that the halos were roughly balanced in the two 1-sided image differences, implying that the threshold on the real-valued prediction was probably close to correct.\n\nIn a very small number of cases, I found that there were gloms apparent in the ground truth images that were completely missing in the predicted image; call these \"false negatives\".\n\nIn a slightly larger number of cases (a few %) I found that there were gloms apparent in the predicted image that had no corresponding glom in the ground truth image; call these \"false positives\".   In some of these cases, the \"hallucinated gloms\" were smaller than the average size for gloms; in a subset of those, the gloms were partial or malformed, not the general size and convex shape of most gloms.\n\nIt's possible we might be able to clean up *some* of these cases in several ways:\n* In the predicted image, perform connectivity analysis to identify \"blobs\" and then threshold these blobs according to area and convexity (perimeter ^ 2 / area)\n* In the predicted image, perform a morphological opening using a structuring element that's approximately the size of the smallest acceptable glom, something like a diameter of 200 pixels; this of course depends on a more or less uniform scale across the images.\n\n",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1199226": "I was curious about what the errors would look like in my solution for identifying glomeruli, so I tried the following:\n\nFirst, I trained a model following the approach suggested by Wojtek Rosa https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train.   Then I used the model to predict glomeruli in the test sets, public and private.   My best submission so far got a public leaderboard score of 0.850   My next question was, what was the nature of the errors in this score?\n\nSince I had the ground truth glom mask images only for the training set, I decided to run predictions using the images in the training set as inputs.   I understand that the resulting scores are not valid, due to leakage from the training test cases, but nevertheless believed that errors I saw between predicted and ground truth would provide some insights as to where my real submissions were deficient.\n\nI created and ran a public notebook \nhttps://www.kaggle.com/markalavin/hubmap-validate-against-training-set.   The notebook performs several comparisons between the predicted and ground truth glom mask images:\n\n*  For each of the 8 training images, I calculated the individual image DICE coefficient, then took the average of the DICE coefficient, weighted by the total number of 1's in the predicted and ground truth images.\n*   I divided each of the images (predicted and ground truth) into 1024 x 1024 windows, and computed the DICE coefficient for each window pair, then plotted histograms of the values.\n\nI realized that looking only at the DICE coefficients, which normalize by the number of 1's in predicted and ground truth, I was not getting a picture of what the magnitude and \"shape\" of the discrepancies was.  So, I tried a different measure I call the XOR coefficient, which is simply the count of the number of pixels that differ between predicted and ground truth.\n\n*   For each predicted/ground truth image pair, I computed the XOR coefficient for each window pair, then thresholded the coefficient so that I only considered window pairs with at least 5000 discrepant pixels out of 1 M.   For each such pair, I plotted the predicted and ground truth window images, along with the two-way differences between each, i.e., predicted and not ground truth, and ground truth and not predicted.   So I end up plotting for each discrepant window four images, which I then examined \"manually\".\n\nI suggest that you take a look at the XOR images yourself, but here's a brief summary of what I saw.   In the vast majority of cases, the set of gloms resembled each other and the differences were thin \"halos\" around the periphery of each glom.   In of these cases, I believe the halo reflects differences because the ground truth ultimately comes from polygonal approximations.   I also observed that the halos were roughly balanced in the two 1-sided image differences, implying that the threshold on the real-valued prediction was probably close to correct.\n\nIn a very small number of cases, I found that there were gloms apparent in the ground truth images that were completely missing in the predicted image; call these \"false negatives\".\n\nIn a slightly larger number of cases (a few %) I found that there were gloms apparent in the predicted image that had no corresponding glom in the ground truth image; call these \"false positives\".   In some of these cases, the \"hallucinated gloms\" were smaller than the average size for gloms; in a subset of those, the gloms were partial or malformed, not the general size and convex shape of most gloms.\n\nIt's possible we might be able to clean up *some* of these cases in several ways:\n* In the predicted image, perform connectivity analysis to identify \"blobs\" and then threshold these blobs according to area and convexity (perimeter ^ 2 / area)\n* In the predicted image, perform a morphological opening using a structuring element that's approximately the size of the smallest acceptable glom, something like a diameter of 200 pixels; this of course depends on a more or less uniform scale across the images.\n\n"
  }
}