{
  "id": 230494,
  "title": "Is it really worthy to use a classifier?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/230494",
  "author_name": "",
  "post_date": "2021-04-04T04:26:10.493255100Z",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Previously I only used folded efficient-unet models for inference, it shows 0.919 in the public LB and the submission can take as long as 8 hours.</p>\n<p>To speed up the submission, I trained a binary classifier to skip some empty masks. Now the submission takes only less than 5 hours, even if I put more and more models in. However, the public LB drops from 0.919 to 0.914.</p>\n<p>I was hoping that a classifier can help me save some time and memory, so I can use more models to do inference, thus increase my public LB score. Now it seems that the result was not that good.</p>\n<p>I was wondering if this is common in this competition, or it is only me(in most cases, I code poorly).</p>",
  "messages": [
    {
      "id": "1262288",
      "postDate": "04/04/2021 04:26:10",
      "content": "<p>Previously I only used folded efficient-unet models for inference, it shows 0.919 in the public LB and the submission can take as long as 8 hours.</p>\n<p>To speed up the submission, I trained a binary classifier to skip some empty masks. Now the submission takes only less than 5 hours, even if I put more and more models in. However, the public LB drops from 0.919 to 0.914.</p>\n<p>I was hoping that a classifier can help me save some time and memory, so I can use more models to do inference, thus increase my public LB score. Now it seems that the result was not that good.</p>\n<p>I was wondering if this is common in this competition, or it is only me(in most cases, I code poorly).</p>",
      "rawMarkdown": "Previously I only used folded efficient-unet models for inference, it shows 0.919 in the public LB and the submission can take as long as 8 hours.\n\nTo speed up the submission, I trained a binary classifier to skip some empty masks. Now the submission takes only less than 5 hours, even if I put more and more models in. However, the public LB drops from 0.919 to 0.914.\n\nI was hoping that a classifier can help me save some time and memory, so I can use more models to do inference, thus increase my public LB score. Now it seems that the result was not that good.\n\nI was wondering if this is common in this competition, or it is only me(in most cases, I code poorly).",
      "votes": null
    },
    {
      "id": "1262650",
      "postDate": "04/04/2021 15:00:41",
      "content": "<p>It’s hard to give advice without knowing much about your code or approach, pruning is always an option. But there’s a trade off between accuracy and speed.</p>",
      "rawMarkdown": "It’s hard to give advice without knowing much about your code or approach, pruning is always an option. But there’s a trade off between accuracy and speed.",
      "votes": null
    },
    {
      "id": "1263053",
      "postDate": "04/05/2021 04:01:03",
      "content": "<p>Thanks for the reply.</p>\n<p>Yes, I do see the trade-off here. The classifier was based on efficientnet b5 and it converges real fast(compared with segmentation models), with good val precision(I used precision as monitor). It seems that this classifier is just getting by on the testing set, it is not a very reliable classifier, yet.</p>\n<p>Also, I am still wondering if the ratio of empty/non-empty masks will affect the performance of classifier(should we use/abandon classifier in balanced/unbalanced scenario). Still lots of things to try…</p>",
      "rawMarkdown": "Thanks for the reply.\n\nYes, I do see the trade-off here. The classifier was based on efficientnet b5 and it converges real fast(compared with segmentation models), with good val precision(I used precision as monitor). It seems that this classifier is just getting by on the testing set, it is not a very reliable classifier, yet.\n\nAlso, I am still wondering if the ratio of empty/non-empty masks will affect the performance of classifier(should we use/abandon classifier in balanced/unbalanced scenario). Still lots of things to try...",
      "votes": null
    },
    {
      "id": "1263820",
      "postDate": "04/05/2021 17:30:06",
      "content": "<p>Those scores 0.919 and 0.914 are within the margin of error from each other. Maybe with more fine tuning the classifier could be more accurate. If you are taking a Bayesian approach there's not much you can do about the time/hardware restrictions of the competition.</p>\n<p>By \"empty/non-empty masks\" do you mean the anatomical mask or the glomeruli masks?</p>",
      "rawMarkdown": "Those scores 0.919 and 0.914 are within the margin of error from each other. Maybe with more fine tuning the classifier could be more accurate. If you are taking a Bayesian approach there's not much you can do about the time/hardware restrictions of the competition.\n\nBy \"empty/non-empty masks\" do you mean the anatomical mask or the glomeruli masks?",
      "votes": null
    },
    {
      "id": "1264273",
      "postDate": "04/06/2021 03:33:00",
      "content": "<p>I have to say sometimes it works but sometimes not. I have a shake of about 0.01 with different principles of classifier.</p>",
      "rawMarkdown": "I have to say sometimes it works but sometimes not. I have a shake of about 0.01 with different principles of classifier.",
      "votes": null
    },
    {
      "id": "1264315",
      "postDate": "04/06/2021 04:31:06",
      "content": "<p>I used Wojtek Rosa's public tfrecords <a href=\"url\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512</a> and <a href=\"url\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512-2</a> as my training/validation set(thanks him very much). I didn't build my own pipeline so I don't know exactly if the masks are anatomical masks or glomeruli masks(sorry).</p>\n<p>I remember there was a notebook showing how the tfrecords are created but I forgot where this notebook can be found. If my memory serves, the masks were decoded from train.csv, but I am not so sure of it…</p>",
      "rawMarkdown": "I used Wojtek Rosa's public tfrecords [https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512](url) and [https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512-2](url) as my training/validation set(thanks him very much). I didn't build my own pipeline so I don't know exactly if the masks are anatomical masks or glomeruli masks(sorry).\n\nI remember there was a notebook showing how the tfrecords are created but I forgot where this notebook can be found. If my memory serves, the masks were decoded from train.csv, but I am not so sure of it...",
      "votes": null
    },
    {
      "id": "1264321",
      "postDate": "04/06/2021 04:38:20",
      "content": "<p>Thanks for the reply. </p>\n<p>Then I think the classification-then-segmentation approach is still worth trying. A shake of 0.01 surely means something in this competition.</p>",
      "rawMarkdown": "Thanks for the reply. \n\nThen I think the classification-then-segmentation approach is still worth trying. A shake of 0.01 surely means something in this competition.",
      "votes": null
    },
    {
      "id": "1264327",
      "postDate": "04/06/2021 04:50:42",
      "content": "<p>Sure. At first I want to keep the original ratio of each samples but it seems to reflect much dffierent results with different qualities of curve worked comparing to using some classifiers. </p>",
      "rawMarkdown": "Sure. At first I want to keep the original ratio of each samples but it seems to reflect much dffierent results with different qualities of curve worked comparing to using some classifiers.",
      "votes": null
    },
    {
      "id": "1264514",
      "postDate": "04/06/2021 08:05:07",
      "content": "<p>The classifier is not perfect. So your segmentator will not see all true positives.<br>\nA clasifier makes sense if you can improve your segmentation models more efficiently than the ratio of the true positives you might lose with your classifier. A simple tradeoff that is really hard to improve at your stage beyond 915+, so good luck :-)</p>",
      "rawMarkdown": "The classifier is not perfect. So your segmentator will not see all true positives.\nA clasifier makes sense if you can improve your segmentation models more efficiently than the ratio of the true positives you might lose with your classifier. A simple tradeoff that is really hard to improve at your stage beyond 915+, so good luck :-)",
      "votes": null
    },
    {
      "id": "1264585",
      "postDate": "04/06/2021 09:24:20",
      "content": "<p>I agreed. The classifier can filter out some of the nearly meaningless samples and make the dataset more balance thus the model perform more stable.</p>",
      "rawMarkdown": "I agreed. The classifier can filter out some of the nearly meaningless samples and make the dataset more balance thus the model perform more stable.",
      "votes": null
    },
    {
      "id": "1264862",
      "postDate": "04/06/2021 13:26:06",
      "content": "<p>if the additional post processing you can do with the extra time at hand can offset the loss of true positives, then it may be worth it. It is a risky proposition, nevertheless, as you don't know what it will do on the private dataset.</p>",
      "rawMarkdown": "if the additional post processing you can do with the extra time at hand can offset the loss of true positives, then it may be worth it. It is a risky proposition, nevertheless, as you don't know what it will do on the private dataset.",
      "votes": null
    },
    {
      "id": "1265675",
      "postDate": "04/07/2021 05:51:30",
      "content": "<p>Thanks for the reply.</p>\n<p>Yes, now I believe that my classifier is far from being perfect, anyway, it is just a roughly trained single model.</p>\n<p>I saw good val precision(0.97+ at threshold 0.5) while training this classifier so I just applied it to my model. However, I am still losing lots of true positives…I was expecting I can counter this loss by ensembling more segmentation models(so the segmentation part will be more precise), but I failed.</p>\n<p>I believe I need a more robust classifier.</p>",
      "rawMarkdown": "Thanks for the reply.\n\nYes, now I believe that my classifier is far from being perfect, anyway, it is just a roughly trained single model.\n\nI saw good val precision(0.97+ at threshold 0.5) while training this classifier so I just applied it to my model. However, I am still losing lots of true positives...I was expecting I can counter this loss by ensembling more segmentation models(so the segmentation part will be more precise), but I failed.\n\nI believe I need a more robust classifier.",
      "votes": null
    },
    {
      "id": "1265679",
      "postDate": "04/07/2021 06:00:23",
      "content": "<p>Thanks for the reply.</p>\n<p>I am still confused about how to do post processing to offset the loss. To make up for the loss, I ensembled more segmentation models with different backbones (effcientnet b4 and b5) in but it doesn't seem to be much useful. </p>\n<p>Anyway, now I see how risky it is to use a classifier…</p>",
      "rawMarkdown": "Thanks for the reply.\n\nI am still confused about how to do post processing to offset the loss. To make up for the loss, I ensembled more segmentation models with different backbones (effcientnet b4 and b5) in but it doesn't seem to be much useful. \n\nAnyway, now I see how risky it is to use a classifier...",
      "votes": null
    },
    {
      "id": "1270302",
      "postDate": "04/11/2021 13:59:59",
      "content": "<p>Thinking about this more, although the classifier may be harmful in leaving out some true positives outside, it may also filter some of the negatives, which might show up as false positives in the segmentation. <br>\nAdditionally as you hinted, you can play with threshold to decrease false negatives, so you can boost your 97+. <br>\nThis might bring in more negatives into false positive, but that is not too bad at all. Because you will still use your segmentation for the results, and it is always better to segment less tiles…<br>\n…At the end it is something worth expertimenting I think..</p>",
      "rawMarkdown": "Thinking about this more, although the classifier may be harmful in leaving out some true positives outside, it may also filter some of the negatives, which might show up as false positives in the segmentation. \nAdditionally as you hinted, you can play with threshold to decrease false negatives, so you can boost your 97+. \nThis might bring in more negatives into false positive, but that is not too bad at all. Because you will still use your segmentation for the results, and it is always better to segment less tiles...\n...At the end it is something worth expertimenting I think..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1262650,
      "author_name": "erikdali",
      "author_url": "",
      "post_date": "04/04/2021 15:00:41",
      "content": "<p>It’s hard to give advice without knowing much about your code or approach, pruning is always an option. But there’s a trade off between accuracy and speed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1263053,
          "author_name": "tlopesma",
          "author_url": "",
          "post_date": "04/05/2021 04:01:03",
          "content": "<p>Thanks for the reply.</p>\n<p>Yes, I do see the trade-off here. The classifier was based on efficientnet b5 and it converges real fast(compared with segmentation models), with good val precision(I used precision as monitor). It seems that this classifier is just getting by on the testing set, it is not a very reliable classifier, yet.</p>\n<p>Also, I am still wondering if the ratio of empty/non-empty masks will affect the performance of classifier(should we use/abandon classifier in balanced/unbalanced scenario). Still lots of things to try…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1263820,
          "author_name": "erikdali",
          "author_url": "",
          "post_date": "04/05/2021 17:30:06",
          "content": "<p>Those scores 0.919 and 0.914 are within the margin of error from each other. Maybe with more fine tuning the classifier could be more accurate. If you are taking a Bayesian approach there's not much you can do about the time/hardware restrictions of the competition.</p>\n<p>By \"empty/non-empty masks\" do you mean the anatomical mask or the glomeruli masks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264315,
          "author_name": "tlopesma",
          "author_url": "",
          "post_date": "04/06/2021 04:31:06",
          "content": "<p>I used Wojtek Rosa's public tfrecords <a href=\"url\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512</a> and <a href=\"url\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512-2</a> as my training/validation set(thanks him very much). I didn't build my own pipeline so I don't know exactly if the masks are anatomical masks or glomeruli masks(sorry).</p>\n<p>I remember there was a notebook showing how the tfrecords are created but I forgot where this notebook can be found. If my memory serves, the masks were decoded from train.csv, but I am not so sure of it…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1264273,
      "author_name": "southsakura",
      "author_url": "",
      "post_date": "04/06/2021 03:33:00",
      "content": "<p>I have to say sometimes it works but sometimes not. I have a shake of about 0.01 with different principles of classifier.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1264321,
          "author_name": "tlopesma",
          "author_url": "",
          "post_date": "04/06/2021 04:38:20",
          "content": "<p>Thanks for the reply. </p>\n<p>Then I think the classification-then-segmentation approach is still worth trying. A shake of 0.01 surely means something in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264327,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "04/06/2021 04:50:42",
          "content": "<p>Sure. At first I want to keep the original ratio of each samples but it seems to reflect much dffierent results with different qualities of curve worked comparing to using some classifiers. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1264514,
      "author_name": "abdulkadirguner",
      "author_url": "",
      "post_date": "04/06/2021 08:05:07",
      "content": "<p>The classifier is not perfect. So your segmentator will not see all true positives.<br>\nA clasifier makes sense if you can improve your segmentation models more efficiently than the ratio of the true positives you might lose with your classifier. A simple tradeoff that is really hard to improve at your stage beyond 915+, so good luck :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1264585,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "04/06/2021 09:24:20",
          "content": "<p>I agreed. The classifier can filter out some of the nearly meaningless samples and make the dataset more balance thus the model perform more stable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1265675,
          "author_name": "tlopesma",
          "author_url": "",
          "post_date": "04/07/2021 05:51:30",
          "content": "<p>Thanks for the reply.</p>\n<p>Yes, now I believe that my classifier is far from being perfect, anyway, it is just a roughly trained single model.</p>\n<p>I saw good val precision(0.97+ at threshold 0.5) while training this classifier so I just applied it to my model. However, I am still losing lots of true positives…I was expecting I can counter this loss by ensembling more segmentation models(so the segmentation part will be more precise), but I failed.</p>\n<p>I believe I need a more robust classifier.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1270302,
          "author_name": "abdulkadirguner",
          "author_url": "",
          "post_date": "04/11/2021 13:59:59",
          "content": "<p>Thinking about this more, although the classifier may be harmful in leaving out some true positives outside, it may also filter some of the negatives, which might show up as false positives in the segmentation. <br>\nAdditionally as you hinted, you can play with threshold to decrease false negatives, so you can boost your 97+. <br>\nThis might bring in more negatives into false positive, but that is not too bad at all. Because you will still use your segmentation for the results, and it is always better to segment less tiles…<br>\n…At the end it is something worth expertimenting I think..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1264862,
      "author_name": "andrasferenczi",
      "author_url": "",
      "post_date": "04/06/2021 13:26:06",
      "content": "<p>if the additional post processing you can do with the extra time at hand can offset the loss of true positives, then it may be worth it. It is a risky proposition, nevertheless, as you don't know what it will do on the private dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1265679,
          "author_name": "tlopesma",
          "author_url": "",
          "post_date": "04/07/2021 06:00:23",
          "content": "<p>Thanks for the reply.</p>\n<p>I am still confused about how to do post processing to offset the loss. To make up for the loss, I ensembled more segmentation models with different backbones (effcientnet b4 and b5) in but it doesn't seem to be much useful. </p>\n<p>Anyway, now I see how risky it is to use a classifier…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1262288": "Previously I only used folded efficient-unet models for inference, it shows 0.919 in the public LB and the submission can take as long as 8 hours.\n\nTo speed up the submission, I trained a binary classifier to skip some empty masks. Now the submission takes only less than 5 hours, even if I put more and more models in. However, the public LB drops from 0.919 to 0.914.\n\nI was hoping that a classifier can help me save some time and memory, so I can use more models to do inference, thus increase my public LB score. Now it seems that the result was not that good.\n\nI was wondering if this is common in this competition, or it is only me(in most cases, I code poorly).",
    "1262650": "It’s hard to give advice without knowing much about your code or approach, pruning is always an option. But there’s a trade off between accuracy and speed.",
    "1263053": "Thanks for the reply.\n\nYes, I do see the trade-off here. The classifier was based on efficientnet b5 and it converges real fast(compared with segmentation models), with good val precision(I used precision as monitor). It seems that this classifier is just getting by on the testing set, it is not a very reliable classifier, yet.\n\nAlso, I am still wondering if the ratio of empty/non-empty masks will affect the performance of classifier(should we use/abandon classifier in balanced/unbalanced scenario). Still lots of things to try...",
    "1263820": "Those scores 0.919 and 0.914 are within the margin of error from each other. Maybe with more fine tuning the classifier could be more accurate. If you are taking a Bayesian approach there's not much you can do about the time/hardware restrictions of the competition.\n\nBy \"empty/non-empty masks\" do you mean the anatomical mask or the glomeruli masks?",
    "1264273": "I have to say sometimes it works but sometimes not. I have a shake of about 0.01 with different principles of classifier.",
    "1264315": "I used Wojtek Rosa's public tfrecords [https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512](url) and [https://www.kaggle.com/wrrosa/hubmap-tfrecords-1024-512-2](url) as my training/validation set(thanks him very much). I didn't build my own pipeline so I don't know exactly if the masks are anatomical masks or glomeruli masks(sorry).\n\nI remember there was a notebook showing how the tfrecords are created but I forgot where this notebook can be found. If my memory serves, the masks were decoded from train.csv, but I am not so sure of it...",
    "1264321": "Thanks for the reply. \n\nThen I think the classification-then-segmentation approach is still worth trying. A shake of 0.01 surely means something in this competition.",
    "1264327": "Sure. At first I want to keep the original ratio of each samples but it seems to reflect much dffierent results with different qualities of curve worked comparing to using some classifiers.",
    "1264514": "The classifier is not perfect. So your segmentator will not see all true positives.\nA clasifier makes sense if you can improve your segmentation models more efficiently than the ratio of the true positives you might lose with your classifier. A simple tradeoff that is really hard to improve at your stage beyond 915+, so good luck :-)",
    "1264585": "I agreed. The classifier can filter out some of the nearly meaningless samples and make the dataset more balance thus the model perform more stable.",
    "1264862": "if the additional post processing you can do with the extra time at hand can offset the loss of true positives, then it may be worth it. It is a risky proposition, nevertheless, as you don't know what it will do on the private dataset.",
    "1265675": "Thanks for the reply.\n\nYes, now I believe that my classifier is far from being perfect, anyway, it is just a roughly trained single model.\n\nI saw good val precision(0.97+ at threshold 0.5) while training this classifier so I just applied it to my model. However, I am still losing lots of true positives...I was expecting I can counter this loss by ensembling more segmentation models(so the segmentation part will be more precise), but I failed.\n\nI believe I need a more robust classifier.",
    "1265679": "Thanks for the reply.\n\nI am still confused about how to do post processing to offset the loss. To make up for the loss, I ensembled more segmentation models with different backbones (effcientnet b4 and b5) in but it doesn't seem to be much useful. \n\nAnyway, now I see how risky it is to use a classifier...",
    "1270302": "Thinking about this more, although the classifier may be harmful in leaving out some true positives outside, it may also filter some of the negatives, which might show up as false positives in the segmentation. \nAdditionally as you hinted, you can play with threshold to decrease false negatives, so you can boost your 97+. \nThis might bring in more negatives into false positive, but that is not too bad at all. Because you will still use your segmentation for the results, and it is always better to segment less tiles...\n...At the end it is something worth expertimenting I think.."
  },
  "source": "meta"
}