{
  "id": 284215,
  "title": "Threshold by class",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/284215",
  "author_name": "Slawek Biel",
  "post_date": "2021-10-30T12:25:17.955000",
  "votes": 85,
  "comment_count": 19,
  "views": 0,
  "content": "<p>When we run inference with a model like Mask R-CNN, as the output we get the masks predictions and also a score associated with each object, representing the confidence of our model in that object. Now we need to decide which of the predictions to include in the submission. We add too few - our false negative count will be high, we add too many low scoring predictions and our false positives count rises. This leads to a unimodal function of score vs threshold as plotted bellow. From this we can select the threshold to use in a submission (either based on the validation set, or probing the public test set)<br>\n<img src=\"https://raw.githubusercontent.com/slawekslex/random/main/threshold_total.png\" alt=\"\"></p>\n<p>However if we zoom closer and break samples by different cell types we can see they have quite different characteristics. Looking at the chart bellow we can see that their maximas are quite far from one another and it would make more sense to set the thresholds separately.</p>\n<p><img src=\"https://raw.githubusercontent.com/slawekslex/random/main/Screenshot%20from%202021-10-30%2013-51-11.png\" alt=\"\"></p>\n<p>That's exactly what I tried in my <a href=\"https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference\" target=\"_blank\">Inference notebook</a> and got a small boost in score with the same model and little effort.</p>\n<p><em>PS: I would appreciate if you respect my work and don't create public forks that do nothing else but tweak these numbers for a slightly better score</em></p>",
  "messages": [
    {
      "id": 1565374,
      "postDate": "2021-10-30T12:25:17.957Z",
      "content": "<p>When we run inference with a model like Mask R-CNN, as the output we get the masks predictions and also a score associated with each object, representing the confidence of our model in that object. Now we need to decide which of the predictions to include in the submission. We add too few - our false negative count will be high, we add too many low scoring predictions and our false positives count rises. This leads to a unimodal function of score vs threshold as plotted bellow. From this we can select the threshold to use in a submission (either based on the validation set, or probing the public test set)<br>\n<img src=\"https://raw.githubusercontent.com/slawekslex/random/main/threshold_total.png\" alt=\"\"></p>\n<p>However if we zoom closer and break samples by different cell types we can see they have quite different characteristics. Looking at the chart bellow we can see that their maximas are quite far from one another and it would make more sense to set the thresholds separately.</p>\n<p><img src=\"https://raw.githubusercontent.com/slawekslex/random/main/Screenshot%20from%202021-10-30%2013-51-11.png\" alt=\"\"></p>\n<p>That's exactly what I tried in my <a href=\"https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference\" target=\"_blank\">Inference notebook</a> and got a small boost in score with the same model and little effort.</p>\n<p><em>PS: I would appreciate if you respect my work and don't create public forks that do nothing else but tweak these numbers for a slightly better score</em></p>",
      "rawMarkdown": "When we run inference with a model like Mask R-CNN, as the output we get the masks predictions and also a score associated with each object, representing the confidence of our model in that object. Now we need to decide which of the predictions to include in the submission. We add too few - our false negative count will be high, we add too many low scoring predictions and our false positives count rises. This leads to a unimodal function of score vs threshold as plotted bellow. From this we can select the threshold to use in a submission (either based on the validation set, or probing the public test set)\n![](https://raw.githubusercontent.com/slawekslex/random/main/threshold_total.png)\n\nHowever if we zoom closer and break samples by different cell types we can see they have quite different characteristics. Looking at the chart bellow we can see that their maximas are quite far from one another and it would make more sense to set the thresholds separately.\n\n![](https://raw.githubusercontent.com/slawekslex/random/main/Screenshot%20from%202021-10-30%2013-51-11.png)\n\nThat's exactly what I tried in my [Inference notebook](https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference) and got a small boost in score with the same model and little effort.\n\n*PS: I would appreciate if you respect my work and don't create public forks that do nothing else but tweak these numbers for a slightly better score*",
      "votes": 85
    },
    {
      "id": 1565492,
      "postDate": "2021-10-30T15:03:15.290Z",
      "content": "<p>It makes perfect sense to use lower thresholds on shsy5y and astro because shsy5y cell lines have lots of object and astro cell lines have very large objects. Their masks contain larger sections compared to cort masks. One problem is the test set doesn't have cell_type column, so how did you manage to separate different cell types? Did you train a classifier?</p>",
      "rawMarkdown": "It makes perfect sense to use lower thresholds on shsy5y and astro because shsy5y cell lines have lots of object and astro cell lines have very large objects. Their masks contain larger sections compared to cort masks. One problem is the test set doesn't have cell_type column, so how did you manage to separate different cell types? Did you train a classifier?",
      "votes": 5,
      "replies": [
        {
          "id": 1565500,
          "postDate": "2021-10-30T15:19:46.503Z",
          "content": "<p>The Mask R-CNN model I use outputs class predictions too.</p>",
          "rawMarkdown": "The Mask R-CNN model I use outputs class predictions too.",
          "votes": 4
        },
        {
          "id": 1565573,
          "postDate": "2021-10-30T17:18:08.417Z",
          "content": "<p>Do you mean you trained mask r-cnn with 3 labels?</p>",
          "rawMarkdown": "Do you mean you trained mask r-cnn with 3 labels?"
        },
        {
          "id": 1565706,
          "postDate": "2021-10-30T21:58:51.677Z",
          "content": "<blockquote>\n  <p>Do you mean you trained mask r-cnn with 3 labels?</p>\n</blockquote>\n<p>Yes</p>",
          "rawMarkdown": "> Do you mean you trained mask r-cnn with 3 labels?\n\nYes",
          "votes": 1
        }
      ]
    },
    {
      "id": 1570853,
      "postDate": "2021-11-04T12:35:20.317Z",
      "content": "<p>Hi there Slawek, just wanted to say that this is a really cool insight into our data! Best of luck moving forward</p>",
      "rawMarkdown": "Hi there Slawek, just wanted to say that this is a really cool insight into our data! Best of luck moving forward",
      "votes": 6
    },
    {
      "id": 1565943,
      "postDate": "2021-10-31T07:30:34.497Z",
      "content": "<p>Good work. Highly Appreciated. </p>",
      "rawMarkdown": "Good work. Highly Appreciated. ",
      "votes": 2
    },
    {
      "id": 1565523,
      "postDate": "2021-10-30T15:54:24.100Z",
      "content": "<p>Appreciate you running tests on different classes and showing the results. Makes total sense, given that there are clear size and shape characteristics associated with different types. I believe in one of his notebooks, <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> also mentioned having benefit to classifying first and using that information to prune / limit the instances. I wouldn't doubt that the top scores are using separate models for each type of cell. Thanks for sharing.</p>",
      "rawMarkdown": "Appreciate you running tests on different classes and showing the results. Makes total sense, given that there are clear size and shape characteristics associated with different types. I believe in one of his notebooks, @julian3833 also mentioned having benefit to classifying first and using that information to prune / limit the instances. I wouldn't doubt that the top scores are using separate models for each type of cell. Thanks for sharing.",
      "votes": 2,
      "replies": [
        {
          "id": 1565711,
          "postDate": "2021-10-30T22:01:16Z",
          "content": "<p>Hi! Thanks for mentioning my work.</p>\n<p>Yes I did play around with a classifier determining the <code>cell_type</code> of an image. I could obtain an improvement of <code>0.007</code> with it. The improvement happened refining masks mostly (I tried putting boundaries to the number of predictions for each <code>cell_type</code>, but I couldn't get it to work better than a simple threshold over the confidence score). </p>\n<p>I limited the masks sizes for each cell_type based on the train dataset statistics.</p>\n<p>The classifier's validation accuracy is 100%. I preferred a separate classifier to assess the <code>cell_type</code> of the full image instead of the type of each instance (that's what Mask R-CNN does I think). My intuition is that it's a simpler task and therefore a good performance is easier to achieve.</p>\n<p>I tried using 3 different Mask R-CNN based on the classifier output, but with no good results till now.</p>\n<p>If you want to give it a shot with a separated classifier, it would be a good experiment.</p>\n<p>In case anyone wants to try:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/julian3833/sartorius-classifier-mask-r-cnn-lb-0-28\" target=\"_blank\">🦠 Sartorius - Classifier + Mask R-CNN [LB=0.28]</a> (check here the function <code>refine_mask</code> and the definition of <code>df_pixels</code></li>\n<li><a href=\"https://www.kaggle.com/julian3833/sartorius-resnet34-classifier\" target=\"_blank\">🦠 Sartorius - Resnet34 Classifier</a></li>\n<li><a href=\"https://www.kaggle.com/julian3833/sartorius-resnet-34-classifier-finetuned\" target=\"_blank\">Resnet34 Classifier weights as a dataset</a></li>\n</ul>",
          "rawMarkdown": "Hi! Thanks for mentioning my work.\n\nYes I did play around with a classifier determining the `cell_type` of an image. I could obtain an improvement of `0.007` with it. The improvement happened refining masks mostly (I tried putting boundaries to the number of predictions for each `cell_type`, but I couldn't get it to work better than a simple threshold over the confidence score). \n\nI limited the masks sizes for each cell_type based on the train dataset statistics.\n\n\nThe classifier's validation accuracy is 100%. I preferred a separate classifier to assess the `cell_type` of the full image instead of the type of each instance (that's what Mask R-CNN does I think). My intuition is that it's a simpler task and therefore a good performance is easier to achieve.\n\nI tried using 3 different Mask R-CNN based on the classifier output, but with no good results till now.\n\nIf you want to give it a shot with a separated classifier, it would be a good experiment.\n\nIn case anyone wants to try:\n* [🦠 Sartorius - Classifier + Mask R-CNN [LB=0.28]](https://www.kaggle.com/julian3833/sartorius-classifier-mask-r-cnn-lb-0-28) (check here the function `refine_mask` and the definition of `df_pixels`\n* [🦠 Sartorius - Resnet34 Classifier](https://www.kaggle.com/julian3833/sartorius-resnet34-classifier)\n* [Resnet34 Classifier weights as a dataset](https://www.kaggle.com/julian3833/sartorius-resnet-34-classifier-finetuned)",
          "votes": 2
        },
        {
          "id": 1565714,
          "postDate": "2021-10-30T22:06:29.070Z",
          "content": "<p>FYI I’m getting 100% validation accuracy too by taking the most common class predicted by Mask RCNN. It’s seems that the cells are different enough that it’s a simple task regardless how you do it.</p>",
          "rawMarkdown": "FYI I’m getting 100% validation accuracy too by taking the most common class predicted by Mask RCNN. It’s seems that the cells are different enough that it’s a simple task regardless how you do it.",
          "votes": 2
        },
        {
          "id": 1566893,
          "postDate": "2021-11-01T11:04:55.680Z",
          "content": "<p>I switch to resnet50 classifer and mask-rcnn using your pipe line, got LB 0.281</p>",
          "rawMarkdown": "I switch to resnet50 classifer and mask-rcnn using your pipe line, got LB 0.281",
          "votes": 2
        },
        {
          "id": 1567178,
          "postDate": "2021-11-01T16:36:39.613Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> . So I did decide to experiment on it. First I averaged the errors based on cell type and saw that \"astro\" type cell had almost double the cross entropy loss as the rest. As you mentioned, classifying based on a simple conv net was straight forward with 100% val accuracy. I made a separate model just for astro type cell, and if the classifier detected astro, used the different model. The score went up from 0.175 to 0.190. I know the scores is low (almost half based on Mask-RCNN), but separating showed an improvement. I'm thinking that the low score could be due to limitations of my model architecture.</p>",
          "rawMarkdown": "Hi @julian3833 . So I did decide to experiment on it. First I averaged the errors based on cell type and saw that \"astro\" type cell had almost double the cross entropy loss as the rest. As you mentioned, classifying based on a simple conv net was straight forward with 100% val accuracy. I made a separate model just for astro type cell, and if the classifier detected astro, used the different model. The score went up from 0.175 to 0.190. I know the scores is low (almost half based on Mask-RCNN), but separating showed an improvement. I'm thinking that the low score could be due to limitations of my model architecture.",
          "votes": 4
        },
        {
          "id": 1567561,
          "postDate": "2021-11-02T02:23:09.177Z",
          "content": "<p>That's a very interesting finding <a href=\"https://www.kaggle.com/dunedinb\" target=\"_blank\">@dunedinb</a>. A <code>0.015</code> increase is a lot for the current LB scores.</p>",
          "rawMarkdown": "That's a very interesting finding @dunedinb. A `0.015` increase is a lot for the current LB scores.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1623730,
      "postDate": "2021-12-20T07:13:15.470Z",
      "content": "<p><a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> Thanks for the work, I am just wondering if x-axis is thresholds range from [0-1] and y-axis is MaP IoU, MaP IoU should be close to 1 at threshold 0 and decreasing monolithic function? but looking at your plot at 0 in x-axis it's around 0.12. It seems I misunderstood something..</p>",
      "rawMarkdown": "@slawekbiel Thanks for the work, I am just wondering if x-axis is thresholds range from [0-1] and y-axis is MaP IoU, MaP IoU should be close to 1 at threshold 0 and decreasing monolithic function? but looking at your plot at 0 in x-axis it's around 0.12. It seems I misunderstood something..",
      "replies": [
        {
          "id": 1623760,
          "postDate": "2021-12-20T08:06:28.167Z",
          "content": "<p>The confusion might be due to different definitions of average precision. I’m using the one from this competition evaluation metric TP/(TP+FP+FN) if it was just TP/(TP+FP) then it would have started close to 1 - but that of course wouldn’t make sense from evaluation standpoint as everyone would just submit a single mask.</p>",
          "rawMarkdown": "The confusion might be due to different definitions of average precision. I’m using the one from this competition evaluation metric TP/(TP+FP+FN) if it was just TP/(TP+FP) then it would have started close to 1 - but that of course wouldn’t make sense from evaluation standpoint as everyone would just submit a single mask.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1589829,
      "postDate": "2021-11-20T16:08:15.070Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1589836,
          "postDate": "2021-11-20T16:20:56.243Z",
          "content": "<p>It has no effect on training. Just on calculating validation score or generating submission.</p>",
          "rawMarkdown": "It has no effect on training. Just on calculating validation score or generating submission."
        },
        {
          "id": 1589843,
          "postDate": "2021-11-20T16:29:24.047Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1589846,
          "postDate": "2021-11-20T16:31:17.143Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1617271,
      "postDate": "2021-12-14T02:14:34.813Z",
      "content": "<p>It's a great work. Thank you!</p>",
      "rawMarkdown": "It's a great work. Thank you!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1565492,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2021-10-30T15:03:15.290000",
      "content": "<p>It makes perfect sense to use lower thresholds on shsy5y and astro because shsy5y cell lines have lots of object and astro cell lines have very large objects. Their masks contain larger sections compared to cort masks. One problem is the test set doesn't have cell_type column, so how did you manage to separate different cell types? Did you train a classifier?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1565500,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-10-30T15:19:46.503000",
          "content": "<p>The Mask R-CNN model I use outputs class predictions too.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1565573,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2021-10-30T17:18:08.417000",
          "content": "<p>Do you mean you trained mask r-cnn with 3 labels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1565706,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-10-30T21:58:51.677000",
          "content": "<blockquote>\n  <p>Do you mean you trained mask r-cnn with 3 labels?</p>\n</blockquote>\n<p>Yes</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1570853,
      "author_name": "CorporateResearchSartorius",
      "author_url": "",
      "post_date": "2021-11-04T12:35:20.317000",
      "content": "<p>Hi there Slawek, just wanted to say that this is a really cool insight into our data! Best of luck moving forward</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1565943,
      "author_name": "Muhammad Maaz",
      "author_url": "",
      "post_date": "2021-10-31T07:30:34.497000",
      "content": "<p>Good work. Highly Appreciated. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1565523,
      "author_name": "Michael Bolton",
      "author_url": "",
      "post_date": "2021-10-30T15:54:24.100000",
      "content": "<p>Appreciate you running tests on different classes and showing the results. Makes total sense, given that there are clear size and shape characteristics associated with different types. I believe in one of his notebooks, <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> also mentioned having benefit to classifying first and using that information to prune / limit the instances. I wouldn't doubt that the top scores are using separate models for each type of cell. Thanks for sharing.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1565711,
          "author_name": "Julián Peller (dataista0)",
          "author_url": "",
          "post_date": "2021-10-30T22:01:16",
          "content": "<p>Hi! Thanks for mentioning my work.</p>\n<p>Yes I did play around with a classifier determining the <code>cell_type</code> of an image. I could obtain an improvement of <code>0.007</code> with it. The improvement happened refining masks mostly (I tried putting boundaries to the number of predictions for each <code>cell_type</code>, but I couldn't get it to work better than a simple threshold over the confidence score). </p>\n<p>I limited the masks sizes for each cell_type based on the train dataset statistics.</p>\n<p>The classifier's validation accuracy is 100%. I preferred a separate classifier to assess the <code>cell_type</code> of the full image instead of the type of each instance (that's what Mask R-CNN does I think). My intuition is that it's a simpler task and therefore a good performance is easier to achieve.</p>\n<p>I tried using 3 different Mask R-CNN based on the classifier output, but with no good results till now.</p>\n<p>If you want to give it a shot with a separated classifier, it would be a good experiment.</p>\n<p>In case anyone wants to try:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/julian3833/sartorius-classifier-mask-r-cnn-lb-0-28\" target=\"_blank\">🦠 Sartorius - Classifier + Mask R-CNN [LB=0.28]</a> (check here the function <code>refine_mask</code> and the definition of <code>df_pixels</code></li>\n<li><a href=\"https://www.kaggle.com/julian3833/sartorius-resnet34-classifier\" target=\"_blank\">🦠 Sartorius - Resnet34 Classifier</a></li>\n<li><a href=\"https://www.kaggle.com/julian3833/sartorius-resnet-34-classifier-finetuned\" target=\"_blank\">Resnet34 Classifier weights as a dataset</a></li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1565714,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-10-30T22:06:29.070000",
          "content": "<p>FYI I’m getting 100% validation accuracy too by taking the most common class predicted by Mask RCNN. It’s seems that the cells are different enough that it’s a simple task regardless how you do it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1566893,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-11-01T11:04:55.680000",
          "content": "<p>I switch to resnet50 classifer and mask-rcnn using your pipe line, got LB 0.281</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1567178,
          "author_name": "Michael Bolton",
          "author_url": "",
          "post_date": "2021-11-01T16:36:39.613000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> . So I did decide to experiment on it. First I averaged the errors based on cell type and saw that \"astro\" type cell had almost double the cross entropy loss as the rest. As you mentioned, classifying based on a simple conv net was straight forward with 100% val accuracy. I made a separate model just for astro type cell, and if the classifier detected astro, used the different model. The score went up from 0.175 to 0.190. I know the scores is low (almost half based on Mask-RCNN), but separating showed an improvement. I'm thinking that the low score could be due to limitations of my model architecture.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1567561,
          "author_name": "Julián Peller (dataista0)",
          "author_url": "",
          "post_date": "2021-11-02T02:23:09.177000",
          "content": "<p>That's a very interesting finding <a href=\"https://www.kaggle.com/dunedinb\" target=\"_blank\">@dunedinb</a>. A <code>0.015</code> increase is a lot for the current LB scores.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1623730,
      "author_name": "enddl22",
      "author_url": "",
      "post_date": "2021-12-20T07:13:15.470000",
      "content": "<p><a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> Thanks for the work, I am just wondering if x-axis is thresholds range from [0-1] and y-axis is MaP IoU, MaP IoU should be close to 1 at threshold 0 and decreasing monolithic function? but looking at your plot at 0 in x-axis it's around 0.12. It seems I misunderstood something..</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1623760,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-12-20T08:06:28.167000",
          "content": "<p>The confusion might be due to different definitions of average precision. I’m using the one from this competition evaluation metric TP/(TP+FP+FN) if it was just TP/(TP+FP) then it would have started close to 1 - but that of course wouldn’t make sense from evaluation standpoint as everyone would just submit a single mask.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1589829,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-20T16:08:15.070000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1589836,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-11-20T16:20:56.243000",
          "content": "<p>It has no effect on training. Just on calculating validation score or generating submission.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1589843,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-11-20T16:29:24.047000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1589846,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-11-20T16:31:17.143000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1617271,
      "author_name": "Wonjun Park",
      "author_url": "",
      "post_date": "2021-12-14T02:14:34.813000",
      "content": "<p>It's a great work. Thank you!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1565374": "When we run inference with a model like Mask R-CNN, as the output we get the masks predictions and also a score associated with each object, representing the confidence of our model in that object. Now we need to decide which of the predictions to include in the submission. We add too few - our false negative count will be high, we add too many low scoring predictions and our false positives count rises. This leads to a unimodal function of score vs threshold as plotted bellow. From this we can select the threshold to use in a submission (either based on the validation set, or probing the public test set)\n![](https://raw.githubusercontent.com/slawekslex/random/main/threshold_total.png)\n\nHowever if we zoom closer and break samples by different cell types we can see they have quite different characteristics. Looking at the chart bellow we can see that their maximas are quite far from one another and it would make more sense to set the thresholds separately.\n\n![](https://raw.githubusercontent.com/slawekslex/random/main/Screenshot%20from%202021-10-30%2013-51-11.png)\n\nThat's exactly what I tried in my [Inference notebook](https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference) and got a small boost in score with the same model and little effort.\n\n*PS: I would appreciate if you respect my work and don't create public forks that do nothing else but tweak these numbers for a slightly better score*",
    "1565492": "It makes perfect sense to use lower thresholds on shsy5y and astro because shsy5y cell lines have lots of object and astro cell lines have very large objects. Their masks contain larger sections compared to cort masks. One problem is the test set doesn't have cell_type column, so how did you manage to separate different cell types? Did you train a classifier?",
    "1570853": "Hi there Slawek, just wanted to say that this is a really cool insight into our data! Best of luck moving forward",
    "1565943": "Good work. Highly Appreciated. ",
    "1565523": "Appreciate you running tests on different classes and showing the results. Makes total sense, given that there are clear size and shape characteristics associated with different types. I believe in one of his notebooks, @julian3833 also mentioned having benefit to classifying first and using that information to prune / limit the instances. I wouldn't doubt that the top scores are using separate models for each type of cell. Thanks for sharing.",
    "1623730": "@slawekbiel Thanks for the work, I am just wondering if x-axis is thresholds range from [0-1] and y-axis is MaP IoU, MaP IoU should be close to 1 at threshold 0 and decreasing monolithic function? but looking at your plot at 0 in x-axis it's around 0.12. It seems I misunderstood something..",
    "1589829": "",
    "1617271": "It's a great work. Thank you!"
  }
}