{
  "id": 285351,
  "title": "CV LB-Correlation and Best Single Model",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/285351",
  "author_name": "Jaideep",
  "post_date": "2021-11-04T10:33:20.338000",
  "votes": 4,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I am not able to find the thread so far. Starting one now.<br>\nI am new entrant so will post once i have some thing available to.</p>\n<p>Please post your Best single model CV-LB score  and N  folds  </p>\n<p>Thanks..</p>",
  "messages": [
    {
      "id": 1579644,
      "postDate": "2021-11-12T04:58:19Z",
      "content": "<p>Mask R-CNN (torchvision). Labels are cort, shsy5y and astro.</p>\n<pre><code>Fold 1 - mAP: 0.263966 ({1: 0.3644, 2: 0.1553, 3: 0.1504})\nFold 2 - mAP: 0.245860 ({1: 0.3531, 2: 0.1250, 3: 0.1258})\nFold 3 - mAP: 0.262556 ({1: 0.3667, 2: 0.1523, 3: 0.1375})\nFold 4 - mAP: 0.259697 ({1: 0.3653, 2: 0.1409, 3: 0.1412})\nFold 5 - mAP: 0.242949 ({1: 0.3345, 2: 0.1492, 3: 0.1292})\n------------------------------\nOOF mAP: 0.25502 ({1: 0.3568, 2: 0.1445, 3: 0.1369})\n------------------------------\n</code></pre>\n<p>Public LB score of this model is 0.283. This is my first instance segmentation project and I noticed torchvision is very outdated. It only has Mask R-CNN with ResNet50 FPN. I'll update this post once I adapt frameworks like detectron2, mmdetection and etc.</p>",
      "rawMarkdown": "Mask R-CNN (torchvision). Labels are cort, shsy5y and astro.\n\n```\nFold 1 - mAP: 0.263966 ({1: 0.3644, 2: 0.1553, 3: 0.1504})\nFold 2 - mAP: 0.245860 ({1: 0.3531, 2: 0.1250, 3: 0.1258})\nFold 3 - mAP: 0.262556 ({1: 0.3667, 2: 0.1523, 3: 0.1375})\nFold 4 - mAP: 0.259697 ({1: 0.3653, 2: 0.1409, 3: 0.1412})\nFold 5 - mAP: 0.242949 ({1: 0.3345, 2: 0.1492, 3: 0.1292})\n------------------------------\nOOF mAP: 0.25502 ({1: 0.3568, 2: 0.1445, 3: 0.1369})\n------------------------------\n```\n\nPublic LB score of this model is 0.283. This is my first instance segmentation project and I noticed torchvision is very outdated. It only has Mask R-CNN with ResNet50 FPN. I'll update this post once I adapt frameworks like detectron2, mmdetection and etc.",
      "votes": 3,
      "replies": [
        {
          "id": 1580287,
          "postDate": "2021-11-12T16:09:02.113Z",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  thanks..<br>\none thing confuses me when i look masks plots, people are color coding the masks of same class. why is there a need. Looking at all image ids i see each image has got only one and one class. So it  is Multiclass Segmentation, Unlike Severstial steel defects where we got more than one defect each sample, but no pixel overlaps</p>",
          "rawMarkdown": "@gunesevitan  thanks..\none thing confuses me when i look masks plots, people are color coding the masks of same class. why is there a need. Looking at all image ids i see each image has got only one and one class. So it  is Multiclass Segmentation, Unlike Severstial steel defects where we got more than one defect each sample, but no pixel overlaps"
        },
        {
          "id": 1580296,
          "postDate": "2021-11-12T16:17:18.753Z",
          "content": "<p>You are right. Each image has only one unique cell type. I don't think there is a \"need\" of what you are saying. It's just a choice of visualization. The reason why they are doing that could be separating borders of overlapping instances.</p>",
          "rawMarkdown": "You are right. Each image has only one unique cell type. I don't think there is a \"need\" of what you are saying. It's just a choice of visualization. The reason why they are doing that could be separating borders of overlapping instances."
        },
        {
          "id": 1583319,
          "postDate": "2021-11-15T17:57:29.520Z",
          "content": "<p>what is the number of validations images for each cell type in each fold?<br>\nare you using the ratio follwing?<br>\nshsy5y                  52286<br>\ncort                     10777<br>\nastro                     10522</p>\n<pre><code>e.g. for fold-1 you report:\n  0.263966 = ?*0.3644 + ?*0.1553 + ?*0.1504\n</code></pre>",
          "rawMarkdown": "what is the number of validations images for each cell type in each fold?\nare you using the ratio follwing?\nshsy5y  \t\t\t\t52286\ncort     \t\t\t\t10777\nastro     \t\t\t\t10522\n```\ne.g. for fold-1 you report:\n  0.263966 = ?*0.3644 + ?*0.1553 + ?*0.1504\n```"
        },
        {
          "id": 1583325,
          "postDate": "2021-11-15T18:05:46.203Z",
          "content": "<p>They are the same folds from this topic.<br>\n<a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285546\" target=\"_blank\">https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285546</a></p>\n<pre><code>Fold 1 (122, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 27}\nFold 2 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 3 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 4 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 5 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\n</code></pre>",
          "rawMarkdown": "They are the same folds from this topic.\nhttps://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285546\n\n```\nFold 1 (122, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 27}\nFold 2 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 3 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 4 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 5 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\n```"
        }
      ]
    },
    {
      "id": 1570737,
      "postDate": "2021-11-04T10:33:20.340Z",
      "content": "<p>I am not able to find the thread so far. Starting one now.<br>\nI am new entrant so will post once i have some thing available to.</p>\n<p>Please post your Best single model CV-LB score  and N  folds  </p>\n<p>Thanks..</p>",
      "rawMarkdown": "I am not able to find the thread so far. Starting one now.\nI am new entrant so will post once i have some thing available to.\n\nPlease post your Best single model CV-LB score  and N  folds  \n\nThanks..\n",
      "votes": 4
    },
    {
      "id": 1575232,
      "postDate": "2021-11-08T09:01:26.633Z",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  the difficult part is that the metric is so sensitive to the way you decide which predictions to include and how to process them that I can have the same model scoring between 2.8 and 3.0 just depending on the postprocessing. It's meaningless to compare CV scores if we don't have a common way to score them (and that doesn't depend on hand picked thresholds)</p>\n<p>If you want to create a benchmark notebook for that I'll be happy to run my models through it and report back.</p>",
      "rawMarkdown": "@jaideepvalani  the difficult part is that the metric is so sensitive to the way you decide which predictions to include and how to process them that I can have the same model scoring between 2.8 and 3.0 just depending on the postprocessing. It's meaningless to compare CV scores if we don't have a common way to score them (and that doesn't depend on hand picked thresholds)\n\nIf you want to create a benchmark notebook for that I'll be happy to run my models through it and report back.",
      "votes": 2,
      "replies": [
        {
          "id": 1576099,
          "postDate": "2021-11-08T22:09:06.487Z",
          "content": "<p>i suggest report on the train mAP as well (in addition to cv mAP and public lb)<br>\nyou can also report ground-truth mAP, i.e this will show  \" perfect network output = target \" + \"post processing\"</p>\n<p>if you are using mask-rcnn, the ground truth mask is 28x28. This will remove some pixel when the details are too small. If you are using watershed/Unet, the component labelling may not get back to the ground truth if your seed/valley are not perfect enough</p>",
          "rawMarkdown": "i suggest report on the train mAP as well (in addition to cv mAP and public lb)\nyou can also report ground-truth mAP, i.e this will show  \" perfect network output = target \" + \"post processing\"\n\nif you are using mask-rcnn, the ground truth mask is 28x28. This will remove some pixel when the details are too small. If you are using watershed/Unet, the component labelling may not get back to the ground truth if your seed/valley are not perfect enough",
          "votes": 2
        },
        {
          "id": 1576363,
          "postDate": "2021-11-09T05:51:38.023Z",
          "content": "<p><a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> I actually tried to split the dataset into 5 folds using different criteria (so I have 3x5-fold splits), and train a baseline copied from <a href=\"https://www.kaggle.com/slawekbiel/positive-score-with-detectron-2-3-training\" target=\"_blank\">your notebook</a> on them. Here is what I found:</p>\n<ul>\n<li>With the same training parameters, <code>MaP IoU</code> on validation set on each fold can be different by up to <code>0.017</code>!!!</li>\n<li>With the same training setting, some training folds produces better LB results than others (0.004 difference, from <code>0.29</code> to <code>0.294</code>).</li>\n</ul>",
          "rawMarkdown": "@slawekbiel I actually tried to split the dataset into 5 folds using different criteria (so I have 3x5-fold splits), and train a baseline copied from [your notebook](https://www.kaggle.com/slawekbiel/positive-score-with-detectron-2-3-training) on them. Here is what I found:\n- With the same training parameters, `MaP IoU` on validation set on each fold can be different by up to `0.017`!!!\n- With the same training setting, some training folds produces better LB results than others (0.004 difference, from `0.29` to `0.294`).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1576308,
      "postDate": "2021-11-09T04:34:49.660Z",
      "content": "<p>Up until now it doesn't correlate so much.</p>",
      "rawMarkdown": "Up until now it doesn't correlate so much."
    },
    {
      "id": 1575056,
      "postDate": "2021-11-08T06:14:40.730Z",
      "content": "<p>It still quite early for the competition, but figure this information could be useful. </p>\n<p>Used 5 folds <code>StratifiedGroupKFold</code> by id.<br>\nMaskrcnn has good consistent CV/LB correlation; SOLOv2 has better CV but no correlation with LB.</p>\n<p>Best single model is R50-Maskrcnn (no postprocessing, not ensembled): 0.287/0.289 (CV/LB)</p>",
      "rawMarkdown": "It still quite early for the competition, but figure this information could be useful. \n\nUsed 5 folds `StratifiedGroupKFold` by id.\nMaskrcnn has good consistent CV/LB correlation; SOLOv2 has better CV but no correlation with LB.\n\nBest single model is R50-Maskrcnn (no postprocessing, not ensembled): 0.287/0.289 (CV/LB)"
    },
    {
      "id": 1583737,
      "postDate": "2021-11-16T03:13:37.873Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1579644,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2021-11-12T04:58:19",
      "content": "<p>Mask R-CNN (torchvision). Labels are cort, shsy5y and astro.</p>\n<pre><code>Fold 1 - mAP: 0.263966 ({1: 0.3644, 2: 0.1553, 3: 0.1504})\nFold 2 - mAP: 0.245860 ({1: 0.3531, 2: 0.1250, 3: 0.1258})\nFold 3 - mAP: 0.262556 ({1: 0.3667, 2: 0.1523, 3: 0.1375})\nFold 4 - mAP: 0.259697 ({1: 0.3653, 2: 0.1409, 3: 0.1412})\nFold 5 - mAP: 0.242949 ({1: 0.3345, 2: 0.1492, 3: 0.1292})\n------------------------------\nOOF mAP: 0.25502 ({1: 0.3568, 2: 0.1445, 3: 0.1369})\n------------------------------\n</code></pre>\n<p>Public LB score of this model is 0.283. This is my first instance segmentation project and I noticed torchvision is very outdated. It only has Mask R-CNN with ResNet50 FPN. I'll update this post once I adapt frameworks like detectron2, mmdetection and etc.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1580287,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-11-12T16:09:02.113000",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  thanks..<br>\none thing confuses me when i look masks plots, people are color coding the masks of same class. why is there a need. Looking at all image ids i see each image has got only one and one class. So it  is Multiclass Segmentation, Unlike Severstial steel defects where we got more than one defect each sample, but no pixel overlaps</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1580296,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2021-11-12T16:17:18.753000",
          "content": "<p>You are right. Each image has only one unique cell type. I don't think there is a \"need\" of what you are saying. It's just a choice of visualization. The reason why they are doing that could be separating borders of overlapping instances.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1583319,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-11-15T17:57:29.520000",
          "content": "<p>what is the number of validations images for each cell type in each fold?<br>\nare you using the ratio follwing?<br>\nshsy5y                  52286<br>\ncort                     10777<br>\nastro                     10522</p>\n<pre><code>e.g. for fold-1 you report:\n  0.263966 = ?*0.3644 + ?*0.1553 + ?*0.1504\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1583325,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2021-11-15T18:05:46.203000",
          "content": "<p>They are the same folds from this topic.<br>\n<a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285546\" target=\"_blank\">https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285546</a></p>\n<pre><code>Fold 1 (122, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 27}\nFold 2 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 3 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 4 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 5 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1575232,
      "author_name": "Slawek Biel",
      "author_url": "",
      "post_date": "2021-11-08T09:01:26.633000",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  the difficult part is that the metric is so sensitive to the way you decide which predictions to include and how to process them that I can have the same model scoring between 2.8 and 3.0 just depending on the postprocessing. It's meaningless to compare CV scores if we don't have a common way to score them (and that doesn't depend on hand picked thresholds)</p>\n<p>If you want to create a benchmark notebook for that I'll be happy to run my models through it and report back.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1576099,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-11-08T22:09:06.487000",
          "content": "<p>i suggest report on the train mAP as well (in addition to cv mAP and public lb)<br>\nyou can also report ground-truth mAP, i.e this will show  \" perfect network output = target \" + \"post processing\"</p>\n<p>if you are using mask-rcnn, the ground truth mask is 28x28. This will remove some pixel when the details are too small. If you are using watershed/Unet, the component labelling may not get back to the ground truth if your seed/valley are not perfect enough</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1576363,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2021-11-09T05:51:38.023000",
          "content": "<p><a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> I actually tried to split the dataset into 5 folds using different criteria (so I have 3x5-fold splits), and train a baseline copied from <a href=\"https://www.kaggle.com/slawekbiel/positive-score-with-detectron-2-3-training\" target=\"_blank\">your notebook</a> on them. Here is what I found:</p>\n<ul>\n<li>With the same training parameters, <code>MaP IoU</code> on validation set on each fold can be different by up to <code>0.017</code>!!!</li>\n<li>With the same training setting, some training folds produces better LB results than others (0.004 difference, from <code>0.29</code> to <code>0.294</code>).</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1576308,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-11-09T04:34:49.660000",
      "content": "<p>Up until now it doesn't correlate so much.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1575056,
      "author_name": "JunYong Tong",
      "author_url": "",
      "post_date": "2021-11-08T06:14:40.730000",
      "content": "<p>It still quite early for the competition, but figure this information could be useful. </p>\n<p>Used 5 folds <code>StratifiedGroupKFold</code> by id.<br>\nMaskrcnn has good consistent CV/LB correlation; SOLOv2 has better CV but no correlation with LB.</p>\n<p>Best single model is R50-Maskrcnn (no postprocessing, not ensembled): 0.287/0.289 (CV/LB)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1583737,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-16T03:13:37.873000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1579644": "Mask R-CNN (torchvision). Labels are cort, shsy5y and astro.\n\n```\nFold 1 - mAP: 0.263966 ({1: 0.3644, 2: 0.1553, 3: 0.1504})\nFold 2 - mAP: 0.245860 ({1: 0.3531, 2: 0.1250, 3: 0.1258})\nFold 3 - mAP: 0.262556 ({1: 0.3667, 2: 0.1523, 3: 0.1375})\nFold 4 - mAP: 0.259697 ({1: 0.3653, 2: 0.1409, 3: 0.1412})\nFold 5 - mAP: 0.242949 ({1: 0.3345, 2: 0.1492, 3: 0.1292})\n------------------------------\nOOF mAP: 0.25502 ({1: 0.3568, 2: 0.1445, 3: 0.1369})\n------------------------------\n```\n\nPublic LB score of this model is 0.283. This is my first instance segmentation project and I noticed torchvision is very outdated. It only has Mask R-CNN with ResNet50 FPN. I'll update this post once I adapt frameworks like detectron2, mmdetection and etc.",
    "1570737": "I am not able to find the thread so far. Starting one now.\nI am new entrant so will post once i have some thing available to.\n\nPlease post your Best single model CV-LB score  and N  folds  \n\nThanks..\n",
    "1575232": "@jaideepvalani  the difficult part is that the metric is so sensitive to the way you decide which predictions to include and how to process them that I can have the same model scoring between 2.8 and 3.0 just depending on the postprocessing. It's meaningless to compare CV scores if we don't have a common way to score them (and that doesn't depend on hand picked thresholds)\n\nIf you want to create a benchmark notebook for that I'll be happy to run my models through it and report back.",
    "1576308": "Up until now it doesn't correlate so much.",
    "1575056": "It still quite early for the competition, but figure this information could be useful. \n\nUsed 5 folds `StratifiedGroupKFold` by id.\nMaskrcnn has good consistent CV/LB correlation; SOLOv2 has better CV but no correlation with LB.\n\nBest single model is R50-Maskrcnn (no postprocessing, not ensembled): 0.287/0.289 (CV/LB)",
    "1583737": ""
  }
}