{
  "id": 285504,
  "title": "Validation strategy",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/285504",
  "author_name": "",
  "post_date": "2021-11-05T00:03:50.853690Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The first thing I usually try to do in a new competition is getting a validation scheme working, and this one is no exception - as in this kernel:</p>\n<p><a href=\"https://www.kaggle.com/konradb/validation-score\" target=\"_blank\">https://www.kaggle.com/konradb/validation-score</a></p>\n<p>The weird thing is, I took a relatively simple setup (copied from the excellent trilogy by <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> <a href=\"https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference\" target=\"_blank\">https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference</a>) and built around it:</p>\n<ol>\n<li>I created a five-fold split of the data (<code>GroupKFold</code> on id)</li>\n<li>fitted a model on folds 1,2,3,4</li>\n<li>used it for evaluation on fold 0 data (the kernel above).</li>\n</ol>\n<p>The part I'm struggling with is that during training my last MaP IoU was 0.2554, the LB score is 0.24 and the calculated metric is 0.056. Either there is something really strange in the data, or there is an embarrassing hole in my logic - whichever it is, I would love some feedback.</p>",
  "messages": [
    {
      "id": "1571515",
      "postDate": "11/05/2021 00:03:50",
      "content": "<p>The first thing I usually try to do in a new competition is getting a validation scheme working, and this one is no exception - as in this kernel:</p>\n<p><a href=\"https://www.kaggle.com/konradb/validation-score\" target=\"_blank\">https://www.kaggle.com/konradb/validation-score</a></p>\n<p>The weird thing is, I took a relatively simple setup (copied from the excellent trilogy by <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> <a href=\"https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference\" target=\"_blank\">https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference</a>) and built around it:</p>\n<ol>\n<li>I created a five-fold split of the data (<code>GroupKFold</code> on id)</li>\n<li>fitted a model on folds 1,2,3,4</li>\n<li>used it for evaluation on fold 0 data (the kernel above).</li>\n</ol>\n<p>The part I'm struggling with is that during training my last MaP IoU was 0.2554, the LB score is 0.24 and the calculated metric is 0.056. Either there is something really strange in the data, or there is an embarrassing hole in my logic - whichever it is, I would love some feedback.</p>",
      "rawMarkdown": "The first thing I usually try to do in a new competition is getting a validation scheme working, and this one is no exception - as in this kernel:\n\nhttps://www.kaggle.com/konradb/validation-score\n\nThe weird thing is, I took a relatively simple setup (copied from the excellent trilogy by @slawekbiel https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference) and built around it:\n1. I created a five-fold split of the data (`GroupKFold` on id)\n2. fitted a model on folds 1,2,3,4\n3. used it for evaluation on fold 0 data (the kernel above).\n\nThe part I'm struggling with is that during training my last MaP IoU was 0.2554, the LB score is 0.24 and the calculated metric is 0.056. Either there is something really strange in the data, or there is an embarrassing hole in my logic - whichever it is, I would love some feedback.",
      "votes": null
    },
    {
      "id": "1571778",
      "postDate": "11/05/2021 06:43:35",
      "content": "<p>I have the same problem but I haven't solved it yet. My best bet is on we are calculating error on background too.</p>",
      "rawMarkdown": "I have the same problem but I haven't solved it yet. My best bet is on we are calculating error on background too.",
      "votes": null
    },
    {
      "id": "1571984",
      "postDate": "11/05/2021 10:26:17",
      "content": "<p>It looks to me that the problem is here:<br>\n<code>masks_valid_agg = masks_valid_agg.groupby('id').sum().reset_index()</code><br>\nSumming all the masks together looses information about individual instances. The <code>iou_map</code> input states</p>\n<blockquote>\n  <p>Masks contain the segmented pixels where each object has one value associated</p>\n</blockquote>\n<p>Wouldn't your sum just give 0 for background and 1 for every object (since overlaps are removed).<br>\nI haven't actually run it, so sorry If I'm misunderstanding what you do.</p>",
      "rawMarkdown": "It looks to me that the problem is here:\n```masks_valid_agg = masks_valid_agg.groupby('id').sum().reset_index()```\nSumming all the masks together looses information about individual instances. The `iou_map` input states\n> Masks contain the segmented pixels where each object has one value associated\n\nWouldn't your sum just give 0 for background and 1 for every object (since overlaps are removed).\nI haven't actually run it, so sorry If I'm misunderstanding what you do.",
      "votes": null
    },
    {
      "id": "1572009",
      "postDate": "11/05/2021 10:57:12",
      "content": "<p>It will do just that, yes - and that might indeed be a problem. It's just that histogram2d needs the samples to be the same length, and aggregating per id seemed like the best idea.</p>\n<p>I may have mentioned this before, but thanks for the fantastic work you shared - I learned a lot.</p>",
      "rawMarkdown": "It will do just that, yes - and that might indeed be a problem. It's just that histogram2d needs the samples to be the same length, and aggregating per id seemed like the best idea.\n\nI may have mentioned this before, but thanks for the fantastic work you shared - I learned a lot.",
      "votes": null
    },
    {
      "id": "1572042",
      "postDate": "11/05/2021 11:21:39",
      "content": "<p>You can first multiply each mask by a different number and then aggregate with max. This should give the format you want.</p>",
      "rawMarkdown": "You can first multiply each mask by a different number and then aggregate with max. This should give the format you want.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1571778,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "11/05/2021 06:43:35",
      "content": "<p>I have the same problem but I haven't solved it yet. My best bet is on we are calculating error on background too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1571984,
      "author_name": "slawekbiel",
      "author_url": "",
      "post_date": "11/05/2021 10:26:17",
      "content": "<p>It looks to me that the problem is here:<br>\n<code>masks_valid_agg = masks_valid_agg.groupby('id').sum().reset_index()</code><br>\nSumming all the masks together looses information about individual instances. The <code>iou_map</code> input states</p>\n<blockquote>\n  <p>Masks contain the segmented pixels where each object has one value associated</p>\n</blockquote>\n<p>Wouldn't your sum just give 0 for background and 1 for every object (since overlaps are removed).<br>\nI haven't actually run it, so sorry If I'm misunderstanding what you do.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1572009,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "11/05/2021 10:57:12",
          "content": "<p>It will do just that, yes - and that might indeed be a problem. It's just that histogram2d needs the samples to be the same length, and aggregating per id seemed like the best idea.</p>\n<p>I may have mentioned this before, but thanks for the fantastic work you shared - I learned a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1572042,
          "author_name": "slawekbiel",
          "author_url": "",
          "post_date": "11/05/2021 11:21:39",
          "content": "<p>You can first multiply each mask by a different number and then aggregate with max. This should give the format you want.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1571515": "The first thing I usually try to do in a new competition is getting a validation scheme working, and this one is no exception - as in this kernel:\n\nhttps://www.kaggle.com/konradb/validation-score\n\nThe weird thing is, I took a relatively simple setup (copied from the excellent trilogy by @slawekbiel https://www.kaggle.com/slawekbiel/positive-score-with-detectron-3-3-inference) and built around it:\n1. I created a five-fold split of the data (`GroupKFold` on id)\n2. fitted a model on folds 1,2,3,4\n3. used it for evaluation on fold 0 data (the kernel above).\n\nThe part I'm struggling with is that during training my last MaP IoU was 0.2554, the LB score is 0.24 and the calculated metric is 0.056. Either there is something really strange in the data, or there is an embarrassing hole in my logic - whichever it is, I would love some feedback.",
    "1571778": "I have the same problem but I haven't solved it yet. My best bet is on we are calculating error on background too.",
    "1571984": "It looks to me that the problem is here:\n```masks_valid_agg = masks_valid_agg.groupby('id').sum().reset_index()```\nSumming all the masks together looses information about individual instances. The `iou_map` input states\n> Masks contain the segmented pixels where each object has one value associated\n\nWouldn't your sum just give 0 for background and 1 for every object (since overlaps are removed).\nI haven't actually run it, so sorry If I'm misunderstanding what you do.",
    "1572009": "It will do just that, yes - and that might indeed be a problem. It's just that histogram2d needs the samples to be the same length, and aggregating per id seemed like the best idea.\n\nI may have mentioned this before, but thanks for the fantastic work you shared - I learned a lot.",
    "1572042": "You can first multiply each mask by a different number and then aggregate with max. This should give the format you want."
  },
  "source": "meta"
}