{
  "id": 291639,
  "title": "First experiment with cleaned astro masks",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/291639",
  "author_name": "Slawek Biel",
  "post_date": "2021-11-30T10:27:13.476000",
  "votes": 55,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I took the <a href=\"https://www.kaggle.com/hengck23/clean-astro-mask\" target=\"_blank\">cleaned masks</a> from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> (see <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/291371\" target=\"_blank\">discussion</a>) and run it through my pipeline, comparing to previous results with same split and hyperparameters. I then scored it on the original masks as is and with fillConvexPoly postprocessing similar to what I've done in <a href=\"https://www.kaggle.com/slawekbiel/broken-mask-example\" target=\"_blank\">this notebook</a></p>\n<p>During training, unsurprisingly it did improve astro performance, without affecting the other two. The orange line is AP with original train and val and teal line is with both cleaned.<br>\n<img src=\"https://raw.githubusercontent.com/slawekslex/random/main/astro.png\" alt=\"\"></p>\n<p>On inference however I didn't see an improvement so far:</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>val score</th>\n<th>LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>original</td>\n<td>.312</td>\n<td>.320</td>\n</tr>\n<tr>\n<td>cleaned without fillConvex</td>\n<td>.308</td>\n<td>.319</td>\n</tr>\n<tr>\n<td>cleaned with fillConvex</td>\n<td>.309</td>\n<td>.320</td>\n</tr>\n</tbody>\n</table>\n<p>This is of course only comparing two single runs, the difference might as well be just variance.</p>",
  "messages": [
    {
      "id": 1600355,
      "postDate": "2021-11-30T10:27:13.477Z",
      "content": "<p>I took the <a href=\"https://www.kaggle.com/hengck23/clean-astro-mask\" target=\"_blank\">cleaned masks</a> from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> (see <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/291371\" target=\"_blank\">discussion</a>) and run it through my pipeline, comparing to previous results with same split and hyperparameters. I then scored it on the original masks as is and with fillConvexPoly postprocessing similar to what I've done in <a href=\"https://www.kaggle.com/slawekbiel/broken-mask-example\" target=\"_blank\">this notebook</a></p>\n<p>During training, unsurprisingly it did improve astro performance, without affecting the other two. The orange line is AP with original train and val and teal line is with both cleaned.<br>\n<img src=\"https://raw.githubusercontent.com/slawekslex/random/main/astro.png\" alt=\"\"></p>\n<p>On inference however I didn't see an improvement so far:</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>val score</th>\n<th>LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>original</td>\n<td>.312</td>\n<td>.320</td>\n</tr>\n<tr>\n<td>cleaned without fillConvex</td>\n<td>.308</td>\n<td>.319</td>\n</tr>\n<tr>\n<td>cleaned with fillConvex</td>\n<td>.309</td>\n<td>.320</td>\n</tr>\n</tbody>\n</table>\n<p>This is of course only comparing two single runs, the difference might as well be just variance.</p>",
      "rawMarkdown": "I took the [cleaned masks](https://www.kaggle.com/hengck23/clean-astro-mask) from @hengck23 (see [discussion](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/291371)) and run it through my pipeline, comparing to previous results with same split and hyperparameters. I then scored it on the original masks as is and with fillConvexPoly postprocessing similar to what I've done in [this notebook](https://www.kaggle.com/slawekbiel/broken-mask-example)\n\nDuring training, unsurprisingly it did improve astro performance, without affecting the other two. The orange line is AP with original train and val and teal line is with both cleaned.\n![](https://raw.githubusercontent.com/slawekslex/random/main/astro.png)\n\nOn inference however I didn't see an improvement so far:\n| model |val score  |LB score\n| --- | --- | --- |\n| original | .312 | .320\n| cleaned without fillConvex | .308 | .319\n| cleaned with fillConvex | .309 | .320\n\nThis is of course only comparing two single runs, the difference might as well be just variance.",
      "votes": 55
    },
    {
      "id": 1600378,
      "postDate": "2021-11-30T10:51:39.057Z",
      "content": "<p>It does confirm that the same mask issue is present in the test data, since intentionally breaking my masks improved the LB score.</p>",
      "rawMarkdown": "It does confirm that the same mask issue is present in the test data, since intentionally breaking my masks improved the LB score.",
      "votes": 6,
      "replies": [
        {
          "id": 1600646,
          "postDate": "2021-11-30T15:30:16.830Z",
          "content": "<p>Intentionally breaking?</p>\n<p>(sorry, couldn't help it)</p>",
          "rawMarkdown": "Intentionally breaking?\n\n(sorry, couldn't help it)"
        },
        {
          "id": 1600727,
          "postDate": "2021-11-30T17:04:20.403Z",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> I mean making predictions worse in the hope of better matching incorrectly encoded masks in the test set.</p>",
          "rawMarkdown": "@authman I mean making predictions worse in the hope of better matching incorrectly encoded masks in the test set."
        },
        {
          "id": 1600864,
          "postDate": "2021-11-30T19:13:38.560Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1601840,
          "postDate": "2021-12-01T16:07:21.420Z",
          "content": "<p><a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a>  which one in above list is the one intentionally broken part. <br>\nsorry i am trying to understand.</p>",
          "rawMarkdown": "@slawekbiel  which one in above list is the one intentionally broken part. \nsorry i am trying to understand."
        },
        {
          "id": 1601899,
          "postDate": "2021-12-01T16:17:30.203Z",
          "content": "<blockquote>\n  <p>cleaned with fillConvex</p>\n</blockquote>",
          "rawMarkdown": "> cleaned with fillConvex"
        },
        {
          "id": 1601912,
          "postDate": "2021-12-01T16:28:09.377Z",
          "content": "<p>Thanks..<br>\nwhat is diff between cleaned with and without.. <br>\nWhich one is same as hengs ?</p>",
          "rawMarkdown": "Thanks..\nwhat is diff between cleaned with and without.. \nWhich one is same as hengs ?"
        },
        {
          "id": 1602049,
          "postDate": "2021-12-01T18:04:03.057Z",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> please read the post and the notebook I've linked in the top post, it should make everything clear.</p>",
          "rawMarkdown": "@jaideepvalani please read the post and the notebook I've linked in the top post, it should make everything clear."
        }
      ]
    },
    {
      "id": 1611992,
      "postDate": "2021-12-08T12:57:58.367Z",
      "content": "<p>Thanks you you great job，i want to ask how to slove erro [Errno 2] No such file or directory: '../input/sample-mask/mask.npy'</p>",
      "rawMarkdown": "Thanks you you great job，i want to ask how to slove erro [Errno 2] No such file or directory: '../input/sample-mask/mask.npy'"
    },
    {
      "id": 1603406,
      "postDate": "2021-12-02T13:52:55.670Z",
      "content": "<p><a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> what kind of AP metric you visualized? 0.24 mAP IoU for asto looks extremely good to me.</p>",
      "rawMarkdown": "@slawekbiel what kind of AP metric you visualized? 0.24 mAP IoU for asto looks extremely good to me.",
      "replies": [
        {
          "id": 1603575,
          "postDate": "2021-12-02T15:27:16.430Z",
          "content": "<p><a href=\"https://www.kaggle.com/rednikotin\" target=\"_blank\">@rednikotin</a> I use the evaluation code from the LIVECell repo: <a href=\"https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py\" target=\"_blank\">https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py</a></p>",
          "rawMarkdown": "@rednikotin I use the evaluation code from the LIVECell repo: https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py",
          "votes": 1
        }
      ]
    },
    {
      "id": 1604184,
      "postDate": "2021-12-03T06:59:00.567Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1603607,
      "postDate": "2021-12-02T16:13:17.970Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 1606644,
      "postDate": "2021-12-05T06:41:16.337Z",
      "content": "<p>Thank you!<br>\nGreat work!</p>",
      "rawMarkdown": "Thank you!\nGreat work!"
    },
    {
      "id": 1606115,
      "postDate": "2021-12-04T21:07:36.583Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!"
    },
    {
      "id": 1604745,
      "postDate": "2021-12-03T17:04:06.757Z",
      "content": "<p>wwooo, thank you so much</p>",
      "rawMarkdown": "wwooo, thank you so much"
    }
  ],
  "comments": [
    {
      "id": 1600378,
      "author_name": "Slawek Biel",
      "author_url": "",
      "post_date": "2021-11-30T10:51:39.057000",
      "content": "<p>It does confirm that the same mask issue is present in the test data, since intentionally breaking my masks improved the LB score.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1600646,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-11-30T15:30:16.830000",
          "content": "<p>Intentionally breaking?</p>\n<p>(sorry, couldn't help it)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1600727,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-11-30T17:04:20.403000",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> I mean making predictions worse in the hope of better matching incorrectly encoded masks in the test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1600864,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-11-30T19:13:38.560000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1601840,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-12-01T16:07:21.420000",
          "content": "<p><a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a>  which one in above list is the one intentionally broken part. <br>\nsorry i am trying to understand.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1601899,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-12-01T16:17:30.203000",
          "content": "<blockquote>\n  <p>cleaned with fillConvex</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1601912,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-12-01T16:28:09.377000",
          "content": "<p>Thanks..<br>\nwhat is diff between cleaned with and without.. <br>\nWhich one is same as hengs ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1602049,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-12-01T18:04:03.057000",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> please read the post and the notebook I've linked in the top post, it should make everything clear.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1611992,
      "author_name": "Rongye Ye",
      "author_url": "",
      "post_date": "2021-12-08T12:57:58.367000",
      "content": "<p>Thanks you you great job，i want to ask how to slove erro [Errno 2] No such file or directory: '../input/sample-mask/mask.npy'</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1603406,
      "author_name": "Valentin Nikotin",
      "author_url": "",
      "post_date": "2021-12-02T13:52:55.670000",
      "content": "<p><a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> what kind of AP metric you visualized? 0.24 mAP IoU for asto looks extremely good to me.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1603575,
          "author_name": "Slawek Biel",
          "author_url": "",
          "post_date": "2021-12-02T15:27:16.430000",
          "content": "<p><a href=\"https://www.kaggle.com/rednikotin\" target=\"_blank\">@rednikotin</a> I use the evaluation code from the LIVECell repo: <a href=\"https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py\" target=\"_blank\">https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1604184,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-12-03T06:59:00.567000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1603607,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-12-02T16:13:17.970000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 1606644,
      "author_name": "John Doe",
      "author_url": "",
      "post_date": "2021-12-05T06:41:16.337000",
      "content": "<p>Thank you!<br>\nGreat work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1606115,
      "author_name": "KAZI SHAMIM SHAHAREAR ISLAM",
      "author_url": "",
      "post_date": "2021-12-04T21:07:36.583000",
      "content": "<p>Thanks a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1604745,
      "author_name": "Tư Mã Trọng Đạt",
      "author_url": "",
      "post_date": "2021-12-03T17:04:06.757000",
      "content": "<p>wwooo, thank you so much</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1600355": "I took the [cleaned masks](https://www.kaggle.com/hengck23/clean-astro-mask) from @hengck23 (see [discussion](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/291371)) and run it through my pipeline, comparing to previous results with same split and hyperparameters. I then scored it on the original masks as is and with fillConvexPoly postprocessing similar to what I've done in [this notebook](https://www.kaggle.com/slawekbiel/broken-mask-example)\n\nDuring training, unsurprisingly it did improve astro performance, without affecting the other two. The orange line is AP with original train and val and teal line is with both cleaned.\n![](https://raw.githubusercontent.com/slawekslex/random/main/astro.png)\n\nOn inference however I didn't see an improvement so far:\n| model |val score  |LB score\n| --- | --- | --- |\n| original | .312 | .320\n| cleaned without fillConvex | .308 | .319\n| cleaned with fillConvex | .309 | .320\n\nThis is of course only comparing two single runs, the difference might as well be just variance.",
    "1600378": "It does confirm that the same mask issue is present in the test data, since intentionally breaking my masks improved the LB score.",
    "1611992": "Thanks you you great job，i want to ask how to slove erro [Errno 2] No such file or directory: '../input/sample-mask/mask.npy'",
    "1603406": "@slawekbiel what kind of AP metric you visualized? 0.24 mAP IoU for asto looks extremely good to me.",
    "1604184": "",
    "1603607": "",
    "1606644": "Thank you!\nGreat work!",
    "1606115": "Thanks a lot!",
    "1604745": "wwooo, thank you so much"
  }
}