{
  "id": 281205,
  "title": "Annotations Are too Noisy for the Metric",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/281205",
  "author_name": "",
  "post_date": "2021-10-23T21:02:47.300643400Z",
  "votes": 174,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Here, we aim at segmenting cells, such task can be divided in two steps :</p>\n<ol>\n<li>Detecting individual cells</li>\n<li>Correctly predicting their boundaries</li>\n</ol>\n<p>I'm focusing here on the 2nd step : <br>\nThe mAP IoU evaluates the correctness of the boundaries by penalizing low IoUs using different thresholds (0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95).</p>\n<p>However, I think this is a mistake, let's take two cort cells as example, this also works with shsy5y cells.</p>\n<p><a href=\"https://ibb.co/B6NDQVS\"><img src=\"https://i.ibb.co/vdY5TxK/ious.png\" alt=\"ious\"></a></p>\n<ul>\n<li>Cells are small : about 10x10 px each, because of the resolution the data is acquired at</li>\n<li>Annotations are not pixel perfect : there is a lot of ambiguity on where the cell actually stops, sometimes the brighter pixels are included, sometimes excluded</li>\n<li>Lets say such ambiguity results in a ~20px variation : this will result in a ~0.2 IoU variation for the same prediction</li>\n<li>Which means reaching the 0.8, 0.85, 0.9 and 0.95 thresholds is basically luck. I.e 40% of the evaluation metric is random</li>\n<li>The two predictions here score IoU 0.76 and 0.79 which only gives a 0.6 mAP despite them being decent enough to be considered as good as the labels</li>\n</ul>\n<p><strong>Conclusion :</strong> because cells are smalls and labels not pixel perfect, the metric fluctuates a lot.<br>\nIt makes no sense to evaluate the IoU at high thresholds for small cells. </p>\n<p>An easy fix would be to consider the mAP at thresholds ranging from 0.5 to 0.75 (or 0.8).</p>\n<p>A more complicated (but better) fix would be to choose threshold according to the ground truth size, for instance :</p>\n<ul>\n<li>&lt; 200 px cells should use 0.5 -&gt; 0.75</li>\n<li>&lt; 400 px cells should use 0.5 -&gt; 0.85</li>\n<li>&gt;= 400 px cells should use 0.5 -&gt; 0.95</li>\n</ul>\n<p>High thresholds are useful to make sure boundaries are correct on big cells (astro) as well.</p>",
  "messages": [
    {
      "id": "1555493",
      "postDate": "10/23/2021 21:02:47",
      "content": "<p>Here, we aim at segmenting cells, such task can be divided in two steps :</p>\n<ol>\n<li>Detecting individual cells</li>\n<li>Correctly predicting their boundaries</li>\n</ol>\n<p>I'm focusing here on the 2nd step : <br>\nThe mAP IoU evaluates the correctness of the boundaries by penalizing low IoUs using different thresholds (0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95).</p>\n<p>However, I think this is a mistake, let's take two cort cells as example, this also works with shsy5y cells.</p>\n<p><a href=\"https://ibb.co/B6NDQVS\"><img src=\"https://i.ibb.co/vdY5TxK/ious.png\" alt=\"ious\"></a></p>\n<ul>\n<li>Cells are small : about 10x10 px each, because of the resolution the data is acquired at</li>\n<li>Annotations are not pixel perfect : there is a lot of ambiguity on where the cell actually stops, sometimes the brighter pixels are included, sometimes excluded</li>\n<li>Lets say such ambiguity results in a ~20px variation : this will result in a ~0.2 IoU variation for the same prediction</li>\n<li>Which means reaching the 0.8, 0.85, 0.9 and 0.95 thresholds is basically luck. I.e 40% of the evaluation metric is random</li>\n<li>The two predictions here score IoU 0.76 and 0.79 which only gives a 0.6 mAP despite them being decent enough to be considered as good as the labels</li>\n</ul>\n<p><strong>Conclusion :</strong> because cells are smalls and labels not pixel perfect, the metric fluctuates a lot.<br>\nIt makes no sense to evaluate the IoU at high thresholds for small cells. </p>\n<p>An easy fix would be to consider the mAP at thresholds ranging from 0.5 to 0.75 (or 0.8).</p>\n<p>A more complicated (but better) fix would be to choose threshold according to the ground truth size, for instance :</p>\n<ul>\n<li>&lt; 200 px cells should use 0.5 -&gt; 0.75</li>\n<li>&lt; 400 px cells should use 0.5 -&gt; 0.85</li>\n<li>&gt;= 400 px cells should use 0.5 -&gt; 0.95</li>\n</ul>\n<p>High thresholds are useful to make sure boundaries are correct on big cells (astro) as well.</p>",
      "rawMarkdown": "Here, we aim at segmenting cells, such task can be divided in two steps :\n1. Detecting individual cells\n2. Correctly predicting their boundaries\n\nI'm focusing here on the 2nd step : \nThe mAP IoU evaluates the correctness of the boundaries by penalizing low IoUs using different thresholds (0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95).\n\nHowever, I think this is a mistake, let's take two cort cells as example, this also works with shsy5y cells.\n\n<a href=\"https://ibb.co/B6NDQVS\"><img src=\"https://i.ibb.co/vdY5TxK/ious.png\" alt=\"ious\" border=\"0\"></a>\n\n- Cells are small : about 10x10 px each, because of the resolution the data is acquired at\n- Annotations are not pixel perfect : there is a lot of ambiguity on where the cell actually stops, sometimes the brighter pixels are included, sometimes excluded\n- Lets say such ambiguity results in a ~20px variation : this will result in a ~0.2 IoU variation for the same prediction\n- Which means reaching the 0.8, 0.85, 0.9 and 0.95 thresholds is basically luck. I.e 40% of the evaluation metric is random\n- The two predictions here score IoU 0.76 and 0.79 which only gives a 0.6 mAP despite them being decent enough to be considered as good as the labels\n\n**Conclusion :** because cells are smalls and labels not pixel perfect, the metric fluctuates a lot.\nIt makes no sense to evaluate the IoU at high thresholds for small cells. \n\nAn easy fix would be to consider the mAP at thresholds ranging from 0.5 to 0.75 (or 0.8).\n\nA more complicated (but better) fix would be to choose threshold according to the ground truth size, for instance :\n- < 200 px cells should use 0.5 -> 0.75\n- < 400 px cells should use 0.5 -> 0.85\n- >= 400 px cells should use 0.5 -> 0.95\n\nHigh thresholds are useful to make sure boundaries are correct on big cells (astro) as well.",
      "votes": null
    },
    {
      "id": "1555683",
      "postDate": "10/24/2021 05:33:50",
      "content": "<p>I completely agree, looks like the metric may introduce quite a significant amount of unnecessary noise to the score.</p>\n<p>Another option would be to shift the threshold ranges for smaller cells, like:</p>\n<ul>\n<li>&lt; 200 px cells should use 0.3 -&gt; 0.75</li>\n<li>&lt; 400 px cells should use 0.4 -&gt; 0.85</li>\n<li>&gt;= 400 px cells should use 0.5 -&gt; 0.95</li>\n</ul>\n<p>Even 0.3-04 IOU sounds like a useful but less accurate detection.</p>",
      "rawMarkdown": "I completely agree, looks like the metric may introduce quite a significant amount of unnecessary noise to the score.\n\nAnother option would be to shift the threshold ranges for smaller cells, like:\n\n- < 200 px cells should use 0.3 -> 0.75\n- < 400 px cells should use 0.4 -> 0.85\n- >= 400 px cells should use 0.5 -> 0.95\n\nEven 0.3-04 IOU sounds like a useful but less accurate detection.",
      "votes": null
    },
    {
      "id": "1555795",
      "postDate": "10/24/2021 08:04:31",
      "content": "<p>I thought the same thing. Thresholds &gt; 0.8 can be unrealistic with such noisy annotations.</p>",
      "rawMarkdown": "I thought the same thing. Thresholds > 0.8 can be unrealistic with such noisy annotations.",
      "votes": null
    },
    {
      "id": "1555801",
      "postDate": "10/24/2021 08:14:15",
      "content": "<p>Agreed, shifting thresholds is even better</p>",
      "rawMarkdown": "Agreed, shifting thresholds is even better",
      "votes": null
    },
    {
      "id": "1555823",
      "postDate": "10/24/2021 08:51:37",
      "content": "<p>Add to this the <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/280250\" target=\"_blank\">disambiguity of the overlaps</a> that <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281084\" target=\"_blank\">averages to 10-20 pixels</a> on an image (depending on cell type).</p>",
      "rawMarkdown": "Add to this the [disambiguity of the overlaps](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/280250) that [averages to 10-20 pixels](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281084) on an image (depending on cell type).",
      "votes": null
    },
    {
      "id": "1556026",
      "postDate": "10/24/2021 12:58:45",
      "content": "<p>I guess that perfectly simulates the real world data</p>",
      "rawMarkdown": "I guess that perfectly simulates the real world data",
      "votes": null
    },
    {
      "id": "1557248",
      "postDate": "10/25/2021 14:25:51",
      "content": "<p>I'm pretty new to kaggle. Is it pretty common that the hosts of the competitions actually modify the metrics in the middle of the competitions? because that sounds pretty confusing sometimes…</p>",
      "rawMarkdown": "I'm pretty new to kaggle. Is it pretty common that the hosts of the competitions actually modify the metrics in the middle of the competitions? because that sounds pretty confusing sometimes...",
      "votes": null
    },
    {
      "id": "1558201",
      "postDate": "10/26/2021 06:08:47",
      "content": "<p><a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> <br>\nCould you please give some comment on that? </p>",
      "rawMarkdown": "christoffersartorius \nCould you please give some comment on that?",
      "votes": null
    },
    {
      "id": "1558954",
      "postDate": "10/26/2021 15:27:00",
      "content": "<p>It is easy to see on these samples that solution possibly could predict boundaries better than ground truth but still being punished for it. </p>",
      "rawMarkdown": "It is easy to see on these samples that solution possibly could predict boundaries better than ground truth but still being punished for it.",
      "votes": null
    },
    {
      "id": "1563279",
      "postDate": "10/28/2021 06:53:26",
      "content": "<p>Agreed, I was looking through the annotation masks and respective images and came to the same conclusion. It does seem as though there's a large amount of variability between annotative markings, as well as determinable variance in the quality of each marking.</p>",
      "rawMarkdown": "Agreed, I was looking through the annotation masks and respective images and came to the same conclusion. It does seem as though there's a large amount of variability between annotative markings, as well as determinable variance in the quality of each marking.",
      "votes": null
    },
    {
      "id": "1564428",
      "postDate": "10/29/2021 08:03:36",
      "content": "<p>Would it be safe to assume this is the case for the test dataset as well or the private dataset might be more properly labelled?</p>",
      "rawMarkdown": "Would it be safe to assume this is the case for the test dataset as well or the private dataset might be more properly labelled?",
      "votes": null
    },
    {
      "id": "1567462",
      "postDate": "11/01/2021 22:42:01",
      "content": "<p>Can you explain annotation column in the train.csv <br>\nHow run length coding can be interpreted from annotation column and there is one more integer end of the run length code?Thanks</p>",
      "rawMarkdown": "Can you explain annotation column in the train.csv \nHow run length coding can be interpreted from annotation column and there is one more integer end of the run length code?Thanks",
      "votes": null
    },
    {
      "id": "1570576",
      "postDate": "11/04/2021 08:41:48",
      "content": "<p>after training and comparing for a while, Now I understand this topic.</p>",
      "rawMarkdown": "after training and comparing for a while, Now I understand this topic.",
      "votes": null
    },
    {
      "id": "1570845",
      "postDate": "11/04/2021 12:26:56",
      "content": "<p>Hi, this is a really interesting and relevant observation. </p>\n<p>I really like your suggestion of adjusting the relevant thresholds based on object size, that makes a lot of sense and will definitely take that with us moving forward when working with this type of data.</p>\n<p>I will also inquire about what can be done in the context of a running competition.</p>",
      "rawMarkdown": "Hi, this is a really interesting and relevant observation. \n\nI really like your suggestion of adjusting the relevant thresholds based on object size, that makes a lot of sense and will definitely take that with us moving forward when working with this type of data.\n\nI will also inquire about what can be done in the context of a running competition.",
      "votes": null
    },
    {
      "id": "1571993",
      "postDate": "11/05/2021 10:36:30",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> !</p>",
      "rawMarkdown": "Thanks a lot @christoffersartorius !",
      "votes": null
    },
    {
      "id": "1595916",
      "postDate": "11/26/2021 04:55:11",
      "content": "<p><a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> is there any final decision on this recommendation. Should we need to follow this one or the one stated on the evaluation page?</p>",
      "rawMarkdown": "christoffersartorius is there any final decision on this recommendation. Should we need to follow this one or the one stated on the evaluation page?",
      "votes": null
    },
    {
      "id": "1608567",
      "postDate": "12/06/2021 13:00:46",
      "content": "<p>Hi there <a href=\"https://www.kaggle.com/projdev\" target=\"_blank\">@projdev</a>, for this competition we will stick with what is described on the evaluation page. </p>",
      "rawMarkdown": "Hi there @projdev, for this competition we will stick with what is described on the evaluation page.",
      "votes": null
    },
    {
      "id": "1608597",
      "postDate": "12/06/2021 13:38:25",
      "content": "<p>sad.         </p>",
      "rawMarkdown": "sad.",
      "votes": null
    },
    {
      "id": "1625506",
      "postDate": "12/21/2021 21:29:58",
      "content": "<p>So guys does anyone have an idea about % of 'failed' GT in hidden set? Could it be from the same dataset as train or val, so the distribution of corrupted and good masks more or less the same?</p>",
      "rawMarkdown": "So guys does anyone have an idea about % of 'failed' GT in hidden set? Could it be from the same dataset as train or val, so the distribution of corrupted and good masks more or less the same?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1555683,
      "author_name": "dmytropoplavskiy",
      "author_url": "",
      "post_date": "10/24/2021 05:33:50",
      "content": "<p>I completely agree, looks like the metric may introduce quite a significant amount of unnecessary noise to the score.</p>\n<p>Another option would be to shift the threshold ranges for smaller cells, like:</p>\n<ul>\n<li>&lt; 200 px cells should use 0.3 -&gt; 0.75</li>\n<li>&lt; 400 px cells should use 0.4 -&gt; 0.85</li>\n<li>&gt;= 400 px cells should use 0.5 -&gt; 0.95</li>\n</ul>\n<p>Even 0.3-04 IOU sounds like a useful but less accurate detection.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1555801,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "10/24/2021 08:14:15",
          "content": "<p>Agreed, shifting thresholds is even better</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1570576,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "11/04/2021 08:41:48",
          "content": "<p>after training and comparing for a while, Now I understand this topic.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1555795,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "10/24/2021 08:04:31",
      "content": "<p>I thought the same thing. Thresholds &gt; 0.8 can be unrealistic with such noisy annotations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1555823,
      "author_name": "sandorkonya",
      "author_url": "",
      "post_date": "10/24/2021 08:51:37",
      "content": "<p>Add to this the <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/280250\" target=\"_blank\">disambiguity of the overlaps</a> that <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281084\" target=\"_blank\">averages to 10-20 pixels</a> on an image (depending on cell type).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1556026,
      "author_name": "pranshu154",
      "author_url": "",
      "post_date": "10/24/2021 12:58:45",
      "content": "<p>I guess that perfectly simulates the real world data</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1557248,
      "author_name": "lilkoke",
      "author_url": "",
      "post_date": "10/25/2021 14:25:51",
      "content": "<p>I'm pretty new to kaggle. Is it pretty common that the hosts of the competitions actually modify the metrics in the middle of the competitions? because that sounds pretty confusing sometimes…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1558201,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "10/26/2021 06:08:47",
      "content": "<p><a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> <br>\nCould you please give some comment on that? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1558954,
      "author_name": "ilyazhenin",
      "author_url": "",
      "post_date": "10/26/2021 15:27:00",
      "content": "<p>It is easy to see on these samples that solution possibly could predict boundaries better than ground truth but still being punished for it. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1563279,
      "author_name": "ferrariic",
      "author_url": "",
      "post_date": "10/28/2021 06:53:26",
      "content": "<p>Agreed, I was looking through the annotation masks and respective images and came to the same conclusion. It does seem as though there's a large amount of variability between annotative markings, as well as determinable variance in the quality of each marking.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1564428,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "10/29/2021 08:03:36",
      "content": "<p>Would it be safe to assume this is the case for the test dataset as well or the private dataset might be more properly labelled?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1567462,
      "author_name": "bogivinaykumar",
      "author_url": "",
      "post_date": "11/01/2021 22:42:01",
      "content": "<p>Can you explain annotation column in the train.csv <br>\nHow run length coding can be interpreted from annotation column and there is one more integer end of the run length code?Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1570845,
      "author_name": "christoffersartorius",
      "author_url": "",
      "post_date": "11/04/2021 12:26:56",
      "content": "<p>Hi, this is a really interesting and relevant observation. </p>\n<p>I really like your suggestion of adjusting the relevant thresholds based on object size, that makes a lot of sense and will definitely take that with us moving forward when working with this type of data.</p>\n<p>I will also inquire about what can be done in the context of a running competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1571993,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "11/05/2021 10:36:30",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1595916,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "11/26/2021 04:55:11",
      "content": "<p><a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> is there any final decision on this recommendation. Should we need to follow this one or the one stated on the evaluation page?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1608567,
          "author_name": "christoffersartorius",
          "author_url": "",
          "post_date": "12/06/2021 13:00:46",
          "content": "<p>Hi there <a href=\"https://www.kaggle.com/projdev\" target=\"_blank\">@projdev</a>, for this competition we will stick with what is described on the evaluation page. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1608597,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/06/2021 13:38:25",
          "content": "<p>sad.         </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1625506,
      "author_name": "someshadydude",
      "author_url": "",
      "post_date": "12/21/2021 21:29:58",
      "content": "<p>So guys does anyone have an idea about % of 'failed' GT in hidden set? Could it be from the same dataset as train or val, so the distribution of corrupted and good masks more or less the same?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1555493": "Here, we aim at segmenting cells, such task can be divided in two steps :\n1. Detecting individual cells\n2. Correctly predicting their boundaries\n\nI'm focusing here on the 2nd step : \nThe mAP IoU evaluates the correctness of the boundaries by penalizing low IoUs using different thresholds (0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95).\n\nHowever, I think this is a mistake, let's take two cort cells as example, this also works with shsy5y cells.\n\n<a href=\"https://ibb.co/B6NDQVS\"><img src=\"https://i.ibb.co/vdY5TxK/ious.png\" alt=\"ious\" border=\"0\"></a>\n\n- Cells are small : about 10x10 px each, because of the resolution the data is acquired at\n- Annotations are not pixel perfect : there is a lot of ambiguity on where the cell actually stops, sometimes the brighter pixels are included, sometimes excluded\n- Lets say such ambiguity results in a ~20px variation : this will result in a ~0.2 IoU variation for the same prediction\n- Which means reaching the 0.8, 0.85, 0.9 and 0.95 thresholds is basically luck. I.e 40% of the evaluation metric is random\n- The two predictions here score IoU 0.76 and 0.79 which only gives a 0.6 mAP despite them being decent enough to be considered as good as the labels\n\n**Conclusion :** because cells are smalls and labels not pixel perfect, the metric fluctuates a lot.\nIt makes no sense to evaluate the IoU at high thresholds for small cells. \n\nAn easy fix would be to consider the mAP at thresholds ranging from 0.5 to 0.75 (or 0.8).\n\nA more complicated (but better) fix would be to choose threshold according to the ground truth size, for instance :\n- < 200 px cells should use 0.5 -> 0.75\n- < 400 px cells should use 0.5 -> 0.85\n- >= 400 px cells should use 0.5 -> 0.95\n\nHigh thresholds are useful to make sure boundaries are correct on big cells (astro) as well.",
    "1555683": "I completely agree, looks like the metric may introduce quite a significant amount of unnecessary noise to the score.\n\nAnother option would be to shift the threshold ranges for smaller cells, like:\n\n- < 200 px cells should use 0.3 -> 0.75\n- < 400 px cells should use 0.4 -> 0.85\n- >= 400 px cells should use 0.5 -> 0.95\n\nEven 0.3-04 IOU sounds like a useful but less accurate detection.",
    "1555795": "I thought the same thing. Thresholds > 0.8 can be unrealistic with such noisy annotations.",
    "1555801": "Agreed, shifting thresholds is even better",
    "1555823": "Add to this the [disambiguity of the overlaps](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/280250) that [averages to 10-20 pixels](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281084) on an image (depending on cell type).",
    "1556026": "I guess that perfectly simulates the real world data",
    "1557248": "I'm pretty new to kaggle. Is it pretty common that the hosts of the competitions actually modify the metrics in the middle of the competitions? because that sounds pretty confusing sometimes...",
    "1558201": "christoffersartorius \nCould you please give some comment on that?",
    "1558954": "It is easy to see on these samples that solution possibly could predict boundaries better than ground truth but still being punished for it.",
    "1563279": "Agreed, I was looking through the annotation masks and respective images and came to the same conclusion. It does seem as though there's a large amount of variability between annotative markings, as well as determinable variance in the quality of each marking.",
    "1564428": "Would it be safe to assume this is the case for the test dataset as well or the private dataset might be more properly labelled?",
    "1567462": "Can you explain annotation column in the train.csv \nHow run length coding can be interpreted from annotation column and there is one more integer end of the run length code?Thanks",
    "1570576": "after training and comparing for a while, Now I understand this topic.",
    "1570845": "Hi, this is a really interesting and relevant observation. \n\nI really like your suggestion of adjusting the relevant thresholds based on object size, that makes a lot of sense and will definitely take that with us moving forward when working with this type of data.\n\nI will also inquire about what can be done in the context of a running competition.",
    "1571993": "Thanks a lot @christoffersartorius !",
    "1595916": "christoffersartorius is there any final decision on this recommendation. Should we need to follow this one or the one stated on the evaluation page?",
    "1608567": "Hi there @projdev, for this competition we will stick with what is described on the evaluation page.",
    "1608597": "sad.",
    "1625506": "So guys does anyone have an idea about % of 'failed' GT in hidden set? Could it be from the same dataset as train or val, so the distribution of corrupted and good masks more or less the same?"
  },
  "source": "meta"
}