{
  "id": 221099,
  "title": "Does removing borderline cells matter any more with mAP?",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/221099",
  "author_name": "Darek Kłeczek",
  "post_date": "2021-02-21T06:32:47.049000",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>There is a comment by host that cells at the border are likely to not be annotated:</p>\n<blockquote>\n  <p>A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it! </p>\n</blockquote>\n<p>This might have been a big deal with the original F1 metric, but based on my current understanding this should not matter with mAP. I've experimented a bit with removing cells below certain threshold from my predictions and it didn't influence the score. Am I right on this? </p>",
  "messages": [
    {
      "id": 1212344,
      "postDate": "2021-02-21T06:32:47.050Z",
      "content": "<p>There is a comment by host that cells at the border are likely to not be annotated:</p>\n<blockquote>\n  <p>A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it! </p>\n</blockquote>\n<p>This might have been a big deal with the original F1 metric, but based on my current understanding this should not matter with mAP. I've experimented a bit with removing cells below certain threshold from my predictions and it didn't influence the score. Am I right on this? </p>",
      "rawMarkdown": "There is a comment by host that cells at the border are likely to not be annotated:\n\n> A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it! \n\nThis might have been a big deal with the original F1 metric, but based on my current understanding this should not matter with mAP. I've experimented a bit with removing cells below certain threshold from my predictions and it didn't influence the score. Am I right on this? ",
      "votes": 6
    },
    {
      "id": 1212361,
      "postDate": "2021-02-21T06:45:42.767Z",
      "content": "<p>Not played with this yet - but my understanding is that the score flow runs thru the truth segments and looks for a 0.6 matching segment in our submission.  If the truth segments do not include that edge segment than it appears No Harm No Foul.</p>\n<p>But an annotators 1/2 and my 1/2 might not match - so I believe it's OK to make a prediction on an edge cell without mAP harm if that single edge cell not really a part of the ground truth.</p>",
      "rawMarkdown": "Not played with this yet - but my understanding is that the score flow runs thru the truth segments and looks for a 0.6 matching segment in our submission.  If the truth segments do not include that edge segment than it appears No Harm No Foul.\n\nBut an annotators 1/2 and my 1/2 might not match - so I believe it's OK to make a prediction on an edge cell without mAP harm if that single edge cell not really a part of the ground truth.\n",
      "votes": 2
    },
    {
      "id": 1214192,
      "postDate": "2021-02-22T16:39:49.340Z",
      "content": "<p>Hello, can you please tell me how you removed the border cells while experimenting? Like, how to know whether more than half the cell is present or not? Sorry if it's a silly question 😅 </p>",
      "rawMarkdown": "Hello, can you please tell me how you removed the border cells while experimenting? Like, how to know whether more than half the cell is present or not? Sorry if it's a silly question 😅 ",
      "replies": [
        {
          "id": 1214206,
          "postDate": "2021-02-22T16:53:21.650Z",
          "content": "<p>That's a good question. There could be different ways. I computed the cell and the nucleus segments, with the \"label_cell\" function, then computed the centers for  all the nuclei.<br>\nThen check if there is such a center inside of your cell. If no nucleus discard the cell. If there is a nucleus you can further filter by checking if the center of  nucleus is close to the boundary of the image (say closer than 50 pixels). </p>",
          "rawMarkdown": "That's a good question. There could be different ways. I computed the cell and the nucleus segments, with the \"label_cell\" function, then computed the centers for  all the nuclei.\nThen check if there is such a center inside of your cell. If no nucleus discard the cell. If there is a nucleus you can further filter by checking if the center of  nucleus is close to the boundary of the image (say closer than 50 pixels). ",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 1214223,
          "postDate": "2021-02-22T17:05:56.623Z",
          "content": "<p>I see, so there is some sort of heuristic (&lt;50px) involved afterall</p>",
          "rawMarkdown": "I see, so there is some sort of heuristic (<50px) involved afterall"
        },
        {
          "id": 1214757,
          "postDate": "2021-02-23T05:29:05.157Z",
          "content": "<p>Thanks for sharing your approach Zoltan! My approach was much more naive, I calculated average size of cell per image and discarded cells that had significantly smaller size. Thinking about it now, this probably wasn't the best heuristic. </p>",
          "rawMarkdown": "Thanks for sharing your approach Zoltan! My approach was much more naive, I calculated average size of cell per image and discarded cells that had significantly smaller size. Thinking about it now, this probably wasn't the best heuristic. "
        }
      ]
    },
    {
      "id": 1213238,
      "postDate": "2021-02-22T01:43:36.307Z",
      "content": "<p>I have also experimented with this. So far it didn't change my scores.</p>",
      "rawMarkdown": "I have also experimented with this. So far it didn't change my scores.",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1212361,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-02-21T06:45:42.767000",
      "content": "<p>Not played with this yet - but my understanding is that the score flow runs thru the truth segments and looks for a 0.6 matching segment in our submission.  If the truth segments do not include that edge segment than it appears No Harm No Foul.</p>\n<p>But an annotators 1/2 and my 1/2 might not match - so I believe it's OK to make a prediction on an edge cell without mAP harm if that single edge cell not really a part of the ground truth.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1214192,
      "author_name": "Arka Saha",
      "author_url": "",
      "post_date": "2021-02-22T16:39:49.340000",
      "content": "<p>Hello, can you please tell me how you removed the border cells while experimenting? Like, how to know whether more than half the cell is present or not? Sorry if it's a silly question 😅 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1214206,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-22T16:53:21.650000",
          "content": "<p>That's a good question. There could be different ways. I computed the cell and the nucleus segments, with the \"label_cell\" function, then computed the centers for  all the nuclei.<br>\nThen check if there is such a center inside of your cell. If no nucleus discard the cell. If there is a nucleus you can further filter by checking if the center of  nucleus is close to the boundary of the image (say closer than 50 pixels). </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1214223,
          "author_name": "Arka Saha",
          "author_url": "",
          "post_date": "2021-02-22T17:05:56.623000",
          "content": "<p>I see, so there is some sort of heuristic (&lt;50px) involved afterall</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1214757,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-02-23T05:29:05.157000",
          "content": "<p>Thanks for sharing your approach Zoltan! My approach was much more naive, I calculated average size of cell per image and discarded cells that had significantly smaller size. Thinking about it now, this probably wasn't the best heuristic. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1213238,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-22T01:43:36.307000",
      "content": "<p>I have also experimented with this. So far it didn't change my scores.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1212344": "There is a comment by host that cells at the border are likely to not be annotated:\n\n> A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it! \n\nThis might have been a big deal with the original F1 metric, but based on my current understanding this should not matter with mAP. I've experimented a bit with removing cells below certain threshold from my predictions and it didn't influence the score. Am I right on this? ",
    "1212361": "Not played with this yet - but my understanding is that the score flow runs thru the truth segments and looks for a 0.6 matching segment in our submission.  If the truth segments do not include that edge segment than it appears No Harm No Foul.\n\nBut an annotators 1/2 and my 1/2 might not match - so I believe it's OK to make a prediction on an edge cell without mAP harm if that single edge cell not really a part of the ground truth.\n",
    "1214192": "Hello, can you please tell me how you removed the border cells while experimenting? Like, how to know whether more than half the cell is present or not? Sorry if it's a silly question 😅 ",
    "1213238": "I have also experimented with this. So far it didn't change my scores."
  }
}