{
  "id": 230089,
  "title": "Question on the evaluation metric (edge vs blob)",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/230089",
  "author_name": "",
  "post_date": "2021-04-01T22:52:08.959288200Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>We know that the precision (positional accuracy) of the label edges in the training set is not extremely accurate, and probably also in the private set - there is always a small deviation between model predictions and the labels provided by the host. Would it be more appropriate to penalize the presence/absence of glomeruli more than the edges when evaluating on the private set? </p>\n<p>This can be easily achieved if the final evaluation will be weighted lower within a small buffer around edges. For example, by adding a distance transform starting with let's say 0.1 on edges and increasing to 1 within some buffer, let's say 15 pixels around edges and then use that weight to weigh the dice coefficient.</p>\n<p><img src=\"https://i.ibb.co/HqRS657/image.png\" alt=\"image\"></p>\n<p>This way we can avoid randomness in the final evaluations and models that identify glomeruli best will have higher scores than those that will just guess edges better by co-incidence.</p>\n<p>An alternative is the host needs to really pay much more attention to the final labels, making sure they are of a much higher quality.</p>",
  "messages": [
    {
      "id": "1260168",
      "postDate": "04/01/2021 22:52:08",
      "content": "<p>We know that the precision (positional accuracy) of the label edges in the training set is not extremely accurate, and probably also in the private set - there is always a small deviation between model predictions and the labels provided by the host. Would it be more appropriate to penalize the presence/absence of glomeruli more than the edges when evaluating on the private set? </p>\n<p>This can be easily achieved if the final evaluation will be weighted lower within a small buffer around edges. For example, by adding a distance transform starting with let's say 0.1 on edges and increasing to 1 within some buffer, let's say 15 pixels around edges and then use that weight to weigh the dice coefficient.</p>\n<p><img src=\"https://i.ibb.co/HqRS657/image.png\" alt=\"image\"></p>\n<p>This way we can avoid randomness in the final evaluations and models that identify glomeruli best will have higher scores than those that will just guess edges better by co-incidence.</p>\n<p>An alternative is the host needs to really pay much more attention to the final labels, making sure they are of a much higher quality.</p>",
      "rawMarkdown": "We know that the precision (positional accuracy) of the label edges in the training set is not extremely accurate, and probably also in the private set - there is always a small deviation between model predictions and the labels provided by the host. Would it be more appropriate to penalize the presence/absence of glomeruli more than the edges when evaluating on the private set? \n\nThis can be easily achieved if the final evaluation will be weighted lower within a small buffer around edges. For example, by adding a distance transform starting with let's say 0.1 on edges and increasing to 1 within some buffer, let's say 15 pixels around edges and then use that weight to weigh the dice coefficient.\n\n![image](https://i.ibb.co/HqRS657/image.png)\n\nThis way we can avoid randomness in the final evaluations and models that identify glomeruli best will have higher scores than those that will just guess edges better by co-incidence.\n\nAn alternative is the host needs to really pay much more attention to the final labels, making sure they are of a much higher quality.",
      "votes": null
    },
    {
      "id": "1260244",
      "postDate": "04/02/2021 00:49:31",
      "content": "<p>My hope is that the private test set is different enough that these slight differences won’t effect the outcome. As an experiment, I split the training data into a few sets based on their overall similarity. Then I trained my model with one of those sets (without augmentation) and used the other sets as test samples. Because of the significant difference between the training set and the test set, the model performed worse enough that the over/underestimation of boundaries became inconsequential.</p>\n<p>Anyway if the private test set is challenging enough, what might make the outcome random is if several teams are using the same model.</p>",
      "rawMarkdown": "My hope is that the private test set is different enough that these slight differences won’t effect the outcome. As an experiment, I split the training data into a few sets based on their overall similarity. Then I trained my model with one of those sets (without augmentation) and used the other sets as test samples. Because of the significant difference between the training set and the test set, the model performed worse enough that the over/underestimation of boundaries became inconsequential.\n\nAnyway if the private test set is challenging enough, what might make the outcome random is if several teams are using the same model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1260244,
      "author_name": "erikdali",
      "author_url": "",
      "post_date": "04/02/2021 00:49:31",
      "content": "<p>My hope is that the private test set is different enough that these slight differences won’t effect the outcome. As an experiment, I split the training data into a few sets based on their overall similarity. Then I trained my model with one of those sets (without augmentation) and used the other sets as test samples. Because of the significant difference between the training set and the test set, the model performed worse enough that the over/underestimation of boundaries became inconsequential.</p>\n<p>Anyway if the private test set is challenging enough, what might make the outcome random is if several teams are using the same model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1260168": "We know that the precision (positional accuracy) of the label edges in the training set is not extremely accurate, and probably also in the private set - there is always a small deviation between model predictions and the labels provided by the host. Would it be more appropriate to penalize the presence/absence of glomeruli more than the edges when evaluating on the private set? \n\nThis can be easily achieved if the final evaluation will be weighted lower within a small buffer around edges. For example, by adding a distance transform starting with let's say 0.1 on edges and increasing to 1 within some buffer, let's say 15 pixels around edges and then use that weight to weigh the dice coefficient.\n\n![image](https://i.ibb.co/HqRS657/image.png)\n\nThis way we can avoid randomness in the final evaluations and models that identify glomeruli best will have higher scores than those that will just guess edges better by co-incidence.\n\nAn alternative is the host needs to really pay much more attention to the final labels, making sure they are of a much higher quality.",
    "1260244": "My hope is that the private test set is different enough that these slight differences won’t effect the outcome. As an experiment, I split the training data into a few sets based on their overall similarity. Then I trained my model with one of those sets (without augmentation) and used the other sets as test samples. Because of the significant difference between the training set and the test set, the model performed worse enough that the over/underestimation of boundaries became inconsequential.\n\nAnyway if the private test set is challenging enough, what might make the outcome random is if several teams are using the same model."
  },
  "source": "meta"
}