{
  "id": 242253,
  "title": "Is a score of 0.56 enough ?",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/242253",
  "author_name": "Philip Hucklesby",
  "post_date": "2021-05-28T07:26:13.221000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>The scoring method has been excellently described, see for instance the input from Tito:<br>\n     <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217158\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217158</a> </p>\n<p>But this still leaves me wondering what the score \"means\" … Is the 0.56 produced by the best efforts of our admirable best competitors now good, bad or indifferent ?</p>\n<p>My immediate subjective impression is that for software to be useful it should have an accuracy of at least somewhere around 95%.</p>\n<p>Therefore, not meaning any offence and in a genuine spirit of desire to learn, I would like to ask:</p>\n<p>a) What score would be necessary for a cell-labelling software to be \"fit for purpose\" ?</p>\n<p>b) Can the scores be improved much by even better machine learning, or are we near the best possible solution ?</p>\n<p>c) What are the limiting factors that stop us getting 95%:</p>\n<ul>\n<li>Problems with quality of the images (unwanted staining, poor illumination etc.) ? </li>\n<li>Do different organelles simply not differ enough in the images, so that the problem is undecidable with any method ?</li>\n<li>Does further domain knowledge about the cells and imaging need to be incorporated ?</li>\n<li>Does the scoring method make it inherently difficult to get a high score ?</li>\n</ul>",
  "messages": [
    {
      "id": 1331042,
      "postDate": "2021-06-01T08:35:06.983Z",
      "content": "<p>Hey!</p>\n<p>Excellent questions!</p>\n<p>Part of  the reason the score is low is indeed that the method to calculate it is inherently punishing against missing rarer classes which lack enough data to be easily trained to be identified. While we haven't done a full analysis of the results of this challenge yet, we have seen in the previous challenge that the performance on different classes can be very different.</p>\n<p>If the models are great at recognizing the \"easy\" and common classes, they will still be able help us immensely. The model may be able to classify all those images and leave the ones it is unsure on to humans, with the relevant domain knowledge, to manually annotate.</p>\n<p>We don't know if we are reaching the limit of machine learning power for this problem. Models for this problem has come extremely far just in the last few years (much thanks to the Kaggle community!) and there is no reason for us to believe that the problem is undecidable yet.</p>",
      "rawMarkdown": "Hey!\n\nExcellent questions!\n\nPart of  the reason the score is low is indeed that the method to calculate it is inherently punishing against missing rarer classes which lack enough data to be easily trained to be identified. While we haven't done a full analysis of the results of this challenge yet, we have seen in the previous challenge that the performance on different classes can be very different.\n\nIf the models are great at recognizing the \"easy\" and common classes, they will still be able help us immensely. The model may be able to classify all those images and leave the ones it is unsure on to humans, with the relevant domain knowledge, to manually annotate.\n\nWe don't know if we are reaching the limit of machine learning power for this problem. Models for this problem has come extremely far just in the last few years (much thanks to the Kaggle community!) and there is no reason for us to believe that the problem is undecidable yet.",
      "votes": 6,
      "replies": [
        {
          "id": 1357724,
          "postDate": "2021-06-19T23:23:38.287Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> </p>\n<p>Thank you for your response and apologies for not acknowledging sooner - I have been thinking about your answer.</p>\n<p>I see two levels to understanding the mAP score:</p>\n<pre><code>1. Effect on the score by building the mean AP from the AP scores of the classes.\n2. Intuitive understanding of the single AP score of each class.\n</code></pre>\n<p>I guess we all have a pretty good intuitive grasp of the first point since it's just averages and outliers.</p>\n<p>Still 0.56 seemed to me surprisingly low, so I had a closer look at point 2.</p>\n<p>I am sharing this notebook with my examination of the effect:</p>\n<pre><code>https://www.kaggle.com/philiphucklesby/intuition-of-ap-score\n</code></pre>\n<p>My conclusion is:</p>\n<p>The AP score is significantly dependent on the confidence scores attributed to false positives and true positives.</p>\n<p>If the confidence is successfully allocated such that false positives are less confident than true positives, then AP will correspond to the intuitive notion of the accuracy of the test (the 95% mentioned above).</p>\n<p>However, if the confidence is not successfully allocated, then particularly for rare labels, the AP score may be heavily penalised</p>",
          "rawMarkdown": "Hi @cwinsnes \n\nThank you for your response and apologies for not acknowledging sooner - I have been thinking about your answer.\n\nI see two levels to understanding the mAP score:\n\n\t1. Effect on the score by building the mean AP from the AP scores of the classes.\n\t2. Intuitive understanding of the single AP score of each class.\n\nI guess we all have a pretty good intuitive grasp of the first point since it's just averages and outliers.\n\nStill 0.56 seemed to me surprisingly low, so I had a closer look at point 2.\n\nI am sharing this notebook with my examination of the effect:\n\n\thttps://www.kaggle.com/philiphucklesby/intuition-of-ap-score\n\nMy conclusion is:\n\nThe AP score is significantly dependent on the confidence scores attributed to false positives and true positives.\n\nIf the confidence is successfully allocated such that false positives are less confident than true positives, then AP will correspond to the intuitive notion of the accuracy of the test (the 95% mentioned above).\n\nHowever, if the confidence is not successfully allocated, then particularly for rare labels, the AP score may be heavily penalised\n\n"
        }
      ]
    },
    {
      "id": 1326049,
      "postDate": "2021-05-28T07:26:13.223Z",
      "content": "<p>Hi all,</p>\n<p>The scoring method has been excellently described, see for instance the input from Tito:<br>\n     <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217158\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217158</a> </p>\n<p>But this still leaves me wondering what the score \"means\" … Is the 0.56 produced by the best efforts of our admirable best competitors now good, bad or indifferent ?</p>\n<p>My immediate subjective impression is that for software to be useful it should have an accuracy of at least somewhere around 95%.</p>\n<p>Therefore, not meaning any offence and in a genuine spirit of desire to learn, I would like to ask:</p>\n<p>a) What score would be necessary for a cell-labelling software to be \"fit for purpose\" ?</p>\n<p>b) Can the scores be improved much by even better machine learning, or are we near the best possible solution ?</p>\n<p>c) What are the limiting factors that stop us getting 95%:</p>\n<ul>\n<li>Problems with quality of the images (unwanted staining, poor illumination etc.) ? </li>\n<li>Do different organelles simply not differ enough in the images, so that the problem is undecidable with any method ?</li>\n<li>Does further domain knowledge about the cells and imaging need to be incorporated ?</li>\n<li>Does the scoring method make it inherently difficult to get a high score ?</li>\n</ul>",
      "rawMarkdown": "Hi all,\n\nThe scoring method has been excellently described, see for instance the input from Tito:\n     https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217158 \n\nBut this still leaves me wondering what the score \"means\" ... Is the 0.56 produced by the best efforts of our admirable best competitors now good, bad or indifferent ?\n\nMy immediate subjective impression is that for software to be useful it should have an accuracy of at least somewhere around 95%.\n\nTherefore, not meaning any offence and in a genuine spirit of desire to learn, I would like to ask:\n\na) What score would be necessary for a cell-labelling software to be \"fit for purpose\" ?\n\nb) Can the scores be improved much by even better machine learning, or are we near the best possible solution ?\n\nc) What are the limiting factors that stop us getting 95%:\n- Problems with quality of the images (unwanted staining, poor illumination etc.) ? \n- Do different organelles simply not differ enough in the images, so that the problem is undecidable with any method ?\n- Does further domain knowledge about the cells and imaging need to be incorporated ?\n- Does the scoring method make it inherently difficult to get a high score ?\n\n",
      "votes": 3
    },
    {
      "id": 1370465,
      "postDate": "2021-06-30T07:24:08.573Z",
      "content": "<p>The more interesting question in my opinion is not the absolute score but the comparison against the human level performance. <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> Do you have more insights, how humans performed on this task and how consistent the classification is?</p>",
      "rawMarkdown": "The more interesting question in my opinion is not the absolute score but the comparison against the human level performance. @cwinsnes Do you have more insights, how humans performed on this task and how consistent the classification is?",
      "votes": 1,
      "replies": [
        {
          "id": 1370730,
          "postDate": "2021-06-30T11:13:08.720Z",
          "content": "<p>We tested how consistent humans are on image level annotation in our <a href=\"https://www.nature.com/articles/nbt.4225\" target=\"_blank\">Nature Biotechnology paper</a> a while back. </p>\n<p>Specifically, you can see the precision and recall of experts in <a href=\"https://www.nature.com/articles/nbt.4225/figures/6\" target=\"_blank\">this</a> figure. The humans experts were pretty good, doing much better than the feature based neural net or the gamers of Project Discovery, but as you can see they're not perfectly consistent.    <br>\n<strong>Note</strong> that this was tested with experts doing a single pass on annotating. While the precision and recall of our experts was less than perfect in that test, we do many control checks to make sure everything is correct before publishing our data.</p>\n<p>I don't have any similar analysis for single cells, like in this challenge.</p>",
          "rawMarkdown": "We tested how consistent humans are on image level annotation in our [Nature Biotechnology paper](https://www.nature.com/articles/nbt.4225) a while back. \n\nSpecifically, you can see the precision and recall of experts in [this](https://www.nature.com/articles/nbt.4225/figures/6) figure. The humans experts were pretty good, doing much better than the feature based neural net or the gamers of Project Discovery, but as you can see they're not perfectly consistent.    \n**Note** that this was tested with experts doing a single pass on annotating. While the precision and recall of our experts was less than perfect in that test, we do many control checks to make sure everything is correct before publishing our data.\n\nI don't have any similar analysis for single cells, like in this challenge.",
          "votes": 3
        },
        {
          "id": 1371262,
          "postDate": "2021-06-30T20:00:24.607Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/Dieter\" target=\"_blank\">@Dieter</a>,</p>\n<p>I agree, having seen the effect of confidence values, that the absolute mAP score does not seem of much interest (unless maybe, if you specifically attach importance to the quality of the confidence allocations). Maybe this effect is just useful for competitions, so the scores don't all cluster together in the 90's ?.</p>\n<p>The underlying precision and recall figures as plotted in the figure cited by <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> seem a more useful comparison.</p>\n<p>Thank you both for the input</p>",
          "rawMarkdown": "Hi @Dieter,\n\nI agree, having seen the effect of confidence values, that the absolute mAP score does not seem of much interest (unless maybe, if you specifically attach importance to the quality of the confidence allocations). Maybe this effect is just useful for competitions, so the scores don't all cluster together in the 90's ?.\n\nThe underlying precision and recall figures as plotted in the figure cited by @cwinsnes seem a more useful comparison.\n\nThank you both for the input"
        },
        {
          "id": 1371991,
          "postDate": "2021-07-01T10:46:39.190Z",
          "content": "<p>Wouldn't the requirement of having good confidence distribution rather be a good thing for a metric like this?</p>",
          "rawMarkdown": "Wouldn't the requirement of having good confidence distribution rather be a good thing for a metric like this?"
        },
        {
          "id": 1372382,
          "postDate": "2021-07-01T16:38:36.143Z",
          "content": "<p>Certainly the confidence distributions are useful functionality, since if well allocated they allow you to do things like identifying which images would be good to have checked by human experts.</p>\n<p>But outside a competition context I guess you could use two separate metrics: one for the correctness of labelling and a second for the quality of the confidence allocations.  Rolling both into one metric, you don't know from the absolute score whether you are looking at an excellent labelling performance with a bad confidence allocation or a poorer labelling performance with a more consistent confidence allocation.</p>",
          "rawMarkdown": "Certainly the confidence distributions are useful functionality, since if well allocated they allow you to do things like identifying which images would be good to have checked by human experts.\n\nBut outside a competition context I guess you could use two separate metrics: one for the correctness of labelling and a second for the quality of the confidence allocations.  Rolling both into one metric, you don't know from the absolute score whether you are looking at an excellent labelling performance with a bad confidence allocation or a poorer labelling performance with a more consistent confidence allocation."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1331042,
      "author_name": "Casper Winsnes",
      "author_url": "",
      "post_date": "2021-06-01T08:35:06.983000",
      "content": "<p>Hey!</p>\n<p>Excellent questions!</p>\n<p>Part of  the reason the score is low is indeed that the method to calculate it is inherently punishing against missing rarer classes which lack enough data to be easily trained to be identified. While we haven't done a full analysis of the results of this challenge yet, we have seen in the previous challenge that the performance on different classes can be very different.</p>\n<p>If the models are great at recognizing the \"easy\" and common classes, they will still be able help us immensely. The model may be able to classify all those images and leave the ones it is unsure on to humans, with the relevant domain knowledge, to manually annotate.</p>\n<p>We don't know if we are reaching the limit of machine learning power for this problem. Models for this problem has come extremely far just in the last few years (much thanks to the Kaggle community!) and there is no reason for us to believe that the problem is undecidable yet.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1357724,
          "author_name": "Philip Hucklesby",
          "author_url": "",
          "post_date": "2021-06-19T23:23:38.287000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> </p>\n<p>Thank you for your response and apologies for not acknowledging sooner - I have been thinking about your answer.</p>\n<p>I see two levels to understanding the mAP score:</p>\n<pre><code>1. Effect on the score by building the mean AP from the AP scores of the classes.\n2. Intuitive understanding of the single AP score of each class.\n</code></pre>\n<p>I guess we all have a pretty good intuitive grasp of the first point since it's just averages and outliers.</p>\n<p>Still 0.56 seemed to me surprisingly low, so I had a closer look at point 2.</p>\n<p>I am sharing this notebook with my examination of the effect:</p>\n<pre><code>https://www.kaggle.com/philiphucklesby/intuition-of-ap-score\n</code></pre>\n<p>My conclusion is:</p>\n<p>The AP score is significantly dependent on the confidence scores attributed to false positives and true positives.</p>\n<p>If the confidence is successfully allocated such that false positives are less confident than true positives, then AP will correspond to the intuitive notion of the accuracy of the test (the 95% mentioned above).</p>\n<p>However, if the confidence is not successfully allocated, then particularly for rare labels, the AP score may be heavily penalised</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1370465,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2021-06-30T07:24:08.573000",
      "content": "<p>The more interesting question in my opinion is not the absolute score but the comparison against the human level performance. <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> Do you have more insights, how humans performed on this task and how consistent the classification is?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1370730,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-06-30T11:13:08.720000",
          "content": "<p>We tested how consistent humans are on image level annotation in our <a href=\"https://www.nature.com/articles/nbt.4225\" target=\"_blank\">Nature Biotechnology paper</a> a while back. </p>\n<p>Specifically, you can see the precision and recall of experts in <a href=\"https://www.nature.com/articles/nbt.4225/figures/6\" target=\"_blank\">this</a> figure. The humans experts were pretty good, doing much better than the feature based neural net or the gamers of Project Discovery, but as you can see they're not perfectly consistent.    <br>\n<strong>Note</strong> that this was tested with experts doing a single pass on annotating. While the precision and recall of our experts was less than perfect in that test, we do many control checks to make sure everything is correct before publishing our data.</p>\n<p>I don't have any similar analysis for single cells, like in this challenge.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1371262,
          "author_name": "Philip Hucklesby",
          "author_url": "",
          "post_date": "2021-06-30T20:00:24.607000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/Dieter\" target=\"_blank\">@Dieter</a>,</p>\n<p>I agree, having seen the effect of confidence values, that the absolute mAP score does not seem of much interest (unless maybe, if you specifically attach importance to the quality of the confidence allocations). Maybe this effect is just useful for competitions, so the scores don't all cluster together in the 90's ?.</p>\n<p>The underlying precision and recall figures as plotted in the figure cited by <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> seem a more useful comparison.</p>\n<p>Thank you both for the input</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1371991,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-07-01T10:46:39.190000",
          "content": "<p>Wouldn't the requirement of having good confidence distribution rather be a good thing for a metric like this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1372382,
          "author_name": "Philip Hucklesby",
          "author_url": "",
          "post_date": "2021-07-01T16:38:36.143000",
          "content": "<p>Certainly the confidence distributions are useful functionality, since if well allocated they allow you to do things like identifying which images would be good to have checked by human experts.</p>\n<p>But outside a competition context I guess you could use two separate metrics: one for the correctness of labelling and a second for the quality of the confidence allocations.  Rolling both into one metric, you don't know from the absolute score whether you are looking at an excellent labelling performance with a bad confidence allocation or a poorer labelling performance with a more consistent confidence allocation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1331042": "Hey!\n\nExcellent questions!\n\nPart of  the reason the score is low is indeed that the method to calculate it is inherently punishing against missing rarer classes which lack enough data to be easily trained to be identified. While we haven't done a full analysis of the results of this challenge yet, we have seen in the previous challenge that the performance on different classes can be very different.\n\nIf the models are great at recognizing the \"easy\" and common classes, they will still be able help us immensely. The model may be able to classify all those images and leave the ones it is unsure on to humans, with the relevant domain knowledge, to manually annotate.\n\nWe don't know if we are reaching the limit of machine learning power for this problem. Models for this problem has come extremely far just in the last few years (much thanks to the Kaggle community!) and there is no reason for us to believe that the problem is undecidable yet.",
    "1326049": "Hi all,\n\nThe scoring method has been excellently described, see for instance the input from Tito:\n     https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217158 \n\nBut this still leaves me wondering what the score \"means\" ... Is the 0.56 produced by the best efforts of our admirable best competitors now good, bad or indifferent ?\n\nMy immediate subjective impression is that for software to be useful it should have an accuracy of at least somewhere around 95%.\n\nTherefore, not meaning any offence and in a genuine spirit of desire to learn, I would like to ask:\n\na) What score would be necessary for a cell-labelling software to be \"fit for purpose\" ?\n\nb) Can the scores be improved much by even better machine learning, or are we near the best possible solution ?\n\nc) What are the limiting factors that stop us getting 95%:\n- Problems with quality of the images (unwanted staining, poor illumination etc.) ? \n- Do different organelles simply not differ enough in the images, so that the problem is undecidable with any method ?\n- Does further domain knowledge about the cells and imaging need to be incorporated ?\n- Does the scoring method make it inherently difficult to get a high score ?\n\n",
    "1370465": "The more interesting question in my opinion is not the absolute score but the comparison against the human level performance. @cwinsnes Do you have more insights, how humans performed on this task and how consistent the classification is?"
  }
}