{
  "id": 514812,
  "title": "Evaluation Metric",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514812",
  "author_name": "",
  "post_date": "2024-06-25T16:24:18.497479200Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Someone please explain evaluation metric as if you are explaining it to 5-year old.<br>\nHow we calculate error for a single study_id. What is significance of 'any_severe' ?</p>",
  "messages": [
    {
      "id": "2889705",
      "postDate": "06/25/2024 16:24:18",
      "content": "<p>Someone please explain evaluation metric as if you are explaining it to 5-year old.<br>\nHow we calculate error for a single study_id. What is significance of 'any_severe' ?</p>",
      "rawMarkdown": "Someone please explain evaluation metric as if you are explaining it to 5-year old.\nHow we calculate error for a single study_id. What is significance of 'any_severe' ?",
      "votes": null
    },
    {
      "id": "2891487",
      "postDate": "06/26/2024 16:51:50",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3548771%2Ffa5dfd33f0f01490846600e99c67e9e0%2F90149Capture0.png?generation=1719420784611884&amp;alt=media\"></p>\n<p>This competition uses weighted log loss to score.  For each prediction, you will assign a probability to each classification.  A perfect response is to assign 100% to the correct label and 0% to the other two.  This would result in a loss of 0 as you answered it perfectly.  We assign weights to these values, meaning correctly matching a Severe prediction is worth 4 points because there are less of them while predicting Normal gives 1 point as that is the majority.  The second part of the loss calculation that is summed to this is a log loss metric on whether you predicted that some part of the spine had Severe Degeneration.  So if the label says some part of the spine is Severe but you do not say any part of the spine is Severe then you get 0 but if you both say there exists a Severe then you get 1.  These two calculations (25 per sample for part 1 and 1 per sample for part 2) are added together.  I wrote this fairly quickly so I am sure I got something wrong but I wanted to help the best I could :D</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3548771%2Ffa5dfd33f0f01490846600e99c67e9e0%2F90149Capture0.png?generation=1719420784611884&alt=media)\n\nThis competition uses weighted log loss to score.  For each prediction, you will assign a probability to each classification.  A perfect response is to assign 100% to the correct label and 0% to the other two.  This would result in a loss of 0 as you answered it perfectly.  We assign weights to these values, meaning correctly matching a Severe prediction is worth 4 points because there are less of them while predicting Normal gives 1 point as that is the majority.  The second part of the loss calculation that is summed to this is a log loss metric on whether you predicted that some part of the spine had Severe Degeneration.  So if the label says some part of the spine is Severe but you do not say any part of the spine is Severe then you get 0 but if you both say there exists a Severe then you get 1.  These two calculations (25 per sample for part 1 and 1 per sample for part 2) are added together.  I wrote this fairly quickly so I am sure I got something wrong but I wanted to help the best I could :D",
      "votes": null
    },
    {
      "id": "2891564",
      "postDate": "06/26/2024 17:31:49",
      "content": "<p>But we do not have explicit any_severe column so how to account for this in actual loss function and I also read some discussion that there are 10,5,10 labels and all are weighted differently. Can you throw some light in that direction ?</p>",
      "rawMarkdown": "But we do not have explicit any_severe column so how to account for this in actual loss function and I also read some discussion that there are 10,5,10 labels and all are weighted differently. Can you throw some light in that direction ?",
      "votes": null
    },
    {
      "id": "2893663",
      "postDate": "06/28/2024 01:27:15",
      "content": "<p>I am not sure what the 10,5,10 labels mean but for the first part of your question I can answer that.  How I interpret it is we have two loss functions for this competition, the first is log loss to evaluate your prediction for each of the 25 tasks which checks how well you classify those.  The second part is to check if there is ANY severe degeneration in the sample and to also check that you predicted that there is ANY severe degeneration.  This is not done through a column but instead it looks through the 25 predictions you made and checks for your max probability in the severe column.  It is slightly flawed as some other Kagglers pointed out but that is how the second part is calculated.  The two scores are then combined.  I hope that helps :)</p>",
      "rawMarkdown": "I am not sure what the 10,5,10 labels mean but for the first part of your question I can answer that.  How I interpret it is we have two loss functions for this competition, the first is log loss to evaluate your prediction for each of the 25 tasks which checks how well you classify those.  The second part is to check if there is ANY severe degeneration in the sample and to also check that you predicted that there is ANY severe degeneration.  This is not done through a column but instead it looks through the 25 predictions you made and checks for your max probability in the severe column.  It is slightly flawed as some other Kagglers pointed out but that is how the second part is calculated.  The two scores are then combined.  I hope that helps :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2891487,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "06/26/2024 16:51:50",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3548771%2Ffa5dfd33f0f01490846600e99c67e9e0%2F90149Capture0.png?generation=1719420784611884&amp;alt=media\"></p>\n<p>This competition uses weighted log loss to score.  For each prediction, you will assign a probability to each classification.  A perfect response is to assign 100% to the correct label and 0% to the other two.  This would result in a loss of 0 as you answered it perfectly.  We assign weights to these values, meaning correctly matching a Severe prediction is worth 4 points because there are less of them while predicting Normal gives 1 point as that is the majority.  The second part of the loss calculation that is summed to this is a log loss metric on whether you predicted that some part of the spine had Severe Degeneration.  So if the label says some part of the spine is Severe but you do not say any part of the spine is Severe then you get 0 but if you both say there exists a Severe then you get 1.  These two calculations (25 per sample for part 1 and 1 per sample for part 2) are added together.  I wrote this fairly quickly so I am sure I got something wrong but I wanted to help the best I could :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 2891564,
          "author_name": "rohitchaudhari25",
          "author_url": "",
          "post_date": "06/26/2024 17:31:49",
          "content": "<p>But we do not have explicit any_severe column so how to account for this in actual loss function and I also read some discussion that there are 10,5,10 labels and all are weighted differently. Can you throw some light in that direction ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2893663,
              "author_name": "connorjd",
              "author_url": "",
              "post_date": "06/28/2024 01:27:15",
              "content": "<p>I am not sure what the 10,5,10 labels mean but for the first part of your question I can answer that.  How I interpret it is we have two loss functions for this competition, the first is log loss to evaluate your prediction for each of the 25 tasks which checks how well you classify those.  The second part is to check if there is ANY severe degeneration in the sample and to also check that you predicted that there is ANY severe degeneration.  This is not done through a column but instead it looks through the 25 predictions you made and checks for your max probability in the severe column.  It is slightly flawed as some other Kagglers pointed out but that is how the second part is calculated.  The two scores are then combined.  I hope that helps :)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2889705": "Someone please explain evaluation metric as if you are explaining it to 5-year old.\nHow we calculate error for a single study_id. What is significance of 'any_severe' ?",
    "2891487": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3548771%2Ffa5dfd33f0f01490846600e99c67e9e0%2F90149Capture0.png?generation=1719420784611884&alt=media)\n\nThis competition uses weighted log loss to score.  For each prediction, you will assign a probability to each classification.  A perfect response is to assign 100% to the correct label and 0% to the other two.  This would result in a loss of 0 as you answered it perfectly.  We assign weights to these values, meaning correctly matching a Severe prediction is worth 4 points because there are less of them while predicting Normal gives 1 point as that is the majority.  The second part of the loss calculation that is summed to this is a log loss metric on whether you predicted that some part of the spine had Severe Degeneration.  So if the label says some part of the spine is Severe but you do not say any part of the spine is Severe then you get 0 but if you both say there exists a Severe then you get 1.  These two calculations (25 per sample for part 1 and 1 per sample for part 2) are added together.  I wrote this fairly quickly so I am sure I got something wrong but I wanted to help the best I could :D",
    "2891564": "But we do not have explicit any_severe column so how to account for this in actual loss function and I also read some discussion that there are 10,5,10 labels and all are weighted differently. Can you throw some light in that direction ?",
    "2893663": "I am not sure what the 10,5,10 labels mean but for the first part of your question I can answer that.  How I interpret it is we have two loss functions for this competition, the first is log loss to evaluate your prediction for each of the 25 tasks which checks how well you classify those.  The second part is to check if there is ANY severe degeneration in the sample and to also check that you predicted that there is ANY severe degeneration.  This is not done through a column but instead it looks through the 25 predictions you made and checks for your max probability in the severe column.  It is slightly flawed as some other Kagglers pointed out but that is how the second part is calculated.  The two scores are then combined.  I hope that helps :)"
  },
  "source": "meta"
}