{
  "id": 414165,
  "title": "Metric could be the Key!!",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/414165",
  "author_name": "Murugesan Narayanaswamy",
  "post_date": "2023-05-31T15:52:23.967000",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Understanding the metric seems very crucial in this competition. In the evaluation page, it is mentioned that <em>\"In the ground truth, <strong>at most one event class</strong> has a non-zero value for each Id\"</em>…</p>\n<p>So, I was trying to build models that are not multilabel -means whatever model I constructed, I tried to ensure that it is normalized so that there is at most only one event class that has label 1.</p>\n<p>However, the metric, average_precision_score functions in a different way</p>\n<p>Consider the following example of ground truth and prediction scores: </p>\n<p>*ground_truth = np.array([[0,0,1], [1,0,0], [0,1,0], [0,0,0]])<br>\ny_true=pd.DataFrame(ground_truth, columns=['StartHesitation','Turn','Walking'])</p>\n<p>scores=np.array([[0.98,0.97,0.51],  [1,0.1,.1], [0.3,1,0], [0.4,0.5,.49]])<br>\ny_scores=pd.DataFrame(scores, columns=['StartHesitation','Turn','Walking'])</p>\n<p>print(\"y_true\\n\",y_true)<br>\nprint(\"y_scores\\n\",y_scores)<br>\naverage_precision_score(y_true, y_scores)*</p>\n<p>The average precision score for the above case is 1</p>\n<p>For the first sample, the ground truth says the true label is 'Walking'. But the prediction scores indicate that probability of 'StartHesitation' is 0.98 and probability of 'Turn' is 0.97 and the probability of true label 'Walking' is only 0.51 !! But still the precision score is 100%!!</p>\n<p>If we normalize the above probabilities, we will get completely wrong scores!</p>\n<p>Suppose the prediction of our model for the first sample for 'Walking' event is 0.5 instead of 0.51 - then guess what will be the score! The score will be reduced 20% - the average precision score will be 83%</p>\n<p>Though the requirement is that there could be at most only one event possible for a sample, and even though our model above predicted more than one, and wrong events, at above 95% probability, the average precision score will still be 100%!!</p>\n<p>Similarly, the ground truth says that fourth sample is None - there is no event that occurred with that reading, but our model above predicts around 50% probability of any of the three events - still we get 100% score! (as long as score is less than 0.5, it does not matter)</p>\n<p>So, even though ground truth is that at most only one event can be non-zero, our model should ignore that fact for predictions!</p>\n<p><a href=\"https://www.kaggle.com/code/murugesann/average-precision-score/notebook\" target=\"_blank\">https://www.kaggle.com/code/murugesann/average-precision-score/notebook</a></p>",
  "messages": [
    {
      "id": 2282484,
      "postDate": "2023-05-31T15:52:23.967Z",
      "content": "<p>Understanding the metric seems very crucial in this competition. In the evaluation page, it is mentioned that <em>\"In the ground truth, <strong>at most one event class</strong> has a non-zero value for each Id\"</em>…</p>\n<p>So, I was trying to build models that are not multilabel -means whatever model I constructed, I tried to ensure that it is normalized so that there is at most only one event class that has label 1.</p>\n<p>However, the metric, average_precision_score functions in a different way</p>\n<p>Consider the following example of ground truth and prediction scores: </p>\n<p>*ground_truth = np.array([[0,0,1], [1,0,0], [0,1,0], [0,0,0]])<br>\ny_true=pd.DataFrame(ground_truth, columns=['StartHesitation','Turn','Walking'])</p>\n<p>scores=np.array([[0.98,0.97,0.51],  [1,0.1,.1], [0.3,1,0], [0.4,0.5,.49]])<br>\ny_scores=pd.DataFrame(scores, columns=['StartHesitation','Turn','Walking'])</p>\n<p>print(\"y_true\\n\",y_true)<br>\nprint(\"y_scores\\n\",y_scores)<br>\naverage_precision_score(y_true, y_scores)*</p>\n<p>The average precision score for the above case is 1</p>\n<p>For the first sample, the ground truth says the true label is 'Walking'. But the prediction scores indicate that probability of 'StartHesitation' is 0.98 and probability of 'Turn' is 0.97 and the probability of true label 'Walking' is only 0.51 !! But still the precision score is 100%!!</p>\n<p>If we normalize the above probabilities, we will get completely wrong scores!</p>\n<p>Suppose the prediction of our model for the first sample for 'Walking' event is 0.5 instead of 0.51 - then guess what will be the score! The score will be reduced 20% - the average precision score will be 83%</p>\n<p>Though the requirement is that there could be at most only one event possible for a sample, and even though our model above predicted more than one, and wrong events, at above 95% probability, the average precision score will still be 100%!!</p>\n<p>Similarly, the ground truth says that fourth sample is None - there is no event that occurred with that reading, but our model above predicts around 50% probability of any of the three events - still we get 100% score! (as long as score is less than 0.5, it does not matter)</p>\n<p>So, even though ground truth is that at most only one event can be non-zero, our model should ignore that fact for predictions!</p>\n<p><a href=\"https://www.kaggle.com/code/murugesann/average-precision-score/notebook\" target=\"_blank\">https://www.kaggle.com/code/murugesann/average-precision-score/notebook</a></p>",
      "rawMarkdown": "Understanding the metric seems very crucial in this competition. In the evaluation page, it is mentioned that *\"In the ground truth, **at most one event class** has a non-zero value for each Id\"*...\n\nSo, I was trying to build models that are not multilabel -means whatever model I constructed, I tried to ensure that it is normalized so that there is at most only one event class that has label 1.\n\nHowever, the metric, average_precision_score functions in a different way\n\nConsider the following example of ground truth and prediction scores: \n\n*ground_truth = np.array([[0,0,1], [1,0,0], [0,1,0], [0,0,0]])\ny_true=pd.DataFrame(ground_truth, columns=['StartHesitation','Turn','Walking'])\n\nscores=np.array([[0.98,0.97,0.51],  [1,0.1,.1], [0.3,1,0], [0.4,0.5,.49]])\ny_scores=pd.DataFrame(scores, columns=['StartHesitation','Turn','Walking'])\n \n\nprint(\"y_true\\n\",y_true)\nprint(\"y_scores\\n\",y_scores)\naverage_precision_score(y_true, y_scores)*\n\nThe average precision score for the above case is 1\n\nFor the first sample, the ground truth says the true label is 'Walking'. But the prediction scores indicate that probability of 'StartHesitation' is 0.98 and probability of 'Turn' is 0.97 and the probability of true label 'Walking' is only 0.51 !! But still the precision score is 100%!!\n\nIf we normalize the above probabilities, we will get completely wrong scores!\n \nSuppose the prediction of our model for the first sample for 'Walking' event is 0.5 instead of 0.51 - then guess what will be the score! The score will be reduced 20% - the average precision score will be 83%\n\nThough the requirement is that there could be at most only one event possible for a sample, and even though our model above predicted more than one, and wrong events, at above 95% probability, the average precision score will still be 100%!!\n\nSimilarly, the ground truth says that fourth sample is None - there is no event that occurred with that reading, but our model above predicts around 50% probability of any of the three events - still we get 100% score! (as long as score is less than 0.5, it does not matter)\n\nSo, even though ground truth is that at most only one event can be non-zero, our model should ignore that fact for predictions!\n\nhttps://www.kaggle.com/code/murugesann/average-precision-score/notebook\n",
      "votes": 1
    },
    {
      "id": 2282727,
      "postDate": "2023-05-31T19:27:19.590Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2283084,
          "postDate": "2023-06-01T04:04:01.970Z",
          "content": "<p>This is a multilabel, multiclass problem. However, the statement in the evaluation page that says 'at most one event shall be non-zero' leads to confusion. There can be more than one event that can be non-zero seems to be the fact. Also, when it is mentioned that the probability scores between  1 and 0 can also be submitted, it led to the usual assumption that these are multi-class probabilities as against multi-label. The implication is that I can predict two events (depending on the data - assumption is that there can be two identical rows occurring in the dataset - one for each event) for a sample reading as against the construction of dataset and sample submission files.</p>",
          "rawMarkdown": "This is a multilabel, multiclass problem. However, the statement in the evaluation page that says 'at most one event shall be non-zero' leads to confusion. There can be more than one event that can be non-zero seems to be the fact. Also, when it is mentioned that the probability scores between  1 and 0 can also be submitted, it led to the usual assumption that these are multi-class probabilities as against multi-label. The implication is that I can predict two events (depending on the data - assumption is that there can be two identical rows occurring in the dataset - one for each event) for a sample reading as against the construction of dataset and sample submission files.\n",
          "votes": -1,
          "replies": [
            {
              "id": 2283534,
              "postDate": "2023-06-01T10:24:43.653Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 2285742,
              "postDate": "2023-06-03T00:25:21.240Z",
              "content": "<p>Hi, sorry for delay in replying. </p>\n<p>First of all, if this is a multi-class problem, then we can't use the evaluation metric - average_precision_score. This is stated in the sklearn documentation as I have already pointed out in a different post.</p>\n<p>The evaluation page says that ground truth is with at most one event greater than zero, implying multiclass. But it allows probabilities that does not sum up to one across events, thereby usage of the given metric. For the given metric to be used, we have to train separate estimators for each label. When we train separate estimators for each label, then it is a multi-label problem and thereby we can use given metric.</p>\n<p>When the ground truth is labeled with at most one event greater than zero, how it can be multilabel? Because, though the objective is multiclass, the data used for predictions are not amenable for multiclass. For example, I was analyzing the results of a sample - the 'None\" label was predicted with 75% probability while another event was predicted just 25% probability. But the ground truth is 1 for that event for that  sample. For values ranging from 25% probabilities to 75% probabilities, the ground truth is one. This means those readings can be labelled either 'None' or 'Turn' for example.  In other words, a same set of reading can occur in any of the events. For that reading, the labelling should strictly be [1,1,1,0] for example! Though this is not labelled that way, since we have \"nearly\" duplicate rows with different lablels, and since we train separate estimators for each label, it leads to multilabel problem.</p>\n<p>With regard to using regressors (with which I reached top 6%), I don't think it is correct. The competition host should not accept a winning model with regression - though I have explained why there are reasons that lead to its usage in another post. </p>\n<p>We cannot use regressor for this problem because we are not interested in predicting the degree of each event (from zero to one) as far as the evaluation metric is concerned. The evaluation metric expects probabilities and not regression scores.  With  just regression scores, even ignoring the metric, the scores are not interpretable since we are not predicting the classes as stated in evaluation page of this competition, instead we are predicting degree of each FOG event for each reading. Hence, my assumption is that competition hosts will reject notebooks with regression scores that are directly presented to the metric without converting them to probabilities</p>\n<p>We can use regression scores as used in the published notebook ONLY if these regression scores are converted into probabilities through some other post processing. The process of converting regression scores into probabilities is what classification is all about :-) </p>\n<p>The logistic regression first calculates the regression scores (as done in these notebooks) and then converts them to probabilities using logloss optimization.</p>\n<p>(Note: Even though sklearn documentation says that logistic regression is well calibrated when compared to ensembles, calibration will further improve logistic regression scores - if the dataset size is small, calibration is required even for logistic regression)</p>",
              "rawMarkdown": "Hi, sorry for delay in replying. \n\nFirst of all, if this is a multi-class problem, then we can't use the evaluation metric - average_precision_score. This is stated in the sklearn documentation as I have already pointed out in a different post.\n\nThe evaluation page says that ground truth is with at most one event greater than zero, implying multiclass. But it allows probabilities that does not sum up to one across events, thereby usage of the given metric. For the given metric to be used, we have to train separate estimators for each label. When we train separate estimators for each label, then it is a multi-label problem and thereby we can use given metric.\n\nWhen the ground truth is labeled with at most one event greater than zero, how it can be multilabel? Because, though the objective is multiclass, the data used for predictions are not amenable for multiclass. For example, I was analyzing the results of a sample - the 'None\" label was predicted with 75% probability while another event was predicted just 25% probability. But the ground truth is 1 for that event for that  sample. For values ranging from 25% probabilities to 75% probabilities, the ground truth is one. This means those readings can be labelled either 'None' or 'Turn' for example.  In other words, a same set of reading can occur in any of the events. For that reading, the labelling should strictly be [1,1,1,0] for example! Though this is not labelled that way, since we have \"nearly\" duplicate rows with different lablels, and since we train separate estimators for each label, it leads to multilabel problem.\n\nWith regard to using regressors (with which I reached top 6%), I don't think it is correct. The competition host should not accept a winning model with regression - though I have explained why there are reasons that lead to its usage in another post. \n\nWe cannot use regressor for this problem because we are not interested in predicting the degree of each event (from zero to one) as far as the evaluation metric is concerned. The evaluation metric expects probabilities and not regression scores.  With  just regression scores, even ignoring the metric, the scores are not interpretable since we are not predicting the classes as stated in evaluation page of this competition, instead we are predicting degree of each FOG event for each reading. Hence, my assumption is that competition hosts will reject notebooks with regression scores that are directly presented to the metric without converting them to probabilities\n\nWe can use regression scores as used in the published notebook ONLY if these regression scores are converted into probabilities through some other post processing. The process of converting regression scores into probabilities is what classification is all about :-) \n\nThe logistic regression first calculates the regression scores (as done in these notebooks) and then converts them to probabilities using logloss optimization.\n\n(Note: Even though sklearn documentation says that logistic regression is well calibrated when compared to ensembles, calibration will further improve logistic regression scores - if the dataset size is small, calibration is required even for logistic regression)"
            },
            {
              "id": 2287946,
              "postDate": "2023-06-05T04:15:39.457Z",
              "content": "<p>The whole point and challenge behind this competition seems to be the nature of dataset. Unlike normal classification datasets, here the label values are continuous in nature..</p>\n<p>I had asked competition host whether and will they accept regression based models, but no response. My point is that the multioutput based raw regression scores cannot give any information about the probability of which class, any particular reading belongs to….which I think they want us to predict.  If top scores are based on these models, they must be doing some post processing to convert the scores to probabilities or may be they are based on deep learning based lstm, transformer models.</p>\n<p>Anyway, I could achieve same scores with multiclass models - the metric does  not differentiate the model from which the scores have arisen. AP metric cannot restrict that at most one alone can be positive label. </p>",
              "rawMarkdown": "The whole point and challenge behind this competition seems to be the nature of dataset. Unlike normal classification datasets, here the label values are continuous in nature..\n\nI had asked competition host whether and will they accept regression based models, but no response. My point is that the multioutput based raw regression scores cannot give any information about the probability of which class, any particular reading belongs to....which I think they want us to predict.  If top scores are based on these models, they must be doing some post processing to convert the scores to probabilities or may be they are based on deep learning based lstm, transformer models.\n\nAnyway, I could achieve same scores with multiclass models - the metric does  not differentiate the model from which the scores have arisen. AP metric cannot restrict that at most one alone can be positive label. "
            },
            {
              "id": 2287973,
              "postDate": "2023-06-05T04:51:57.983Z",
              "content": "<p>Hi, interesting discussions.  From my perspective, the goal is to achieve the best average P-R score (with the most generalizable and least overfitting). If that can be done using regression based models, that is totally fine.</p>",
              "rawMarkdown": "Hi, interesting discussions.  From my perspective, the goal is to achieve the best average P-R score (with the most generalizable and least overfitting). If that can be done using regression based models, that is totally fine.",
              "votes": 1
            },
            {
              "id": 2288008,
              "postDate": "2023-06-05T05:50:09.373Z",
              "content": "<p>Hi, the multi-output regression models are giving linear scores that indicate degree of any FOG event,  independent of other events, and they are not confidence scores. These scores cannot be directly fed into Average Precision Metric which expects confidence / probability scores. For a particular accelerometer reading, if we want to know whether it indicates onset of an event, and if so, what is the probability of that event being any of the three, then regression scores do not convey any meaningful information. Metric does not know whether it is probability scores or regression scores, and so, even if we get high AP score from a regression model, how can that model output be used in practice? The output of regression models are different from the outputs of classification model. Moreover, when similar but more meaningful information can be obtained by designing a multiclass model, why resort to regression models? And hence my question: if a regression model gives higher AP score than a multiclass classifier model, will that model be accepted?</p>\n<p>Note: Sklearn says outputs of multiclass models cannot be used for calculating Average Precision Score. But that is because the implementation does not restrict multiclass outputs - it does not check whether the probabilities sum to one. But in our case, we don't require the probabilities to sum to  one - a very peculiar use case which neither truly belongs to multiclass nor to multilabel</p>\n<p>Yesterday, an analysis of predictions of  train data by a model showed this case: A particular reading was predicted to be 'None' - not an event, with a probability of 75% while the probability predicted for StartHesitation was around 10% and Turn was 15% or so. But the ground truth labelled this event as Start Hesitation while my model predicted it to be Turn event  (I added some post processing to my model which resulted in both StartHesitation and Turn labelled as 1:-)  </p>\n<p>For the above case, if I give scores as one for both StartHesitation and Turn events, the metric won't know - it processes each event separately. This is what sklearn documentation says I guess (final score after post processing does not improve anyway:-( )</p>\n<p>The above case does not seem to be a data anomaly or outlier - it could be a valid case because the readings at the start of the events are not much distinguishable by models</p>\n<p>If the above readings (0.10 and 0.15) are from regression scores, the interpretation will be that this reading has 0.1 degree of StartHesitation event and 0.15 degree of Turn event. And I don't think that would be meaningful</p>",
              "rawMarkdown": "Hi, the multi-output regression models are giving linear scores that indicate degree of any FOG event,  independent of other events, and they are not confidence scores. These scores cannot be directly fed into Average Precision Metric which expects confidence / probability scores. For a particular accelerometer reading, if we want to know whether it indicates onset of an event, and if so, what is the probability of that event being any of the three, then regression scores do not convey any meaningful information. Metric does not know whether it is probability scores or regression scores, and so, even if we get high AP score from a regression model, how can that model output be used in practice? The output of regression models are different from the outputs of classification model. Moreover, when similar but more meaningful information can be obtained by designing a multiclass model, why resort to regression models? And hence my question: if a regression model gives higher AP score than a multiclass classifier model, will that model be accepted?\n\nNote: Sklearn says outputs of multiclass models cannot be used for calculating Average Precision Score. But that is because the implementation does not restrict multiclass outputs - it does not check whether the probabilities sum to one. But in our case, we don't require the probabilities to sum to  one - a very peculiar use case which neither truly belongs to multiclass nor to multilabel\n\nYesterday, an analysis of predictions of  train data by a model showed this case: A particular reading was predicted to be 'None' - not an event, with a probability of 75% while the probability predicted for StartHesitation was around 10% and Turn was 15% or so. But the ground truth labelled this event as Start Hesitation while my model predicted it to be Turn event  (I added some post processing to my model which resulted in both StartHesitation and Turn labelled as 1:-)  \n\nFor the above case, if I give scores as one for both StartHesitation and Turn events, the metric won't know - it processes each event separately. This is what sklearn documentation says I guess (final score after post processing does not improve anyway:-( )\n\nThe above case does not seem to be a data anomaly or outlier - it could be a valid case because the readings at the start of the events are not much distinguishable by models\n\nIf the above readings (0.10 and 0.15) are from regression scores, the interpretation will be that this reading has 0.1 degree of StartHesitation event and 0.15 degree of Turn event. And I don't think that would be meaningful"
            },
            {
              "id": 2289971,
              "postDate": "2023-06-06T13:20:42.873Z",
              "content": "<p>Do you think calibration of regression scores converts them into probabilities? I don't think so..!!<br>\n (by the way, I am sure the top notebooks - prize contenders of this competition are not using regression scores - else competition host would have taken this seriously)</p>",
              "rawMarkdown": "Do you think calibration of regression scores converts them into probabilities? I don't think so..!!\n (by the way, I am sure the top notebooks - prize contenders of this competition are not using regression scores - else competition host would have taken this seriously)"
            },
            {
              "id": 2304727,
              "postDate": "2023-06-16T07:42:14.803Z",
              "content": "<p>Regression based scored are too easy to interpret - they are from 0 to 1 and thereby they show probability of some classes (events). If regression gives numbers above 1 or lower than 0, they can be rounded to 1 and 0 respectively, I think.</p>",
              "rawMarkdown": "Regression based scored are too easy to interpret - they are from 0 to 1 and thereby they show probability of some classes (events). If regression gives numbers above 1 or lower than 0, they can be rounded to 1 and 0 respectively, I think."
            }
          ]
        },
        {
          "id": 2283594,
          "postDate": "2023-06-01T11:33:04.447Z",
          "content": "<p>Can you elaborate what do you mean by \"plotting a histogram for each score type\", and how does it related to the calibrator? Never used the calibratedclassifier, and finding it hard to understand its relationship with this problem. Regarding your point on normalization, I did notice it had no effect, but what would you suggest to introduce to it?</p>",
          "rawMarkdown": "Can you elaborate what do you mean by \"plotting a histogram for each score type\", and how does it related to the calibrator? Never used the calibratedclassifier, and finding it hard to understand its relationship with this problem. Regarding your point on normalization, I did notice it had no effect, but what would you suggest to introduce to it?",
          "replies": [
            {
              "id": 2283821,
              "postDate": "2023-06-01T14:16:09.380Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2282727,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-31T19:27:19.590000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2283084,
          "author_name": "Murugesan Narayanaswamy",
          "author_url": "",
          "post_date": "2023-06-01T04:04:01.970000",
          "content": "<p>This is a multilabel, multiclass problem. However, the statement in the evaluation page that says 'at most one event shall be non-zero' leads to confusion. There can be more than one event that can be non-zero seems to be the fact. Also, when it is mentioned that the probability scores between  1 and 0 can also be submitted, it led to the usual assumption that these are multi-class probabilities as against multi-label. The implication is that I can predict two events (depending on the data - assumption is that there can be two identical rows occurring in the dataset - one for each event) for a sample reading as against the construction of dataset and sample submission files.</p>",
          "votes": -1,
          "replies": [
            {
              "id": 2283534,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-01T10:24:43.653000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2285742,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-06-03T00:25:21.240000",
              "content": "<p>Hi, sorry for delay in replying. </p>\n<p>First of all, if this is a multi-class problem, then we can't use the evaluation metric - average_precision_score. This is stated in the sklearn documentation as I have already pointed out in a different post.</p>\n<p>The evaluation page says that ground truth is with at most one event greater than zero, implying multiclass. But it allows probabilities that does not sum up to one across events, thereby usage of the given metric. For the given metric to be used, we have to train separate estimators for each label. When we train separate estimators for each label, then it is a multi-label problem and thereby we can use given metric.</p>\n<p>When the ground truth is labeled with at most one event greater than zero, how it can be multilabel? Because, though the objective is multiclass, the data used for predictions are not amenable for multiclass. For example, I was analyzing the results of a sample - the 'None\" label was predicted with 75% probability while another event was predicted just 25% probability. But the ground truth is 1 for that event for that  sample. For values ranging from 25% probabilities to 75% probabilities, the ground truth is one. This means those readings can be labelled either 'None' or 'Turn' for example.  In other words, a same set of reading can occur in any of the events. For that reading, the labelling should strictly be [1,1,1,0] for example! Though this is not labelled that way, since we have \"nearly\" duplicate rows with different lablels, and since we train separate estimators for each label, it leads to multilabel problem.</p>\n<p>With regard to using regressors (with which I reached top 6%), I don't think it is correct. The competition host should not accept a winning model with regression - though I have explained why there are reasons that lead to its usage in another post. </p>\n<p>We cannot use regressor for this problem because we are not interested in predicting the degree of each event (from zero to one) as far as the evaluation metric is concerned. The evaluation metric expects probabilities and not regression scores.  With  just regression scores, even ignoring the metric, the scores are not interpretable since we are not predicting the classes as stated in evaluation page of this competition, instead we are predicting degree of each FOG event for each reading. Hence, my assumption is that competition hosts will reject notebooks with regression scores that are directly presented to the metric without converting them to probabilities</p>\n<p>We can use regression scores as used in the published notebook ONLY if these regression scores are converted into probabilities through some other post processing. The process of converting regression scores into probabilities is what classification is all about :-) </p>\n<p>The logistic regression first calculates the regression scores (as done in these notebooks) and then converts them to probabilities using logloss optimization.</p>\n<p>(Note: Even though sklearn documentation says that logistic regression is well calibrated when compared to ensembles, calibration will further improve logistic regression scores - if the dataset size is small, calibration is required even for logistic regression)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2287946,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-06-05T04:15:39.457000",
              "content": "<p>The whole point and challenge behind this competition seems to be the nature of dataset. Unlike normal classification datasets, here the label values are continuous in nature..</p>\n<p>I had asked competition host whether and will they accept regression based models, but no response. My point is that the multioutput based raw regression scores cannot give any information about the probability of which class, any particular reading belongs to….which I think they want us to predict.  If top scores are based on these models, they must be doing some post processing to convert the scores to probabilities or may be they are based on deep learning based lstm, transformer models.</p>\n<p>Anyway, I could achieve same scores with multiclass models - the metric does  not differentiate the model from which the scores have arisen. AP metric cannot restrict that at most one alone can be positive label. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2287973,
              "author_name": "Jeff Hausdorff",
              "author_url": "",
              "post_date": "2023-06-05T04:51:57.983000",
              "content": "<p>Hi, interesting discussions.  From my perspective, the goal is to achieve the best average P-R score (with the most generalizable and least overfitting). If that can be done using regression based models, that is totally fine.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2288008,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-06-05T05:50:09.373000",
              "content": "<p>Hi, the multi-output regression models are giving linear scores that indicate degree of any FOG event,  independent of other events, and they are not confidence scores. These scores cannot be directly fed into Average Precision Metric which expects confidence / probability scores. For a particular accelerometer reading, if we want to know whether it indicates onset of an event, and if so, what is the probability of that event being any of the three, then regression scores do not convey any meaningful information. Metric does not know whether it is probability scores or regression scores, and so, even if we get high AP score from a regression model, how can that model output be used in practice? The output of regression models are different from the outputs of classification model. Moreover, when similar but more meaningful information can be obtained by designing a multiclass model, why resort to regression models? And hence my question: if a regression model gives higher AP score than a multiclass classifier model, will that model be accepted?</p>\n<p>Note: Sklearn says outputs of multiclass models cannot be used for calculating Average Precision Score. But that is because the implementation does not restrict multiclass outputs - it does not check whether the probabilities sum to one. But in our case, we don't require the probabilities to sum to  one - a very peculiar use case which neither truly belongs to multiclass nor to multilabel</p>\n<p>Yesterday, an analysis of predictions of  train data by a model showed this case: A particular reading was predicted to be 'None' - not an event, with a probability of 75% while the probability predicted for StartHesitation was around 10% and Turn was 15% or so. But the ground truth labelled this event as Start Hesitation while my model predicted it to be Turn event  (I added some post processing to my model which resulted in both StartHesitation and Turn labelled as 1:-)  </p>\n<p>For the above case, if I give scores as one for both StartHesitation and Turn events, the metric won't know - it processes each event separately. This is what sklearn documentation says I guess (final score after post processing does not improve anyway:-( )</p>\n<p>The above case does not seem to be a data anomaly or outlier - it could be a valid case because the readings at the start of the events are not much distinguishable by models</p>\n<p>If the above readings (0.10 and 0.15) are from regression scores, the interpretation will be that this reading has 0.1 degree of StartHesitation event and 0.15 degree of Turn event. And I don't think that would be meaningful</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2289971,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-06-06T13:20:42.873000",
              "content": "<p>Do you think calibration of regression scores converts them into probabilities? I don't think so..!!<br>\n (by the way, I am sure the top notebooks - prize contenders of this competition are not using regression scores - else competition host would have taken this seriously)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2304727,
              "author_name": "Sdelnikov Aleksandr",
              "author_url": "",
              "post_date": "2023-06-16T07:42:14.803000",
              "content": "<p>Regression based scored are too easy to interpret - they are from 0 to 1 and thereby they show probability of some classes (events). If regression gives numbers above 1 or lower than 0, they can be rounded to 1 and 0 respectively, I think.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2283594,
          "author_name": "Yijie Xu",
          "author_url": "",
          "post_date": "2023-06-01T11:33:04.447000",
          "content": "<p>Can you elaborate what do you mean by \"plotting a histogram for each score type\", and how does it related to the calibrator? Never used the calibratedclassifier, and finding it hard to understand its relationship with this problem. Regarding your point on normalization, I did notice it had no effect, but what would you suggest to introduce to it?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2283821,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-01T14:16:09.380000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2282484": "Understanding the metric seems very crucial in this competition. In the evaluation page, it is mentioned that *\"In the ground truth, **at most one event class** has a non-zero value for each Id\"*...\n\nSo, I was trying to build models that are not multilabel -means whatever model I constructed, I tried to ensure that it is normalized so that there is at most only one event class that has label 1.\n\nHowever, the metric, average_precision_score functions in a different way\n\nConsider the following example of ground truth and prediction scores: \n\n*ground_truth = np.array([[0,0,1], [1,0,0], [0,1,0], [0,0,0]])\ny_true=pd.DataFrame(ground_truth, columns=['StartHesitation','Turn','Walking'])\n\nscores=np.array([[0.98,0.97,0.51],  [1,0.1,.1], [0.3,1,0], [0.4,0.5,.49]])\ny_scores=pd.DataFrame(scores, columns=['StartHesitation','Turn','Walking'])\n \n\nprint(\"y_true\\n\",y_true)\nprint(\"y_scores\\n\",y_scores)\naverage_precision_score(y_true, y_scores)*\n\nThe average precision score for the above case is 1\n\nFor the first sample, the ground truth says the true label is 'Walking'. But the prediction scores indicate that probability of 'StartHesitation' is 0.98 and probability of 'Turn' is 0.97 and the probability of true label 'Walking' is only 0.51 !! But still the precision score is 100%!!\n\nIf we normalize the above probabilities, we will get completely wrong scores!\n \nSuppose the prediction of our model for the first sample for 'Walking' event is 0.5 instead of 0.51 - then guess what will be the score! The score will be reduced 20% - the average precision score will be 83%\n\nThough the requirement is that there could be at most only one event possible for a sample, and even though our model above predicted more than one, and wrong events, at above 95% probability, the average precision score will still be 100%!!\n\nSimilarly, the ground truth says that fourth sample is None - there is no event that occurred with that reading, but our model above predicts around 50% probability of any of the three events - still we get 100% score! (as long as score is less than 0.5, it does not matter)\n\nSo, even though ground truth is that at most only one event can be non-zero, our model should ignore that fact for predictions!\n\nhttps://www.kaggle.com/code/murugesann/average-precision-score/notebook\n",
    "2282727": ""
  }
}