{
  "id": 413638,
  "title": "In Top 6% of LB with a wrong model ?l!  - Classification or Regression?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638",
  "author_name": "",
  "post_date": "2023-05-29T14:48:50.642576200Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>My notebook is currently in 78th position out of 1242 teams! But I wonder whether the approach of my notebook is correct!</p>\n<p>I joined this competition 10 days ago and started with the top scoring notebook published (thanks to the notebook by Nicholas Gray <a href=\"https://www.kaggle.com/nickcgray\" target=\"_blank\">@nickcgray</a> - <a href=\"https://www.kaggle.com/code/nickcgray/gait-prediction\" target=\"_blank\">https://www.kaggle.com/code/nickcgray/gait-prediction</a> - it has saved tremendous amount of time). After some finetuning, I could reach a score of 0.311 putting me into top 6% !!</p>\n<p>But the problem is that the notebook uses a multi-output regression model. Is this right to use a regression model for solving this problem of FOG event prediction? Though it might look using a regression model out of place for this kind of problem, there are strong reasons why regression model gives good scores for a classification problem :-)</p>\n<p>Here are some points:</p>\n<ol>\n<li>We are supposed to predict the events of FOG episodes based on training data of 3D accelerometer readings. These events are distinct from each other and are based on manually induced protocols in labs in some cases that there is no doubt that this is a classification problem - is there an event? If there is an event, which one of the three?</li>\n</ol>\n<p>However, a classification problem is one where the data is not continuous. If we are classifying flowers or animals like cat and dog, the data won't be continuous. There is a clear distinction between two flowers or between cat and dog. But in this problem we are determining an event based on <strong><em>individual accelerometer readings</em></strong>. If the accelerometer readings are very distinct for each of these events, then this is amenable to classification. But what if the event is a 'continuous one'?. By that I mean, the event slowly starts and reaches its peak and then subsides. <em>If we consider entire set of events as marker, then we can distinguish between events</em>. But in this training dataset, the classification is not done for a chunk of accelerometer readings, but for each individual accelerometer reading. What this means is the first few rows of a 'StartHesitation' event and first few rows of 'Turn' event may not be indistinguishable. This is like there exists some animals which are both cat and dog - or whether something is a cat or dog is a matter of degree..</p>\n<p>This is how this problem looks like. The event is a continuous one - in the sense, for a patient without an event and with event, the difference in accelerometer readings may not be very distinct and it can also be only a matter of degree. That is a particular reading might imply a FOG episode of degree 0.1 and as it progresses, the event increases in degrees till it reaches a peak and subsides. This means, for say a 'StartHesitation' event, there are several classes like 0.1 StartHesitation to 0.5 Start Hesitation and may even be equivalent to 0.3 Turn event..</p>\n<p>This is the reason why a regression model is giving good score!</p>\n<ol>\n<li>What about a classification model? This seems a good multi-class classification model with four classes - No event + three events. But is this a multi-class or multi-label problem? Since the evaluation parameters clearly says that at most one of the event alone is possible - this cannot be treated as multi-label problem. When we use independent estimators for each of the events, it implies a possibility of multi-label. But the metric chosen for evaluation is 'average precision score'. The sklearn documentation says that average_precision_score is relevant only for binary or multilabel classification and not for multi-class classification. </li>\n</ol>\n<p>Since regression scores are similar to multi-label scores, the average precision score metric is working fine!</p>\n<p>With respect to approaching this as a classification problem, irrespective of multi-class or multi-label, the training data and also the metric should be evaluated using four classes. How can using three classes alone and their prediction labels or probabilities be used for average precision score which anyway does not have an implementation for this? It looks wrong to use average precision score for a multi-class problem ignoring one of the classes.  Since only atmost one event alone is possible, it is not multi-label problem. It is neither a binary classification - using individual estimators for each event is similar to multilabel problem</p>\n<ol>\n<li>So, why not use regression model?</li>\n</ol>\n<p>I also wonder whether we can use a regression model at all  - the regression assumes the problem is continuous one and predicts 'a degree of occurrence of \"each event\" between zero and 1. This is fine, but from the perspective of evaluation metric, it may not be correct. </p>\n<p>The evaluation is done based on predicting which of the event a particular reading imply. The evaluation metric is clearly a classification metric which imply that losses are to be optimized based on which among the four classes the reading might belong. So, the sample evaluation test file shows a 1 in only one of the three columns. If the problem is addressed through a regression model which assumes continuity, then the target labels should also be continuous  - it cannot be either one or zero alone. Moreover, the very objective of the competition is to create a model that predicts which one of the three events a particular reading might belong (though strictly it should be chunk of readings and not individual readings) - how the regression score and unrelated average precision score can be used?</p>\n<p>I have created a multi-class classifiction model and the scores are nearer to the regression model - these probabilities are more relevant than the regression linear scores from the problem objectives. However, the metric of average precision score can not to be used for multiclass problem. We cannot approach it as a multilabel problem either as explained above. Having separate estimators and not optimizing loss considering all events does not correlate with the sample evaluation target variable file. Moreover, after predicting for four classes, we have to use probabilities of three classes for metric - this cannot be correct. </p>\n<p>If the competition host wants a classification and average precision score as metric, then the problem and metric should be for four classes and not three. Can metric be evaluated based on three classes alone leaving one out? If regression scores are ok, then it may not serve their purpose.</p>\n<p>Assuming a  regression notebook reaches top, will it be ok for the competition hosts? A clarification on this would be great !) Hi <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> !! - some clarifications can be every helpful!!</p>",
  "messages": [
    {
      "id": "2279711",
      "postDate": "05/29/2023 14:48:50",
      "content": "<p>My notebook is currently in 78th position out of 1242 teams! But I wonder whether the approach of my notebook is correct!</p>\n<p>I joined this competition 10 days ago and started with the top scoring notebook published (thanks to the notebook by Nicholas Gray <a href=\"https://www.kaggle.com/nickcgray\" target=\"_blank\">@nickcgray</a> - <a href=\"https://www.kaggle.com/code/nickcgray/gait-prediction\" target=\"_blank\">https://www.kaggle.com/code/nickcgray/gait-prediction</a> - it has saved tremendous amount of time). After some finetuning, I could reach a score of 0.311 putting me into top 6% !!</p>\n<p>But the problem is that the notebook uses a multi-output regression model. Is this right to use a regression model for solving this problem of FOG event prediction? Though it might look using a regression model out of place for this kind of problem, there are strong reasons why regression model gives good scores for a classification problem :-)</p>\n<p>Here are some points:</p>\n<ol>\n<li>We are supposed to predict the events of FOG episodes based on training data of 3D accelerometer readings. These events are distinct from each other and are based on manually induced protocols in labs in some cases that there is no doubt that this is a classification problem - is there an event? If there is an event, which one of the three?</li>\n</ol>\n<p>However, a classification problem is one where the data is not continuous. If we are classifying flowers or animals like cat and dog, the data won't be continuous. There is a clear distinction between two flowers or between cat and dog. But in this problem we are determining an event based on <strong><em>individual accelerometer readings</em></strong>. If the accelerometer readings are very distinct for each of these events, then this is amenable to classification. But what if the event is a 'continuous one'?. By that I mean, the event slowly starts and reaches its peak and then subsides. <em>If we consider entire set of events as marker, then we can distinguish between events</em>. But in this training dataset, the classification is not done for a chunk of accelerometer readings, but for each individual accelerometer reading. What this means is the first few rows of a 'StartHesitation' event and first few rows of 'Turn' event may not be indistinguishable. This is like there exists some animals which are both cat and dog - or whether something is a cat or dog is a matter of degree..</p>\n<p>This is how this problem looks like. The event is a continuous one - in the sense, for a patient without an event and with event, the difference in accelerometer readings may not be very distinct and it can also be only a matter of degree. That is a particular reading might imply a FOG episode of degree 0.1 and as it progresses, the event increases in degrees till it reaches a peak and subsides. This means, for say a 'StartHesitation' event, there are several classes like 0.1 StartHesitation to 0.5 Start Hesitation and may even be equivalent to 0.3 Turn event..</p>\n<p>This is the reason why a regression model is giving good score!</p>\n<ol>\n<li>What about a classification model? This seems a good multi-class classification model with four classes - No event + three events. But is this a multi-class or multi-label problem? Since the evaluation parameters clearly says that at most one of the event alone is possible - this cannot be treated as multi-label problem. When we use independent estimators for each of the events, it implies a possibility of multi-label. But the metric chosen for evaluation is 'average precision score'. The sklearn documentation says that average_precision_score is relevant only for binary or multilabel classification and not for multi-class classification. </li>\n</ol>\n<p>Since regression scores are similar to multi-label scores, the average precision score metric is working fine!</p>\n<p>With respect to approaching this as a classification problem, irrespective of multi-class or multi-label, the training data and also the metric should be evaluated using four classes. How can using three classes alone and their prediction labels or probabilities be used for average precision score which anyway does not have an implementation for this? It looks wrong to use average precision score for a multi-class problem ignoring one of the classes.  Since only atmost one event alone is possible, it is not multi-label problem. It is neither a binary classification - using individual estimators for each event is similar to multilabel problem</p>\n<ol>\n<li>So, why not use regression model?</li>\n</ol>\n<p>I also wonder whether we can use a regression model at all  - the regression assumes the problem is continuous one and predicts 'a degree of occurrence of \"each event\" between zero and 1. This is fine, but from the perspective of evaluation metric, it may not be correct. </p>\n<p>The evaluation is done based on predicting which of the event a particular reading imply. The evaluation metric is clearly a classification metric which imply that losses are to be optimized based on which among the four classes the reading might belong. So, the sample evaluation test file shows a 1 in only one of the three columns. If the problem is addressed through a regression model which assumes continuity, then the target labels should also be continuous  - it cannot be either one or zero alone. Moreover, the very objective of the competition is to create a model that predicts which one of the three events a particular reading might belong (though strictly it should be chunk of readings and not individual readings) - how the regression score and unrelated average precision score can be used?</p>\n<p>I have created a multi-class classifiction model and the scores are nearer to the regression model - these probabilities are more relevant than the regression linear scores from the problem objectives. However, the metric of average precision score can not to be used for multiclass problem. We cannot approach it as a multilabel problem either as explained above. Having separate estimators and not optimizing loss considering all events does not correlate with the sample evaluation target variable file. Moreover, after predicting for four classes, we have to use probabilities of three classes for metric - this cannot be correct. </p>\n<p>If the competition host wants a classification and average precision score as metric, then the problem and metric should be for four classes and not three. Can metric be evaluated based on three classes alone leaving one out? If regression scores are ok, then it may not serve their purpose.</p>\n<p>Assuming a  regression notebook reaches top, will it be ok for the competition hosts? A clarification on this would be great !) Hi <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> !! - some clarifications can be every helpful!!</p>",
      "rawMarkdown": "My notebook is currently in 78th position out of 1242 teams! But I wonder whether the approach of my notebook is correct!\n\nI joined this competition 10 days ago and started with the top scoring notebook published (thanks to the notebook by Nicholas Gray @nickcgray - https://www.kaggle.com/code/nickcgray/gait-prediction - it has saved tremendous amount of time). After some finetuning, I could reach a score of 0.311 putting me into top 6% !!\n\nBut the problem is that the notebook uses a multi-output regression model. Is this right to use a regression model for solving this problem of FOG event prediction? Though it might look using a regression model out of place for this kind of problem, there are strong reasons why regression model gives good scores for a classification problem :-)\n\nHere are some points:\n\n1. We are supposed to predict the events of FOG episodes based on training data of 3D accelerometer readings. These events are distinct from each other and are based on manually induced protocols in labs in some cases that there is no doubt that this is a classification problem - is there an event? If there is an event, which one of the three?\n\nHowever, a classification problem is one where the data is not continuous. If we are classifying flowers or animals like cat and dog, the data won't be continuous. There is a clear distinction between two flowers or between cat and dog. But in this problem we are determining an event based on ***individual accelerometer readings***. If the accelerometer readings are very distinct for each of these events, then this is amenable to classification. But what if the event is a 'continuous one'?. By that I mean, the event slowly starts and reaches its peak and then subsides. *If we consider entire set of events as marker, then we can distinguish between events*. But in this training dataset, the classification is not done for a chunk of accelerometer readings, but for each individual accelerometer reading. What this means is the first few rows of a 'StartHesitation' event and first few rows of 'Turn' event may not be indistinguishable. This is like there exists some animals which are both cat and dog - or whether something is a cat or dog is a matter of degree..\n\nThis is how this problem looks like. The event is a continuous one - in the sense, for a patient without an event and with event, the difference in accelerometer readings may not be very distinct and it can also be only a matter of degree. That is a particular reading might imply a FOG episode of degree 0.1 and as it progresses, the event increases in degrees till it reaches a peak and subsides. This means, for say a 'StartHesitation' event, there are several classes like 0.1 StartHesitation to 0.5 Start Hesitation and may even be equivalent to 0.3 Turn event..\n\nThis is the reason why a regression model is giving good score!\n\n2. What about a classification model? This seems a good multi-class classification model with four classes - No event + three events. But is this a multi-class or multi-label problem? Since the evaluation parameters clearly says that at most one of the event alone is possible - this cannot be treated as multi-label problem. When we use independent estimators for each of the events, it implies a possibility of multi-label. But the metric chosen for evaluation is 'average precision score'. The sklearn documentation says that average_precision_score is relevant only for binary or multilabel classification and not for multi-class classification. \n\nSince regression scores are similar to multi-label scores, the average precision score metric is working fine!\n\nWith respect to approaching this as a classification problem, irrespective of multi-class or multi-label, the training data and also the metric should be evaluated using four classes. How can using three classes alone and their prediction labels or probabilities be used for average precision score which anyway does not have an implementation for this? It looks wrong to use average precision score for a multi-class problem ignoring one of the classes.  Since only atmost one event alone is possible, it is not multi-label problem. It is neither a binary classification - using individual estimators for each event is similar to multilabel problem\n\n3. So, why not use regression model?\n\nI also wonder whether we can use a regression model at all  - the regression assumes the problem is continuous one and predicts 'a degree of occurrence of \"each event\" between zero and 1. This is fine, but from the perspective of evaluation metric, it may not be correct. \n\nThe evaluation is done based on predicting which of the event a particular reading imply. The evaluation metric is clearly a classification metric which imply that losses are to be optimized based on which among the four classes the reading might belong. So, the sample evaluation test file shows a 1 in only one of the three columns. If the problem is addressed through a regression model which assumes continuity, then the target labels should also be continuous  - it cannot be either one or zero alone. Moreover, the very objective of the competition is to create a model that predicts which one of the three events a particular reading might belong (though strictly it should be chunk of readings and not individual readings) - how the regression score and unrelated average precision score can be used?\n\nI have created a multi-class classifiction model and the scores are nearer to the regression model - these probabilities are more relevant than the regression linear scores from the problem objectives. However, the metric of average precision score can not to be used for multiclass problem. We cannot approach it as a multilabel problem either as explained above. Having separate estimators and not optimizing loss considering all events does not correlate with the sample evaluation target variable file. Moreover, after predicting for four classes, we have to use probabilities of three classes for metric - this cannot be correct. \n\nIf the competition host wants a classification and average precision score as metric, then the problem and metric should be for four classes and not three. Can metric be evaluated based on three classes alone leaving one out? If regression scores are ok, then it may not serve their purpose.\n\nAssuming a  regression notebook reaches top, will it be ok for the competition hosts? A clarification on this would be great !) Hi @ryanholbrook !! - some clarifications can be every helpful!!",
      "votes": null
    },
    {
      "id": "2285421",
      "postDate": "06/02/2023 17:40:32",
      "content": "<p>Thank you for sharing your thoughts. These insights indeed will be useful to design a better model. </p>",
      "rawMarkdown": "Thank you for sharing your thoughts. These insights indeed will be useful to design a better model.",
      "votes": null
    },
    {
      "id": "2292410",
      "postDate": "06/08/2023 10:17:41",
      "content": "<p>I think regression is ok, since Evaluation contains this remark: \"The predicted scores may, but are not required, to take the form of a probability (to be in the range 0.0 to 1.0, that is).\"</p>",
      "rawMarkdown": "I think regression is ok, since Evaluation contains this remark: \"The predicted scores may, but are not required, to take the form of a probability (to be in the range 0.0 to 1.0, that is).\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2285421,
      "author_name": "kshitizkumarjangir",
      "author_url": "",
      "post_date": "06/02/2023 17:40:32",
      "content": "<p>Thank you for sharing your thoughts. These insights indeed will be useful to design a better model. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2292410,
      "author_name": "sdelnikovaleksandr",
      "author_url": "",
      "post_date": "06/08/2023 10:17:41",
      "content": "<p>I think regression is ok, since Evaluation contains this remark: \"The predicted scores may, but are not required, to take the form of a probability (to be in the range 0.0 to 1.0, that is).\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2279711": "My notebook is currently in 78th position out of 1242 teams! But I wonder whether the approach of my notebook is correct!\n\nI joined this competition 10 days ago and started with the top scoring notebook published (thanks to the notebook by Nicholas Gray @nickcgray - https://www.kaggle.com/code/nickcgray/gait-prediction - it has saved tremendous amount of time). After some finetuning, I could reach a score of 0.311 putting me into top 6% !!\n\nBut the problem is that the notebook uses a multi-output regression model. Is this right to use a regression model for solving this problem of FOG event prediction? Though it might look using a regression model out of place for this kind of problem, there are strong reasons why regression model gives good scores for a classification problem :-)\n\nHere are some points:\n\n1. We are supposed to predict the events of FOG episodes based on training data of 3D accelerometer readings. These events are distinct from each other and are based on manually induced protocols in labs in some cases that there is no doubt that this is a classification problem - is there an event? If there is an event, which one of the three?\n\nHowever, a classification problem is one where the data is not continuous. If we are classifying flowers or animals like cat and dog, the data won't be continuous. There is a clear distinction between two flowers or between cat and dog. But in this problem we are determining an event based on ***individual accelerometer readings***. If the accelerometer readings are very distinct for each of these events, then this is amenable to classification. But what if the event is a 'continuous one'?. By that I mean, the event slowly starts and reaches its peak and then subsides. *If we consider entire set of events as marker, then we can distinguish between events*. But in this training dataset, the classification is not done for a chunk of accelerometer readings, but for each individual accelerometer reading. What this means is the first few rows of a 'StartHesitation' event and first few rows of 'Turn' event may not be indistinguishable. This is like there exists some animals which are both cat and dog - or whether something is a cat or dog is a matter of degree..\n\nThis is how this problem looks like. The event is a continuous one - in the sense, for a patient without an event and with event, the difference in accelerometer readings may not be very distinct and it can also be only a matter of degree. That is a particular reading might imply a FOG episode of degree 0.1 and as it progresses, the event increases in degrees till it reaches a peak and subsides. This means, for say a 'StartHesitation' event, there are several classes like 0.1 StartHesitation to 0.5 Start Hesitation and may even be equivalent to 0.3 Turn event..\n\nThis is the reason why a regression model is giving good score!\n\n2. What about a classification model? This seems a good multi-class classification model with four classes - No event + three events. But is this a multi-class or multi-label problem? Since the evaluation parameters clearly says that at most one of the event alone is possible - this cannot be treated as multi-label problem. When we use independent estimators for each of the events, it implies a possibility of multi-label. But the metric chosen for evaluation is 'average precision score'. The sklearn documentation says that average_precision_score is relevant only for binary or multilabel classification and not for multi-class classification. \n\nSince regression scores are similar to multi-label scores, the average precision score metric is working fine!\n\nWith respect to approaching this as a classification problem, irrespective of multi-class or multi-label, the training data and also the metric should be evaluated using four classes. How can using three classes alone and their prediction labels or probabilities be used for average precision score which anyway does not have an implementation for this? It looks wrong to use average precision score for a multi-class problem ignoring one of the classes.  Since only atmost one event alone is possible, it is not multi-label problem. It is neither a binary classification - using individual estimators for each event is similar to multilabel problem\n\n3. So, why not use regression model?\n\nI also wonder whether we can use a regression model at all  - the regression assumes the problem is continuous one and predicts 'a degree of occurrence of \"each event\" between zero and 1. This is fine, but from the perspective of evaluation metric, it may not be correct. \n\nThe evaluation is done based on predicting which of the event a particular reading imply. The evaluation metric is clearly a classification metric which imply that losses are to be optimized based on which among the four classes the reading might belong. So, the sample evaluation test file shows a 1 in only one of the three columns. If the problem is addressed through a regression model which assumes continuity, then the target labels should also be continuous  - it cannot be either one or zero alone. Moreover, the very objective of the competition is to create a model that predicts which one of the three events a particular reading might belong (though strictly it should be chunk of readings and not individual readings) - how the regression score and unrelated average precision score can be used?\n\nI have created a multi-class classifiction model and the scores are nearer to the regression model - these probabilities are more relevant than the regression linear scores from the problem objectives. However, the metric of average precision score can not to be used for multiclass problem. We cannot approach it as a multilabel problem either as explained above. Having separate estimators and not optimizing loss considering all events does not correlate with the sample evaluation target variable file. Moreover, after predicting for four classes, we have to use probabilities of three classes for metric - this cannot be correct. \n\nIf the competition host wants a classification and average precision score as metric, then the problem and metric should be for four classes and not three. Can metric be evaluated based on three classes alone leaving one out? If regression scores are ok, then it may not serve their purpose.\n\nAssuming a  regression notebook reaches top, will it be ok for the competition hosts? A clarification on this would be great !) Hi @ryanholbrook !! - some clarifications can be every helpful!!",
    "2285421": "Thank you for sharing your thoughts. These insights indeed will be useful to design a better model.",
    "2292410": "I think regression is ok, since Evaluation contains this remark: \"The predicted scores may, but are not required, to take the form of a probability (to be in the range 0.0 to 1.0, that is).\""
  },
  "source": "meta"
}