{
  "id": 547105,
  "title": "classification or regression?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/547105",
  "author_name": "",
  "post_date": "2024-11-19T20:06:05.903273800Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Why do most public notebooks and solution approaches for this competition use regression followed by rounding the results? Why not use a classifier directly, optimized with quadratic weighted kappa for evaluation?</p>",
  "messages": [
    {
      "id": "3050091",
      "postDate": "11/19/2024 20:06:05",
      "content": "<p>Why do most public notebooks and solution approaches for this competition use regression followed by rounding the results? Why not use a classifier directly, optimized with quadratic weighted kappa for evaluation?</p>",
      "rawMarkdown": "Why do most public notebooks and solution approaches for this competition use regression followed by rounding the results? Why not use a classifier directly, optimized with quadratic weighted kappa for evaluation?",
      "votes": null
    },
    {
      "id": "3050134",
      "postDate": "11/19/2024 21:10:30",
      "content": "<p>It's because we are handling a multi-class classification where there is an inherent ordering between the classes.<br>\nMost classification algorithms can't handle that well. They only learn if they predicted the target right or wrong. They can't judge, that '0' is worse than '2' if the target is '3'.</p>",
      "rawMarkdown": "It's because we are handling a multi-class classification where there is an inherent ordering between the classes.\nMost classification algorithms can't handle that well. They only learn if they predicted the target right or wrong. They can't judge, that '0' is worse than '2' if the target is '3'.",
      "votes": null
    },
    {
      "id": "3052124",
      "postDate": "11/22/2024 04:16:51",
      "content": "<p>The choice between treating this problem as classification or regression depends on how the labels are derived and the nature of the evaluation metric. Since the labels are determined based on thresholds applied to the total questionnaire scores, it suggests that the problem has a continuous underlying structure. This makes regression a natural choice because:</p>\n<p>Continuous Nature of Scores: Regression models are better suited to predict continuous scores. By training a regression model, you leverage this continuity, potentially capturing more information about the relationship between the features and the target variable.</p>\n<p>Flexibility in Label Mapping: Once you have a continuous output, you can easily map it back to discrete labels using rounding or applying the predefined thresholds. This approach allows you to use the regression model's full potential and adjust thresholds dynamically if needed.</p>\n<p>Smooth Optimization: Regression models optimize loss functions like Mean Squared Error (MSE) or Mean Absolute Error (MAE), which are smooth and differentiable, making the training process stable and efficient. On the other hand, optimizing directly for metrics like quadratic weighted kappa (QWK) can be more challenging as it is not differentiable.</p>\n<p>Classification Challenges: If you treat the problem as a classification task, the labels need to be pre-binned, potentially losing the nuance of the underlying continuous structure. Furthermore, training a classifier optimized for QWK directly is complex, and public tools or frameworks supporting this are limited. Regression followed by rounding is a simpler and more practical approach.</p>\n<p>In this competition, regression followed by rounding is popular because it balances predictive accuracy and ease of implementation, while still aligning with the quadratic weighted kappa evaluation metric after the labels are converted.</p>",
      "rawMarkdown": "The choice between treating this problem as classification or regression depends on how the labels are derived and the nature of the evaluation metric. Since the labels are determined based on thresholds applied to the total questionnaire scores, it suggests that the problem has a continuous underlying structure. This makes regression a natural choice because:\n\nContinuous Nature of Scores: Regression models are better suited to predict continuous scores. By training a regression model, you leverage this continuity, potentially capturing more information about the relationship between the features and the target variable.\n\nFlexibility in Label Mapping: Once you have a continuous output, you can easily map it back to discrete labels using rounding or applying the predefined thresholds. This approach allows you to use the regression model's full potential and adjust thresholds dynamically if needed.\n\nSmooth Optimization: Regression models optimize loss functions like Mean Squared Error (MSE) or Mean Absolute Error (MAE), which are smooth and differentiable, making the training process stable and efficient. On the other hand, optimizing directly for metrics like quadratic weighted kappa (QWK) can be more challenging as it is not differentiable.\n\nClassification Challenges: If you treat the problem as a classification task, the labels need to be pre-binned, potentially losing the nuance of the underlying continuous structure. Furthermore, training a classifier optimized for QWK directly is complex, and public tools or frameworks supporting this are limited. Regression followed by rounding is a simpler and more practical approach.\n\nIn this competition, regression followed by rounding is popular because it balances predictive accuracy and ease of implementation, while still aligning with the quadratic weighted kappa evaluation metric after the labels are converted.",
      "votes": null
    },
    {
      "id": "3052212",
      "postDate": "11/22/2024 07:27:47",
      "content": "<p>I think it should be used regression</p>",
      "rawMarkdown": "I think it should be used regression",
      "votes": null
    },
    {
      "id": "3062700",
      "postDate": "12/03/2024 20:29:43",
      "content": "<p>As something of a \"Regress-ification\" hybrid, I'm fitting PCIAT Total but modified it to have gaps at the threshold values: the idea is to provide an MSE signal for the model to pay more attention at the sii boundaries. It does help the score a little (0.005 on LB)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fe1c980078b9013bab750934e40c26052%2FTotal_for_regressification.png?generation=1733257482416179&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "As something of a \"Regress-ification\" hybrid, I'm fitting PCIAT Total but modified it to have gaps at the threshold values: the idea is to provide an MSE signal for the model to pay more attention at the sii boundaries. It does help the score a little (0.005 on LB)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fe1c980078b9013bab750934e40c26052%2FTotal_for_regressification.png?generation=1733257482416179&alt=media)",
      "votes": null
    },
    {
      "id": "3075379",
      "postDate": "12/18/2024 18:55:40",
      "content": "<p>I personally perceive this as a regression problem since the target values 0,1,2,3 resulted directly from the <strong>continuous</strong> test scores, which are also missing from the test data, so I do believe that using regression corresponds to the nature of this problem!</p>",
      "rawMarkdown": "I personally perceive this as a regression problem since the target values 0,1,2,3 resulted directly from the **continuous** test scores, which are also missing from the test data, so I do believe that using regression corresponds to the nature of this problem!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3050134,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "11/19/2024 21:10:30",
      "content": "<p>It's because we are handling a multi-class classification where there is an inherent ordering between the classes.<br>\nMost classification algorithms can't handle that well. They only learn if they predicted the target right or wrong. They can't judge, that '0' is worse than '2' if the target is '3'.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3052124,
      "author_name": "wayne127",
      "author_url": "",
      "post_date": "11/22/2024 04:16:51",
      "content": "<p>The choice between treating this problem as classification or regression depends on how the labels are derived and the nature of the evaluation metric. Since the labels are determined based on thresholds applied to the total questionnaire scores, it suggests that the problem has a continuous underlying structure. This makes regression a natural choice because:</p>\n<p>Continuous Nature of Scores: Regression models are better suited to predict continuous scores. By training a regression model, you leverage this continuity, potentially capturing more information about the relationship between the features and the target variable.</p>\n<p>Flexibility in Label Mapping: Once you have a continuous output, you can easily map it back to discrete labels using rounding or applying the predefined thresholds. This approach allows you to use the regression model's full potential and adjust thresholds dynamically if needed.</p>\n<p>Smooth Optimization: Regression models optimize loss functions like Mean Squared Error (MSE) or Mean Absolute Error (MAE), which are smooth and differentiable, making the training process stable and efficient. On the other hand, optimizing directly for metrics like quadratic weighted kappa (QWK) can be more challenging as it is not differentiable.</p>\n<p>Classification Challenges: If you treat the problem as a classification task, the labels need to be pre-binned, potentially losing the nuance of the underlying continuous structure. Furthermore, training a classifier optimized for QWK directly is complex, and public tools or frameworks supporting this are limited. Regression followed by rounding is a simpler and more practical approach.</p>\n<p>In this competition, regression followed by rounding is popular because it balances predictive accuracy and ease of implementation, while still aligning with the quadratic weighted kappa evaluation metric after the labels are converted.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3052212,
      "author_name": "dtleez",
      "author_url": "",
      "post_date": "11/22/2024 07:27:47",
      "content": "<p>I think it should be used regression</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3062700,
      "author_name": "dan3dewey",
      "author_url": "",
      "post_date": "12/03/2024 20:29:43",
      "content": "<p>As something of a \"Regress-ification\" hybrid, I'm fitting PCIAT Total but modified it to have gaps at the threshold values: the idea is to provide an MSE signal for the model to pay more attention at the sii boundaries. It does help the score a little (0.005 on LB)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fe1c980078b9013bab750934e40c26052%2FTotal_for_regressification.png?generation=1733257482416179&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3075379,
      "author_name": "viktoriamelkumyan",
      "author_url": "",
      "post_date": "12/18/2024 18:55:40",
      "content": "<p>I personally perceive this as a regression problem since the target values 0,1,2,3 resulted directly from the <strong>continuous</strong> test scores, which are also missing from the test data, so I do believe that using regression corresponds to the nature of this problem!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3050091": "Why do most public notebooks and solution approaches for this competition use regression followed by rounding the results? Why not use a classifier directly, optimized with quadratic weighted kappa for evaluation?",
    "3050134": "It's because we are handling a multi-class classification where there is an inherent ordering between the classes.\nMost classification algorithms can't handle that well. They only learn if they predicted the target right or wrong. They can't judge, that '0' is worse than '2' if the target is '3'.",
    "3052124": "The choice between treating this problem as classification or regression depends on how the labels are derived and the nature of the evaluation metric. Since the labels are determined based on thresholds applied to the total questionnaire scores, it suggests that the problem has a continuous underlying structure. This makes regression a natural choice because:\n\nContinuous Nature of Scores: Regression models are better suited to predict continuous scores. By training a regression model, you leverage this continuity, potentially capturing more information about the relationship between the features and the target variable.\n\nFlexibility in Label Mapping: Once you have a continuous output, you can easily map it back to discrete labels using rounding or applying the predefined thresholds. This approach allows you to use the regression model's full potential and adjust thresholds dynamically if needed.\n\nSmooth Optimization: Regression models optimize loss functions like Mean Squared Error (MSE) or Mean Absolute Error (MAE), which are smooth and differentiable, making the training process stable and efficient. On the other hand, optimizing directly for metrics like quadratic weighted kappa (QWK) can be more challenging as it is not differentiable.\n\nClassification Challenges: If you treat the problem as a classification task, the labels need to be pre-binned, potentially losing the nuance of the underlying continuous structure. Furthermore, training a classifier optimized for QWK directly is complex, and public tools or frameworks supporting this are limited. Regression followed by rounding is a simpler and more practical approach.\n\nIn this competition, regression followed by rounding is popular because it balances predictive accuracy and ease of implementation, while still aligning with the quadratic weighted kappa evaluation metric after the labels are converted.",
    "3052212": "I think it should be used regression",
    "3062700": "As something of a \"Regress-ification\" hybrid, I'm fitting PCIAT Total but modified it to have gaps at the threshold values: the idea is to provide an MSE signal for the model to pay more attention at the sii boundaries. It does help the score a little (0.005 on LB)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fe1c980078b9013bab750934e40c26052%2FTotal_for_regressification.png?generation=1733257482416179&alt=media)",
    "3075379": "I personally perceive this as a regression problem since the target values 0,1,2,3 resulted directly from the **continuous** test scores, which are also missing from the test data, so I do believe that using regression corresponds to the nature of this problem!"
  },
  "source": "meta"
}