{
  "id": 97893,
  "title": "Ordinal Regression",
  "url": "/competitions/aptos2019-blindness-detection/discussion/97893",
  "author_name": "",
  "post_date": "2019-06-29T14:15:49.445761500Z",
  "votes": 34,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This problem is one of ordinal regression, i.e., the labels are ordered 0, 1, ..., 4, but the difference between, say, 0 and 1 is not the same as the difference between 2 and 3. Often machine learning practioners take the easy route and treat this either as a multi-class classification problem or as a regression problem. As a classification problem, we throw away the ordering of the labels. As a regression problem we impose the restriction that the differences between 0 and 1, 1 and 2, etc. are the same. </p>\n\n<p>If you want to treat the problem correctly as one of ordinal regression then you will need to use a custom loss function in Pytorch / Tensorflow / etc. training. An alternative is to use this package: <a href=\"https://www.ethanrosenthal.com/2018/12/06/spacecutter-ordinal-regression/\">spacecutter</a>, which wraps ordinal regression around a Pytorch neural network. Here is another <a href=\"https://github.com/gspell/TF-OrdinalRegression\">package</a> that does something similar with Tensorflow.</p>\n\n<p>A quicker approach which can be easily added to your algorithm of choice is to train a regression model against the integer labels, then estimate an ordinal regression from the real-valued predictions to the integer labels.</p>\n\n<p>And finally, although this is more computationally expensive, one can generate all pairs of training examples and label them 0 or 1 according to whether the first or second of the pair has a higher label. Then, train a binary classifier which can be used to order the combined train and test data, from which test label predictions can be extracted.</p>",
  "messages": [
    {
      "id": "564505",
      "postDate": "06/29/2019 14:15:49",
      "content": "<p>This problem is one of ordinal regression, i.e., the labels are ordered 0, 1, ..., 4, but the difference between, say, 0 and 1 is not the same as the difference between 2 and 3. Often machine learning practioners take the easy route and treat this either as a multi-class classification problem or as a regression problem. As a classification problem, we throw away the ordering of the labels. As a regression problem we impose the restriction that the differences between 0 and 1, 1 and 2, etc. are the same. </p>\n\n<p>If you want to treat the problem correctly as one of ordinal regression then you will need to use a custom loss function in Pytorch / Tensorflow / etc. training. An alternative is to use this package: <a href=\"https://www.ethanrosenthal.com/2018/12/06/spacecutter-ordinal-regression/\">spacecutter</a>, which wraps ordinal regression around a Pytorch neural network. Here is another <a href=\"https://github.com/gspell/TF-OrdinalRegression\">package</a> that does something similar with Tensorflow.</p>\n\n<p>A quicker approach which can be easily added to your algorithm of choice is to train a regression model against the integer labels, then estimate an ordinal regression from the real-valued predictions to the integer labels.</p>\n\n<p>And finally, although this is more computationally expensive, one can generate all pairs of training examples and label them 0 or 1 according to whether the first or second of the pair has a higher label. Then, train a binary classifier which can be used to order the combined train and test data, from which test label predictions can be extracted.</p>",
      "rawMarkdown": "This problem is one of ordinal regression, i.e., the labels are ordered 0, 1, ..., 4, but the difference between, say, 0 and 1 is not the same as the difference between 2 and 3. Often machine learning practioners take the easy route and treat this either as a multi-class classification problem or as a regression problem. As a classification problem, we throw away the ordering of the labels. As a regression problem we impose the restriction that the differences between 0 and 1, 1 and 2, etc. are the same. \n\nIf you want to treat the problem correctly as one of ordinal regression then you will need to use a custom loss function in Pytorch / Tensorflow / etc. training. An alternative is to use this package: [spacecutter](https://www.ethanrosenthal.com/2018/12/06/spacecutter-ordinal-regression/), which wraps ordinal regression around a Pytorch neural network. Here is another [package](https://github.com/gspell/TF-OrdinalRegression) that does something similar with Tensorflow.\n\nA quicker approach which can be easily added to your algorithm of choice is to train a regression model against the integer labels, then estimate an ordinal regression from the real-valued predictions to the integer labels.\n\nAnd finally, although this is more computationally expensive, one can generate all pairs of training examples and label them 0 or 1 according to whether the first or second of the pair has a higher label. Then, train a binary classifier which can be used to order the combined train and test data, from which test label predictions can be extracted.",
      "votes": null
    },
    {
      "id": "564801",
      "postDate": "06/30/2019 02:22:13",
      "content": "<p>Thanks for sharing! Spacecutter looks pretty handy.</p>",
      "rawMarkdown": "Thanks for sharing! Spacecutter looks pretty handy.",
      "votes": null
    },
    {
      "id": "564826",
      "postDate": "06/30/2019 04:01:39",
      "content": "<p>Or this; <a href=\"https://stats.stackexchange.com/a/324879/226626\">https://stats.stackexchange.com/a/324879/226626</a>.</p>\n\n<p>I suppose this approach results in ambiguities though. How do you treat a prediction that doesn't match the assumed ordering, such as '0 1 0 1 0'.</p>",
      "rawMarkdown": "Or this; https://stats.stackexchange.com/a/324879/226626.\n\nI suppose this approach results in ambiguities though. How do you treat a prediction that doesn't match the assumed ordering, such as '0 1 0 1 0'.",
      "votes": null
    },
    {
      "id": "565395",
      "postDate": "06/30/2019 22:16:18",
      "content": "<p>You could add an extra component to your pipeline that uses a standard statistical ordinal regression to map the neural network outputs to a ranking.</p>",
      "rawMarkdown": "You could add an extra component to your pipeline that uses a standard statistical ordinal regression to map the neural network outputs to a ranking.",
      "votes": null
    },
    {
      "id": "566307",
      "postDate": "07/02/2019 03:21:44",
      "content": "<p>I've made a <a href=\"https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets\">kernel</a> public implementing this technique. It seems to help a little bit.</p>\n\n<p>Basically, I make the target look like that of a multi-label problem, by turning 1 into <code>0,1</code> and 2 into <code>0,1,2</code> etc (Fast.ai automatically detects a multi label problem doing this) then I wrote a custom prediction parser which takes the sigmoid of the model's output thresholded to 0.5, finds the first position that's not zero and returns that as the prediction.</p>\n\n<p>```\ndef get_preds(arr):</p>\n\n<pre><code>mask = arr == 0\nreturn np.clip(np.where(mask.any(1), mask.argmax(1), 5) - 1, 0, 4)\n</code></pre>\n\n<p>val_preds = get_preds((torch.sigmoid(last_output) &gt; 0.5).numpy())\n```</p>\n\n<p>Let me know if you have any questions about it.</p>",
      "rawMarkdown": "I've made a [kernel](https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets) public implementing this technique. It seems to help a little bit.\n\nBasically, I make the target look like that of a multi-label problem, by turning 1 into `0,1` and 2 into `0,1,2` etc (Fast.ai automatically detects a multi label problem doing this) then I wrote a custom prediction parser which takes the sigmoid of the model's output thresholded to 0.5, finds the first position that's not zero and returns that as the prediction.\n\n```\ndef get_preds(arr):\n  \n    mask = arr == 0\n    return np.clip(np.where(mask.any(1), mask.argmax(1), 5) - 1, 0, 4)\n\nval_preds = get_preds((torch.sigmoid(last_output) &gt; 0.5).numpy())\n```\n\nLet me know if you have any questions about it.",
      "votes": null
    },
    {
      "id": "601257",
      "postDate": "08/17/2019 09:36:14",
      "content": "<p>Thank you for introducing spacecutter! I tried spacecutter and here is a <a href=\"https://www.kaggle.com/yosefardhito/aptos-pytorch-ordinal-regression-with-spacecutter\">kernel</a> that utilizes it.</p>\n\n<p>I really think this problem should be a regression instead of classification, since it is clear that, for example, class 3 is less severe than class 4 and more severe than class 2. However, regression with pre-defined threshold does not really capture the difference in gap between the labels (is it correct to assume that going from label 1 to label 2 is the same as going from label 3 to label 4?). This threshold problem might not be a big deal due to the capability of neural network models to establish a non-linear boundary, but ordinal regression is the natural way to go for me. </p>\n\n<p>I know, in the end, all that matters is the performance, so use whatever that works for you! it is just boggling me that none of the public kernels approach this with ordinal regression.</p>",
      "rawMarkdown": "Thank you for introducing spacecutter! I tried spacecutter and here is a [kernel](https://www.kaggle.com/yosefardhito/aptos-pytorch-ordinal-regression-with-spacecutter) that utilizes it.\n\nI really think this problem should be a regression instead of classification, since it is clear that, for example, class 3 is less severe than class 4 and more severe than class 2. However, regression with pre-defined threshold does not really capture the difference in gap between the labels (is it correct to assume that going from label 1 to label 2 is the same as going from label 3 to label 4?). This threshold problem might not be a big deal due to the capability of neural network models to establish a non-linear boundary, but ordinal regression is the natural way to go for me. \n\nI know, in the end, all that matters is the performance, so use whatever that works for you! it is just boggling me that none of the public kernels approach this with ordinal regression.",
      "votes": null
    },
    {
      "id": "1646002",
      "postDate": "01/11/2022 12:13:33",
      "content": "<p>Ordinal regression predicts ranking values. The method works well with ordinal dependent variables. Ordinal <a href=\"https://www.learnbay.co/data-science-course/introduction-to-simple-linear-regression-in-machine-learning/\" target=\"_blank\">regression </a>includes ordered logit and ordered probit.</p>",
      "rawMarkdown": "Ordinal regression predicts ranking values. The method works well with ordinal dependent variables. Ordinal [regression ](https://www.learnbay.co/data-science-course/introduction-to-simple-linear-regression-in-machine-learning/)includes ordered logit and ordered probit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1646002,
      "author_name": "datamachinelearning",
      "author_url": "",
      "post_date": "01/11/2022 12:13:33",
      "content": "<p>Ordinal regression predicts ranking values. The method works well with ordinal dependent variables. Ordinal <a href=\"https://www.learnbay.co/data-science-course/introduction-to-simple-linear-regression-in-machine-learning/\" target=\"_blank\">regression </a>includes ordered logit and ordered probit.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 564801,
      "author_name": "puremath86",
      "author_url": "",
      "post_date": "06/30/2019 02:22:13",
      "content": "<p>Thanks for sharing! Spacecutter looks pretty handy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 564826,
      "author_name": "roman99",
      "author_url": "",
      "post_date": "06/30/2019 04:01:39",
      "content": "<p>Or this; <a href=\"https://stats.stackexchange.com/a/324879/226626\">https://stats.stackexchange.com/a/324879/226626</a>.</p>\n\n<p>I suppose this approach results in ambiguities though. How do you treat a prediction that doesn't match the assumed ordering, such as '0 1 0 1 0'.</p>",
      "votes": null,
      "replies": [
        {
          "id": 565395,
          "author_name": "robertburbidge",
          "author_url": "",
          "post_date": "06/30/2019 22:16:18",
          "content": "<p>You could add an extra component to your pipeline that uses a standard statistical ordinal regression to map the neural network outputs to a ranking.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 566307,
      "author_name": "lextoumbourou",
      "author_url": "",
      "post_date": "07/02/2019 03:21:44",
      "content": "<p>I've made a <a href=\"https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets\">kernel</a> public implementing this technique. It seems to help a little bit.</p>\n\n<p>Basically, I make the target look like that of a multi-label problem, by turning 1 into <code>0,1</code> and 2 into <code>0,1,2</code> etc (Fast.ai automatically detects a multi label problem doing this) then I wrote a custom prediction parser which takes the sigmoid of the model's output thresholded to 0.5, finds the first position that's not zero and returns that as the prediction.</p>\n\n<p>```\ndef get_preds(arr):</p>\n\n<pre><code>mask = arr == 0\nreturn np.clip(np.where(mask.any(1), mask.argmax(1), 5) - 1, 0, 4)\n</code></pre>\n\n<p>val_preds = get_preds((torch.sigmoid(last_output) &gt; 0.5).numpy())\n```</p>\n\n<p>Let me know if you have any questions about it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 601257,
      "author_name": "yosefardhito",
      "author_url": "",
      "post_date": "08/17/2019 09:36:14",
      "content": "<p>Thank you for introducing spacecutter! I tried spacecutter and here is a <a href=\"https://www.kaggle.com/yosefardhito/aptos-pytorch-ordinal-regression-with-spacecutter\">kernel</a> that utilizes it.</p>\n\n<p>I really think this problem should be a regression instead of classification, since it is clear that, for example, class 3 is less severe than class 4 and more severe than class 2. However, regression with pre-defined threshold does not really capture the difference in gap between the labels (is it correct to assume that going from label 1 to label 2 is the same as going from label 3 to label 4?). This threshold problem might not be a big deal due to the capability of neural network models to establish a non-linear boundary, but ordinal regression is the natural way to go for me. </p>\n\n<p>I know, in the end, all that matters is the performance, so use whatever that works for you! it is just boggling me that none of the public kernels approach this with ordinal regression.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "564505": "This problem is one of ordinal regression, i.e., the labels are ordered 0, 1, ..., 4, but the difference between, say, 0 and 1 is not the same as the difference between 2 and 3. Often machine learning practioners take the easy route and treat this either as a multi-class classification problem or as a regression problem. As a classification problem, we throw away the ordering of the labels. As a regression problem we impose the restriction that the differences between 0 and 1, 1 and 2, etc. are the same. \n\nIf you want to treat the problem correctly as one of ordinal regression then you will need to use a custom loss function in Pytorch / Tensorflow / etc. training. An alternative is to use this package: [spacecutter](https://www.ethanrosenthal.com/2018/12/06/spacecutter-ordinal-regression/), which wraps ordinal regression around a Pytorch neural network. Here is another [package](https://github.com/gspell/TF-OrdinalRegression) that does something similar with Tensorflow.\n\nA quicker approach which can be easily added to your algorithm of choice is to train a regression model against the integer labels, then estimate an ordinal regression from the real-valued predictions to the integer labels.\n\nAnd finally, although this is more computationally expensive, one can generate all pairs of training examples and label them 0 or 1 according to whether the first or second of the pair has a higher label. Then, train a binary classifier which can be used to order the combined train and test data, from which test label predictions can be extracted.",
    "564801": "Thanks for sharing! Spacecutter looks pretty handy.",
    "564826": "Or this; https://stats.stackexchange.com/a/324879/226626.\n\nI suppose this approach results in ambiguities though. How do you treat a prediction that doesn't match the assumed ordering, such as '0 1 0 1 0'.",
    "565395": "You could add an extra component to your pipeline that uses a standard statistical ordinal regression to map the neural network outputs to a ranking.",
    "566307": "I've made a [kernel](https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets) public implementing this technique. It seems to help a little bit.\n\nBasically, I make the target look like that of a multi-label problem, by turning 1 into `0,1` and 2 into `0,1,2` etc (Fast.ai automatically detects a multi label problem doing this) then I wrote a custom prediction parser which takes the sigmoid of the model's output thresholded to 0.5, finds the first position that's not zero and returns that as the prediction.\n\n```\ndef get_preds(arr):\n  \n    mask = arr == 0\n    return np.clip(np.where(mask.any(1), mask.argmax(1), 5) - 1, 0, 4)\n\nval_preds = get_preds((torch.sigmoid(last_output) &gt; 0.5).numpy())\n```\n\nLet me know if you have any questions about it.",
    "601257": "Thank you for introducing spacecutter! I tried spacecutter and here is a [kernel](https://www.kaggle.com/yosefardhito/aptos-pytorch-ordinal-regression-with-spacecutter) that utilizes it.\n\nI really think this problem should be a regression instead of classification, since it is clear that, for example, class 3 is less severe than class 4 and more severe than class 2. However, regression with pre-defined threshold does not really capture the difference in gap between the labels (is it correct to assume that going from label 1 to label 2 is the same as going from label 3 to label 4?). This threshold problem might not be a big deal due to the capability of neural network models to establish a non-linear boundary, but ordinal regression is the natural way to go for me. \n\nI know, in the end, all that matters is the performance, so use whatever that works for you! it is just boggling me that none of the public kernels approach this with ordinal regression.",
    "1646002": "Ordinal regression predicts ranking values. The method works well with ordinal dependent variables. Ordinal [regression ](https://www.learnbay.co/data-science-course/introduction-to-simple-linear-regression-in-machine-learning/)includes ordered logit and ordered probit."
  },
  "source": "meta"
}