{
  "id": 13115,
  "title": "Paper on Using ANN for Ordinal Problems",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/13115",
  "author_name": "",
  "post_date": "2015-03-29T07:45:20.330Z",
  "votes": 19,
  "comment_count": 12,
  "views": 4036,
  "content": "<p>As has already been pointed out, this is an ordinal problem, rather than a&nbsp;classification problem. I found the following paper, which describes a fairly simple way to set up&nbsp;ordinal problems on a neural net, helpful in setting up the inputs and outputs to my net. Enjoy:</p>\n<p>https://web.missouri.edu/~zwyw6/files/rank.pdf</p>",
  "messages": [
    {
      "id": "68823",
      "postDate": "03/29/2015 07:45:20",
      "content": "<p>As has already been pointed out, this is an ordinal problem, rather than a&nbsp;classification problem. I found the following paper, which describes a fairly simple way to set up&nbsp;ordinal problems on a neural net, helpful in setting up the inputs and outputs to my net. Enjoy:</p>\n<p>https://web.missouri.edu/~zwyw6/files/rank.pdf</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72642",
      "postDate": "04/20/2015 14:57:43",
      "content": "<p>I was also thinking about ordinal classification. Have you tried this approach?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72668",
      "postDate": "04/20/2015 16:49:01",
      "content": "<p>Yes. I started off using this approach exactly. Since then I'm made some super-secret modifications to the way the output is decoded&nbsp;that improve Kappa, but general idea is the same. &nbsp;This has been good for a kappa&nbsp;of &nbsp;~0.63 with a single net. If I decoded&nbsp;the output of this same net through the decoding method&nbsp;they propose in the paper it's only good for a kappa of ~0.55. So there is considerable room for improvement on the decoding side. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72672",
      "postDate": "04/20/2015 17:05:09",
      "content": "<p>I'm sorry for stupid question, but what is kappa?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72673",
      "postDate": "04/20/2015 17:10:40",
      "content": "<p>No worries. Kappa is how the contest is being scored. &nbsp;Specifically quadratic kappa (https://www.kaggle.com/c/diabetic-retinopathy-detection/details/evaluation). &nbsp;As it turns out, optimizing against kappa tends to reduce the accuracy as bit since you end up emphasizing the rare cases more. &nbsp;If you are using Python, you can install skll and use the kappa function from there (http://skll.readthedocs.org/en/latest/api/skll.html?highlight=kappa#skll.kappa), this is what I've been doing. &nbsp;If you are using some other environment, I can't help you, but I imagine there is something out there.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72700",
      "postDate": "04/20/2015 18:20:44",
      "content": "<p>cool, thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75579",
      "postDate": "04/30/2015 07:19:09",
      "content": "<p>Hi Tim,</p>\n<p>Firstly congrats for your new achievement! :)</p>\n<p>And thank you for both link you provide to us.Just to clarify: under decoding you meant manipulating the output of CNN?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75621",
      "postDate": "04/30/2015 12:45:17",
      "content": "<p>[quote=George Solymosi;75579]</p>\n<p>Hi Tim,</p>\n<p>Firstly congrats for your new achievement! :)</p>\n<p>[/quote]</p>\n<p>Thanks!</p>\n<p>[quote=George Solymosi;75579]</p>\n<p>And thank you for both link you provide to us.Just to clarify: under decoding you meant manipulating the output of CNN?</p>\n<p>[/quote]</p>\n<p>The output of the CNN is just some number of floating point numbers. &nbsp;In the paper that I referenced they use a CNN with 4 outputs and map the five cases map to &quot;0000&quot;, &quot;0001&quot;, &quot;0011&quot;, &quot;0111&quot; and &quot;1111&quot;. &nbsp;However the output values aren't going to be exactly 0 and 1, although they use a sigmoid nonlinearity on the output to&nbsp;force them to be <em>between&nbsp;</em>0 and 1 . So the problem of decoding the output, in the case of this paper, &nbsp;is one of mapping four floating point values between, 0 and 1, onto the integers 0 to 4. &nbsp;The paper proposes a very simple and easy to understand scheme for doing this, which I'll let you read for yourself, but it turns out that one can do better.</p>\n<p><em>*EDIT*</em></p>\n<p><em>When I said above that in the paper they use a CNN with four outputs, etc, etc. I meant that is what one would do if applying that paper to THIS problem. In general the number of outputs would vary on the number of ordinal categories. Specifically there would be one fewer outputs than categories. Sorry for any confusion.</em></p>\n<p><em>*EDIT*</em></p>\n<p>One way to think of the way they set up the mapping in the paper is that four bits correspond to &quot;at least case 1&quot;, &quot;at least case 2&quot;, &quot;at least case 3&quot; and &quot;at least case 4&quot;. &nbsp;This is only slightly different than using a standard categorical setup were there would be 5 outputs &quot;00001&quot;, &quot;00010&quot;, &quot;00100&quot;, &quot;01000&quot; and &quot;10000&quot; which correspond to &quot;exactly case 0&quot;, &quot;exactly case 1&quot;, etc. However, in the categorical case, a softmax function is typically applied to the output to force one of the values to dominate and decoding is typically just a manner of applying argmax to the outputs.&nbsp;</p>\n<p>I hope that helps.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "132360",
      "postDate": "08/24/2016 21:33:33",
      "content": "<p>A current link for <em>A Neural Network Approach to Ordinal Regression</em>:</p>\n\n<p><a href=\"http://orca.st.usm.edu/~zwang/files/rank.pdf\">http://orca.st.usm.edu/~zwang/files/rank.pdf</a></p>",
      "rawMarkdown": "A current link for _A Neural Network Approach to Ordinal Regression_:\r\n \r\nhttp://orca.st.usm.edu/~zwang/files/rank.pdf",
      "votes": null
    },
    {
      "id": "554983",
      "postDate": "06/18/2019 09:25:23",
      "content": "<p>Hi <a href=\"/bitsofbits\">@bitsofbits</a> ,</p>\n\n<p>I assume that this competition is over. I am trying to apply ordinal classification for my data for age classification. Could you please share the super secret modification you mentioned.</p>\n\n<p>Srini</p>",
      "rawMarkdown": "Hi @bitsofbits ,\n\nI assume that this competition is over. I am trying to apply ordinal classification for my data for age classification. Could you please share the super secret modification you mentioned.\n\nSrini",
      "votes": null
    },
    {
      "id": "638305",
      "postDate": "10/01/2019 18:11:21",
      "content": "<p>Thanks for the idea. Would you mind to let me know if there is an existing python package function to perform such encoding? I didn't find the exactly same function from the latest sklearn.</p>\n\n<p>Thanks again!</p>",
      "rawMarkdown": "Thanks for the idea. Would you mind to let me know if there is an existing python package function to perform such encoding? I didn't find the exactly same function from the latest sklearn.\n\nThanks again!",
      "votes": null
    },
    {
      "id": "641627",
      "postDate": "10/04/2019 21:26:20",
      "content": "<p><a href=\"/yantinghuang\">@yantinghuang</a>, not that I know of. It's been a while since I looked at this, but you should be able to get the correct effect using one-hot encoding followed by something involving cumsum.  Maybe '1 - cumsum(one_hot(x))'? But check that before you try it.</p>",
      "rawMarkdown": "yantinghuang, not that I know of. It's been a while since I looked at this, but you should be able to get the correct effect using one-hot encoding followed by something involving cumsum.  Maybe '1 - cumsum(one_hot(x))'? But check that before you try it.",
      "votes": null
    },
    {
      "id": "2327149",
      "postDate": "07/02/2023 17:24:41",
      "content": "<p>I'm trying to apply this to a problem of my own and I really like the approach. I know this thread is way old but I'd really appreciate some advice on your superior decoding method for the ordinal encoded output! Thanks. </p>",
      "rawMarkdown": "I'm trying to apply this to a problem of my own and I really like the approach. I know this thread is way old but I'd really appreciate some advice on your superior decoding method for the ordinal encoded output! Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 72642,
      "author_name": "",
      "author_url": "",
      "post_date": "04/20/2015 14:57:43",
      "content": "<p>I was also thinking about ordinal classification. Have you tried this approach?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72668,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "04/20/2015 16:49:01",
      "content": "<p>Yes. I started off using this approach exactly. Since then I'm made some super-secret modifications to the way the output is decoded&nbsp;that improve Kappa, but general idea is the same. &nbsp;This has been good for a kappa&nbsp;of &nbsp;~0.63 with a single net. If I decoded&nbsp;the output of this same net through the decoding method&nbsp;they propose in the paper it's only good for a kappa of ~0.55. So there is considerable room for improvement on the decoding side. &nbsp;</p>",
      "votes": null,
      "replies": [
        {
          "id": 554983,
          "author_name": "sreenivasaupadhyaya",
          "author_url": "",
          "post_date": "06/18/2019 09:25:23",
          "content": "<p>Hi <a href=\"/bitsofbits\">@bitsofbits</a> ,</p>\n\n<p>I assume that this competition is over. I am trying to apply ordinal classification for my data for age classification. Could you please share the super secret modification you mentioned.</p>\n\n<p>Srini</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2327149,
          "author_name": "abalter85",
          "author_url": "",
          "post_date": "07/02/2023 17:24:41",
          "content": "<p>I'm trying to apply this to a problem of my own and I really like the approach. I know this thread is way old but I'd really appreciate some advice on your superior decoding method for the ordinal encoded output! Thanks. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 72672,
      "author_name": "",
      "author_url": "",
      "post_date": "04/20/2015 17:05:09",
      "content": "<p>I'm sorry for stupid question, but what is kappa?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72673,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "04/20/2015 17:10:40",
      "content": "<p>No worries. Kappa is how the contest is being scored. &nbsp;Specifically quadratic kappa (https://www.kaggle.com/c/diabetic-retinopathy-detection/details/evaluation). &nbsp;As it turns out, optimizing against kappa tends to reduce the accuracy as bit since you end up emphasizing the rare cases more. &nbsp;If you are using Python, you can install skll and use the kappa function from there (http://skll.readthedocs.org/en/latest/api/skll.html?highlight=kappa#skll.kappa), this is what I've been doing. &nbsp;If you are using some other environment, I can't help you, but I imagine there is something out there.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72700,
      "author_name": "",
      "author_url": "",
      "post_date": "04/20/2015 18:20:44",
      "content": "<p>cool, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75579,
      "author_name": "",
      "author_url": "",
      "post_date": "04/30/2015 07:19:09",
      "content": "<p>Hi Tim,</p>\n<p>Firstly congrats for your new achievement! :)</p>\n<p>And thank you for both link you provide to us.Just to clarify: under decoding you meant manipulating the output of CNN?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75621,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "04/30/2015 12:45:17",
      "content": "<p>[quote=George Solymosi;75579]</p>\n<p>Hi Tim,</p>\n<p>Firstly congrats for your new achievement! :)</p>\n<p>[/quote]</p>\n<p>Thanks!</p>\n<p>[quote=George Solymosi;75579]</p>\n<p>And thank you for both link you provide to us.Just to clarify: under decoding you meant manipulating the output of CNN?</p>\n<p>[/quote]</p>\n<p>The output of the CNN is just some number of floating point numbers. &nbsp;In the paper that I referenced they use a CNN with 4 outputs and map the five cases map to &quot;0000&quot;, &quot;0001&quot;, &quot;0011&quot;, &quot;0111&quot; and &quot;1111&quot;. &nbsp;However the output values aren't going to be exactly 0 and 1, although they use a sigmoid nonlinearity on the output to&nbsp;force them to be <em>between&nbsp;</em>0 and 1 . So the problem of decoding the output, in the case of this paper, &nbsp;is one of mapping four floating point values between, 0 and 1, onto the integers 0 to 4. &nbsp;The paper proposes a very simple and easy to understand scheme for doing this, which I'll let you read for yourself, but it turns out that one can do better.</p>\n<p><em>*EDIT*</em></p>\n<p><em>When I said above that in the paper they use a CNN with four outputs, etc, etc. I meant that is what one would do if applying that paper to THIS problem. In general the number of outputs would vary on the number of ordinal categories. Specifically there would be one fewer outputs than categories. Sorry for any confusion.</em></p>\n<p><em>*EDIT*</em></p>\n<p>One way to think of the way they set up the mapping in the paper is that four bits correspond to &quot;at least case 1&quot;, &quot;at least case 2&quot;, &quot;at least case 3&quot; and &quot;at least case 4&quot;. &nbsp;This is only slightly different than using a standard categorical setup were there would be 5 outputs &quot;00001&quot;, &quot;00010&quot;, &quot;00100&quot;, &quot;01000&quot; and &quot;10000&quot; which correspond to &quot;exactly case 0&quot;, &quot;exactly case 1&quot;, etc. However, in the categorical case, a softmax function is typically applied to the output to force one of the values to dominate and decoding is typically just a manner of applying argmax to the outputs.&nbsp;</p>\n<p>I hope that helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 132360,
      "author_name": "zygmunt",
      "author_url": "",
      "post_date": "08/24/2016 21:33:33",
      "content": "<p>A current link for <em>A Neural Network Approach to Ordinal Regression</em>:</p>\n\n<p><a href=\"http://orca.st.usm.edu/~zwang/files/rank.pdf\">http://orca.st.usm.edu/~zwang/files/rank.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 638305,
      "author_name": "yantinghuang",
      "author_url": "",
      "post_date": "10/01/2019 18:11:21",
      "content": "<p>Thanks for the idea. Would you mind to let me know if there is an existing python package function to perform such encoding? I didn't find the exactly same function from the latest sklearn.</p>\n\n<p>Thanks again!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 641627,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "10/04/2019 21:26:20",
      "content": "<p><a href=\"/yantinghuang\">@yantinghuang</a>, not that I know of. It's been a while since I looked at this, but you should be able to get the correct effect using one-hot encoding followed by something involving cumsum.  Maybe '1 - cumsum(one_hot(x))'? But check that before you try it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "68823": "",
    "72642": "",
    "72668": "",
    "72672": "",
    "72673": "",
    "72700": "",
    "75579": "",
    "75621": "",
    "132360": "A current link for _A Neural Network Approach to Ordinal Regression_:\r\n \r\nhttp://orca.st.usm.edu/~zwang/files/rank.pdf",
    "554983": "Hi @bitsofbits ,\n\nI assume that this competition is over. I am trying to apply ordinal classification for my data for age classification. Could you please share the super secret modification you mentioned.\n\nSrini",
    "638305": "Thanks for the idea. Would you mind to let me know if there is an existing python package function to perform such encoding? I didn't find the exactly same function from the latest sklearn.\n\nThanks again!",
    "641627": "yantinghuang, not that I know of. It's been a while since I looked at this, but you should be able to get the correct effect using one-hot encoding followed by something involving cumsum.  Maybe '1 - cumsum(one_hot(x))'? But check that before you try it.",
    "2327149": "I'm trying to apply this to a problem of my own and I really like the approach. I know this thread is way old but I'd really appreciate some advice on your superior decoding method for the ordinal encoded output! Thanks."
  },
  "source": "meta"
}