{
  "id": 552753,
  "title": "Private LB: 0.460. Ensemble Model WIth Ordinal Target Encoding ",
  "url": "/competitions/child-mind-institute-problematic-internet-use/writeups/abhisek-dash-private-lb-0-460-ensemble-model-with-",
  "author_name": "",
  "post_date": "2024-12-25T04:29:47.773Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello.</p>\n<p>I did not select this notebook as my final solutions, but this ended up as the highest scoring. (LOL)</p>\n<p>Salient Features of the modelling process:</p>\n<ol>\n<li>Ensemble of XGBoost, Random Forest and LightGBM with cross validation.</li>\n<li>Descriptive statistics of enmo and light from actigraphy data. </li>\n<li>Ordinal encoding of target sii. Ordinal encoding is based on the paper <a href=\"https://researchcommons.waikato.ac.nz/server/api/core/bitstreams/d4decc50-dcf2-486b-9860-0c650b92b173/content\" target=\"_blank\">A Simple Approach to Ordinal Classification</a><br>\nThis type of encoding preserves the ordinal relation between the different sii values. <br>\nThe model tries to predict  P(sii) &gt; 0, P(sii) &gt; 1 and P(sii) &gt;2 </li>\n</ol>\n<p>For e.g. in the ordinal encoding section of the notebook,<br>\n[0,0,0] represents sii=0<br>\n[1,0,0] represents sii = 1<br>\n[1,1,0] represents sii = 2<br>\n[1,1,1] represents sii = 3</p>\n<p>[1,1,0] means that P(sii&gt;0) = 1, P(sii&gt;1) = 1 and P(sii&gt;2)=0. So naturally the datapoint has sii value =2.<br>\n[0,0,0] means that P(sii&gt;0) = 0, P(sii&gt;1) = 0 and P(sii&gt;2) = 0. So naturally the datapoint has sii=0.<br>\n[1,1,1] means that P(sii&gt;0)=1,P(sii&gt;1)=1 and P(sii&gt;2)=1, So naturally the datapoint has sii=3.  </p>\n<p>The probability of sii is calculated by the following formulae. (Note that this is slightly different than the method described in the paper)</p>\n<p>P(sii=0) = 1-P(sii&gt;0)</p>\n<p>P(sii=1) = P(sii&gt;0)*(1-P(sii&gt;1)) (In plain english this means probability of sii &gt; 0 *and* probability of sii &lt;=1)</p>\n<p>P(sii=2) = P(sii&gt;1)*(1-P(sii&gt;2))</p>\n<p>P(sii=3) = P(sii&gt;2)*(1-P(sii&gt;3))</p>\n<p>The classification of sii is the one with highest probability as per the above formulae. For e.g. if P(sii=2) is highest, the datapoint is classified as 2</p>\n<p>Link to notebook: <a href=\"https://www.kaggle.com/code/abhisekdash37/ensemble-learning-problematic-internet-use?scriptVersionId=205432552\" target=\"_blank\">https://www.kaggle.com/code/abhisekdash37/ensemble-learning-problematic-internet-use?scriptVersionId=205432552</a></p>",
  "messages": [
    {
      "id": "3077724",
      "postDate": "12/21/2024 10:36:38",
      "content": "<p>Hello.</p>\n<p>I did not select this notebook as my final solutions, but this ended up as the highest scoring. (LOL)</p>\n<p>Salient Features of the modelling process:</p>\n<ol>\n<li>Ensemble of XGBoost, Random Forest and LightGBM with cross validation.</li>\n<li>Descriptive statistics of enmo and light from actigraphy data. </li>\n<li>Ordinal encoding of target sii. Ordinal encoding is based on the paper <a href=\"https://researchcommons.waikato.ac.nz/server/api/core/bitstreams/d4decc50-dcf2-486b-9860-0c650b92b173/content\" target=\"_blank\">A Simple Approach to Ordinal Classification</a><br>\nThis type of encoding preserves the ordinal relation between the different sii values. <br>\nThe model tries to predict  P(sii) &gt; 0, P(sii) &gt; 1 and P(sii) &gt;2 </li>\n</ol>\n<p>For e.g. in the ordinal encoding section of the notebook,<br>\n[0,0,0] represents sii=0<br>\n[1,0,0] represents sii = 1<br>\n[1,1,0] represents sii = 2<br>\n[1,1,1] represents sii = 3</p>\n<p>[1,1,0] means that P(sii&gt;0) = 1, P(sii&gt;1) = 1 and P(sii&gt;2)=0. So naturally the datapoint has sii value =2.<br>\n[0,0,0] means that P(sii&gt;0) = 0, P(sii&gt;1) = 0 and P(sii&gt;2) = 0. So naturally the datapoint has sii=0.<br>\n[1,1,1] means that P(sii&gt;0)=1,P(sii&gt;1)=1 and P(sii&gt;2)=1, So naturally the datapoint has sii=3.  </p>\n<p>The probability of sii is calculated by the following formulae. (Note that this is slightly different than the method described in the paper)</p>\n<p>P(sii=0) = 1-P(sii&gt;0)</p>\n<p>P(sii=1) = P(sii&gt;0)*(1-P(sii&gt;1)) (In plain english this means probability of sii &gt; 0 *and* probability of sii &lt;=1)</p>\n<p>P(sii=2) = P(sii&gt;1)*(1-P(sii&gt;2))</p>\n<p>P(sii=3) = P(sii&gt;2)*(1-P(sii&gt;3))</p>\n<p>The classification of sii is the one with highest probability as per the above formulae. For e.g. if P(sii=2) is highest, the datapoint is classified as 2</p>\n<p>Link to notebook: <a href=\"https://www.kaggle.com/code/abhisekdash37/ensemble-learning-problematic-internet-use?scriptVersionId=205432552\" target=\"_blank\">https://www.kaggle.com/code/abhisekdash37/ensemble-learning-problematic-internet-use?scriptVersionId=205432552</a></p>",
      "rawMarkdown": "Hello.\n\nI did not select this notebook as my final solutions, but this ended up as the highest scoring. (LOL)\n\nSalient Features of the modelling process:\n\n1. Ensemble of XGBoost, Random Forest and LightGBM with cross validation.\n2. Descriptive statistics of enmo and light from actigraphy data. \n3. Ordinal encoding of target sii. Ordinal encoding is based on the paper [A Simple Approach to Ordinal Classification](https://researchcommons.waikato.ac.nz/server/api/core/bitstreams/d4decc50-dcf2-486b-9860-0c650b92b173/content)\nThis type of encoding preserves the ordinal relation between the different sii values. \nThe model tries to predict  P(sii) > 0, P(sii) > 1 and P(sii) >2 \n\n\nFor e.g. in the ordinal encoding section of the notebook,\n[0,0,0] represents sii=0\n[1,0,0] represents sii = 1\n[1,1,0] represents sii = 2\n[1,1,1] represents sii = 3\n\n[1,1,0] means that P(sii>0) = 1, P(sii>1) = 1 and P(sii>2)=0. So naturally the datapoint has sii value =2.\n[0,0,0] means that P(sii>0) = 0, P(sii>1) = 0 and P(sii>2) = 0. So naturally the datapoint has sii=0.\n[1,1,1] means that P(sii>0)=1,P(sii>1)=1 and P(sii>2)=1, So naturally the datapoint has sii=3.  \n\nThe probability of sii is calculated by the following formulae. (Note that this is slightly different than the method described in the paper)\n\nP(sii=0) = 1-P(sii>0)\n\nP(sii=1) = P(sii>0)*(1-P(sii>1)) (In plain english this means probability of sii > 0 *and* probability of sii <=1)\n\nP(sii=2) = P(sii>1)*(1-P(sii>2))\n\nP(sii=3) = P(sii>2)*(1-P(sii>3))\n\n\nThe classification of sii is the one with highest probability as per the above formulae. For e.g. if P(sii=2) is highest, the datapoint is classified as 2\n\nLink to notebook: https://www.kaggle.com/code/abhisekdash37/ensemble-learning-problematic-internet-use?scriptVersionId=205432552",
      "votes": null
    },
    {
      "id": "3077736",
      "postDate": "12/21/2024 10:53:40",
      "content": "<p>This looks like 4 binary classifiers. Am I right <a href=\"https://www.kaggle.com/abhisekdash37\" target=\"_blank\">@abhisekdash37</a>?</p>",
      "rawMarkdown": "This looks like 4 binary classifiers. Am I right @abhisekdash37?",
      "votes": null
    },
    {
      "id": "3077767",
      "postDate": "12/21/2024 11:32:30",
      "content": "<p>Hello Ravi. Thanks for the upvote. </p>\n<p>All the models in the ensemble predict 3 logistic values (between 0 and 1). So I essentially converted the modelling to a multilabel logistic regression problem. (See the ordinal encoding section of the shared notebook)</p>\n<p>The first value represents P(sii&gt;0), second value represents P(sii&gt;1), third value represents P(sii&gt;2).</p>\n<p>So for example the encoding [1,1,0] represents sii=2. <br>\nThat is because in this encoding scheme P(sii&gt;0)=1 (first index) and P(sii&gt;1)=1 (second index) and P(sii&gt;2)=0 (third index)</p>\n<p>So what the encoding scheme is saying in this particular case  is that P(sii&gt;0)=1 and P(sii&gt;1) = 1 but P(sii&gt;2)=0. Since we are dealing with discrete values here, this means that the sii = 2.</p>\n<p>Similarly [1,1,1] represents 3 because P(sii&gt;0)=1, P(sii&gt;1)=1 and P(sii&gt;2)=1. This can only mean that sii=3</p>\n<p>The model tries to predict this encoding scheme. So for e.g. if the model predicts [0.9,0.7,0.5] for a particular datapoint, the probability for sii=1 or 2 or 3 is evaluated as follows:</p>\n<p>P(sii=0) = 1-P(sii&gt;0) = 1-0.9 = 0.1 <br>\nP(sii=1) = P(sii&gt;0) and P(sii&lt;=1) = P(sii&gt;0)<em>(1-P(sii&gt;1)) = 0.9</em>(1-0.7)<br>\nP(sii=2) = P(sii&gt;1 and P(sii&lt;=2)) = P(sii&gt;1)<em>(1-P(sii&gt;2)) = 0.7</em>(1-0.5)<br>\nP(sii=3) = P(sii&gt;2) = 0.5</p>\n<p>The datapoint is assigned the class with the highest probability<br>\nNote that this probability calculation is different than the one mentioned in the paper that I have tagged.</p>\n<p>Please let me know if you need further clarification.</p>",
      "rawMarkdown": "Hello Ravi. Thanks for the upvote. \n\nAll the models in the ensemble predict 3 logistic values (between 0 and 1). So I essentially converted the modelling to a multilabel logistic regression problem. (See the ordinal encoding section of the shared notebook)\n\nThe first value represents P(sii>0), second value represents P(sii>1), third value represents P(sii>2).\n\nSo for example the encoding [1,1,0] represents sii=2. \nThat is because in this encoding scheme P(sii>0)=1 (first index) and P(sii>1)=1 (second index) and P(sii>2)=0 (third index)\n\nSo what the encoding scheme is saying in this particular case  is that P(sii>0)=1 and P(sii>1) = 1 but P(sii>2)=0. Since we are dealing with discrete values here, this means that the sii = 2.\n\nSimilarly [1,1,1] represents 3 because P(sii>0)=1, P(sii>1)=1 and P(sii>2)=1. This can only mean that sii=3\n\nThe model tries to predict this encoding scheme. So for e.g. if the model predicts [0.9,0.7,0.5] for a particular datapoint, the probability for sii=1 or 2 or 3 is evaluated as follows:\n\nP(sii=0) = 1-P(sii>0) = 1-0.9 = 0.1 \nP(sii=1) = P(sii>0) and P(sii<=1) = P(sii>0)*(1-P(sii>1)) = 0.9*(1-0.7)\nP(sii=2) = P(sii>1 and P(sii<=2)) = P(sii>1)*(1-P(sii>2)) = 0.7*(1-0.5)\nP(sii=3) = P(sii>2) = 0.5\n\nThe datapoint is assigned the class with the highest probability\nNote that this probability calculation is different than the one mentioned in the paper that I have tagged.\n\nPlease let me know if you need further clarification.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3077736,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/21/2024 10:53:40",
      "content": "<p>This looks like 4 binary classifiers. Am I right <a href=\"https://www.kaggle.com/abhisekdash37\" target=\"_blank\">@abhisekdash37</a>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3077767,
          "author_name": "abhisekdash37",
          "author_url": "",
          "post_date": "12/21/2024 11:32:30",
          "content": "<p>Hello Ravi. Thanks for the upvote. </p>\n<p>All the models in the ensemble predict 3 logistic values (between 0 and 1). So I essentially converted the modelling to a multilabel logistic regression problem. (See the ordinal encoding section of the shared notebook)</p>\n<p>The first value represents P(sii&gt;0), second value represents P(sii&gt;1), third value represents P(sii&gt;2).</p>\n<p>So for example the encoding [1,1,0] represents sii=2. <br>\nThat is because in this encoding scheme P(sii&gt;0)=1 (first index) and P(sii&gt;1)=1 (second index) and P(sii&gt;2)=0 (third index)</p>\n<p>So what the encoding scheme is saying in this particular case  is that P(sii&gt;0)=1 and P(sii&gt;1) = 1 but P(sii&gt;2)=0. Since we are dealing with discrete values here, this means that the sii = 2.</p>\n<p>Similarly [1,1,1] represents 3 because P(sii&gt;0)=1, P(sii&gt;1)=1 and P(sii&gt;2)=1. This can only mean that sii=3</p>\n<p>The model tries to predict this encoding scheme. So for e.g. if the model predicts [0.9,0.7,0.5] for a particular datapoint, the probability for sii=1 or 2 or 3 is evaluated as follows:</p>\n<p>P(sii=0) = 1-P(sii&gt;0) = 1-0.9 = 0.1 <br>\nP(sii=1) = P(sii&gt;0) and P(sii&lt;=1) = P(sii&gt;0)<em>(1-P(sii&gt;1)) = 0.9</em>(1-0.7)<br>\nP(sii=2) = P(sii&gt;1 and P(sii&lt;=2)) = P(sii&gt;1)<em>(1-P(sii&gt;2)) = 0.7</em>(1-0.5)<br>\nP(sii=3) = P(sii&gt;2) = 0.5</p>\n<p>The datapoint is assigned the class with the highest probability<br>\nNote that this probability calculation is different than the one mentioned in the paper that I have tagged.</p>\n<p>Please let me know if you need further clarification.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3077724": "Hello.\n\nI did not select this notebook as my final solutions, but this ended up as the highest scoring. (LOL)\n\nSalient Features of the modelling process:\n\n1. Ensemble of XGBoost, Random Forest and LightGBM with cross validation.\n2. Descriptive statistics of enmo and light from actigraphy data. \n3. Ordinal encoding of target sii. Ordinal encoding is based on the paper [A Simple Approach to Ordinal Classification](https://researchcommons.waikato.ac.nz/server/api/core/bitstreams/d4decc50-dcf2-486b-9860-0c650b92b173/content)\nThis type of encoding preserves the ordinal relation between the different sii values. \nThe model tries to predict  P(sii) > 0, P(sii) > 1 and P(sii) >2 \n\n\nFor e.g. in the ordinal encoding section of the notebook,\n[0,0,0] represents sii=0\n[1,0,0] represents sii = 1\n[1,1,0] represents sii = 2\n[1,1,1] represents sii = 3\n\n[1,1,0] means that P(sii>0) = 1, P(sii>1) = 1 and P(sii>2)=0. So naturally the datapoint has sii value =2.\n[0,0,0] means that P(sii>0) = 0, P(sii>1) = 0 and P(sii>2) = 0. So naturally the datapoint has sii=0.\n[1,1,1] means that P(sii>0)=1,P(sii>1)=1 and P(sii>2)=1, So naturally the datapoint has sii=3.  \n\nThe probability of sii is calculated by the following formulae. (Note that this is slightly different than the method described in the paper)\n\nP(sii=0) = 1-P(sii>0)\n\nP(sii=1) = P(sii>0)*(1-P(sii>1)) (In plain english this means probability of sii > 0 *and* probability of sii <=1)\n\nP(sii=2) = P(sii>1)*(1-P(sii>2))\n\nP(sii=3) = P(sii>2)*(1-P(sii>3))\n\n\nThe classification of sii is the one with highest probability as per the above formulae. For e.g. if P(sii=2) is highest, the datapoint is classified as 2\n\nLink to notebook: https://www.kaggle.com/code/abhisekdash37/ensemble-learning-problematic-internet-use?scriptVersionId=205432552",
    "3077736": "This looks like 4 binary classifiers. Am I right @abhisekdash37?",
    "3077767": "Hello Ravi. Thanks for the upvote. \n\nAll the models in the ensemble predict 3 logistic values (between 0 and 1). So I essentially converted the modelling to a multilabel logistic regression problem. (See the ordinal encoding section of the shared notebook)\n\nThe first value represents P(sii>0), second value represents P(sii>1), third value represents P(sii>2).\n\nSo for example the encoding [1,1,0] represents sii=2. \nThat is because in this encoding scheme P(sii>0)=1 (first index) and P(sii>1)=1 (second index) and P(sii>2)=0 (third index)\n\nSo what the encoding scheme is saying in this particular case  is that P(sii>0)=1 and P(sii>1) = 1 but P(sii>2)=0. Since we are dealing with discrete values here, this means that the sii = 2.\n\nSimilarly [1,1,1] represents 3 because P(sii>0)=1, P(sii>1)=1 and P(sii>2)=1. This can only mean that sii=3\n\nThe model tries to predict this encoding scheme. So for e.g. if the model predicts [0.9,0.7,0.5] for a particular datapoint, the probability for sii=1 or 2 or 3 is evaluated as follows:\n\nP(sii=0) = 1-P(sii>0) = 1-0.9 = 0.1 \nP(sii=1) = P(sii>0) and P(sii<=1) = P(sii>0)*(1-P(sii>1)) = 0.9*(1-0.7)\nP(sii=2) = P(sii>1 and P(sii<=2)) = P(sii>1)*(1-P(sii>2)) = 0.7*(1-0.5)\nP(sii=3) = P(sii>2) = 0.5\n\nThe datapoint is assigned the class with the highest probability\nNote that this probability calculation is different than the one mentioned in the paper that I have tagged.\n\nPlease let me know if you need further clarification."
  },
  "source": "meta"
}