{
  "id": 536407,
  "title": "PCIAT variables ??",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/536407",
  "author_name": "",
  "post_date": "2024-09-27T10:22:49.873648100Z",
  "votes": 19,
  "comment_count": 16,
  "views": 0,
  "content": "<p>As PCIAT variables data is missing from test set what should we do?<br>\ndo we remove all PCIAT from train as well ?</p>",
  "messages": [
    {
      "id": "3000125",
      "postDate": "09/27/2024 10:22:49",
      "content": "<p>As PCIAT variables data is missing from test set what should we do?<br>\ndo we remove all PCIAT from train as well ?</p>",
      "rawMarkdown": "As PCIAT variables data is missing from test set what should we do?\ndo we remove all PCIAT from train as well ?",
      "votes": null
    },
    {
      "id": "3000226",
      "postDate": "09/27/2024 12:16:23",
      "content": "<p>This is probably the way most of us do it😭</p>",
      "rawMarkdown": "This is probably the way most of us do it😭",
      "votes": null
    },
    {
      "id": "3000246",
      "postDate": "09/27/2024 12:39:04",
      "content": "<p>My current belief is that test set true sii is based on PCIAT_total, but of course PCIAT_total in test set is hidden before disclosed to us.<br>\nI'm now trying to generalize PCIAT_total to all train data, and then using the train data to predict a PCIAT_total in test set, and convert it to sii for submission. Basically a LGBM on top of another LGBM.<br>\nOf course the flow is tedious, and the result is not promising yet.<br>\nGood luck!</p>",
      "rawMarkdown": "My current belief is that test set true sii is based on PCIAT_total, but of course PCIAT_total in test set is hidden before disclosed to us.\nI'm now trying to generalize PCIAT_total to all train data, and then using the train data to predict a PCIAT_total in test set, and convert it to sii for submission. Basically a LGBM on top of another LGBM.\nOf course the flow is tedious, and the result is not promising yet.\nGood luck!",
      "votes": null
    },
    {
      "id": "3000248",
      "postDate": "09/27/2024 12:41:45",
      "content": "<p>This is the best option if you don't wish to structure <strong>sii</strong> from <strong>PCIAT-Total</strong> <a href=\"https://www.kaggle.com/chetan8007\" target=\"_blank\">@chetan8007</a> </p>",
      "rawMarkdown": "This is the best option if you don't wish to structure **sii** from **PCIAT-Total** @chetan8007",
      "votes": null
    },
    {
      "id": "3000472",
      "postDate": "09/27/2024 16:48:50",
      "content": "<p>Hi,</p>\n<p>'sii' comes from PCIAT-total and PCIAT-total is the sum of the 20 PCIAT-PCiAT-xx.</p>\n<p>Instead of fitting directly a model which predict 'sii', we can try to fit models to predict each of the 20 PCIAT-PCIAT, then compute PCIAT-total as the sum of PCIAT-PCIAT's predictions and then compute 'sii' predictions.</p>\n<p>Because there is several ways to have a high PCIAT-total value (or a high \"sii' value)</p>\n<pre><code>display((train[] - train[[  i  ()]].(axis = )).value_counts(dropna=))\n\n_ = {i:  i  ()}\n_.update({i:  i  (, )})\n_.update({i:  i  (, )})\n_.update({i:  i  (, )})\n\ndisplay((train[].(_) - train[]).value_counts(dropna=))\n</code></pre>\n<pre><code>    \n    \n   \n\n    \n    \n   \n</code></pre>",
      "rawMarkdown": "Hi,\n\n'sii' comes from PCIAT-total and PCIAT-total is the sum of the 20 PCIAT-PCiAT-xx.\n\nInstead of fitting directly a model which predict 'sii', we can try to fit models to predict each of the 20 PCIAT-PCIAT, then compute PCIAT-total as the sum of PCIAT-PCIAT's predictions and then compute 'sii' predictions.\n\nBecause there is several ways to have a high PCIAT-total value (or a high \"sii' value)\n\n```python\ndisplay((train[\"PCIAT-PCIAT_Total\"] - train[[f\"PCIAT-PCIAT_{i+1:02}\" for i in range(20)]].sum(axis = 1)).value_counts(dropna=False))\n\n_map = {i:0 for i in range(31)}\n_map.update({i:1 for i in range(31, 50)})\n_map.update({i:2 for i in range(50, 80)})\n_map.update({i:3 for i in range(80, 101)})\n\ndisplay((train[\"PCIAT-PCIAT_Total\"].map(_map) - train['sii']).value_counts(dropna=False))\n```\n```\n0.0    2736\nNaN    1224\nName: count, dtype: int64\n\n0.0    2736\nNaN    1224\nName: count, dtype: int64\n```",
      "votes": null
    },
    {
      "id": "3000602",
      "postDate": "09/27/2024 19:51:34",
      "content": "<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> I tried this to no good effect. LB score is quite tepid.</p>",
      "rawMarkdown": "adaubas I tried this to no good effect. LB score is quite tepid.",
      "votes": null
    },
    {
      "id": "3000620",
      "postDate": "09/27/2024 20:45:06",
      "content": "<p><code>PCIAT-Season</code> is NaN if and only if <code>sii</code> is NaN. </p>\n<pre><code>np.(train[].isna() ^ train[].isna())\n</code></pre>\n<pre><code>0\n</code></pre>\n<p>However, even if <code>PCIAT-Season</code> is one of the valid values (<code>Spring</code>, <code>Fall</code>, <code>Summer</code>, <code>Winter</code>), some of the summands can be NaN, and <code>PCIAT-PCIAT_Total</code> is the <code>nansum</code> of them. If <code>PCIAT-Season</code> is NaN, then all the summands would be NaN, but <code>PCIAT-PCIAT_Total</code> would be NaN too, not the <code>nansum</code> (which would be 0).</p>\n<p>So, when modeling as a multi-output regression/classification problem, one possible approach is, first remove all the rows with NaN <code>PCIAT-Season</code>, and then <code>fillna</code> all the summand columns in the remaining rows with 0 so that the multi-output target would not have any missing values. This is assuming that in the (hidden) test set, the latent <code>PCIAT-Season</code> variable would never be NaN, since <code>sii</code> is not supposed to be NaN in the test set.</p>\n<pre><code>X = train[~train[].isna()]\ndf = {}\ncols = [  i  ()]\n col  cols:\n    df.update({col:X[col].value_counts(dropna=)})\npd.DataFrame(df)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1048539%2F04e5de4a4a9a337f5c34f444fb5fcd8b%2FCapture.PNG?generation=1727469823753177&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "`PCIAT-Season` is NaN if and only if `sii` is NaN. \n```python\nnp.sum(train['PCIAT-Season'].isna() ^ train['sii'].isna())\n```\n```text\n0\n```\nHowever, even if `PCIAT-Season` is one of the valid values (`Spring`, `Fall`, `Summer`, `Winter`), some of the summands can be NaN, and `PCIAT-PCIAT_Total` is the `nansum` of them. If `PCIAT-Season` is NaN, then all the summands would be NaN, but `PCIAT-PCIAT_Total` would be NaN too, not the `nansum` (which would be 0).\n\nSo, when modeling as a multi-output regression/classification problem, one possible approach is, first remove all the rows with NaN `PCIAT-Season`, and then `fillna` all the summand columns in the remaining rows with 0 so that the multi-output target would not have any missing values. This is assuming that in the (hidden) test set, the latent `PCIAT-Season` variable would never be NaN, since `sii` is not supposed to be NaN in the test set.\n```python\nX = train[~train['PCIAT-Season'].isna()]\ndf = {}\ncols = [F'PCIAT-PCIAT_{i+1:02d}' for i in range(20)]\nfor col in cols:\n    df.update({col:X[col].value_counts(dropna=False)})\npd.DataFrame(df)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1048539%2F04e5de4a4a9a337f5c34f444fb5fcd8b%2FCapture.PNG?generation=1727469823753177&alt=media)",
      "votes": null
    },
    {
      "id": "3000864",
      "postDate": "09/28/2024 06:55:24",
      "content": "<p>Thank you for feedback <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "Thank you for feedback @ravi20076",
      "votes": null
    },
    {
      "id": "3000928",
      "postDate": "09/28/2024 09:14:20",
      "content": "<p>Wow, nice catch <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a> ! That might explain some of the discrepancies between SII values and other features, I will check (<em>UPD: i did, 17 rows have incorrect SII, but nothing explained, discrepancies persist</em>)!  It would be really helpful to know the reason why some questions of the Parent-Child Internet Addiction Test can be ignored…</p>",
      "rawMarkdown": "Wow, nice catch @siukeitin ! That might explain some of the discrepancies between SII values and other features, I will check (*UPD: i did, 17 rows have incorrect SII, but nothing explained, discrepancies persist*)!  It would be really helpful to know the reason why some questions of the Parent-Child Internet Addiction Test can be ignored...",
      "votes": null
    },
    {
      "id": "3001407",
      "postDate": "09/28/2024 20:00:58",
      "content": "<p>Would the best option not be to use the PCIAT-Total score as the target as it can provide more information to the model about what scores are on the borders of the SII classes?</p>",
      "rawMarkdown": "Would the best option not be to use the PCIAT-Total score as the target as it can provide more information to the model about what scores are on the borders of the SII classes?",
      "votes": null
    },
    {
      "id": "3002100",
      "postDate": "09/29/2024 15:58:01",
      "content": "<p>As \"sii\" we have many missing values, so colud be the possiblity to take PICAT columns average value as a target value..  what you people think..</p>",
      "rawMarkdown": "As \"sii\" we have many missing values, so colud be the possiblity to take PICAT columns average value as a target value..  what you people think..",
      "votes": null
    },
    {
      "id": "3002413",
      "postDate": "09/30/2024 02:39:58",
      "content": "<p>No that won’t be possible, as usually both the target and PCIAT variables are missing.</p>",
      "rawMarkdown": "No that won’t be possible, as usually both the target and PCIAT variables are missing.",
      "votes": null
    },
    {
      "id": "3003402",
      "postDate": "10/01/2024 01:24:48",
      "content": "<p>Hi Antonina, can you share which discrepancy you found? I may be missing smth but find the sii corresponds to the sum converted. thanks!</p>",
      "rawMarkdown": "Hi Antonina, can you share which discrepancy you found? I may be missing smth but find the sii corresponds to the sum converted. thanks!",
      "votes": null
    },
    {
      "id": "3003634",
      "postDate": "10/01/2024 06:16:51",
      "content": "<p>Sorry if I didn't get your point, but just in case:</p>\n<p>PCIAT-PCIAT_Total is calculated as the sum of the scores for the answers to 20 questions (stored in PCIAT-PCIAT_01 to PCIAT-PCIAT_20). Some of the questions may be ignored by a respondent (missing values in the PCIAT columns), but the SII score is still calculated as the sum of the non-NA values, resulting in potentially invalid SII values:</p>\n<p>e.g. PCIAT-PCIAT_Total between 31 and 49 corresponds to SII = 1 and between 50 and 79 corresponds to SII = 2. A respondent has 47 points and did not answer 1 question. If he did, his PCIAT-PCIAT_Total could be 47 to 52 (max 5 points per question), so his actual SII could be 1 or 2. But in the data we have 1. </p>",
      "rawMarkdown": "Sorry if I didn't get your point, but just in case:\n\nPCIAT-PCIAT_Total is calculated as the sum of the scores for the answers to 20 questions (stored in PCIAT-PCIAT_01 to PCIAT-PCIAT_20). Some of the questions may be ignored by a respondent (missing values in the PCIAT columns), but the SII score is still calculated as the sum of the non-NA values, resulting in potentially invalid SII values:\n\ne.g. PCIAT-PCIAT_Total between 31 and 49 corresponds to SII = 1 and between 50 and 79 corresponds to SII = 2. A respondent has 47 points and did not answer 1 question. If he did, his PCIAT-PCIAT_Total could be 47 to 52 (max 5 points per question), so his actual SII could be 1 or 2. But in the data we have 1.",
      "votes": null
    },
    {
      "id": "3006687",
      "postDate": "10/04/2024 12:28:58",
      "content": "<p><a href=\"https://www.kaggle.com/chetan8007\" target=\"_blank\">@chetan8007</a> I've reviewed all the comments and haven't found any additional insights. If you've drawn any conclusions or developed a method for handling the PCIAT data or imputing the sii values, I would appreciate your guidance.</p>",
      "rawMarkdown": "chetan8007 I've reviewed all the comments and haven't found any additional insights. If you've drawn any conclusions or developed a method for handling the PCIAT data or imputing the sii values, I would appreciate your guidance.",
      "votes": null
    },
    {
      "id": "3008224",
      "postDate": "10/06/2024 11:35:01",
      "content": "<p>Hello,I'm doing this job ,and I want to know what't your LB score by predicting every PCIAT?</p>",
      "rawMarkdown": "Hello,I'm doing this job ,and I want to know what't your LB score by predicting every PCIAT?",
      "votes": null
    },
    {
      "id": "3073986",
      "postDate": "12/17/2024 07:24:22",
      "content": "<p>Thank you for your reply</p>",
      "rawMarkdown": "Thank you for your reply",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3000226,
      "author_name": "wayne127",
      "author_url": "",
      "post_date": "09/27/2024 12:16:23",
      "content": "<p>This is probably the way most of us do it😭</p>",
      "votes": null,
      "replies": [
        {
          "id": 3002100,
          "author_name": "sadiasadique",
          "author_url": "",
          "post_date": "09/29/2024 15:58:01",
          "content": "<p>As \"sii\" we have many missing values, so colud be the possiblity to take PICAT columns average value as a target value..  what you people think..</p>",
          "votes": null,
          "replies": [
            {
              "id": 3002413,
              "author_name": "peterhopkinson",
              "author_url": "",
              "post_date": "09/30/2024 02:39:58",
              "content": "<p>No that won’t be possible, as usually both the target and PCIAT variables are missing.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3000246,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "09/27/2024 12:39:04",
      "content": "<p>My current belief is that test set true sii is based on PCIAT_total, but of course PCIAT_total in test set is hidden before disclosed to us.<br>\nI'm now trying to generalize PCIAT_total to all train data, and then using the train data to predict a PCIAT_total in test set, and convert it to sii for submission. Basically a LGBM on top of another LGBM.<br>\nOf course the flow is tedious, and the result is not promising yet.<br>\nGood luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3000248,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "09/27/2024 12:41:45",
      "content": "<p>This is the best option if you don't wish to structure <strong>sii</strong> from <strong>PCIAT-Total</strong> <a href=\"https://www.kaggle.com/chetan8007\" target=\"_blank\">@chetan8007</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3001407,
          "author_name": "peterhopkinson",
          "author_url": "",
          "post_date": "09/28/2024 20:00:58",
          "content": "<p>Would the best option not be to use the PCIAT-Total score as the target as it can provide more information to the model about what scores are on the borders of the SII classes?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3000472,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "09/27/2024 16:48:50",
      "content": "<p>Hi,</p>\n<p>'sii' comes from PCIAT-total and PCIAT-total is the sum of the 20 PCIAT-PCiAT-xx.</p>\n<p>Instead of fitting directly a model which predict 'sii', we can try to fit models to predict each of the 20 PCIAT-PCIAT, then compute PCIAT-total as the sum of PCIAT-PCIAT's predictions and then compute 'sii' predictions.</p>\n<p>Because there is several ways to have a high PCIAT-total value (or a high \"sii' value)</p>\n<pre><code>display((train[] - train[[  i  ()]].(axis = )).value_counts(dropna=))\n\n_ = {i:  i  ()}\n_.update({i:  i  (, )})\n_.update({i:  i  (, )})\n_.update({i:  i  (, )})\n\ndisplay((train[].(_) - train[]).value_counts(dropna=))\n</code></pre>\n<pre><code>    \n    \n   \n\n    \n    \n   \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3000602,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "09/27/2024 19:51:34",
          "content": "<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> I tried this to no good effect. LB score is quite tepid.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3000864,
              "author_name": "adaubas",
              "author_url": "",
              "post_date": "09/28/2024 06:55:24",
              "content": "<p>Thank you for feedback <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3008224,
              "author_name": "yashi003",
              "author_url": "",
              "post_date": "10/06/2024 11:35:01",
              "content": "<p>Hello,I'm doing this job ,and I want to know what't your LB score by predicting every PCIAT?</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3000620,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "09/27/2024 20:45:06",
          "content": "<p><code>PCIAT-Season</code> is NaN if and only if <code>sii</code> is NaN. </p>\n<pre><code>np.(train[].isna() ^ train[].isna())\n</code></pre>\n<pre><code>0\n</code></pre>\n<p>However, even if <code>PCIAT-Season</code> is one of the valid values (<code>Spring</code>, <code>Fall</code>, <code>Summer</code>, <code>Winter</code>), some of the summands can be NaN, and <code>PCIAT-PCIAT_Total</code> is the <code>nansum</code> of them. If <code>PCIAT-Season</code> is NaN, then all the summands would be NaN, but <code>PCIAT-PCIAT_Total</code> would be NaN too, not the <code>nansum</code> (which would be 0).</p>\n<p>So, when modeling as a multi-output regression/classification problem, one possible approach is, first remove all the rows with NaN <code>PCIAT-Season</code>, and then <code>fillna</code> all the summand columns in the remaining rows with 0 so that the multi-output target would not have any missing values. This is assuming that in the (hidden) test set, the latent <code>PCIAT-Season</code> variable would never be NaN, since <code>sii</code> is not supposed to be NaN in the test set.</p>\n<pre><code>X = train[~train[].isna()]\ndf = {}\ncols = [  i  ()]\n col  cols:\n    df.update({col:X[col].value_counts(dropna=)})\npd.DataFrame(df)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1048539%2F04e5de4a4a9a337f5c34f444fb5fcd8b%2FCapture.PNG?generation=1727469823753177&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 3000928,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "09/28/2024 09:14:20",
              "content": "<p>Wow, nice catch <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a> ! That might explain some of the discrepancies between SII values and other features, I will check (<em>UPD: i did, 17 rows have incorrect SII, but nothing explained, discrepancies persist</em>)!  It would be really helpful to know the reason why some questions of the Parent-Child Internet Addiction Test can be ignored…</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3003402,
                  "author_name": "ibabic",
                  "author_url": "",
                  "post_date": "10/01/2024 01:24:48",
                  "content": "<p>Hi Antonina, can you share which discrepancy you found? I may be missing smth but find the sii corresponds to the sum converted. thanks!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3003634,
                      "author_name": "antoninadolgorukova",
                      "author_url": "",
                      "post_date": "10/01/2024 06:16:51",
                      "content": "<p>Sorry if I didn't get your point, but just in case:</p>\n<p>PCIAT-PCIAT_Total is calculated as the sum of the scores for the answers to 20 questions (stored in PCIAT-PCIAT_01 to PCIAT-PCIAT_20). Some of the questions may be ignored by a respondent (missing values in the PCIAT columns), but the SII score is still calculated as the sum of the non-NA values, resulting in potentially invalid SII values:</p>\n<p>e.g. PCIAT-PCIAT_Total between 31 and 49 corresponds to SII = 1 and between 50 and 79 corresponds to SII = 2. A respondent has 47 points and did not answer 1 question. If he did, his PCIAT-PCIAT_Total could be 47 to 52 (max 5 points per question), so his actual SII could be 1 or 2. But in the data we have 1. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3073986,
          "author_name": "prosepoem",
          "author_url": "",
          "post_date": "12/17/2024 07:24:22",
          "content": "<p>Thank you for your reply</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3006687,
      "author_name": "nassimsfaxi",
      "author_url": "",
      "post_date": "10/04/2024 12:28:58",
      "content": "<p><a href=\"https://www.kaggle.com/chetan8007\" target=\"_blank\">@chetan8007</a> I've reviewed all the comments and haven't found any additional insights. If you've drawn any conclusions or developed a method for handling the PCIAT data or imputing the sii values, I would appreciate your guidance.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3000125": "As PCIAT variables data is missing from test set what should we do?\ndo we remove all PCIAT from train as well ?",
    "3000226": "This is probably the way most of us do it😭",
    "3000246": "My current belief is that test set true sii is based on PCIAT_total, but of course PCIAT_total in test set is hidden before disclosed to us.\nI'm now trying to generalize PCIAT_total to all train data, and then using the train data to predict a PCIAT_total in test set, and convert it to sii for submission. Basically a LGBM on top of another LGBM.\nOf course the flow is tedious, and the result is not promising yet.\nGood luck!",
    "3000248": "This is the best option if you don't wish to structure **sii** from **PCIAT-Total** @chetan8007",
    "3000472": "Hi,\n\n'sii' comes from PCIAT-total and PCIAT-total is the sum of the 20 PCIAT-PCiAT-xx.\n\nInstead of fitting directly a model which predict 'sii', we can try to fit models to predict each of the 20 PCIAT-PCIAT, then compute PCIAT-total as the sum of PCIAT-PCIAT's predictions and then compute 'sii' predictions.\n\nBecause there is several ways to have a high PCIAT-total value (or a high \"sii' value)\n\n```python\ndisplay((train[\"PCIAT-PCIAT_Total\"] - train[[f\"PCIAT-PCIAT_{i+1:02}\" for i in range(20)]].sum(axis = 1)).value_counts(dropna=False))\n\n_map = {i:0 for i in range(31)}\n_map.update({i:1 for i in range(31, 50)})\n_map.update({i:2 for i in range(50, 80)})\n_map.update({i:3 for i in range(80, 101)})\n\ndisplay((train[\"PCIAT-PCIAT_Total\"].map(_map) - train['sii']).value_counts(dropna=False))\n```\n```\n0.0    2736\nNaN    1224\nName: count, dtype: int64\n\n0.0    2736\nNaN    1224\nName: count, dtype: int64\n```",
    "3000602": "adaubas I tried this to no good effect. LB score is quite tepid.",
    "3000620": "`PCIAT-Season` is NaN if and only if `sii` is NaN. \n```python\nnp.sum(train['PCIAT-Season'].isna() ^ train['sii'].isna())\n```\n```text\n0\n```\nHowever, even if `PCIAT-Season` is one of the valid values (`Spring`, `Fall`, `Summer`, `Winter`), some of the summands can be NaN, and `PCIAT-PCIAT_Total` is the `nansum` of them. If `PCIAT-Season` is NaN, then all the summands would be NaN, but `PCIAT-PCIAT_Total` would be NaN too, not the `nansum` (which would be 0).\n\nSo, when modeling as a multi-output regression/classification problem, one possible approach is, first remove all the rows with NaN `PCIAT-Season`, and then `fillna` all the summand columns in the remaining rows with 0 so that the multi-output target would not have any missing values. This is assuming that in the (hidden) test set, the latent `PCIAT-Season` variable would never be NaN, since `sii` is not supposed to be NaN in the test set.\n```python\nX = train[~train['PCIAT-Season'].isna()]\ndf = {}\ncols = [F'PCIAT-PCIAT_{i+1:02d}' for i in range(20)]\nfor col in cols:\n    df.update({col:X[col].value_counts(dropna=False)})\npd.DataFrame(df)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1048539%2F04e5de4a4a9a337f5c34f444fb5fcd8b%2FCapture.PNG?generation=1727469823753177&alt=media)",
    "3000864": "Thank you for feedback @ravi20076",
    "3000928": "Wow, nice catch @siukeitin ! That might explain some of the discrepancies between SII values and other features, I will check (*UPD: i did, 17 rows have incorrect SII, but nothing explained, discrepancies persist*)!  It would be really helpful to know the reason why some questions of the Parent-Child Internet Addiction Test can be ignored...",
    "3001407": "Would the best option not be to use the PCIAT-Total score as the target as it can provide more information to the model about what scores are on the borders of the SII classes?",
    "3002100": "As \"sii\" we have many missing values, so colud be the possiblity to take PICAT columns average value as a target value..  what you people think..",
    "3002413": "No that won’t be possible, as usually both the target and PCIAT variables are missing.",
    "3003402": "Hi Antonina, can you share which discrepancy you found? I may be missing smth but find the sii corresponds to the sum converted. thanks!",
    "3003634": "Sorry if I didn't get your point, but just in case:\n\nPCIAT-PCIAT_Total is calculated as the sum of the scores for the answers to 20 questions (stored in PCIAT-PCIAT_01 to PCIAT-PCIAT_20). Some of the questions may be ignored by a respondent (missing values in the PCIAT columns), but the SII score is still calculated as the sum of the non-NA values, resulting in potentially invalid SII values:\n\ne.g. PCIAT-PCIAT_Total between 31 and 49 corresponds to SII = 1 and between 50 and 79 corresponds to SII = 2. A respondent has 47 points and did not answer 1 question. If he did, his PCIAT-PCIAT_Total could be 47 to 52 (max 5 points per question), so his actual SII could be 1 or 2. But in the data we have 1.",
    "3006687": "chetan8007 I've reviewed all the comments and haven't found any additional insights. If you've drawn any conclusions or developed a method for handling the PCIAT data or imputing the sii values, I would appreciate your guidance.",
    "3008224": "Hello,I'm doing this job ,and I want to know what't your LB score by predicting every PCIAT?",
    "3073986": "Thank you for your reply"
  },
  "source": "meta"
}