{
  "id": 546731,
  "title": "Any insight from the 20-item scale (sii) for this ordinal classification problem?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/546731",
  "author_name": "Fang Zitao",
  "post_date": "2024-11-17T17:08:36.189000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The task for this competition is to predict <code>sii</code>, where <code>0</code> means <code>none</code>, <code>1</code> means <code>mild</code>, <code>2</code> means <code>moderate</code>, and <code>3</code> means <code>severe</code>, an <strong>ordinal classification problem</strong>. The <code>sii</code> is calculated from the <em>Parent-Child Internet Addiction Test (PCIAT)</em>, a 20-item scale used to measure characteristics and behaviors associated with compulsive use of the Internet. The following are the scale questions organized according to <code>data_dictionary.csv</code>:</p>\n<blockquote>\n  <ol>\n  <li>How often does your child disobey time limits you set for online use?</li>\n  <li>How often does your child neglect household chores to spend more time online?</li>\n  <li>How often does your child prefer to spend time online rather than with the rest of your family?</li>\n  <li>How often does your child form new relationships with fellow online users?</li>\n  <li>How often do you complain about the amount of time your child spends online?</li>\n  <li>How often do your child's grades suffer because of the amount of time he or she spends online?</li>\n  <li>How often does your child check his or her e-mail before doing something else?</li>\n  <li>How often does your child seem withdrawn from others since discovering the Internet?</li>\n  <li>How often does your child become defensive or secretive when asked what he or she does online?</li>\n  <li>How often have you caught your child sneaking online against your wishes?</li>\n  <li>How often does your child spend time along in his or her room playing on the computer?</li>\n  <li>How often does your child receive strange phone calls from new \"online\" friends?</li>\n  <li>How often does your child snap, yell, or act annoyed if bothered while online?</li>\n  <li>How often does your child seem more tired and fatigued than he or she did before the Internet came along?</li>\n  <li>How often does your child seem preoccupied with being back online when off-line?</li>\n  <li>How often does your child throw tantrums with your interference about how long he or she spends online?</li>\n  <li>How often does your child choose to spend time online rather than doing once enjoyed hobbies and/or outside interests?</li>\n  <li>How often does your child become angry or belligerent when your place time limits on how much time he or shes is allowed to spend online?</li>\n  <li>How often does your child choose to spend more time online than going out with friends?</li>\n  <li>How often does your child feel depressed, moody, or nervous when off-line which seems to go away once back online?</li>\n  </ol>\n</blockquote>\n<p>where <code>0</code> is <code>does not apply</code>, <code>1</code> is <code>rarely</code>, <code>2</code> is <code>occasionally</code>, <code>3</code> is <code>frequently</code>, <code>4</code> is <code>often</code>, <code>5</code> is <code>always</code>, then the final <code>sii</code> is measured from:</p>\n<table>\n<thead>\n<tr>\n<th>Score Range</th>\n<th>Description</th>\n<th>Label</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 ~ 30</td>\n<td>None</td>\n<td>0</td>\n</tr>\n<tr>\n<td>31 ~ 49</td>\n<td>Mild</td>\n<td>1</td>\n</tr>\n<tr>\n<td>50 ~ 79</td>\n<td>Moderate</td>\n<td>2</td>\n</tr>\n<tr>\n<td>80-100</td>\n<td>Severe</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>\n<p>Above is the information we can get from the official. And thanks to <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a> and <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>, they found the problem in <code>sii</code> calculation <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536407#3000620\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536407#3000620\" target=\"_blank\">here</a>. Specifically, there are several rows where the <code>sii</code> scores are calculated from some <em>PCIAT</em> results with <code>nan</code> value, and a total of <strong>17 rows</strong> will have different <code>sii</code> due to different <code>nan</code> results.</p>\n<p>Quoting <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> from <a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda?scriptVersionId=206130660&amp;cellId=4\" target=\"_blank\">here</a>:</p>\n<blockquote>\n  <p>Type of Machine Learning Problem we can use with sii as a target:</p>\n  <ol>\n  <li>Ordinal classification (ordinal logistic regression, models with custom ordinal loss functions);</li>\n  <li>Multiclass classification (treat sii as a nominal categorical variable without considering the order);</li>\n  <li>Regression (ignore the discrete nature of categories and treat sii as a continuous variable, then round prediction);</li>\n  <li>Custom (e.g. loss functions that penalize errors based on the distance between categories);</li>\n  <li>use <code>PCIAT-PCIAT_Total</code> as a continuous target variable, and implement regression on <code>PCIAT-PCIAT_Total</code> and then map predictions to <code>sii</code> categories.</li>\n  </ol>\n</blockquote>\n<p>The above scoring mechanism and everyone's discussion together lead to some interesting thoughts:</p>\n<ul>\n<li><strong>Considering ordinary during training</strong>: In the current public code, many traditional machine learning methods choose not to consider the ordinal situation during training, and then design a separate model to find the boundary at the end, which may have a negative impact on the model effect;</li>\n<li><strong>Consider 20 sub-questions independently</strong>: Is it necessary to consider the twenty questions separately, or just refer to the final summation? To be more specific, perhaps we could even consider some groups of questions independently;</li>\n<li><strong>Re-scoring</strong>: Some of the above 20 questions are more subjective, and others are relatively objective. For those subjective questions, scoring may introduce greater errors and bring uncertainty to the final score. Especially when our dataset is moderately small, will reducing this error be more helpful for fitting private test sets with unknown distributions?</li>\n<li><strong>Pseudo labels</strong>: Decouple the task into two stages: <ol>\n<li>Use existing data to predict a large amount of missing data and labels; </li>\n<li>Use the generated pseudo data and pseudo labels to train again (pseudo labels can be assigned a smaller weight than the real labels);</li></ol></li>\n<li>…</li>\n</ul>\n<p>Do you have any more in-depth exploration or insights💡? Looking forward to seeing everyone's imaginative ideas and results~</p>",
  "messages": [
    {
      "id": 3048240,
      "postDate": "2024-11-17T17:08:36.190Z",
      "content": "<p>The task for this competition is to predict <code>sii</code>, where <code>0</code> means <code>none</code>, <code>1</code> means <code>mild</code>, <code>2</code> means <code>moderate</code>, and <code>3</code> means <code>severe</code>, an <strong>ordinal classification problem</strong>. The <code>sii</code> is calculated from the <em>Parent-Child Internet Addiction Test (PCIAT)</em>, a 20-item scale used to measure characteristics and behaviors associated with compulsive use of the Internet. The following are the scale questions organized according to <code>data_dictionary.csv</code>:</p>\n<blockquote>\n  <ol>\n  <li>How often does your child disobey time limits you set for online use?</li>\n  <li>How often does your child neglect household chores to spend more time online?</li>\n  <li>How often does your child prefer to spend time online rather than with the rest of your family?</li>\n  <li>How often does your child form new relationships with fellow online users?</li>\n  <li>How often do you complain about the amount of time your child spends online?</li>\n  <li>How often do your child's grades suffer because of the amount of time he or she spends online?</li>\n  <li>How often does your child check his or her e-mail before doing something else?</li>\n  <li>How often does your child seem withdrawn from others since discovering the Internet?</li>\n  <li>How often does your child become defensive or secretive when asked what he or she does online?</li>\n  <li>How often have you caught your child sneaking online against your wishes?</li>\n  <li>How often does your child spend time along in his or her room playing on the computer?</li>\n  <li>How often does your child receive strange phone calls from new \"online\" friends?</li>\n  <li>How often does your child snap, yell, or act annoyed if bothered while online?</li>\n  <li>How often does your child seem more tired and fatigued than he or she did before the Internet came along?</li>\n  <li>How often does your child seem preoccupied with being back online when off-line?</li>\n  <li>How often does your child throw tantrums with your interference about how long he or she spends online?</li>\n  <li>How often does your child choose to spend time online rather than doing once enjoyed hobbies and/or outside interests?</li>\n  <li>How often does your child become angry or belligerent when your place time limits on how much time he or shes is allowed to spend online?</li>\n  <li>How often does your child choose to spend more time online than going out with friends?</li>\n  <li>How often does your child feel depressed, moody, or nervous when off-line which seems to go away once back online?</li>\n  </ol>\n</blockquote>\n<p>where <code>0</code> is <code>does not apply</code>, <code>1</code> is <code>rarely</code>, <code>2</code> is <code>occasionally</code>, <code>3</code> is <code>frequently</code>, <code>4</code> is <code>often</code>, <code>5</code> is <code>always</code>, then the final <code>sii</code> is measured from:</p>\n<table>\n<thead>\n<tr>\n<th>Score Range</th>\n<th>Description</th>\n<th>Label</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 ~ 30</td>\n<td>None</td>\n<td>0</td>\n</tr>\n<tr>\n<td>31 ~ 49</td>\n<td>Mild</td>\n<td>1</td>\n</tr>\n<tr>\n<td>50 ~ 79</td>\n<td>Moderate</td>\n<td>2</td>\n</tr>\n<tr>\n<td>80-100</td>\n<td>Severe</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>\n<p>Above is the information we can get from the official. And thanks to <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a> and <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>, they found the problem in <code>sii</code> calculation <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536407#3000620\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536407#3000620\" target=\"_blank\">here</a>. Specifically, there are several rows where the <code>sii</code> scores are calculated from some <em>PCIAT</em> results with <code>nan</code> value, and a total of <strong>17 rows</strong> will have different <code>sii</code> due to different <code>nan</code> results.</p>\n<p>Quoting <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> from <a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda?scriptVersionId=206130660&amp;cellId=4\" target=\"_blank\">here</a>:</p>\n<blockquote>\n  <p>Type of Machine Learning Problem we can use with sii as a target:</p>\n  <ol>\n  <li>Ordinal classification (ordinal logistic regression, models with custom ordinal loss functions);</li>\n  <li>Multiclass classification (treat sii as a nominal categorical variable without considering the order);</li>\n  <li>Regression (ignore the discrete nature of categories and treat sii as a continuous variable, then round prediction);</li>\n  <li>Custom (e.g. loss functions that penalize errors based on the distance between categories);</li>\n  <li>use <code>PCIAT-PCIAT_Total</code> as a continuous target variable, and implement regression on <code>PCIAT-PCIAT_Total</code> and then map predictions to <code>sii</code> categories.</li>\n  </ol>\n</blockquote>\n<p>The above scoring mechanism and everyone's discussion together lead to some interesting thoughts:</p>\n<ul>\n<li><strong>Considering ordinary during training</strong>: In the current public code, many traditional machine learning methods choose not to consider the ordinal situation during training, and then design a separate model to find the boundary at the end, which may have a negative impact on the model effect;</li>\n<li><strong>Consider 20 sub-questions independently</strong>: Is it necessary to consider the twenty questions separately, or just refer to the final summation? To be more specific, perhaps we could even consider some groups of questions independently;</li>\n<li><strong>Re-scoring</strong>: Some of the above 20 questions are more subjective, and others are relatively objective. For those subjective questions, scoring may introduce greater errors and bring uncertainty to the final score. Especially when our dataset is moderately small, will reducing this error be more helpful for fitting private test sets with unknown distributions?</li>\n<li><strong>Pseudo labels</strong>: Decouple the task into two stages: <ol>\n<li>Use existing data to predict a large amount of missing data and labels; </li>\n<li>Use the generated pseudo data and pseudo labels to train again (pseudo labels can be assigned a smaller weight than the real labels);</li></ol></li>\n<li>…</li>\n</ul>\n<p>Do you have any more in-depth exploration or insights💡? Looking forward to seeing everyone's imaginative ideas and results~</p>",
      "rawMarkdown": "The task for this competition is to predict `sii`, where `0` means `none`, `1` means `mild`, `2` means `moderate`, and `3` means `severe`, an **ordinal classification problem**. The `sii` is calculated from the *Parent-Child Internet Addiction Test (PCIAT)*, a 20-item scale used to measure characteristics and behaviors associated with compulsive use of the Internet. The following are the scale questions organized according to `data_dictionary.csv`:\n\n> 1. How often does your child disobey time limits you set for online use?\n> 2. How often does your child neglect household chores to spend more time online?\n> 3. How often does your child prefer to spend time online rather than with the rest of your family?\n> 4. How often does your child form new relationships with fellow online users?\n> 5. How often do you complain about the amount of time your child spends online?\n> 6. How often do your child's grades suffer because of the amount of time he or she spends online?\n> 7. How often does your child check his or her e-mail before doing something else?\n> 8. How often does your child seem withdrawn from others since discovering the Internet?\n> 9. How often does your child become defensive or secretive when asked what he or she does online?\n> 10. How often have you caught your child sneaking online against your wishes?\n> 11. How often does your child spend time along in his or her room playing on the computer?\n> 12. How often does your child receive strange phone calls from new \"online\" friends?\n> 13. How often does your child snap, yell, or act annoyed if bothered while online?\n> 14. How often does your child seem more tired and fatigued than he or she did before the Internet came along?\n> 15. How often does your child seem preoccupied with being back online when off-line?\n> 16. How often does your child throw tantrums with your interference about how long he or she spends online?\n> 17. How often does your child choose to spend time online rather than doing once enjoyed hobbies and/or outside interests?\n> 18. How often does your child become angry or belligerent when your place time limits on how much time he or shes is allowed to spend online?\n> 19. How often does your child choose to spend more time online than going out with friends?\n> 20. How often does your child feel depressed, moody, or nervous when off-line which seems to go away once back online?\n\nwhere `0` is `does not apply`, `1` is `rarely`, `2` is `occasionally`, `3` is `frequently`, `4` is `often`, `5` is `always`, then the final `sii` is measured from:\n\n| Score Range | Description | Label |\n| :---: | :---: | :---: |\n| 0 ~ 30 | None | 0 |\n| 31 ~ 49 | Mild | 1 |\n| 50 ~ 79 | Moderate | 2 |\n| 80-100 | Severe | 3 |\n\nAbove is the information we can get from the official. And thanks to @siukeitin and @antoninadolgorukova, they found the problem in `sii` calculation [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536407#3000620) and [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536407#3000620). Specifically, there are several rows where the `sii` scores are calculated from some *PCIAT* results with `nan` value, and a total of **17 rows** will have different `sii` due to different `nan` results.\n\nQuoting @antoninadolgorukova from [here](https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda?scriptVersionId=206130660&cellId=4):\n\n> Type of Machine Learning Problem we can use with sii as a target:\n> 1. Ordinal classification (ordinal logistic regression, models with custom ordinal loss functions);\n> 2. Multiclass classification (treat sii as a nominal categorical variable without considering the order);\n> 3. Regression (ignore the discrete nature of categories and treat sii as a continuous variable, then round prediction);\n> 4. Custom (e.g. loss functions that penalize errors based on the distance between categories);\n> 5. use `PCIAT-PCIAT_Total` as a continuous target variable, and implement regression on `PCIAT-PCIAT_Total` and then map predictions to `sii` categories.\n\nThe above scoring mechanism and everyone's discussion together lead to some interesting thoughts:\n- **Considering ordinary during training**: In the current public code, many traditional machine learning methods choose not to consider the ordinal situation during training, and then design a separate model to find the boundary at the end, which may have a negative impact on the model effect;\n- **Consider 20 sub-questions independently**: Is it necessary to consider the twenty questions separately, or just refer to the final summation? To be more specific, perhaps we could even consider some groups of questions independently;\n- **Re-scoring**: Some of the above 20 questions are more subjective, and others are relatively objective. For those subjective questions, scoring may introduce greater errors and bring uncertainty to the final score. Especially when our dataset is moderately small, will reducing this error be more helpful for fitting private test sets with unknown distributions?\n- **Pseudo labels**: Decouple the task into two stages: \n  1. Use existing data to predict a large amount of missing data and labels; \n  2. Use the generated pseudo data and pseudo labels to train again (pseudo labels can be assigned a smaller weight than the real labels);\n- ...\n\nDo you have any more in-depth exploration or insights💡? Looking forward to seeing everyone's imaginative ideas and results~",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3048240": "The task for this competition is to predict `sii`, where `0` means `none`, `1` means `mild`, `2` means `moderate`, and `3` means `severe`, an **ordinal classification problem**. The `sii` is calculated from the *Parent-Child Internet Addiction Test (PCIAT)*, a 20-item scale used to measure characteristics and behaviors associated with compulsive use of the Internet. The following are the scale questions organized according to `data_dictionary.csv`:\n\n> 1. How often does your child disobey time limits you set for online use?\n> 2. How often does your child neglect household chores to spend more time online?\n> 3. How often does your child prefer to spend time online rather than with the rest of your family?\n> 4. How often does your child form new relationships with fellow online users?\n> 5. How often do you complain about the amount of time your child spends online?\n> 6. How often do your child's grades suffer because of the amount of time he or she spends online?\n> 7. How often does your child check his or her e-mail before doing something else?\n> 8. How often does your child seem withdrawn from others since discovering the Internet?\n> 9. How often does your child become defensive or secretive when asked what he or she does online?\n> 10. How often have you caught your child sneaking online against your wishes?\n> 11. How often does your child spend time along in his or her room playing on the computer?\n> 12. How often does your child receive strange phone calls from new \"online\" friends?\n> 13. How often does your child snap, yell, or act annoyed if bothered while online?\n> 14. How often does your child seem more tired and fatigued than he or she did before the Internet came along?\n> 15. How often does your child seem preoccupied with being back online when off-line?\n> 16. How often does your child throw tantrums with your interference about how long he or she spends online?\n> 17. How often does your child choose to spend time online rather than doing once enjoyed hobbies and/or outside interests?\n> 18. How often does your child become angry or belligerent when your place time limits on how much time he or shes is allowed to spend online?\n> 19. How often does your child choose to spend more time online than going out with friends?\n> 20. How often does your child feel depressed, moody, or nervous when off-line which seems to go away once back online?\n\nwhere `0` is `does not apply`, `1` is `rarely`, `2` is `occasionally`, `3` is `frequently`, `4` is `often`, `5` is `always`, then the final `sii` is measured from:\n\n| Score Range | Description | Label |\n| :---: | :---: | :---: |\n| 0 ~ 30 | None | 0 |\n| 31 ~ 49 | Mild | 1 |\n| 50 ~ 79 | Moderate | 2 |\n| 80-100 | Severe | 3 |\n\nAbove is the information we can get from the official. And thanks to @siukeitin and @antoninadolgorukova, they found the problem in `sii` calculation [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536407#3000620) and [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/536407#3000620). Specifically, there are several rows where the `sii` scores are calculated from some *PCIAT* results with `nan` value, and a total of **17 rows** will have different `sii` due to different `nan` results.\n\nQuoting @antoninadolgorukova from [here](https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda?scriptVersionId=206130660&cellId=4):\n\n> Type of Machine Learning Problem we can use with sii as a target:\n> 1. Ordinal classification (ordinal logistic regression, models with custom ordinal loss functions);\n> 2. Multiclass classification (treat sii as a nominal categorical variable without considering the order);\n> 3. Regression (ignore the discrete nature of categories and treat sii as a continuous variable, then round prediction);\n> 4. Custom (e.g. loss functions that penalize errors based on the distance between categories);\n> 5. use `PCIAT-PCIAT_Total` as a continuous target variable, and implement regression on `PCIAT-PCIAT_Total` and then map predictions to `sii` categories.\n\nThe above scoring mechanism and everyone's discussion together lead to some interesting thoughts:\n- **Considering ordinary during training**: In the current public code, many traditional machine learning methods choose not to consider the ordinal situation during training, and then design a separate model to find the boundary at the end, which may have a negative impact on the model effect;\n- **Consider 20 sub-questions independently**: Is it necessary to consider the twenty questions separately, or just refer to the final summation? To be more specific, perhaps we could even consider some groups of questions independently;\n- **Re-scoring**: Some of the above 20 questions are more subjective, and others are relatively objective. For those subjective questions, scoring may introduce greater errors and bring uncertainty to the final score. Especially when our dataset is moderately small, will reducing this error be more helpful for fitting private test sets with unknown distributions?\n- **Pseudo labels**: Decouple the task into two stages: \n  1. Use existing data to predict a large amount of missing data and labels; \n  2. Use the generated pseudo data and pseudo labels to train again (pseudo labels can be assigned a smaller weight than the real labels);\n- ...\n\nDo you have any more in-depth exploration or insights💡? Looking forward to seeing everyone's imaginative ideas and results~"
  }
}