{
  "id": 553030,
  "title": "11th Place Solution for the Child Mind Institute — Problematic Internet Use Competition",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/553030",
  "author_name": "Hezhi Xie",
  "post_date": "2024-12-23T07:19:29.425000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>We are thrilled to achieve a gold medal in our first Kaggle competition and deeply grateful to everyone at Kaggle for all that you’ve shared, which greatly contributed to our learning journey. In this write-up, we’re excited to share our solution and key takeaways.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/overview\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data</a></li>\n</ul>\n<h1>Overview of the Approach</h1>\n<p>Our model builds on this notebook: <a href=\"https://www.kaggle.com/code/laiyunghwei/lb-0-493\" target=\"_blank\">https://www.kaggle.com/code/laiyunghwei/lb-0-493</a>. We used three gradient boosting models-LightGBM, XGBoost, and CatBoost-and combined them using a voting regressor with equal weights.</p>\n<h1>Details of the Submission</h1>\n<h2>Exploratory Data Analysis (EDA)</h2>\n<h3>Missing values</h3>\n<p>Missing data was a significant challenge. Nearly all features, except demographic ones, contained missing values, with about 10 features having missing rates exceeding 70%. The target variable also had approximately 30% missing data.</p>\n<p>Additionally, features from the same instrument had similar missing rates, and the missing values were randomly distributed, allowing us to handle them through either direct removal or imputation.</p>\n<h3>Outliers</h3>\n<p>We identified multiple outliers after a detailed examination of each feature:</p>\n<table>\n<thead>\n<tr>\n<th>Feature</th>\n<th>Outlier Standard</th>\n<th>Number of Outliers</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>CGAS-CGAS_Score</code></td>\n<td>Extremely high value (e.g., 999)</td>\n<td>1</td>\n</tr>\n<tr>\n<td><code>Physical-Weight</code></td>\n<td>Invalid weight value (e.g., 0)</td>\n<td>61</td>\n</tr>\n<tr>\n<td>BIA related features (such as <code>BIA-BIA_TBW</code>, <code>BIA-BIA_ICW</code>)</td>\n<td>Values significantly higher or lower compared to the normal range</td>\n<td>2</td>\n</tr>\n<tr>\n<td><code>Physical-HeartRate</code></td>\n<td>Abnormally low heart rate (&lt;30 bpm)</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Total Body Water percentage (<code>BIA-BIA_TBW</code>/<code>Physical-Weight</code>)</td>\n<td>&lt;20% or &gt;100%</td>\n<td>10</td>\n</tr>\n</tbody>\n</table>\n<h3>Highly correlated features</h3>\n<p>A correlation heatmap revealed that features from the same instrument were often highly correlated (e.g., Bio-electric Impedance Analysis (BIA) features had correlations &gt; 0.9).</p>\n<p>Additionally, some features, like BMI, could be derived from others (BMI = FFMI + FMI). Reducing such redundancy was crucial for simplifying gradient boosting models.</p>\n<h3>Correlation with age</h3>\n<p>Physical features like height and weight exhibited strong correlations with age. By normalizing these features (e.g., height/age, weight/age), we can extract more meaningful information. And age-based regression seems a reasonable method for imputing missing values.</p>\n<h3>Time series data</h3>\n<p>Time series data had high missing rates and minimal impact on prediction accuracy, so we ultimately excluded it from our solution.</p>\n<h2>Data Preprocessing</h2>\n<h3>Handling outliers</h3>\n<p>We removed entries with outlier values (explained in the EDA part) to avoid skewing the model.</p>\n<h3>Feature engineering and selection</h3>\n<p>We used feature engineering to extract more valuable insights and remove duplicate information, and dropped features with too many missing values, redundancy, or strong correlations. The final set of selected features included:</p>\n<table>\n<thead>\n<tr>\n<th>No.</th>\n<th>Feature</th>\n<th>Explanation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td><code>Basic_Demos-Age</code></td>\n<td>Participant's age</td>\n</tr>\n<tr>\n<td>2</td>\n<td><code>Basic_Demos-Sex</code></td>\n<td>Participant's sex</td>\n</tr>\n<tr>\n<td>3</td>\n<td><code>CGAS-CGAS_Score</code></td>\n<td>Children's Global Assessment Scale (CGAS) score</td>\n</tr>\n<tr>\n<td>4</td>\n<td><code>Physical-Height_per_Age</code></td>\n<td><code>Physical-Height</code>/<code>Basic_Demos-Age</code></td>\n</tr>\n<tr>\n<td>5</td>\n<td><code>Physical-Weight_per_Age</code></td>\n<td><code>Physical-Weight</code>/<code>Basic_Demos-Age</code></td>\n</tr>\n<tr>\n<td>6</td>\n<td><code>FGC_Zone_Total</code></td>\n<td><code>FGC_CU_Zone</code>+<code>FGC_SRL_Zone</code>+<code>FGC_SRR_Zone</code>+<code>FGC_TL_Zone</code></td>\n</tr>\n<tr>\n<td>7</td>\n<td><code>BIA-BIA_Activity_Level_num</code></td>\n<td>Activity level</td>\n</tr>\n<tr>\n<td>8</td>\n<td><code>BIA-BIA_BMR</code></td>\n<td>Basal Metabolic Rate</td>\n</tr>\n<tr>\n<td>9</td>\n<td><code>BIA-BIA_DEE</code></td>\n<td>Daily Energy Expenditure</td>\n</tr>\n<tr>\n<td>10</td>\n<td><code>BIA-BIA_FFMI</code></td>\n<td>Fat-Free Mass Index</td>\n</tr>\n<tr>\n<td>11</td>\n<td><code>BIA-BIA_FMI</code></td>\n<td>Fat Mass Index</td>\n</tr>\n<tr>\n<td>12</td>\n<td><code>BIA-BIA_Frame_num</code></td>\n<td>Frame size</td>\n</tr>\n<tr>\n<td>13</td>\n<td><code>BIA-BIA_SMM</code></td>\n<td>Skeletal Muscle Mass</td>\n</tr>\n<tr>\n<td>14</td>\n<td><code>BIA-BIA_TBW</code></td>\n<td>Total Body Water volume</td>\n</tr>\n<tr>\n<td>15</td>\n<td><code>SDS-SDS_Total_T</code></td>\n<td>Sleep Disturbance Scale total score</td>\n</tr>\n<tr>\n<td>16</td>\n<td><code>PreInt_EduHx-computerinternet_hoursday</code></td>\n<td>Average hours per day spent on computer/internet usage</td>\n</tr>\n<tr>\n<td>17</td>\n<td><code>PAQ_Total</code></td>\n<td>Physical Activity Questionnaire score, combined by <code>PAQ_A-PAQ_A_Total</code> and <code>PAQ_C-PAQ_C_Total</code></td>\n</tr>\n</tbody>\n</table>\n<h3>Imputation</h3>\n<p>We evaluated two imputation methods: age-based regression and a hybrid approach combining K-Nearest Neighbors (KNN) for continuous variables and Random Forest for categorical variables. Both methods are reasonable depending on the data characteristics.</p>\n<p>In this project, most features are closely related to physical development, making linear regression with age a suitable choice for imputing missing values. This approach ensures that the imputed values align with the observed data distribution.</p>\n<p>The hybrid method leverages the assumption that continuous features follow local similarity patterns, which KNN can effectively capture, and categorical features are sufficiently correlated with other features, allowing Random Forest to provide accurate predictions.</p>\n<p>After testing both methods, age-based regression was selected as it demonstrated better performance for this dataset.</p>\n<h2>Modeling and Prediction</h2>\n<p>We tested three methods: 1) Ridge Regression; 2) Multilayer Perceptron (MLP); 3) Gradient Boosting. Gradient boosting consistently outperformed the other methods with a stronger Quadratic Weighted Kappa (QWK) score, so we stick to this method. We optimized hyperparameters using Optuna and validated results with 5-fold cross-validation.</p>\n<h1>Key Takeaways</h1>\n<p>This competition was not only a test of technical skills but also an exercise in handling real-world data challenges. While proud of our results, we see room for improvement and valuable lessons for future projects. </p>\n<ul>\n<li><strong>Feature engineering and selection matters</strong>: Due to significant data noise, carefully removing redundant or irrelevant features can greatly improve the performance of gradient boosting models.</li>\n<li><strong>Data imbalance is also a limitation</strong>: The target variable (SII) was highly imbalanced, with about 60% of values being 0 and 85% less than 2. This imbalance complicated the task of establishing clear relationships between features and the target variable.</li>\n<li><strong>Physical data alone may be insufficient</strong>: Physical data alone may not fully explain problematic internet use (PIU). For example, some children with high SII scores displayed strong fitness metrics, which, while unexpected, is reasonable. Including additional data, such as behavioral patterns, may provide more insights.</li>\n<li><strong>Need for objective measurements</strong>: During the competition, we couldn’t help but question the reliability of using Parent-Child Internet Addiction Test (PCIAT) results as a measure of PIU. Since it’s based on parent-reported answers, it can be subjective and potentially inaccurate. We think it might be possible to develop a more objective way to measure PIU. For example, the feature <code>PreInt_EduHx_computerinternet_hoursday</code>, which directly tracks computer usage time, seems like a more reliable indicator.</li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/laiyunghwei/lb-0-493\" target=\"_blank\">https://www.kaggle.com/code/laiyunghwei/lb-0-493</a></li>\n<li><a href=\"https://www.sciencedirect.com/science/article/pii/S1386505624001047\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1386505624001047</a></li>\n</ul>",
  "messages": [
    {
      "id": 3079059,
      "postDate": "2024-12-23T07:19:29.427Z",
      "content": "<p>We are thrilled to achieve a gold medal in our first Kaggle competition and deeply grateful to everyone at Kaggle for all that you’ve shared, which greatly contributed to our learning journey. In this write-up, we’re excited to share our solution and key takeaways.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/overview\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data</a></li>\n</ul>\n<h1>Overview of the Approach</h1>\n<p>Our model builds on this notebook: <a href=\"https://www.kaggle.com/code/laiyunghwei/lb-0-493\" target=\"_blank\">https://www.kaggle.com/code/laiyunghwei/lb-0-493</a>. We used three gradient boosting models-LightGBM, XGBoost, and CatBoost-and combined them using a voting regressor with equal weights.</p>\n<h1>Details of the Submission</h1>\n<h2>Exploratory Data Analysis (EDA)</h2>\n<h3>Missing values</h3>\n<p>Missing data was a significant challenge. Nearly all features, except demographic ones, contained missing values, with about 10 features having missing rates exceeding 70%. The target variable also had approximately 30% missing data.</p>\n<p>Additionally, features from the same instrument had similar missing rates, and the missing values were randomly distributed, allowing us to handle them through either direct removal or imputation.</p>\n<h3>Outliers</h3>\n<p>We identified multiple outliers after a detailed examination of each feature:</p>\n<table>\n<thead>\n<tr>\n<th>Feature</th>\n<th>Outlier Standard</th>\n<th>Number of Outliers</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>CGAS-CGAS_Score</code></td>\n<td>Extremely high value (e.g., 999)</td>\n<td>1</td>\n</tr>\n<tr>\n<td><code>Physical-Weight</code></td>\n<td>Invalid weight value (e.g., 0)</td>\n<td>61</td>\n</tr>\n<tr>\n<td>BIA related features (such as <code>BIA-BIA_TBW</code>, <code>BIA-BIA_ICW</code>)</td>\n<td>Values significantly higher or lower compared to the normal range</td>\n<td>2</td>\n</tr>\n<tr>\n<td><code>Physical-HeartRate</code></td>\n<td>Abnormally low heart rate (&lt;30 bpm)</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Total Body Water percentage (<code>BIA-BIA_TBW</code>/<code>Physical-Weight</code>)</td>\n<td>&lt;20% or &gt;100%</td>\n<td>10</td>\n</tr>\n</tbody>\n</table>\n<h3>Highly correlated features</h3>\n<p>A correlation heatmap revealed that features from the same instrument were often highly correlated (e.g., Bio-electric Impedance Analysis (BIA) features had correlations &gt; 0.9).</p>\n<p>Additionally, some features, like BMI, could be derived from others (BMI = FFMI + FMI). Reducing such redundancy was crucial for simplifying gradient boosting models.</p>\n<h3>Correlation with age</h3>\n<p>Physical features like height and weight exhibited strong correlations with age. By normalizing these features (e.g., height/age, weight/age), we can extract more meaningful information. And age-based regression seems a reasonable method for imputing missing values.</p>\n<h3>Time series data</h3>\n<p>Time series data had high missing rates and minimal impact on prediction accuracy, so we ultimately excluded it from our solution.</p>\n<h2>Data Preprocessing</h2>\n<h3>Handling outliers</h3>\n<p>We removed entries with outlier values (explained in the EDA part) to avoid skewing the model.</p>\n<h3>Feature engineering and selection</h3>\n<p>We used feature engineering to extract more valuable insights and remove duplicate information, and dropped features with too many missing values, redundancy, or strong correlations. The final set of selected features included:</p>\n<table>\n<thead>\n<tr>\n<th>No.</th>\n<th>Feature</th>\n<th>Explanation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td><code>Basic_Demos-Age</code></td>\n<td>Participant's age</td>\n</tr>\n<tr>\n<td>2</td>\n<td><code>Basic_Demos-Sex</code></td>\n<td>Participant's sex</td>\n</tr>\n<tr>\n<td>3</td>\n<td><code>CGAS-CGAS_Score</code></td>\n<td>Children's Global Assessment Scale (CGAS) score</td>\n</tr>\n<tr>\n<td>4</td>\n<td><code>Physical-Height_per_Age</code></td>\n<td><code>Physical-Height</code>/<code>Basic_Demos-Age</code></td>\n</tr>\n<tr>\n<td>5</td>\n<td><code>Physical-Weight_per_Age</code></td>\n<td><code>Physical-Weight</code>/<code>Basic_Demos-Age</code></td>\n</tr>\n<tr>\n<td>6</td>\n<td><code>FGC_Zone_Total</code></td>\n<td><code>FGC_CU_Zone</code>+<code>FGC_SRL_Zone</code>+<code>FGC_SRR_Zone</code>+<code>FGC_TL_Zone</code></td>\n</tr>\n<tr>\n<td>7</td>\n<td><code>BIA-BIA_Activity_Level_num</code></td>\n<td>Activity level</td>\n</tr>\n<tr>\n<td>8</td>\n<td><code>BIA-BIA_BMR</code></td>\n<td>Basal Metabolic Rate</td>\n</tr>\n<tr>\n<td>9</td>\n<td><code>BIA-BIA_DEE</code></td>\n<td>Daily Energy Expenditure</td>\n</tr>\n<tr>\n<td>10</td>\n<td><code>BIA-BIA_FFMI</code></td>\n<td>Fat-Free Mass Index</td>\n</tr>\n<tr>\n<td>11</td>\n<td><code>BIA-BIA_FMI</code></td>\n<td>Fat Mass Index</td>\n</tr>\n<tr>\n<td>12</td>\n<td><code>BIA-BIA_Frame_num</code></td>\n<td>Frame size</td>\n</tr>\n<tr>\n<td>13</td>\n<td><code>BIA-BIA_SMM</code></td>\n<td>Skeletal Muscle Mass</td>\n</tr>\n<tr>\n<td>14</td>\n<td><code>BIA-BIA_TBW</code></td>\n<td>Total Body Water volume</td>\n</tr>\n<tr>\n<td>15</td>\n<td><code>SDS-SDS_Total_T</code></td>\n<td>Sleep Disturbance Scale total score</td>\n</tr>\n<tr>\n<td>16</td>\n<td><code>PreInt_EduHx-computerinternet_hoursday</code></td>\n<td>Average hours per day spent on computer/internet usage</td>\n</tr>\n<tr>\n<td>17</td>\n<td><code>PAQ_Total</code></td>\n<td>Physical Activity Questionnaire score, combined by <code>PAQ_A-PAQ_A_Total</code> and <code>PAQ_C-PAQ_C_Total</code></td>\n</tr>\n</tbody>\n</table>\n<h3>Imputation</h3>\n<p>We evaluated two imputation methods: age-based regression and a hybrid approach combining K-Nearest Neighbors (KNN) for continuous variables and Random Forest for categorical variables. Both methods are reasonable depending on the data characteristics.</p>\n<p>In this project, most features are closely related to physical development, making linear regression with age a suitable choice for imputing missing values. This approach ensures that the imputed values align with the observed data distribution.</p>\n<p>The hybrid method leverages the assumption that continuous features follow local similarity patterns, which KNN can effectively capture, and categorical features are sufficiently correlated with other features, allowing Random Forest to provide accurate predictions.</p>\n<p>After testing both methods, age-based regression was selected as it demonstrated better performance for this dataset.</p>\n<h2>Modeling and Prediction</h2>\n<p>We tested three methods: 1) Ridge Regression; 2) Multilayer Perceptron (MLP); 3) Gradient Boosting. Gradient boosting consistently outperformed the other methods with a stronger Quadratic Weighted Kappa (QWK) score, so we stick to this method. We optimized hyperparameters using Optuna and validated results with 5-fold cross-validation.</p>\n<h1>Key Takeaways</h1>\n<p>This competition was not only a test of technical skills but also an exercise in handling real-world data challenges. While proud of our results, we see room for improvement and valuable lessons for future projects. </p>\n<ul>\n<li><strong>Feature engineering and selection matters</strong>: Due to significant data noise, carefully removing redundant or irrelevant features can greatly improve the performance of gradient boosting models.</li>\n<li><strong>Data imbalance is also a limitation</strong>: The target variable (SII) was highly imbalanced, with about 60% of values being 0 and 85% less than 2. This imbalance complicated the task of establishing clear relationships between features and the target variable.</li>\n<li><strong>Physical data alone may be insufficient</strong>: Physical data alone may not fully explain problematic internet use (PIU). For example, some children with high SII scores displayed strong fitness metrics, which, while unexpected, is reasonable. Including additional data, such as behavioral patterns, may provide more insights.</li>\n<li><strong>Need for objective measurements</strong>: During the competition, we couldn’t help but question the reliability of using Parent-Child Internet Addiction Test (PCIAT) results as a measure of PIU. Since it’s based on parent-reported answers, it can be subjective and potentially inaccurate. We think it might be possible to develop a more objective way to measure PIU. For example, the feature <code>PreInt_EduHx_computerinternet_hoursday</code>, which directly tracks computer usage time, seems like a more reliable indicator.</li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/laiyunghwei/lb-0-493\" target=\"_blank\">https://www.kaggle.com/code/laiyunghwei/lb-0-493</a></li>\n<li><a href=\"https://www.sciencedirect.com/science/article/pii/S1386505624001047\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1386505624001047</a></li>\n</ul>",
      "rawMarkdown": "We are thrilled to achieve a gold medal in our first Kaggle competition and deeply grateful to everyone at Kaggle for all that you’ve shared, which greatly contributed to our learning journey. In this write-up, we’re excited to share our solution and key takeaways.\n\n# Context\n- Business context: https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/overview\n- Data context: https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data\n\n# Overview of the Approach\nOur model builds on this notebook: https://www.kaggle.com/code/laiyunghwei/lb-0-493. We used three gradient boosting models-LightGBM, XGBoost, and CatBoost-and combined them using a voting regressor with equal weights.\n\n# Details of the Submission\n## Exploratory Data Analysis (EDA)\n### Missing values\nMissing data was a significant challenge. Nearly all features, except demographic ones, contained missing values, with about 10 features having missing rates exceeding 70%. The target variable also had approximately 30% missing data.\n\nAdditionally, features from the same instrument had similar missing rates, and the missing values were randomly distributed, allowing us to handle them through either direct removal or imputation.\n\n### Outliers \nWe identified multiple outliers after a detailed examination of each feature:\n\n| Feature | Outlier Standard | Number of Outliers |\n| --- | --- | --- |\n| `CGAS-CGAS_Score` | Extremely high value (e.g., 999) | 1 |\n| `Physical-Weight` | Invalid weight value (e.g., 0) | 61 |\n| BIA related features (such as `BIA-BIA_TBW`, `BIA-BIA_ICW`) | Values significantly higher or lower compared to the normal range | 2 |\n| `Physical-HeartRate` | Abnormally low heart rate (<30 bpm) | 1 |\n| Total Body Water percentage (`BIA-BIA_TBW`/`Physical-Weight`) | <20% or >100% | 10 |\n\n### Highly correlated features\nA correlation heatmap revealed that features from the same instrument were often highly correlated (e.g., Bio-electric Impedance Analysis (BIA) features had correlations > 0.9).\n\nAdditionally, some features, like BMI, could be derived from others (BMI = FFMI + FMI). Reducing such redundancy was crucial for simplifying gradient boosting models.\n\n### Correlation with age\nPhysical features like height and weight exhibited strong correlations with age. By normalizing these features (e.g., height/age, weight/age), we can extract more meaningful information. And age-based regression seems a reasonable method for imputing missing values.\n\n### Time series data\nTime series data had high missing rates and minimal impact on prediction accuracy, so we ultimately excluded it from our solution.\n\n## Data Preprocessing\n### Handling outliers\nWe removed entries with outlier values (explained in the EDA part) to avoid skewing the model.\n\n### Feature engineering and selection\nWe used feature engineering to extract more valuable insights and remove duplicate information, and dropped features with too many missing values, redundancy, or strong correlations. The final set of selected features included:\n\n| No. | Feature | Explanation\n| --- | --- | --- |\n| 1 | `Basic_Demos-Age` | Participant's age |\n| 2 | `Basic_Demos-Sex` | Participant's sex |\n| 3 | `CGAS-CGAS_Score` | Children's Global Assessment Scale (CGAS) score |\n| 4 | `Physical-Height_per_Age` | `Physical-Height`/`Basic_Demos-Age` | \n| 5 | `Physical-Weight_per_Age` | `Physical-Weight`/`Basic_Demos-Age` |\n| 6 | `FGC_Zone_Total` | `FGC_CU_Zone`+`FGC_SRL_Zone`+`FGC_SRR_Zone`+`FGC_TL_Zone` |\n| 7 | `BIA-BIA_Activity_Level_num` | Activity level |\n| 8 | `BIA-BIA_BMR` | Basal Metabolic Rate |\n| 9 | `BIA-BIA_DEE` | Daily Energy Expenditure |\n| 10 |`BIA-BIA_FFMI` | Fat-Free Mass Index | \n| 11 |`BIA-BIA_FMI` | Fat Mass Index |\n| 12 |`BIA-BIA_Frame_num` | Frame size |\n| 13 |`BIA-BIA_SMM` | Skeletal Muscle Mass | \n| 14 | `BIA-BIA_TBW` | Total Body Water volume |\n| 15 | `SDS-SDS_Total_T` | Sleep Disturbance Scale total score |\n| 16 | `PreInt_EduHx-computerinternet_hoursday` | Average hours per day spent on computer/internet usage |\n| 17 | `PAQ_Total` | Physical Activity Questionnaire score, combined by `PAQ_A-PAQ_A_Total` and `PAQ_C-PAQ_C_Total`  |\n\n### Imputation\nWe evaluated two imputation methods: age-based regression and a hybrid approach combining K-Nearest Neighbors (KNN) for continuous variables and Random Forest for categorical variables. Both methods are reasonable depending on the data characteristics.\n\nIn this project, most features are closely related to physical development, making linear regression with age a suitable choice for imputing missing values. This approach ensures that the imputed values align with the observed data distribution.\n\nThe hybrid method leverages the assumption that continuous features follow local similarity patterns, which KNN can effectively capture, and categorical features are sufficiently correlated with other features, allowing Random Forest to provide accurate predictions.\n\nAfter testing both methods, age-based regression was selected as it demonstrated better performance for this dataset.\n\n## Modeling and Prediction\nWe tested three methods: 1) Ridge Regression; 2) Multilayer Perceptron (MLP); 3) Gradient Boosting. Gradient boosting consistently outperformed the other methods with a stronger Quadratic Weighted Kappa (QWK) score, so we stick to this method. We optimized hyperparameters using Optuna and validated results with 5-fold cross-validation.\n\n#Key Takeaways\nThis competition was not only a test of technical skills but also an exercise in handling real-world data challenges. While proud of our results, we see room for improvement and valuable lessons for future projects. \n- **Feature engineering and selection matters**: Due to significant data noise, carefully removing redundant or irrelevant features can greatly improve the performance of gradient boosting models.\n- **Data imbalance is also a limitation**: The target variable (SII) was highly imbalanced, with about 60% of values being 0 and 85% less than 2. This imbalance complicated the task of establishing clear relationships between features and the target variable.\n- **Physical data alone may be insufficient**: Physical data alone may not fully explain problematic internet use (PIU). For example, some children with high SII scores displayed strong fitness metrics, which, while unexpected, is reasonable. Including additional data, such as behavioral patterns, may provide more insights.\n- **Need for objective measurements**: During the competition, we couldn’t help but question the reliability of using Parent-Child Internet Addiction Test (PCIAT) results as a measure of PIU. Since it’s based on parent-reported answers, it can be subjective and potentially inaccurate. We think it might be possible to develop a more objective way to measure PIU. For example, the feature `PreInt_EduHx_computerinternet_hoursday`, which directly tracks computer usage time, seems like a more reliable indicator.\n\n# Sources\n- https://www.kaggle.com/code/laiyunghwei/lb-0-493\n- https://www.sciencedirect.com/science/article/pii/S1386505624001047",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3079059": "We are thrilled to achieve a gold medal in our first Kaggle competition and deeply grateful to everyone at Kaggle for all that you’ve shared, which greatly contributed to our learning journey. In this write-up, we’re excited to share our solution and key takeaways.\n\n# Context\n- Business context: https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/overview\n- Data context: https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data\n\n# Overview of the Approach\nOur model builds on this notebook: https://www.kaggle.com/code/laiyunghwei/lb-0-493. We used three gradient boosting models-LightGBM, XGBoost, and CatBoost-and combined them using a voting regressor with equal weights.\n\n# Details of the Submission\n## Exploratory Data Analysis (EDA)\n### Missing values\nMissing data was a significant challenge. Nearly all features, except demographic ones, contained missing values, with about 10 features having missing rates exceeding 70%. The target variable also had approximately 30% missing data.\n\nAdditionally, features from the same instrument had similar missing rates, and the missing values were randomly distributed, allowing us to handle them through either direct removal or imputation.\n\n### Outliers \nWe identified multiple outliers after a detailed examination of each feature:\n\n| Feature | Outlier Standard | Number of Outliers |\n| --- | --- | --- |\n| `CGAS-CGAS_Score` | Extremely high value (e.g., 999) | 1 |\n| `Physical-Weight` | Invalid weight value (e.g., 0) | 61 |\n| BIA related features (such as `BIA-BIA_TBW`, `BIA-BIA_ICW`) | Values significantly higher or lower compared to the normal range | 2 |\n| `Physical-HeartRate` | Abnormally low heart rate (<30 bpm) | 1 |\n| Total Body Water percentage (`BIA-BIA_TBW`/`Physical-Weight`) | <20% or >100% | 10 |\n\n### Highly correlated features\nA correlation heatmap revealed that features from the same instrument were often highly correlated (e.g., Bio-electric Impedance Analysis (BIA) features had correlations > 0.9).\n\nAdditionally, some features, like BMI, could be derived from others (BMI = FFMI + FMI). Reducing such redundancy was crucial for simplifying gradient boosting models.\n\n### Correlation with age\nPhysical features like height and weight exhibited strong correlations with age. By normalizing these features (e.g., height/age, weight/age), we can extract more meaningful information. And age-based regression seems a reasonable method for imputing missing values.\n\n### Time series data\nTime series data had high missing rates and minimal impact on prediction accuracy, so we ultimately excluded it from our solution.\n\n## Data Preprocessing\n### Handling outliers\nWe removed entries with outlier values (explained in the EDA part) to avoid skewing the model.\n\n### Feature engineering and selection\nWe used feature engineering to extract more valuable insights and remove duplicate information, and dropped features with too many missing values, redundancy, or strong correlations. The final set of selected features included:\n\n| No. | Feature | Explanation\n| --- | --- | --- |\n| 1 | `Basic_Demos-Age` | Participant's age |\n| 2 | `Basic_Demos-Sex` | Participant's sex |\n| 3 | `CGAS-CGAS_Score` | Children's Global Assessment Scale (CGAS) score |\n| 4 | `Physical-Height_per_Age` | `Physical-Height`/`Basic_Demos-Age` | \n| 5 | `Physical-Weight_per_Age` | `Physical-Weight`/`Basic_Demos-Age` |\n| 6 | `FGC_Zone_Total` | `FGC_CU_Zone`+`FGC_SRL_Zone`+`FGC_SRR_Zone`+`FGC_TL_Zone` |\n| 7 | `BIA-BIA_Activity_Level_num` | Activity level |\n| 8 | `BIA-BIA_BMR` | Basal Metabolic Rate |\n| 9 | `BIA-BIA_DEE` | Daily Energy Expenditure |\n| 10 |`BIA-BIA_FFMI` | Fat-Free Mass Index | \n| 11 |`BIA-BIA_FMI` | Fat Mass Index |\n| 12 |`BIA-BIA_Frame_num` | Frame size |\n| 13 |`BIA-BIA_SMM` | Skeletal Muscle Mass | \n| 14 | `BIA-BIA_TBW` | Total Body Water volume |\n| 15 | `SDS-SDS_Total_T` | Sleep Disturbance Scale total score |\n| 16 | `PreInt_EduHx-computerinternet_hoursday` | Average hours per day spent on computer/internet usage |\n| 17 | `PAQ_Total` | Physical Activity Questionnaire score, combined by `PAQ_A-PAQ_A_Total` and `PAQ_C-PAQ_C_Total`  |\n\n### Imputation\nWe evaluated two imputation methods: age-based regression and a hybrid approach combining K-Nearest Neighbors (KNN) for continuous variables and Random Forest for categorical variables. Both methods are reasonable depending on the data characteristics.\n\nIn this project, most features are closely related to physical development, making linear regression with age a suitable choice for imputing missing values. This approach ensures that the imputed values align with the observed data distribution.\n\nThe hybrid method leverages the assumption that continuous features follow local similarity patterns, which KNN can effectively capture, and categorical features are sufficiently correlated with other features, allowing Random Forest to provide accurate predictions.\n\nAfter testing both methods, age-based regression was selected as it demonstrated better performance for this dataset.\n\n## Modeling and Prediction\nWe tested three methods: 1) Ridge Regression; 2) Multilayer Perceptron (MLP); 3) Gradient Boosting. Gradient boosting consistently outperformed the other methods with a stronger Quadratic Weighted Kappa (QWK) score, so we stick to this method. We optimized hyperparameters using Optuna and validated results with 5-fold cross-validation.\n\n#Key Takeaways\nThis competition was not only a test of technical skills but also an exercise in handling real-world data challenges. While proud of our results, we see room for improvement and valuable lessons for future projects. \n- **Feature engineering and selection matters**: Due to significant data noise, carefully removing redundant or irrelevant features can greatly improve the performance of gradient boosting models.\n- **Data imbalance is also a limitation**: The target variable (SII) was highly imbalanced, with about 60% of values being 0 and 85% less than 2. This imbalance complicated the task of establishing clear relationships between features and the target variable.\n- **Physical data alone may be insufficient**: Physical data alone may not fully explain problematic internet use (PIU). For example, some children with high SII scores displayed strong fitness metrics, which, while unexpected, is reasonable. Including additional data, such as behavioral patterns, may provide more insights.\n- **Need for objective measurements**: During the competition, we couldn’t help but question the reliability of using Parent-Child Internet Addiction Test (PCIAT) results as a measure of PIU. Since it’s based on parent-reported answers, it can be subjective and potentially inaccurate. We think it might be possible to develop a more objective way to measure PIU. For example, the feature `PreInt_EduHx_computerinternet_hoursday`, which directly tracks computer usage time, seems like a more reliable indicator.\n\n# Sources\n- https://www.kaggle.com/code/laiyunghwei/lb-0-493\n- https://www.sciencedirect.com/science/article/pii/S1386505624001047"
  }
}