{
  "id": 552760,
  "title": "8th Place Solution",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552760",
  "author_name": "Kzk Knmt",
  "post_date": "2024-12-21T11:12:02.738000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, I'd like to thank all the hosts and Kaggle staff for organizing such an interesting and practical competition!<br>\nI greatly enjoyed the process of trial and error for handling noisy and missing data!</p>\n<h1>OverView</h1>\n<ul>\n<li>I concentrated on improving CV, not LB due to the small number of data and unclear distribution of missing values.</li>\n<li>I extracted features by Null Importance to improve robustness.</li>\n<li>I imputed numerical missing features by Iterative Imputer (estimator=Bayesian Ridge).</li>\n<li>Ensemble in 3 different models. <ul>\n<li>train.csv + descriptive actigraph features</li>\n<li>tain.csv(+ Feature Engineering) + descriptive actigraph features<br>\n(+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])</li>\n<li>train.csv(only)</li></ul></li>\n<li>Regression models<ul>\n<li>objective: mse</li>\n<li>models: lightgbm, xgboost, catboost</li></ul></li>\n<li>I calculated QWK with 3 Seeds×5 CV(StratifiedKFold) to reduce fluctuations.</li>\n<li>QWK thresholds optimized by 3 Seeds×5 Stratified CV to improve robustness.</li>\n</ul>\n<h1>Feature Extraction</h1>\n<p>I used Null Importance based on lightgbm gain importances to reduce features as much as possible while considering the trade-off CV.<br>\nThe number of features after reduction is shown below.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Feature Nums</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train.csv + descriptive actigraph features</td>\n<td>13</td>\n</tr>\n<tr>\n<td>tain.csv(+ Feature Engineering*) + descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400] **)</td>\n<td>40</td>\n</tr>\n<tr>\n<td>train.csv(only)</td>\n<td>14</td>\n</tr>\n</tbody>\n</table>\n<p>* Feature Engineering is the same as many public notebooks. I also calculated the absolute values of the actigraphy data and <code>df['XYZ']=np.sqrt(df['X']**2 + df['Y']**2 + df['Z']**2)</code>.<br>\n**  I calculated descriptive actigraph features for each divided group. </p>\n<h2>Extracted Features</h2>\n<p><strong>train.csv + descriptive actigraph features</strong><br>\n<code>['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'BIA-BIA_Activity_Level_num', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday', 'light_max']</code></p>\n<p><strong>tain.csv(+ Feature Engineering)+ descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])</strong><br>\n<code>['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'Physical-Waist_Circumference', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSND_Zone', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'PAQ_A-PAQ_A_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday', 'X_25%_0-21600', 'Y_std', 'Z_mean_0-21600', 'Z_min_43200-64800', 'Z_75%_64800-86400', 'enmo_50%_43200-64800', 'non-wear_flag_mean_0-21600', 'light_std', 'light_max_21600-43200', 'XYZ_50%_21600-43200', 'XYZ_mean_43200-64800', 'XYZ_mean_64800-86400', 'XYZ_std_64800-86400', 'XYZ_50%_64800-86400', 'abs_X_50%_21600-43200', 'abs_X_25%_64800-86400', 'abs_X_75%_64800-86400', 'abs_Y_75%_0-21600', 'abs_Y_min_21600-43200', 'abs_Y_75%_64800-86400', 'abs_Z_75%_21600-43200', 'abs_Z_75%_43200-64800', 'abs_anglez_min_43200-64800', 'BMI_Age', 'Internet_Hours_Age', 'Muscle_to_Fat']</code></p>\n<p><strong>train.csv(only)</strong><br>\n<code>['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'Fitness_Endurance-Max_Stage', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'BIA-BIA_Activity_Level_num', 'PAQ_A-PAQ_A_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday']\n</code></p>\n<h1>Imputation of missing values</h1>\n<p>I adopted IterativeImputer and the default Bayesian Ridge as an estimator to impute missing values. I also tried GBDT and RandomForest, but the CV was lower than Bayesian Ridge.<br>\nCV was improved the most when training the imputer on data with 30 or fewer missing values. Also, training on data with missing 'sii' would have reduced CV, so we did not use those data.</p>\n<h1>Models</h1>\n<p>I adopted a simple weighted blend, and weights were determined to maximize CV.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train.csv + descriptive actigraph features</td>\n<td>lgb: 0.2, xgb: 0.2</td>\n</tr>\n<tr>\n<td>tain.csv(+ Feature Engineering) + descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])</td>\n<td>xgb: 0.2</td>\n</tr>\n<tr>\n<td>train.csv(only)</td>\n<td>lgb: 0.2, cat: 0.1, xgb: 0.1</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.475</td>\n</tr>\n<tr>\n<td>Submission Model</td>\n<td>0.501</td>\n</tr>\n</tbody>\n</table>\n<p>The baseline model is a single model of lightgbm without imputation of missing values having all train.csv + descriptive actigraph features.(Hyper parameters of models were optimized)</p>\n<h1>Optimization of Thresholds</h1>\n<p>To ensure robust thresholds, they were calculated by the Nelder-Mead method from the entire 3 Seeds×5CV.</p>\n<h1>What didn't work</h1>\n<ul>\n<li>I trained the imputer on data with missing sii, but CV decreased.</li>\n<li>Huber and MAE were used as objective functions but did not improve CV.</li>\n<li>CNN and LSTM were applied to actigraphy data, but learning did not proceed well.</li>\n</ul>\n<h1>What I didn't try</h1>\n<ul>\n<li><p>pseudo-labeling</p></li>\n<li><p>custom-objective</p>\n<p>　<br></p></li>\n</ul>\n<p>Thank you for reading!</p>\n<hr>\n<p>You can check the solution code from the below links.</p>\n<p>training: <a href=\"https://github.com/beagledeveloper/Child_Mind_Institute-Problematic_Internet_Use_8th_Place\" target=\"_blank\">https://github.com/beagledeveloper/Child_Mind_Institute-Problematic_Internet_Use_8th_Place</a></p>\n<p>inference: <a href=\"https://www.kaggle.com/code/kzkknmt/8th-solution-inference-note\" target=\"_blank\">https://www.kaggle.com/code/kzkknmt/8th-solution-inference-note</a></p>",
  "messages": [
    {
      "id": 3077754,
      "postDate": "2024-12-21T11:12:02.740Z",
      "content": "<p>First, I'd like to thank all the hosts and Kaggle staff for organizing such an interesting and practical competition!<br>\nI greatly enjoyed the process of trial and error for handling noisy and missing data!</p>\n<h1>OverView</h1>\n<ul>\n<li>I concentrated on improving CV, not LB due to the small number of data and unclear distribution of missing values.</li>\n<li>I extracted features by Null Importance to improve robustness.</li>\n<li>I imputed numerical missing features by Iterative Imputer (estimator=Bayesian Ridge).</li>\n<li>Ensemble in 3 different models. <ul>\n<li>train.csv + descriptive actigraph features</li>\n<li>tain.csv(+ Feature Engineering) + descriptive actigraph features<br>\n(+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])</li>\n<li>train.csv(only)</li></ul></li>\n<li>Regression models<ul>\n<li>objective: mse</li>\n<li>models: lightgbm, xgboost, catboost</li></ul></li>\n<li>I calculated QWK with 3 Seeds×5 CV(StratifiedKFold) to reduce fluctuations.</li>\n<li>QWK thresholds optimized by 3 Seeds×5 Stratified CV to improve robustness.</li>\n</ul>\n<h1>Feature Extraction</h1>\n<p>I used Null Importance based on lightgbm gain importances to reduce features as much as possible while considering the trade-off CV.<br>\nThe number of features after reduction is shown below.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Feature Nums</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train.csv + descriptive actigraph features</td>\n<td>13</td>\n</tr>\n<tr>\n<td>tain.csv(+ Feature Engineering*) + descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400] **)</td>\n<td>40</td>\n</tr>\n<tr>\n<td>train.csv(only)</td>\n<td>14</td>\n</tr>\n</tbody>\n</table>\n<p>* Feature Engineering is the same as many public notebooks. I also calculated the absolute values of the actigraphy data and <code>df['XYZ']=np.sqrt(df['X']**2 + df['Y']**2 + df['Z']**2)</code>.<br>\n**  I calculated descriptive actigraph features for each divided group. </p>\n<h2>Extracted Features</h2>\n<p><strong>train.csv + descriptive actigraph features</strong><br>\n<code>['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'BIA-BIA_Activity_Level_num', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday', 'light_max']</code></p>\n<p><strong>tain.csv(+ Feature Engineering)+ descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])</strong><br>\n<code>['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'Physical-Waist_Circumference', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSND_Zone', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'PAQ_A-PAQ_A_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday', 'X_25%_0-21600', 'Y_std', 'Z_mean_0-21600', 'Z_min_43200-64800', 'Z_75%_64800-86400', 'enmo_50%_43200-64800', 'non-wear_flag_mean_0-21600', 'light_std', 'light_max_21600-43200', 'XYZ_50%_21600-43200', 'XYZ_mean_43200-64800', 'XYZ_mean_64800-86400', 'XYZ_std_64800-86400', 'XYZ_50%_64800-86400', 'abs_X_50%_21600-43200', 'abs_X_25%_64800-86400', 'abs_X_75%_64800-86400', 'abs_Y_75%_0-21600', 'abs_Y_min_21600-43200', 'abs_Y_75%_64800-86400', 'abs_Z_75%_21600-43200', 'abs_Z_75%_43200-64800', 'abs_anglez_min_43200-64800', 'BMI_Age', 'Internet_Hours_Age', 'Muscle_to_Fat']</code></p>\n<p><strong>train.csv(only)</strong><br>\n<code>['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'Fitness_Endurance-Max_Stage', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'BIA-BIA_Activity_Level_num', 'PAQ_A-PAQ_A_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday']\n</code></p>\n<h1>Imputation of missing values</h1>\n<p>I adopted IterativeImputer and the default Bayesian Ridge as an estimator to impute missing values. I also tried GBDT and RandomForest, but the CV was lower than Bayesian Ridge.<br>\nCV was improved the most when training the imputer on data with 30 or fewer missing values. Also, training on data with missing 'sii' would have reduced CV, so we did not use those data.</p>\n<h1>Models</h1>\n<p>I adopted a simple weighted blend, and weights were determined to maximize CV.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>train.csv + descriptive actigraph features</td>\n<td>lgb: 0.2, xgb: 0.2</td>\n</tr>\n<tr>\n<td>tain.csv(+ Feature Engineering) + descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])</td>\n<td>xgb: 0.2</td>\n</tr>\n<tr>\n<td>train.csv(only)</td>\n<td>lgb: 0.2, cat: 0.1, xgb: 0.1</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.475</td>\n</tr>\n<tr>\n<td>Submission Model</td>\n<td>0.501</td>\n</tr>\n</tbody>\n</table>\n<p>The baseline model is a single model of lightgbm without imputation of missing values having all train.csv + descriptive actigraph features.(Hyper parameters of models were optimized)</p>\n<h1>Optimization of Thresholds</h1>\n<p>To ensure robust thresholds, they were calculated by the Nelder-Mead method from the entire 3 Seeds×5CV.</p>\n<h1>What didn't work</h1>\n<ul>\n<li>I trained the imputer on data with missing sii, but CV decreased.</li>\n<li>Huber and MAE were used as objective functions but did not improve CV.</li>\n<li>CNN and LSTM were applied to actigraphy data, but learning did not proceed well.</li>\n</ul>\n<h1>What I didn't try</h1>\n<ul>\n<li><p>pseudo-labeling</p></li>\n<li><p>custom-objective</p>\n<p>　<br></p></li>\n</ul>\n<p>Thank you for reading!</p>\n<hr>\n<p>You can check the solution code from the below links.</p>\n<p>training: <a href=\"https://github.com/beagledeveloper/Child_Mind_Institute-Problematic_Internet_Use_8th_Place\" target=\"_blank\">https://github.com/beagledeveloper/Child_Mind_Institute-Problematic_Internet_Use_8th_Place</a></p>\n<p>inference: <a href=\"https://www.kaggle.com/code/kzkknmt/8th-solution-inference-note\" target=\"_blank\">https://www.kaggle.com/code/kzkknmt/8th-solution-inference-note</a></p>",
      "rawMarkdown": "First, I'd like to thank all the hosts and Kaggle staff for organizing such an interesting and practical competition!\nI greatly enjoyed the process of trial and error for handling noisy and missing data!\n\n# OverView\n\n- I concentrated on improving CV, not LB due to the small number of data and unclear distribution of missing values.\n- I extracted features by Null Importance to improve robustness.\n- I imputed numerical missing features by Iterative Imputer (estimator=Bayesian Ridge).\n- Ensemble in 3 different models. \n - train.csv + descriptive actigraph features\n - tain.csv(+ Feature Engineering) + descriptive actigraph features\n(+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])\n - train.csv(only)\n- Regression models\n - objective: mse\n - models: lightgbm, xgboost, catboost\n- I calculated QWK with 3 Seeds×5 CV(StratifiedKFold) to reduce fluctuations.\n- QWK thresholds optimized by 3 Seeds×5 Stratified CV to improve robustness.\n\n# Feature Extraction\n\nI used Null Importance based on lightgbm gain importances to reduce features as much as possible while considering the trade-off CV.\nThe number of features after reduction is shown below.\n\n| Model | Feature Nums |\n| --- | --- |\n| train.csv + descriptive actigraph features | 13 |\n| tain.csv(+ Feature Engineering\\*) + descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400] \\*\\*) | 40 |\n| train.csv(only) | 14 |\n\n\\* Feature Engineering is the same as many public notebooks. I also calculated the absolute values of the actigraphy data and `df['XYZ']=np.sqrt(df['X']**2 + df['Y']**2 + df['Z']**2)`.\n\\*\\*  I calculated descriptive actigraph features for each divided group. \n\n## Extracted Features\n\n**train.csv + descriptive actigraph features**\n` ['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'BIA-BIA_Activity_Level_num', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday', 'light_max']`\n\n**tain.csv(+ Feature Engineering)+ descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])**\n`['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'Physical-Waist_Circumference', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSND_Zone', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'PAQ_A-PAQ_A_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday', 'X_25%_0-21600', 'Y_std', 'Z_mean_0-21600', 'Z_min_43200-64800', 'Z_75%_64800-86400', 'enmo_50%_43200-64800', 'non-wear_flag_mean_0-21600', 'light_std', 'light_max_21600-43200', 'XYZ_50%_21600-43200', 'XYZ_mean_43200-64800', 'XYZ_mean_64800-86400', 'XYZ_std_64800-86400', 'XYZ_50%_64800-86400', 'abs_X_50%_21600-43200', 'abs_X_25%_64800-86400', 'abs_X_75%_64800-86400', 'abs_Y_75%_0-21600', 'abs_Y_min_21600-43200', 'abs_Y_75%_64800-86400', 'abs_Z_75%_21600-43200', 'abs_Z_75%_43200-64800', 'abs_anglez_min_43200-64800', 'BMI_Age', 'Internet_Hours_Age', 'Muscle_to_Fat']`\n\n**train.csv(only)**\n`['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'Fitness_Endurance-Max_Stage', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'BIA-BIA_Activity_Level_num', 'PAQ_A-PAQ_A_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday']\n`\n\n# Imputation of missing values\n\nI adopted IterativeImputer and the default Bayesian Ridge as an estimator to impute missing values. I also tried GBDT and RandomForest, but the CV was lower than Bayesian Ridge.\nCV was improved the most when training the imputer on data with 30 or fewer missing values. Also, training on data with missing 'sii' would have reduced CV, so we did not use those data.\n\n# Models\n\nI adopted a simple weighted blend, and weights were determined to maximize CV.\n\n| Model | weights |\n| --- | --- |\n| train.csv + descriptive actigraph features |  lgb: 0.2, xgb: 0.2 |\n| tain.csv(+ Feature Engineering) + descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400]) | xgb: 0.2 |\n| train.csv(only) | lgb: 0.2, cat: 0.1, xgb: 0.1 |\n  \n<br/>\n  \n| Model | CV |\n| --- | --- |\n| Baseline | 0.475 |\n| Submission Model | 0.501 |\n\nThe baseline model is a single model of lightgbm without imputation of missing values having all train.csv + descriptive actigraph features.(Hyper parameters of models were optimized)\n\n# Optimization of Thresholds\n\nTo ensure robust thresholds, they were calculated by the Nelder-Mead method from the entire 3 Seeds×5CV.\n\n# What didn't work\n\n* I trained the imputer on data with missing sii, but CV decreased.\n* Huber and MAE were used as objective functions but did not improve CV.\n* CNN and LSTM were applied to actigraphy data, but learning did not proceed well.\n\n# What I didn't try\n\n* pseudo-labeling\n* custom-objective\n\n　<br/>\n\nThank you for reading!\n\n---\n\nYou can check the solution code from the below links.\n\ntraining: https://github.com/beagledeveloper/Child_Mind_Institute-Problematic_Internet_Use_8th_Place\n\ninference: https://www.kaggle.com/code/kzkknmt/8th-solution-inference-note",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3077754": "First, I'd like to thank all the hosts and Kaggle staff for organizing such an interesting and practical competition!\nI greatly enjoyed the process of trial and error for handling noisy and missing data!\n\n# OverView\n\n- I concentrated on improving CV, not LB due to the small number of data and unclear distribution of missing values.\n- I extracted features by Null Importance to improve robustness.\n- I imputed numerical missing features by Iterative Imputer (estimator=Bayesian Ridge).\n- Ensemble in 3 different models. \n - train.csv + descriptive actigraph features\n - tain.csv(+ Feature Engineering) + descriptive actigraph features\n(+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])\n - train.csv(only)\n- Regression models\n - objective: mse\n - models: lightgbm, xgboost, catboost\n- I calculated QWK with 3 Seeds×5 CV(StratifiedKFold) to reduce fluctuations.\n- QWK thresholds optimized by 3 Seeds×5 Stratified CV to improve robustness.\n\n# Feature Extraction\n\nI used Null Importance based on lightgbm gain importances to reduce features as much as possible while considering the trade-off CV.\nThe number of features after reduction is shown below.\n\n| Model | Feature Nums |\n| --- | --- |\n| train.csv + descriptive actigraph features | 13 |\n| tain.csv(+ Feature Engineering\\*) + descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400] \\*\\*) | 40 |\n| train.csv(only) | 14 |\n\n\\* Feature Engineering is the same as many public notebooks. I also calculated the absolute values of the actigraphy data and `df['XYZ']=np.sqrt(df['X']**2 + df['Y']**2 + df['Z']**2)`.\n\\*\\*  I calculated descriptive actigraph features for each divided group. \n\n## Extracted Features\n\n**train.csv + descriptive actigraph features**\n` ['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'BIA-BIA_Activity_Level_num', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday', 'light_max']`\n\n**tain.csv(+ Feature Engineering)+ descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400])**\n`['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'Physical-Waist_Circumference', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSND_Zone', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'PAQ_A-PAQ_A_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday', 'X_25%_0-21600', 'Y_std', 'Z_mean_0-21600', 'Z_min_43200-64800', 'Z_75%_64800-86400', 'enmo_50%_43200-64800', 'non-wear_flag_mean_0-21600', 'light_std', 'light_max_21600-43200', 'XYZ_50%_21600-43200', 'XYZ_mean_43200-64800', 'XYZ_mean_64800-86400', 'XYZ_std_64800-86400', 'XYZ_50%_64800-86400', 'abs_X_50%_21600-43200', 'abs_X_25%_64800-86400', 'abs_X_75%_64800-86400', 'abs_Y_75%_0-21600', 'abs_Y_min_21600-43200', 'abs_Y_75%_64800-86400', 'abs_Z_75%_21600-43200', 'abs_Z_75%_43200-64800', 'abs_anglez_min_43200-64800', 'BMI_Age', 'Internet_Hours_Age', 'Muscle_to_Fat']`\n\n**train.csv(only)**\n`['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-Height', 'Physical-Weight', 'Fitness_Endurance-Max_Stage', 'FGC-FGC_CU', 'FGC-FGC_GSND', 'FGC-FGC_GSD', 'FGC-FGC_PU', 'BIA-BIA_Activity_Level_num', 'PAQ_A-PAQ_A_Total', 'SDS-SDS_Total_Raw', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday']\n`\n\n# Imputation of missing values\n\nI adopted IterativeImputer and the default Bayesian Ridge as an estimator to impute missing values. I also tried GBDT and RandomForest, but the CV was lower than Bayesian Ridge.\nCV was improved the most when training the imputer on data with 30 or fewer missing values. Also, training on data with missing 'sii' would have reduced CV, so we did not use those data.\n\n# Models\n\nI adopted a simple weighted blend, and weights were determined to maximize CV.\n\n| Model | weights |\n| --- | --- |\n| train.csv + descriptive actigraph features |  lgb: 0.2, xgb: 0.2 |\n| tain.csv(+ Feature Engineering) + descriptive actigraph features <br> (+ time_of_day divided into 4 groups [0-21600, 21600-43200, 43200-64800, 64800-86400]) | xgb: 0.2 |\n| train.csv(only) | lgb: 0.2, cat: 0.1, xgb: 0.1 |\n  \n<br/>\n  \n| Model | CV |\n| --- | --- |\n| Baseline | 0.475 |\n| Submission Model | 0.501 |\n\nThe baseline model is a single model of lightgbm without imputation of missing values having all train.csv + descriptive actigraph features.(Hyper parameters of models were optimized)\n\n# Optimization of Thresholds\n\nTo ensure robust thresholds, they were calculated by the Nelder-Mead method from the entire 3 Seeds×5CV.\n\n# What didn't work\n\n* I trained the imputer on data with missing sii, but CV decreased.\n* Huber and MAE were used as objective functions but did not improve CV.\n* CNN and LSTM were applied to actigraphy data, but learning did not proceed well.\n\n# What I didn't try\n\n* pseudo-labeling\n* custom-objective\n\n　<br/>\n\nThank you for reading!\n\n---\n\nYou can check the solution code from the below links.\n\ntraining: https://github.com/beagledeveloper/Child_Mind_Institute-Problematic_Internet_Use_8th_Place\n\ninference: https://www.kaggle.com/code/kzkknmt/8th-solution-inference-note"
  }
}