{
  "id": 553305,
  "title": "17th Place Solution",
  "url": "/competitions/child-mind-institute-problematic-internet-use/writeups/agf-17th-place-solution",
  "author_name": "",
  "post_date": "2024-12-25T06:59:50.987Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The high-scoring model combines the Global Model and Local Model with a 1:4 weighting ratio.</p>\n<p><strong>Global Model</strong></p>\n<ol>\n<li>Feature Engineering</li>\n</ol>\n<ul>\n<li>Randomly combine multiple features, then use their correlation with sii to roughly filter and retain around 200 features.</li>\n<li>During testing, the model quickly overfits, and further optimization results in worse performance on the Leaderboard (LB).</li>\n</ul>\n<p>2.NAN Imputation</p>\n<ul>\n<li>Use the NAN features to cluster the training data first. After identifying the optimal clusters, fill in the missing values with the mean of each cluster.</li>\n<li>This approach performs slightly better than directly filling wi with the overall mean.</li>\n</ul>\n<p>3.Target Grouping</p>\n<ul>\n<li>Divide the target variable into groups, build separate models for each group, and aggregate the predictions for the final result.</li>\n<li>Grouping strategy for Group 3 (time, interpersonal, emotion):<br>\nPCIAT_GROUP3_1: PCIAT-PCIAT_01, PCIAT-PCIAT_02, PCIAT-PCIAT_05, PCIAT-PCIAT_06, PCIAT-PCIAT_07, PCIAT-PCIAT_10, PCIAT-PCIAT_14, PCIAT-PCIAT_15<br>\nPCIAT_GROUP3_2: PCIAT-PCIAT_03, PCIAT-PCIAT_08, PCIAT-PCIAT_11, PCIAT-PCIAT_12, PCIAT-PCIAT_17, PCIAT-PCIAT_19<br>\nPCIAT_GROUP3_3: PCIAT-PCIAT_04, PCIAT-PCIAT_09, PCIAT-PCIAT_13, PCIAT-PCIAT_16, PCIAT-PCIAT_18, PCIAT-PCIAT_20</li>\n<li>Various grouping strategies were tested, and the above was the final choice.</li>\n</ul>\n<p>4.Five-Model Stacking<br>\nModels used: LGB, XGB, CAT, Random Forest, Gradient Boosting.</p>\n<p>5.Average Seed</p>\n<p><strong>Local Model</strong><br>\nFor each test point, find the closest data points (~100) and build a model for prediction.<br>\nDetailed discussion is available here(<a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552516\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552516</a>)</p>",
  "messages": [
    {
      "id": "3080400",
      "postDate": "12/25/2024 06:58:26",
      "content": "<p>The high-scoring model combines the Global Model and Local Model with a 1:4 weighting ratio.</p>\n<p><strong>Global Model</strong></p>\n<ol>\n<li>Feature Engineering</li>\n</ol>\n<ul>\n<li>Randomly combine multiple features, then use their correlation with sii to roughly filter and retain around 200 features.</li>\n<li>During testing, the model quickly overfits, and further optimization results in worse performance on the Leaderboard (LB).</li>\n</ul>\n<p>2.NAN Imputation</p>\n<ul>\n<li>Use the NAN features to cluster the training data first. After identifying the optimal clusters, fill in the missing values with the mean of each cluster.</li>\n<li>This approach performs slightly better than directly filling wi with the overall mean.</li>\n</ul>\n<p>3.Target Grouping</p>\n<ul>\n<li>Divide the target variable into groups, build separate models for each group, and aggregate the predictions for the final result.</li>\n<li>Grouping strategy for Group 3 (time, interpersonal, emotion):<br>\nPCIAT_GROUP3_1: PCIAT-PCIAT_01, PCIAT-PCIAT_02, PCIAT-PCIAT_05, PCIAT-PCIAT_06, PCIAT-PCIAT_07, PCIAT-PCIAT_10, PCIAT-PCIAT_14, PCIAT-PCIAT_15<br>\nPCIAT_GROUP3_2: PCIAT-PCIAT_03, PCIAT-PCIAT_08, PCIAT-PCIAT_11, PCIAT-PCIAT_12, PCIAT-PCIAT_17, PCIAT-PCIAT_19<br>\nPCIAT_GROUP3_3: PCIAT-PCIAT_04, PCIAT-PCIAT_09, PCIAT-PCIAT_13, PCIAT-PCIAT_16, PCIAT-PCIAT_18, PCIAT-PCIAT_20</li>\n<li>Various grouping strategies were tested, and the above was the final choice.</li>\n</ul>\n<p>4.Five-Model Stacking<br>\nModels used: LGB, XGB, CAT, Random Forest, Gradient Boosting.</p>\n<p>5.Average Seed</p>\n<p><strong>Local Model</strong><br>\nFor each test point, find the closest data points (~100) and build a model for prediction.<br>\nDetailed discussion is available here(<a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552516\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552516</a>)</p>",
      "rawMarkdown": "The high-scoring model combines the Global Model and Local Model with a 1:4 weighting ratio.\n\n**Global Model**\n\n1. Feature Engineering\n- Randomly combine multiple features, then use their correlation with sii to roughly filter and retain around 200 features.\n- During testing, the model quickly overfits, and further optimization results in worse performance on the Leaderboard (LB).\n\n2.NAN Imputation\n- Use the NAN features to cluster the training data first. After identifying the optimal clusters, fill in the missing values with the mean of each cluster.\n- This approach performs slightly better than directly filling wi with the overall mean.\n\n3.Target Grouping\n- Divide the target variable into groups, build separate models for each group, and aggregate the predictions for the final result.\n- Grouping strategy for Group 3 (time, interpersonal, emotion):\nPCIAT_GROUP3_1: PCIAT-PCIAT_01, PCIAT-PCIAT_02, PCIAT-PCIAT_05, PCIAT-PCIAT_06, PCIAT-PCIAT_07, PCIAT-PCIAT_10, PCIAT-PCIAT_14, PCIAT-PCIAT_15\nPCIAT_GROUP3_2: PCIAT-PCIAT_03, PCIAT-PCIAT_08, PCIAT-PCIAT_11, PCIAT-PCIAT_12, PCIAT-PCIAT_17, PCIAT-PCIAT_19\nPCIAT_GROUP3_3: PCIAT-PCIAT_04, PCIAT-PCIAT_09, PCIAT-PCIAT_13, PCIAT-PCIAT_16, PCIAT-PCIAT_18, PCIAT-PCIAT_20\n- Various grouping strategies were tested, and the above was the final choice.\n\n4.Five-Model Stacking\nModels used: LGB, XGB, CAT, Random Forest, Gradient Boosting.\n\n5.Average Seed\n\n**Local Model**\nFor each test point, find the closest data points (~100) and build a model for prediction.\nDetailed discussion is available here(https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552516)",
      "votes": null
    },
    {
      "id": "3102087",
      "postDate": "01/21/2025 17:04:09",
      "content": "<p>It's a very useful notebook! I'm a beginner, and a little confused about 'Use the NAN features to cluster the training data first'. Does it means that using all of the features to cluster samples? but some features have NaN, will it impede the clustering process? (maybe a stupid question) 😭<br>\nthank u! </p>",
      "rawMarkdown": "It's a very useful notebook! I'm a beginner, and a little confused about 'Use the NAN features to cluster the training data first'. Does it means that using all of the features to cluster samples? but some features have NaN, will it impede the clustering process? (maybe a stupid question) 😭\nthank u!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3102087,
      "author_name": "wanyuezhang",
      "author_url": "",
      "post_date": "01/21/2025 17:04:09",
      "content": "<p>It's a very useful notebook! I'm a beginner, and a little confused about 'Use the NAN features to cluster the training data first'. Does it means that using all of the features to cluster samples? but some features have NaN, will it impede the clustering process? (maybe a stupid question) 😭<br>\nthank u! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3080400": "The high-scoring model combines the Global Model and Local Model with a 1:4 weighting ratio.\n\n**Global Model**\n\n1. Feature Engineering\n- Randomly combine multiple features, then use their correlation with sii to roughly filter and retain around 200 features.\n- During testing, the model quickly overfits, and further optimization results in worse performance on the Leaderboard (LB).\n\n2.NAN Imputation\n- Use the NAN features to cluster the training data first. After identifying the optimal clusters, fill in the missing values with the mean of each cluster.\n- This approach performs slightly better than directly filling wi with the overall mean.\n\n3.Target Grouping\n- Divide the target variable into groups, build separate models for each group, and aggregate the predictions for the final result.\n- Grouping strategy for Group 3 (time, interpersonal, emotion):\nPCIAT_GROUP3_1: PCIAT-PCIAT_01, PCIAT-PCIAT_02, PCIAT-PCIAT_05, PCIAT-PCIAT_06, PCIAT-PCIAT_07, PCIAT-PCIAT_10, PCIAT-PCIAT_14, PCIAT-PCIAT_15\nPCIAT_GROUP3_2: PCIAT-PCIAT_03, PCIAT-PCIAT_08, PCIAT-PCIAT_11, PCIAT-PCIAT_12, PCIAT-PCIAT_17, PCIAT-PCIAT_19\nPCIAT_GROUP3_3: PCIAT-PCIAT_04, PCIAT-PCIAT_09, PCIAT-PCIAT_13, PCIAT-PCIAT_16, PCIAT-PCIAT_18, PCIAT-PCIAT_20\n- Various grouping strategies were tested, and the above was the final choice.\n\n4.Five-Model Stacking\nModels used: LGB, XGB, CAT, Random Forest, Gradient Boosting.\n\n5.Average Seed\n\n**Local Model**\nFor each test point, find the closest data points (~100) and build a model for prediction.\nDetailed discussion is available here(https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552516)",
    "3102087": "It's a very useful notebook! I'm a beginner, and a little confused about 'Use the NAN features to cluster the training data first'. Does it means that using all of the features to cluster samples? but some features have NaN, will it impede the clustering process? (maybe a stupid question) 😭\nthank u!"
  },
  "source": "meta"
}