{
  "id": 552503,
  "title": "Private LB 0.466 using only 6 features (Single multiseed CatBoost)",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552503",
  "author_name": "",
  "post_date": "2024-12-20T02:17:41.064780900Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Note: I only started this competition last week, so I probably have a limited view of the competition compared to those that did it for the whole 3 months</p>\n<p><a href=\"https://www.kaggle.com/code/yeoyunsianggeremie/pb-0-466-single-multiseed-catboost-with-6-feats\" target=\"_blank\">Notebook</a> (note there is data leakage in the KNN code here)</p>\n<p>Upon doing initial analysis on the data, I noticed that lots of features do not provide much signal (from the domain perspective) - despite them having some variation on the training data. The dataset is also extremely small and prone to overfitting if lots of features are used. </p>\n<p>Key ideas</p>\n<ul>\n<li>I dropped all rows where the target is NULL</li>\n<li><strong>PreInt_EduHx-computerinternet_hoursday</strong> (Hours of using computer/internet) was by far the most important feature (akin to a pseudo-target) - it is similar to AdvantageP1 in MCTS competition - strong correlation with target as expected</li>\n<li>The other features I used are the age, sex and various assessment scores. PAQ_C-PAQ_C_Total and PAQ_A-PAQ_A_Total were discarded due to a reduction in the CV when they are used. </li>\n<li>I felt that PIU should not influence the features like Height, Weight and scores of various exercises so I assumed these are noise columns and dropped them. </li>\n<li>I tried extracting features (e.g. average sleep onset and awake) from the actigraphy data, but unfortunately they did not improve CV so I discarded them</li>\n</ul>\n<p>This scores <strong>0.438</strong> without any imputation on the missing values and <strong>0.465-0.466</strong> (30th place) with imputation on the missing values. Unfortunately I didn’t trust the imputation and did not select it as a submission 😐</p>",
  "messages": [
    {
      "id": "3076499",
      "postDate": "12/20/2024 02:17:41",
      "content": "<p>Note: I only started this competition last week, so I probably have a limited view of the competition compared to those that did it for the whole 3 months</p>\n<p><a href=\"https://www.kaggle.com/code/yeoyunsianggeremie/pb-0-466-single-multiseed-catboost-with-6-feats\" target=\"_blank\">Notebook</a> (note there is data leakage in the KNN code here)</p>\n<p>Upon doing initial analysis on the data, I noticed that lots of features do not provide much signal (from the domain perspective) - despite them having some variation on the training data. The dataset is also extremely small and prone to overfitting if lots of features are used. </p>\n<p>Key ideas</p>\n<ul>\n<li>I dropped all rows where the target is NULL</li>\n<li><strong>PreInt_EduHx-computerinternet_hoursday</strong> (Hours of using computer/internet) was by far the most important feature (akin to a pseudo-target) - it is similar to AdvantageP1 in MCTS competition - strong correlation with target as expected</li>\n<li>The other features I used are the age, sex and various assessment scores. PAQ_C-PAQ_C_Total and PAQ_A-PAQ_A_Total were discarded due to a reduction in the CV when they are used. </li>\n<li>I felt that PIU should not influence the features like Height, Weight and scores of various exercises so I assumed these are noise columns and dropped them. </li>\n<li>I tried extracting features (e.g. average sleep onset and awake) from the actigraphy data, but unfortunately they did not improve CV so I discarded them</li>\n</ul>\n<p>This scores <strong>0.438</strong> without any imputation on the missing values and <strong>0.465-0.466</strong> (30th place) with imputation on the missing values. Unfortunately I didn’t trust the imputation and did not select it as a submission 😐</p>",
      "rawMarkdown": "Note: I only started this competition last week, so I probably have a limited view of the competition compared to those that did it for the whole 3 months\n\n[Notebook](https://www.kaggle.com/code/yeoyunsianggeremie/pb-0-466-single-multiseed-catboost-with-6-feats) (note there is data leakage in the KNN code here)\n\nUpon doing initial analysis on the data, I noticed that lots of features do not provide much signal (from the domain perspective) - despite them having some variation on the training data. The dataset is also extremely small and prone to overfitting if lots of features are used. \n\nKey ideas\n- I dropped all rows where the target is NULL\n- **PreInt_EduHx-computerinternet_hoursday** (Hours of using computer/internet) was by far the most important feature (akin to a pseudo-target) - it is similar to AdvantageP1 in MCTS competition - strong correlation with target as expected\n- The other features I used are the age, sex and various assessment scores. PAQ_C-PAQ_C_Total and PAQ_A-PAQ_A_Total were discarded due to a reduction in the CV when they are used. \n- I felt that PIU should not influence the features like Height, Weight and scores of various exercises so I assumed these are noise columns and dropped them. \n- I tried extracting features (e.g. average sleep onset and awake) from the actigraphy data, but unfortunately they did not improve CV so I discarded them\n\nThis scores **0.438** without any imputation on the missing values and **0.465-0.466** (30th place) with imputation on the missing values. Unfortunately I didn’t trust the imputation and did not select it as a submission 😐",
      "votes": null
    },
    {
      "id": "3076506",
      "postDate": "12/20/2024 02:25:12",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": null
    },
    {
      "id": "3076622",
      "postDate": "12/20/2024 05:54:22",
      "content": "<p>Same for me. I also didn't trust the imputation and feature selection solution. It could have won me a silver. :(</p>",
      "rawMarkdown": "Same for me. I also didn't trust the imputation and feature selection solution. It could have won me a silver. :(",
      "votes": null
    },
    {
      "id": "3076713",
      "postDate": "12/20/2024 07:48:57",
      "content": "<p>This is a game of luck my friend <a href=\"https://www.kaggle.com/nikhil1e9\" target=\"_blank\">@nikhil1e9</a> </p>",
      "rawMarkdown": "This is a game of luck my friend @nikhil1e9",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076506,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "12/20/2024 02:25:12",
      "content": "<p>congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076622,
      "author_name": "nikhil1e9",
      "author_url": "",
      "post_date": "12/20/2024 05:54:22",
      "content": "<p>Same for me. I also didn't trust the imputation and feature selection solution. It could have won me a silver. :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 3076713,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/20/2024 07:48:57",
          "content": "<p>This is a game of luck my friend <a href=\"https://www.kaggle.com/nikhil1e9\" target=\"_blank\">@nikhil1e9</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3076499": "Note: I only started this competition last week, so I probably have a limited view of the competition compared to those that did it for the whole 3 months\n\n[Notebook](https://www.kaggle.com/code/yeoyunsianggeremie/pb-0-466-single-multiseed-catboost-with-6-feats) (note there is data leakage in the KNN code here)\n\nUpon doing initial analysis on the data, I noticed that lots of features do not provide much signal (from the domain perspective) - despite them having some variation on the training data. The dataset is also extremely small and prone to overfitting if lots of features are used. \n\nKey ideas\n- I dropped all rows where the target is NULL\n- **PreInt_EduHx-computerinternet_hoursday** (Hours of using computer/internet) was by far the most important feature (akin to a pseudo-target) - it is similar to AdvantageP1 in MCTS competition - strong correlation with target as expected\n- The other features I used are the age, sex and various assessment scores. PAQ_C-PAQ_C_Total and PAQ_A-PAQ_A_Total were discarded due to a reduction in the CV when they are used. \n- I felt that PIU should not influence the features like Height, Weight and scores of various exercises so I assumed these are noise columns and dropped them. \n- I tried extracting features (e.g. average sleep onset and awake) from the actigraphy data, but unfortunately they did not improve CV so I discarded them\n\nThis scores **0.438** without any imputation on the missing values and **0.465-0.466** (30th place) with imputation on the missing values. Unfortunately I didn’t trust the imputation and did not select it as a submission 😐",
    "3076506": "congratulations",
    "3076622": "Same for me. I also didn't trust the imputation and feature selection solution. It could have won me a silver. :(",
    "3076713": "This is a game of luck my friend @nikhil1e9"
  },
  "source": "meta"
}