{
  "id": 552488,
  "title": "Private 0.478 notebook",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552488",
  "author_name": "",
  "post_date": "2024-12-20T00:57:07.359787100Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I would like to share a notebook with cv 0.4794 and lb 0.478. The main idea is addressing missing values and outliers, feature selection and voting.<br>\n<a href=\"https://www.kaggle.com/code/takanashihumbert/cmi-address-missing?scriptVersionId=213316097\" target=\"_blank\">CMI - address missing</a></p>",
  "messages": [
    {
      "id": "3076436",
      "postDate": "12/20/2024 00:57:07",
      "content": "<p>I would like to share a notebook with cv 0.4794 and lb 0.478. The main idea is addressing missing values and outliers, feature selection and voting.<br>\n<a href=\"https://www.kaggle.com/code/takanashihumbert/cmi-address-missing?scriptVersionId=213316097\" target=\"_blank\">CMI - address missing</a></p>",
      "rawMarkdown": "I would like to share a notebook with cv 0.4794 and lb 0.478. The main idea is addressing missing values and outliers, feature selection and voting.\n[CMI - address missing](https://www.kaggle.com/code/takanashihumbert/cmi-address-missing?scriptVersionId=213316097)",
      "votes": null
    },
    {
      "id": "3076439",
      "postDate": "12/20/2024 00:59:03",
      "content": "<p>Exciting cv and lb relationship, sorry you didn't choose this notebook</p>",
      "rawMarkdown": "Exciting cv and lb relationship, sorry you didn't choose this notebook",
      "votes": null
    },
    {
      "id": "3076446",
      "postDate": "12/20/2024 01:04:39",
      "content": "<p><a href=\"https://www.kaggle.com/takanashihumbert\" target=\"_blank\">@takanashihumbert</a> , good cv and LB, why you didn't select this ?</p>",
      "rawMarkdown": "takanashihumbert , good cv and LB, why you didn't select this ?",
      "votes": null
    },
    {
      "id": "3076450",
      "postDate": "12/20/2024 01:08:20",
      "content": "<p>I took a look at your preprocessing, and your feature engineering may be a big reason to make your model stable</p>",
      "rawMarkdown": "I took a look at your preprocessing, and your feature engineering may be a big reason to make your model stable",
      "votes": null
    },
    {
      "id": "3076452",
      "postDate": "12/20/2024 01:08:55",
      "content": "<p>Some issues in the test dataset:<br>\nThe missing percentage of series parquet in test: 80-85%(~60% in train) <br>\nThe missing percentage of <code>FGC</code> features in test: 70-75%(29.8% in train) <br>\nThe missing percentage of <code>BIA</code> features in test: 60-65%(33.7% in train) <br>\nThe missing percentage of <code>PreInt_EduHx-computerinternet_hoursday</code> in test: 40-50%(3% in train) <br>\n……</p>",
      "rawMarkdown": "Some issues in the test dataset:\nThe missing percentage of series parquet in test: 80-85%(~60% in train) \nThe missing percentage of `FGC` features in test: 70-75%(29.8% in train) \nThe missing percentage of `BIA` features in test: 60-65%(33.7% in train) \nThe missing percentage of `PreInt_EduHx-computerinternet_hoursday` in test: 40-50%(3% in train) \n......",
      "votes": null
    },
    {
      "id": "3076455",
      "postDate": "12/20/2024 01:11:20",
      "content": "<p>congratulations🏅</p>",
      "rawMarkdown": "congratulations🏅",
      "votes": null
    },
    {
      "id": "3076461",
      "postDate": "12/20/2024 01:24:02",
      "content": "<p>Very interesting…might be the reason why KNNImputer works so well on the private test set despite decreasing the CV on the train set 😐 - both with leakage and without leakage (much worse)</p>\n<p>PreInt_EduHx-computerinternet_hoursday was one of the most important features in my model</p>",
      "rawMarkdown": "Very interesting…might be the reason why KNNImputer works so well on the private test set despite decreasing the CV on the train set 😐 - both with leakage and without leakage (much worse)\n\nPreInt_EduHx-computerinternet_hoursday was one of the most important features in my model",
      "votes": null
    },
    {
      "id": "3076476",
      "postDate": "12/20/2024 01:38:41",
      "content": "<p>Age and PreInt_EduHx-computerinternet_hoursday were the 2 most important features for all my models too, despite the presence of actigraphy features <a href=\"https://www.kaggle.com/yeoyunsianggeremie\" target=\"_blank\">@yeoyunsianggeremie</a> </p>",
      "rawMarkdown": "Age and PreInt_EduHx-computerinternet_hoursday were the 2 most important features for all my models too, despite the presence of actigraphy features @yeoyunsianggeremie",
      "votes": null
    },
    {
      "id": "3076651",
      "postDate": "12/20/2024 06:39:34",
      "content": "<p>how did you come up with these missing percentages in the test dataset?</p>",
      "rawMarkdown": "how did you come up with these missing percentages in the test dataset?",
      "votes": null
    },
    {
      "id": "3076705",
      "postDate": "12/20/2024 07:45:14",
      "content": "<p>I took a look at your preprocessing, and your feature engineering may be a big reason to make your model stable</p>",
      "rawMarkdown": "I took a look at your preprocessing, and your feature engineering may be a big reason to make your model stable",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076439,
      "author_name": "ruichardliu",
      "author_url": "",
      "post_date": "12/20/2024 00:59:03",
      "content": "<p>Exciting cv and lb relationship, sorry you didn't choose this notebook</p>",
      "votes": null,
      "replies": [
        {
          "id": 3076455,
          "author_name": "takanashihumbert",
          "author_url": "",
          "post_date": "12/20/2024 01:11:20",
          "content": "<p>congratulations🏅</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3076446,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "12/20/2024 01:04:39",
      "content": "<p><a href=\"https://www.kaggle.com/takanashihumbert\" target=\"_blank\">@takanashihumbert</a> , good cv and LB, why you didn't select this ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076450,
      "author_name": "ruichardliu",
      "author_url": "",
      "post_date": "12/20/2024 01:08:20",
      "content": "<p>I took a look at your preprocessing, and your feature engineering may be a big reason to make your model stable</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076452,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "12/20/2024 01:08:55",
      "content": "<p>Some issues in the test dataset:<br>\nThe missing percentage of series parquet in test: 80-85%(~60% in train) <br>\nThe missing percentage of <code>FGC</code> features in test: 70-75%(29.8% in train) <br>\nThe missing percentage of <code>BIA</code> features in test: 60-65%(33.7% in train) <br>\nThe missing percentage of <code>PreInt_EduHx-computerinternet_hoursday</code> in test: 40-50%(3% in train) <br>\n……</p>",
      "votes": null,
      "replies": [
        {
          "id": 3076461,
          "author_name": "yeoyunsianggeremie",
          "author_url": "",
          "post_date": "12/20/2024 01:24:02",
          "content": "<p>Very interesting…might be the reason why KNNImputer works so well on the private test set despite decreasing the CV on the train set 😐 - both with leakage and without leakage (much worse)</p>\n<p>PreInt_EduHx-computerinternet_hoursday was one of the most important features in my model</p>",
          "votes": null,
          "replies": [
            {
              "id": 3076476,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "12/20/2024 01:38:41",
              "content": "<p>Age and PreInt_EduHx-computerinternet_hoursday were the 2 most important features for all my models too, despite the presence of actigraphy features <a href=\"https://www.kaggle.com/yeoyunsianggeremie\" target=\"_blank\">@yeoyunsianggeremie</a> </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3076651,
          "author_name": "diegoiglesias",
          "author_url": "",
          "post_date": "12/20/2024 06:39:34",
          "content": "<p>how did you come up with these missing percentages in the test dataset?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3076705,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "12/20/2024 07:45:14",
      "content": "<p>I took a look at your preprocessing, and your feature engineering may be a big reason to make your model stable</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3076436": "I would like to share a notebook with cv 0.4794 and lb 0.478. The main idea is addressing missing values and outliers, feature selection and voting.\n[CMI - address missing](https://www.kaggle.com/code/takanashihumbert/cmi-address-missing?scriptVersionId=213316097)",
    "3076439": "Exciting cv and lb relationship, sorry you didn't choose this notebook",
    "3076446": "takanashihumbert , good cv and LB, why you didn't select this ?",
    "3076450": "I took a look at your preprocessing, and your feature engineering may be a big reason to make your model stable",
    "3076452": "Some issues in the test dataset:\nThe missing percentage of series parquet in test: 80-85%(~60% in train) \nThe missing percentage of `FGC` features in test: 70-75%(29.8% in train) \nThe missing percentage of `BIA` features in test: 60-65%(33.7% in train) \nThe missing percentage of `PreInt_EduHx-computerinternet_hoursday` in test: 40-50%(3% in train) \n......",
    "3076455": "congratulations🏅",
    "3076461": "Very interesting…might be the reason why KNNImputer works so well on the private test set despite decreasing the CV on the train set 😐 - both with leakage and without leakage (much worse)\n\nPreInt_EduHx-computerinternet_hoursday was one of the most important features in my model",
    "3076476": "Age and PreInt_EduHx-computerinternet_hoursday were the 2 most important features for all my models too, despite the presence of actigraphy features @yeoyunsianggeremie",
    "3076651": "how did you come up with these missing percentages in the test dataset?",
    "3076705": "I took a look at your preprocessing, and your feature engineering may be a big reason to make your model stable"
  },
  "source": "meta"
}