{
  "id": 536069,
  "title": "Is the data real / synthetic ? ",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/536069",
  "author_name": "Sheikh Muhammad Abdullah",
  "post_date": "2024-09-25T19:34:40.998000",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I've conducted many experiments, added new features to the data, and then extracted some of the top features from them. I also optimized the parameters using Optuna and managed to improve my model's CV (cross-validation). However, the results on the leaderboard (LB) are still poor.</p>\n<p>On the other hand, using basic preprocessing and basic features, my CV and LB scores are slightly better. Right now, I'm confused about the data because, while the CV improves, the LB performance gets worse.</p>\n<p>For reference, a good CV score for me typically ranges between <code>0.455 and 0.465</code>on the LB.</p>",
  "messages": [
    {
      "id": 2998583,
      "postDate": "2024-09-25T19:34:40.997Z",
      "content": "<p>I've conducted many experiments, added new features to the data, and then extracted some of the top features from them. I also optimized the parameters using Optuna and managed to improve my model's CV (cross-validation). However, the results on the leaderboard (LB) are still poor.</p>\n<p>On the other hand, using basic preprocessing and basic features, my CV and LB scores are slightly better. Right now, I'm confused about the data because, while the CV improves, the LB performance gets worse.</p>\n<p>For reference, a good CV score for me typically ranges between <code>0.455 and 0.465</code>on the LB.</p>",
      "rawMarkdown": "I've conducted many experiments, added new features to the data, and then extracted some of the top features from them. I also optimized the parameters using Optuna and managed to improve my model's CV (cross-validation). However, the results on the leaderboard (LB) are still poor.\n\nOn the other hand, using basic preprocessing and basic features, my CV and LB scores are slightly better. Right now, I'm confused about the data because, while the CV improves, the LB performance gets worse.\n\nFor reference, a good CV score for me typically ranges between `0.455 and 0.465 `on the LB.",
      "votes": 14
    },
    {
      "id": 2998717,
      "postDate": "2024-09-25T23:27:42.053Z",
      "content": "<p>There're 2 possibilities on the hidden test set.</p>\n<ol>\n<li>the hidden test set have gone through PCIAT and sii is based on the PCIAT result, and of course PCIAT is removed then.</li>\n<li>the hidden test didn't pass through PCIAT, and sii is generated from another ML based on other features.</li>\n</ol>\n<p>In scenario 2, since score is based on Cohen's Kappa, we're competing in making the most similar model as the hidden one.<br>\nAnd I've no clue at all.</p>",
      "rawMarkdown": "There're 2 possibilities on the hidden test set.\n1. the hidden test set have gone through PCIAT and sii is based on the PCIAT result, and of course PCIAT is removed then.\n2. the hidden test didn't pass through PCIAT, and sii is generated from another ML based on other features.\n\nIn scenario 2, since score is based on Cohen's Kappa, we're competing in making the most similar model as the hidden one.\nAnd I've no clue at all.",
      "votes": 3,
      "replies": [
        {
          "id": 2998731,
          "postDate": "2024-09-25T23:56:57.247Z",
          "content": "<p><a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> I think the test set sii is predicted from test set PCIAT-Total. I don't think otherwise.<br>\n<a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> <a href=\"https://www.kaggle.com/mushei\" target=\"_blank\">@mushei</a> may add value as well</p>",
          "rawMarkdown": "@tomyuen I think the test set sii is predicted from test set PCIAT-Total. I don't think otherwise.\n@abdmental01 @mushei may add value as well",
          "votes": 3
        }
      ]
    },
    {
      "id": 2998671,
      "postDate": "2024-09-25T22:01:59.853Z",
      "content": "<p>what was your approach to fill the null values of target variable?</p>",
      "rawMarkdown": "what was your approach to fill the null values of target variable?",
      "votes": 2,
      "replies": [
        {
          "id": 2998677,
          "postDate": "2024-09-25T22:19:33.530Z",
          "content": "<p>I Dropped all the null values in the target.</p>\n<p><code>train = train.dropna(subset='sii')</code></p>",
          "rawMarkdown": "I Dropped all the null values in the target.\n\n`train = train.dropna(subset='sii')`",
          "votes": 2,
          "replies": [
            {
              "id": 2999161,
              "postDate": "2024-09-26T12:37:45.137Z",
              "content": "<p>I was thinking of running a 'K-means' algorithm on it, to find similar points together and then fill the target values with it, and i also noticed that, there are alot of columns missing in test set, how did you handle that? </p>",
              "rawMarkdown": "I was thinking of running a 'K-means' algorithm on it, to find similar points together and then fill the target values with it, and i also noticed that, there are alot of columns missing in test set, how did you handle that? ",
              "votes": 2
            },
            {
              "id": 2999163,
              "postDate": "2024-09-26T12:40:25.257Z",
              "content": "<p>I think it is better to drop the missing values from the target column.</p>",
              "rawMarkdown": "I think it is better to drop the missing values from the target column.",
              "votes": 1
            },
            {
              "id": 2999182,
              "postDate": "2024-09-26T13:07:16.920Z",
              "rawMarkdown": "",
              "votes": -1,
              "isDeleted": true
            },
            {
              "id": 2999695,
              "postDate": "2024-09-26T19:30:27.823Z",
              "content": "<p>Basically , Simply Drop The Null in Target, also some of Cols are missing in Test, so also drop those cols from test. Because we cannot generate the same cols for test.</p>",
              "rawMarkdown": "Basically , Simply Drop The Null in Target, also some of Cols are missing in Test, so also drop those cols from test. Because we cannot generate the same cols for test.",
              "votes": 2
            },
            {
              "id": 3013301,
              "postDate": "2024-10-09T22:22:07.833Z",
              "content": "<p>Look into KNNImputer</p>",
              "rawMarkdown": "Look into KNNImputer",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2998717,
      "author_name": "Tom Yuen",
      "author_url": "",
      "post_date": "2024-09-25T23:27:42.053000",
      "content": "<p>There're 2 possibilities on the hidden test set.</p>\n<ol>\n<li>the hidden test set have gone through PCIAT and sii is based on the PCIAT result, and of course PCIAT is removed then.</li>\n<li>the hidden test didn't pass through PCIAT, and sii is generated from another ML based on other features.</li>\n</ol>\n<p>In scenario 2, since score is based on Cohen's Kappa, we're competing in making the most similar model as the hidden one.<br>\nAnd I've no clue at all.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2998731,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-09-25T23:56:57.247000",
          "content": "<p><a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> I think the test set sii is predicted from test set PCIAT-Total. I don't think otherwise.<br>\n<a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> <a href=\"https://www.kaggle.com/mushei\" target=\"_blank\">@mushei</a> may add value as well</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2998671,
      "author_name": "Mushi",
      "author_url": "",
      "post_date": "2024-09-25T22:01:59.853000",
      "content": "<p>what was your approach to fill the null values of target variable?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2998677,
          "author_name": "Sheikh Muhammad Abdullah",
          "author_url": "",
          "post_date": "2024-09-25T22:19:33.530000",
          "content": "<p>I Dropped all the null values in the target.</p>\n<p><code>train = train.dropna(subset='sii')</code></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2999161,
              "author_name": "Mushi",
              "author_url": "",
              "post_date": "2024-09-26T12:37:45.137000",
              "content": "<p>I was thinking of running a 'K-means' algorithm on it, to find similar points together and then fill the target values with it, and i also noticed that, there are alot of columns missing in test set, how did you handle that? </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2999163,
              "author_name": "Adil nawaz ashrafi",
              "author_url": "",
              "post_date": "2024-09-26T12:40:25.257000",
              "content": "<p>I think it is better to drop the missing values from the target column.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2999182,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-09-26T13:07:16.920000",
              "content": "",
              "votes": -1,
              "replies": []
            },
            {
              "id": 2999695,
              "author_name": "Sheikh Muhammad Abdullah",
              "author_url": "",
              "post_date": "2024-09-26T19:30:27.823000",
              "content": "<p>Basically , Simply Drop The Null in Target, also some of Cols are missing in Test, so also drop those cols from test. Because we cannot generate the same cols for test.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3013301,
              "author_name": "Sam",
              "author_url": "",
              "post_date": "2024-10-09T22:22:07.833000",
              "content": "<p>Look into KNNImputer</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2998583": "I've conducted many experiments, added new features to the data, and then extracted some of the top features from them. I also optimized the parameters using Optuna and managed to improve my model's CV (cross-validation). However, the results on the leaderboard (LB) are still poor.\n\nOn the other hand, using basic preprocessing and basic features, my CV and LB scores are slightly better. Right now, I'm confused about the data because, while the CV improves, the LB performance gets worse.\n\nFor reference, a good CV score for me typically ranges between `0.455 and 0.465 `on the LB.",
    "2998717": "There're 2 possibilities on the hidden test set.\n1. the hidden test set have gone through PCIAT and sii is based on the PCIAT result, and of course PCIAT is removed then.\n2. the hidden test didn't pass through PCIAT, and sii is generated from another ML based on other features.\n\nIn scenario 2, since score is based on Cohen's Kappa, we're competing in making the most similar model as the hidden one.\nAnd I've no clue at all.",
    "2998671": "what was your approach to fill the null values of target variable?"
  }
}