{
  "id": 545388,
  "title": "Are NaN Patterns Consistent Between Train and Test Sets?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/545388",
  "author_name": "",
  "post_date": "2024-11-09T21:16:18.749296800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I noticed that not all IDs in the training set have complete data, i.e time series data, leading to NaN values for certain time points and features. Will the test set exhibit the same pattern of missing data across IDs and other features? Additionally, I've observed that some public solutions handle imputation only for the training set and not for the test set. Is this an oversight, and should the test set receive similar imputation strategies for consistency?</p>",
  "messages": [
    {
      "id": "3041045",
      "postDate": "11/09/2024 21:16:18",
      "content": "<p>I noticed that not all IDs in the training set have complete data, i.e time series data, leading to NaN values for certain time points and features. Will the test set exhibit the same pattern of missing data across IDs and other features? Additionally, I've observed that some public solutions handle imputation only for the training set and not for the test set. Is this an oversight, and should the test set receive similar imputation strategies for consistency?</p>",
      "rawMarkdown": "I noticed that not all IDs in the training set have complete data, i.e time series data, leading to NaN values for certain time points and features. Will the test set exhibit the same pattern of missing data across IDs and other features? Additionally, I've observed that some public solutions handle imputation only for the training set and not for the test set. Is this an oversight, and should the test set receive similar imputation strategies for consistency?",
      "votes": null
    },
    {
      "id": "3041450",
      "postDate": "11/10/2024 11:31:25",
      "content": "<p>As far as the test data is concerned nothing can be said concretely as it is hidden. We know only one thing about test data i.e <code>The public leaderboard is calculated with approximately 38% of the test data. The final results will be based on the other 62%.</code> As currently the public leaderboard is calculated with 38% data so, chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data. 38% means more than 1/4th of the data, which I think is sufficient to test our solution's performance.</p>\n<p>As far as imputation is concerned if it is done for train set then it should also be done for test set also. As our model is being trained on train data which is imputed then it should be given test data which is also imputed.</p>",
      "rawMarkdown": "As far as the test data is concerned nothing can be said concretely as it is hidden. We know only one thing about test data i.e `The public leaderboard is calculated with approximately 38% of the test data. The final results will be based on the other 62%.` As currently the public leaderboard is calculated with 38% data so, chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data. 38% means more than 1/4th of the data, which I think is sufficient to test our solution's performance.\n\nAs far as imputation is concerned if it is done for train set then it should also be done for test set also. As our model is being trained on train data which is imputed then it should be given test data which is also imputed.",
      "votes": null
    },
    {
      "id": "3041673",
      "postDate": "11/10/2024 17:17:17",
      "content": "<p>About <code>chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data</code></p>\n<p>I believe chances are not null that the 62% of test samples used for the private LB will be different from the 38% of test samples used for public leaderboard. It happens sometimes on Kaggle competitions. And on the top on public LB is written ``so the final standings may be different```…</p>",
      "rawMarkdown": "About ```chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data```\n\nI believe chances are not null that the 62% of test samples used for the private LB will be different from the 38% of test samples used for public leaderboard. It happens sometimes on Kaggle competitions. And on the top on public LB is written ``so the final standings may be different```...",
      "votes": null
    },
    {
      "id": "3041756",
      "postDate": "11/10/2024 19:53:28",
      "content": "<p><strong><em>Famous last words..</em></strong></p>\n<p><code>As currently the public leaderboard is calculated with 38% data so, chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data.</code> </p>",
      "rawMarkdown": "***Famous last words..***\n\n`As currently the public leaderboard is calculated with 38% data so, chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data.`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3041450,
      "author_name": "taimour",
      "author_url": "",
      "post_date": "11/10/2024 11:31:25",
      "content": "<p>As far as the test data is concerned nothing can be said concretely as it is hidden. We know only one thing about test data i.e <code>The public leaderboard is calculated with approximately 38% of the test data. The final results will be based on the other 62%.</code> As currently the public leaderboard is calculated with 38% data so, chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data. 38% means more than 1/4th of the data, which I think is sufficient to test our solution's performance.</p>\n<p>As far as imputation is concerned if it is done for train set then it should also be done for test set also. As our model is being trained on train data which is imputed then it should be given test data which is also imputed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3041673,
          "author_name": "adaubas",
          "author_url": "",
          "post_date": "11/10/2024 17:17:17",
          "content": "<p>About <code>chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data</code></p>\n<p>I believe chances are not null that the 62% of test samples used for the private LB will be different from the 38% of test samples used for public leaderboard. It happens sometimes on Kaggle competitions. And on the top on public LB is written ``so the final standings may be different```…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3041756,
          "author_name": "bsmelbs",
          "author_url": "",
          "post_date": "11/10/2024 19:53:28",
          "content": "<p><strong><em>Famous last words..</em></strong></p>\n<p><code>As currently the public leaderboard is calculated with 38% data so, chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data.</code> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3041045": "I noticed that not all IDs in the training set have complete data, i.e time series data, leading to NaN values for certain time points and features. Will the test set exhibit the same pattern of missing data across IDs and other features? Additionally, I've observed that some public solutions handle imputation only for the training set and not for the test set. Is this an oversight, and should the test set receive similar imputation strategies for consistency?",
    "3041450": "As far as the test data is concerned nothing can be said concretely as it is hidden. We know only one thing about test data i.e `The public leaderboard is calculated with approximately 38% of the test data. The final results will be based on the other 62%.` As currently the public leaderboard is calculated with 38% data so, chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data. 38% means more than 1/4th of the data, which I think is sufficient to test our solution's performance.\n\nAs far as imputation is concerned if it is done for train set then it should also be done for test set also. As our model is being trained on train data which is imputed then it should be given test data which is also imputed.",
    "3041673": "About ```chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data```\n\nI believe chances are not null that the 62% of test samples used for the private LB will be different from the 38% of test samples used for public leaderboard. It happens sometimes on Kaggle competitions. And on the top on public LB is written ``so the final standings may be different```...",
    "3041756": "***Famous last words..***\n\n`As currently the public leaderboard is calculated with 38% data so, chances are high that if our solution is performing well with 38% of data then it might also give similar performance with the remaining 62% data.`"
  },
  "source": "meta"
}