{
  "id": 551490,
  "title": "One prediction with full df or multiple with folds?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551490",
  "author_name": "",
  "post_date": "2024-12-13T13:48:58.896166600Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey, this is a general question regarding xgblight, xgb and catboost.<br>\nIs it better to make the prediction on aim_df on the full train_df or is it better to make predictions on the aim_df in each fold and take the mean afterwards?<br>\nFor Tabnet it seems to be better to make predictions with each fold, since the default model, which the public notebooks use, internally makes a train test split.<br>\nSo just fitting on the whole df once, would always create the same split. While folds have multiple splits (eventhough they are smaller).<br>\nDo the boosting models work the same way internally?</p>",
  "messages": [
    {
      "id": "3071216",
      "postDate": "12/13/2024 13:48:58",
      "content": "<p>Hey, this is a general question regarding xgblight, xgb and catboost.<br>\nIs it better to make the prediction on aim_df on the full train_df or is it better to make predictions on the aim_df in each fold and take the mean afterwards?<br>\nFor Tabnet it seems to be better to make predictions with each fold, since the default model, which the public notebooks use, internally makes a train test split.<br>\nSo just fitting on the whole df once, would always create the same split. While folds have multiple splits (eventhough they are smaller).<br>\nDo the boosting models work the same way internally?</p>",
      "rawMarkdown": "Hey, this is a general question regarding xgblight, xgb and catboost.\nIs it better to make the prediction on aim_df on the full train_df or is it better to make predictions on the aim_df in each fold and take the mean afterwards?\nFor Tabnet it seems to be better to make predictions with each fold, since the default model, which the public notebooks use, internally makes a train test split.\nSo just fitting on the whole df once, would always create the same split. While folds have multiple splits (eventhough they are smaller).\nDo the boosting models work the same way internally?",
      "votes": null
    },
    {
      "id": "3071504",
      "postDate": "12/13/2024 20:46:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a>,</p>\n<p>I would suggest trying to experiment on your own first, as hands-on practice can often lead to deeper understanding. And if you encounter specific issues, feel free to share the details !</p>",
      "rawMarkdown": "Hi @mariusheuser,\n\nI would suggest trying to experiment on your own first, as hands-on practice can often lead to deeper understanding. And if you encounter specific issues, feel free to share the details !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3071504,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "12/13/2024 20:46:58",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a>,</p>\n<p>I would suggest trying to experiment on your own first, as hands-on practice can often lead to deeper understanding. And if you encounter specific issues, feel free to share the details !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3071216": "Hey, this is a general question regarding xgblight, xgb and catboost.\nIs it better to make the prediction on aim_df on the full train_df or is it better to make predictions on the aim_df in each fold and take the mean afterwards?\nFor Tabnet it seems to be better to make predictions with each fold, since the default model, which the public notebooks use, internally makes a train test split.\nSo just fitting on the whole df once, would always create the same split. While folds have multiple splits (eventhough they are smaller).\nDo the boosting models work the same way internally?",
    "3071504": "Hi @mariusheuser,\n\nI would suggest trying to experiment on your own first, as hands-on practice can often lead to deeper understanding. And if you encounter specific issues, feel free to share the details !"
  },
  "source": "meta"
}