{
  "id": 540637,
  "title": "cross_val_predict vs manual split",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/540637",
  "author_name": "",
  "post_date": "2024-10-15T13:28:56.601776500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Just wondering why everybody uses manual split and a whole bunch of lines of code if there is sklearn's cross_val_predict that returns all the predictions from cross validation? It's just because almost all notebooks are just copy-paste or do you really care about the score on the training set and not just on the validation set? </p>",
  "messages": [
    {
      "id": "3018038",
      "postDate": "10/15/2024 13:28:56",
      "content": "<p>Just wondering why everybody uses manual split and a whole bunch of lines of code if there is sklearn's cross_val_predict that returns all the predictions from cross validation? It's just because almost all notebooks are just copy-paste or do you really care about the score on the training set and not just on the validation set? </p>",
      "rawMarkdown": "Just wondering why everybody uses manual split and a whole bunch of lines of code if there is sklearn's cross_val_predict that returns all the predictions from cross validation? It's just because almost all notebooks are just copy-paste or do you really care about the score on the training set and not just on the validation set?",
      "votes": null
    },
    {
      "id": "3018505",
      "postDate": "10/15/2024 19:53:21",
      "content": "<p>hi <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a>, people often want to keep the oof predictions for tuning an ensemble and the fitted model for submission. The issue is that cross_val_predict gives you the predictions but not the fitted models, and cross_validate gives you the scores and the fitted models but not the predictions. Furthermore, when running the notebook interactively, I like to see every fold's result as soon as it is ready so that I can interrupt the computation if the first two folds show that the model is bad.</p>",
      "rawMarkdown": "hi @eu1234, people often want to keep the oof predictions for tuning an ensemble and the fitted model for submission. The issue is that cross_val_predict gives you the predictions but not the fitted models, and cross_validate gives you the scores and the fitted models but not the predictions. Furthermore, when running the notebook interactively, I like to see every fold's result as soon as it is ready so that I can interrupt the computation if the first two folds show that the model is bad.",
      "votes": null
    },
    {
      "id": "3027326",
      "postDate": "10/24/2024 17:17:40",
      "content": "<p>Also, manually  performing the splits allows you to tweak and manipulate the data, thresholds, etc during training/prediction. </p>",
      "rawMarkdown": "Also, manually  performing the splits allows you to tweak and manipulate the data, thresholds, etc during training/prediction.",
      "votes": null
    },
    {
      "id": "3027401",
      "postDate": "10/24/2024 18:55:13",
      "content": "<p>I understood that you are actually using this because you are training an ensemble of models inside cross-validation so you need to take these trained models from it and not just the results, thanks for clarification.</p>",
      "rawMarkdown": "I understood that you are actually using this because you are training an ensemble of models inside cross-validation so you need to take these trained models from it and not just the results, thanks for clarification.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3018505,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "10/15/2024 19:53:21",
      "content": "<p>hi <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a>, people often want to keep the oof predictions for tuning an ensemble and the fitted model for submission. The issue is that cross_val_predict gives you the predictions but not the fitted models, and cross_validate gives you the scores and the fitted models but not the predictions. Furthermore, when running the notebook interactively, I like to see every fold's result as soon as it is ready so that I can interrupt the computation if the first two folds show that the model is bad.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3027326,
          "author_name": "vitormeurerbesen",
          "author_url": "",
          "post_date": "10/24/2024 17:17:40",
          "content": "<p>Also, manually  performing the splits allows you to tweak and manipulate the data, thresholds, etc during training/prediction. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3027401,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "10/24/2024 18:55:13",
              "content": "<p>I understood that you are actually using this because you are training an ensemble of models inside cross-validation so you need to take these trained models from it and not just the results, thanks for clarification.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3018038": "Just wondering why everybody uses manual split and a whole bunch of lines of code if there is sklearn's cross_val_predict that returns all the predictions from cross validation? It's just because almost all notebooks are just copy-paste or do you really care about the score on the training set and not just on the validation set?",
    "3018505": "hi @eu1234, people often want to keep the oof predictions for tuning an ensemble and the fitted model for submission. The issue is that cross_val_predict gives you the predictions but not the fitted models, and cross_validate gives you the scores and the fitted models but not the predictions. Furthermore, when running the notebook interactively, I like to see every fold's result as soon as it is ready so that I can interrupt the computation if the first two folds show that the model is bad.",
    "3027326": "Also, manually  performing the splits allows you to tweak and manipulate the data, thresholds, etc during training/prediction.",
    "3027401": "I understood that you are actually using this because you are training an ensemble of models inside cross-validation so you need to take these trained models from it and not just the results, thanks for clarification."
  },
  "source": "meta"
}