{
  "id": 543257,
  "title": "How do you use a notebook for training and then take those models into another for inference?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543257",
  "author_name": "",
  "post_date": "2024-10-29T14:43:13.595985600Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello, it's my first competition and I'm trying to figure out how people do things. I've always just done 1 notebook anytime did anything ML related, so I was surprised to see that a lot of people actually have multiple notebooks. Some people, it looks like, are making a notebook where they train the models, and then another where they take those models from that notebook and then use them in an inference notebook. </p>\n<p>I see that a lot of people are taking their trained models, putting it into a dataset, and then using that dataset to load in the models, but is there a way to directly use the outputted models from the training notebook in the inference one just from adding it as an input?</p>",
  "messages": [
    {
      "id": "3031291",
      "postDate": "10/29/2024 14:43:13",
      "content": "<p>Hello, it's my first competition and I'm trying to figure out how people do things. I've always just done 1 notebook anytime did anything ML related, so I was surprised to see that a lot of people actually have multiple notebooks. Some people, it looks like, are making a notebook where they train the models, and then another where they take those models from that notebook and then use them in an inference notebook. </p>\n<p>I see that a lot of people are taking their trained models, putting it into a dataset, and then using that dataset to load in the models, but is there a way to directly use the outputted models from the training notebook in the inference one just from adding it as an input?</p>",
      "rawMarkdown": "Hello, it's my first competition and I'm trying to figure out how people do things. I've always just done 1 notebook anytime did anything ML related, so I was surprised to see that a lot of people actually have multiple notebooks. Some people, it looks like, are making a notebook where they train the models, and then another where they take those models from that notebook and then use them in an inference notebook. \n\nI see that a lot of people are taking their trained models, putting it into a dataset, and then using that dataset to load in the models, but is there a way to directly use the outputted models from the training notebook in the inference one just from adding it as an input?",
      "votes": null
    },
    {
      "id": "3031302",
      "postDate": "10/29/2024 14:53:06",
      "content": "<pre><code>\n     ():\n        \n         (path, mode=)  f:\n            dill.dump(obj, f, protocol=)\n    \n     ():\n        \n         (path, mode=)  f:\n            data = dill.load(f)\n             data\n</code></pre>",
      "rawMarkdown": "```python\n#save models after training\n    def pickle_dump(self,obj, path):\n        #open path,binary write\n        with open(path, mode=\"wb\") as f:\n            dill.dump(obj, f, protocol=4)\n    #load models when inference\n    def pickle_load(self,path):\n        #open path,binary read\n        with open(path, mode=\"rb\") as f:\n            data = dill.load(f)\n            return data\n```",
      "votes": null
    },
    {
      "id": "3031309",
      "postDate": "10/29/2024 14:59:01",
      "content": "<p>Not to sound dumb, but where would the models get saved to?</p>",
      "rawMarkdown": "Not to sound dumb, but where would the models get saved to?",
      "votes": null
    },
    {
      "id": "3031319",
      "postDate": "10/29/2024 15:08:20",
      "content": "<p>Yes, you can save your notebook (the <code>save version</code> button top right) and let it write outputs. Once the notebook is successfully committed, you can add the notebook as a dataset in another notebook and use the output files. The fact that many people upload models to a dataset in this competition is simply because they trained their model on a local machine. </p>",
      "rawMarkdown": "Yes, you can save your notebook (the `save version` button top right) and let it write outputs. Once the notebook is successfully committed, you can add the notebook as a dataset in another notebook and use the output files. The fact that many people upload models to a dataset in this competition is simply because they trained their model on a local machine.",
      "votes": null
    },
    {
      "id": "3031337",
      "postDate": "10/29/2024 15:27:34",
      "content": "<p>OHH okay that makes a lot of sense thank you!</p>",
      "rawMarkdown": "OHH okay that makes a lot of sense thank you!",
      "votes": null
    },
    {
      "id": "3031516",
      "postDate": "10/29/2024 19:12:39",
      "content": "<p>You need to save your trained models in a joblib file and export them to your submission notebook <a href=\"https://www.kaggle.com/coreymichaud\" target=\"_blank\">@coreymichaud</a> </p>",
      "rawMarkdown": "You need to save your trained models in a joblib file and export them to your submission notebook @coreymichaud",
      "votes": null
    },
    {
      "id": "3031936",
      "postDate": "10/30/2024 10:21:01",
      "content": "<p>As simple as: joblib.dump(trained_model, 'trained_model.pkl') and it will be saved in your /kaggle/working directory - on the right side panel you can find it and download locally or re-upload then into another notebook as a dataset and use trained_model = joblib.load('path to your dataset with trained_model.pkl')</p>",
      "rawMarkdown": "As simple as: joblib.dump(trained_model, 'trained_model.pkl') and it will be saved in your /kaggle/working directory - on the right side panel you can find it and download locally or re-upload then into another notebook as a dataset and use trained_model = joblib.load('path to your dataset with trained_model.pkl')",
      "votes": null
    },
    {
      "id": "3032574",
      "postDate": "10/31/2024 05:22:17",
      "content": "<p>save as a json or pkl and use load_model method</p>",
      "rawMarkdown": "save as a json or pkl and use load_model method",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3031302,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "10/29/2024 14:53:06",
      "content": "<pre><code>\n     ():\n        \n         (path, mode=)  f:\n            dill.dump(obj, f, protocol=)\n    \n     ():\n        \n         (path, mode=)  f:\n            data = dill.load(f)\n             data\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3031309,
          "author_name": "coreymichaud",
          "author_url": "",
          "post_date": "10/29/2024 14:59:01",
          "content": "<p>Not to sound dumb, but where would the models get saved to?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3031936,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "10/30/2024 10:21:01",
              "content": "<p>As simple as: joblib.dump(trained_model, 'trained_model.pkl') and it will be saved in your /kaggle/working directory - on the right side panel you can find it and download locally or re-upload then into another notebook as a dataset and use trained_model = joblib.load('path to your dataset with trained_model.pkl')</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3031319,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/29/2024 15:08:20",
      "content": "<p>Yes, you can save your notebook (the <code>save version</code> button top right) and let it write outputs. Once the notebook is successfully committed, you can add the notebook as a dataset in another notebook and use the output files. The fact that many people upload models to a dataset in this competition is simply because they trained their model on a local machine. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3031337,
          "author_name": "coreymichaud",
          "author_url": "",
          "post_date": "10/29/2024 15:27:34",
          "content": "<p>OHH okay that makes a lot of sense thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3031516,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/29/2024 19:12:39",
      "content": "<p>You need to save your trained models in a joblib file and export them to your submission notebook <a href=\"https://www.kaggle.com/coreymichaud\" target=\"_blank\">@coreymichaud</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3032574,
      "author_name": "priyanshu1712",
      "author_url": "",
      "post_date": "10/31/2024 05:22:17",
      "content": "<p>save as a json or pkl and use load_model method</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3031291": "Hello, it's my first competition and I'm trying to figure out how people do things. I've always just done 1 notebook anytime did anything ML related, so I was surprised to see that a lot of people actually have multiple notebooks. Some people, it looks like, are making a notebook where they train the models, and then another where they take those models from that notebook and then use them in an inference notebook. \n\nI see that a lot of people are taking their trained models, putting it into a dataset, and then using that dataset to load in the models, but is there a way to directly use the outputted models from the training notebook in the inference one just from adding it as an input?",
    "3031302": "```python\n#save models after training\n    def pickle_dump(self,obj, path):\n        #open path,binary write\n        with open(path, mode=\"wb\") as f:\n            dill.dump(obj, f, protocol=4)\n    #load models when inference\n    def pickle_load(self,path):\n        #open path,binary read\n        with open(path, mode=\"rb\") as f:\n            data = dill.load(f)\n            return data\n```",
    "3031309": "Not to sound dumb, but where would the models get saved to?",
    "3031319": "Yes, you can save your notebook (the `save version` button top right) and let it write outputs. Once the notebook is successfully committed, you can add the notebook as a dataset in another notebook and use the output files. The fact that many people upload models to a dataset in this competition is simply because they trained their model on a local machine.",
    "3031337": "OHH okay that makes a lot of sense thank you!",
    "3031516": "You need to save your trained models in a joblib file and export them to your submission notebook @coreymichaud",
    "3031936": "As simple as: joblib.dump(trained_model, 'trained_model.pkl') and it will be saved in your /kaggle/working directory - on the right side panel you can find it and download locally or re-upload then into another notebook as a dataset and use trained_model = joblib.load('path to your dataset with trained_model.pkl')",
    "3032574": "save as a json or pkl and use load_model method"
  },
  "source": "meta"
}