{
  "id": 550649,
  "title": "Re-train a model ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/550649",
  "author_name": "",
  "post_date": "2024-12-08T17:54:08.491348600Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, I was wondering if you have succesfullly retrained a model with the test data during inference.<br>\nI gave a try with this notebook: <a href=\"https://www.kaggle.com/code/simonedegasperis/test-cache\" target=\"_blank\">https://www.kaggle.com/code/simonedegasperis/test-cache</a></p>\n<p>In this notebook I've basically set up a cache where I concatenate each served batch at inference time.<br>\nEvery 30 batches I'm retraining a simple lgbm model but the final score is worse, from 0.0044 (without re'training) to -0.014.</p>\n<p>In order to build the ground truth responder_6, I'm shifting 1 datetime back the responder_6_lags because we know they are the values for each symbol at the end of the previous datetime.</p>\n<p>See:</p>\n<pre><code>    \n     batch_count %  == :\n        ()\n        labels = cache[[, , ]]\n        labels = labels.group_by([, ], maintain_order=).last()  \n        lag_cols_rename = {: }\n        labels = labels.rename(lag_cols_rename)\n        \n        labels = labels.with_columns(\n            date_id = pl.col() - ,  \n        )\n        pl_train = cache.group_by([, ], maintain_order=).last()  \n        pl_train = pl_train.join(labels, on=[, ],  how=)\n        \n        X_train = pl_train[CONFIG.feature_cols].to_numpy()\n        y_train = pl_train.select(CONFIG.target_col).to_numpy().flatten()\n\n        train_data = lgb.Dataset(X_train, label=y_train)\n</code></pre>\n<p>I'm not sure if I'm doing some conceptual error.</p>\n<p>I think the big difficulty of this competition is to use the evaluation API to retrain the models or to add additional information to, in case, model the problem as a time series.</p>\n<p>I have seen that the solutions posted mainly use an ensemble of already trained models but this strategy is limited and will hardly have a high score.</p>\n<p>Do you have any suggestions or idea?</p>\n<p>UPDATE</p>\n<p>I've worked a little bit this idea and I've achieved a score of 0.006 on the new test set</p>\n<p><a href=\"https://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700\" target=\"_blank\">https://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700</a></p>",
  "messages": [
    {
      "id": "3067009",
      "postDate": "12/08/2024 17:54:08",
      "content": "<p>Hi, I was wondering if you have succesfullly retrained a model with the test data during inference.<br>\nI gave a try with this notebook: <a href=\"https://www.kaggle.com/code/simonedegasperis/test-cache\" target=\"_blank\">https://www.kaggle.com/code/simonedegasperis/test-cache</a></p>\n<p>In this notebook I've basically set up a cache where I concatenate each served batch at inference time.<br>\nEvery 30 batches I'm retraining a simple lgbm model but the final score is worse, from 0.0044 (without re'training) to -0.014.</p>\n<p>In order to build the ground truth responder_6, I'm shifting 1 datetime back the responder_6_lags because we know they are the values for each symbol at the end of the previous datetime.</p>\n<p>See:</p>\n<pre><code>    \n     batch_count %  == :\n        ()\n        labels = cache[[, , ]]\n        labels = labels.group_by([, ], maintain_order=).last()  \n        lag_cols_rename = {: }\n        labels = labels.rename(lag_cols_rename)\n        \n        labels = labels.with_columns(\n            date_id = pl.col() - ,  \n        )\n        pl_train = cache.group_by([, ], maintain_order=).last()  \n        pl_train = pl_train.join(labels, on=[, ],  how=)\n        \n        X_train = pl_train[CONFIG.feature_cols].to_numpy()\n        y_train = pl_train.select(CONFIG.target_col).to_numpy().flatten()\n\n        train_data = lgb.Dataset(X_train, label=y_train)\n</code></pre>\n<p>I'm not sure if I'm doing some conceptual error.</p>\n<p>I think the big difficulty of this competition is to use the evaluation API to retrain the models or to add additional information to, in case, model the problem as a time series.</p>\n<p>I have seen that the solutions posted mainly use an ensemble of already trained models but this strategy is limited and will hardly have a high score.</p>\n<p>Do you have any suggestions or idea?</p>\n<p>UPDATE</p>\n<p>I've worked a little bit this idea and I've achieved a score of 0.006 on the new test set</p>\n<p><a href=\"https://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700\" target=\"_blank\">https://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700</a></p>",
      "rawMarkdown": "Hi, I was wondering if you have succesfullly retrained a model with the test data during inference.\nI gave a try with this notebook: https://www.kaggle.com/code/simonedegasperis/test-cache\n\nIn this notebook I've basically set up a cache where I concatenate each served batch at inference time.\nEvery 30 batches I'm retraining a simple lgbm model but the final score is worse, from 0.0044 (without re'training) to -0.014.\n\nIn order to build the ground truth responder_6, I'm shifting 1 datetime back the responder_6_lags because we know they are the values for each symbol at the end of the previous datetime.\n\nSee:\n\n```python\n    # re-train a model on the fly every 30 batches\n    if batch_count % 30 == 0:\n        print(\"Using cache data\")\n        labels = cache[['date_id', 'symbol_id', 'responder_6_lag_1']]\n        labels = labels.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).last()  # pick up last record of previous date\n        lag_cols_rename = {\"responder_6_lag_1\": \"responder_6\"}\n        labels = labels.rename(lag_cols_rename)\n        # I shift 1 day back because we know that responder_6_lag_1 correspond to the last recrd of the previous day\n        labels = labels.with_columns(\n            date_id = pl.col('date_id') - 1,  # lagged by 1 day\n        )\n        pl_train = cache.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).last()  # pick up last record of previous date\n        pl_train = pl_train.join(labels, on=[\"date_id\", \"symbol_id\"],  how=\"left\")\n        # after this process we will obtain the labels\n        X_train = pl_train[CONFIG.feature_cols].to_numpy()\n        y_train = pl_train.select(CONFIG.target_col).to_numpy().flatten()\n\n        train_data = lgb.Dataset(X_train, label=y_train)\n```\n\nI'm not sure if I'm doing some conceptual error.\n\nI think the big difficulty of this competition is to use the evaluation API to retrain the models or to add additional information to, in case, model the problem as a time series.\n\nI have seen that the solutions posted mainly use an ensemble of already trained models but this strategy is limited and will hardly have a high score.\n\nDo you have any suggestions or idea?\n\n\nUPDATE\n\nI've worked a little bit this idea and I've achieved a score of 0.006 on the new test set\n\nhttps://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700",
      "votes": null
    },
    {
      "id": "3067036",
      "postDate": "12/08/2024 19:07:50",
      "content": "<p>How do you use the retrained model? </p>\n<p>I would presume that your 0.0044 lgbm is trained on a few years of data, while the \"new\" one is only trained on 30 days. If you simply replace the old with the new, your score should significantly suffer, even if your code creates all features correctly.</p>\n<p>Other than that, for debugging, I think it is quite helpful to create a small offline version of the API, so you can check whether you end up with the correct responder values. Debugging by tripple checking the code only gets you so far, tests work best. </p>",
      "rawMarkdown": "How do you use the retrained model? \n\nI would presume that your 0.0044 lgbm is trained on a few years of data, while the \"new\" one is only trained on 30 days. If you simply replace the old with the new, your score should significantly suffer, even if your code creates all features correctly.\n\n\nOther than that, for debugging, I think it is quite helpful to create a small offline version of the API, so you can check whether you end up with the correct responder values. Debugging by tripple checking the code only gets you so far, tests work best.",
      "votes": null
    },
    {
      "id": "3067042",
      "postDate": "12/08/2024 19:29:05",
      "content": "<p>Thank you for your answer. This looks a good suggestion. Creating an offline version of the API can be really helpful, however I've tested that the code generates the expected features and labels as expected (or at least as I expect) in a playground notebook.</p>\n<p>For the initial model it has been trained for about the last 500 datetimes present in the train set by subsampling a fraction of data for each datetime (due to memory limitations).</p>\n<p>For the retrained model, I'm retraining on each 30 batches incrementally 30, 60, 90,… so you're right, first retrained models will suffer from that. This was a sort of quick experiment to see if I was going in the right direction considering that there's very small work or conversation on the subject of retraining in the competition.<br>\nIt is true that financial markets often follow cycles but very often past data is not significant for future data. This is also the reason why I tried a simple approach to see a possible outcome of retrain on new data even if small subset.</p>",
      "rawMarkdown": "Thank you for your answer. This looks a good suggestion. Creating an offline version of the API can be really helpful, however I've tested that the code generates the expected features and labels as expected (or at least as I expect) in a playground notebook.\n\nFor the initial model it has been trained for about the last 500 datetimes present in the train set by subsampling a fraction of data for each datetime (due to memory limitations).\n\nFor the retrained model, I'm retraining on each 30 batches incrementally 30, 60, 90,... so you're right, first retrained models will suffer from that. This was a sort of quick experiment to see if I was going in the right direction considering that there's very small work or conversation on the subject of retraining in the competition.\nIt is true that financial markets often follow cycles but very often past data is not significant for future data. This is also the reason why I tried a simple approach to see a possible outcome of retrain on new data even if small subset.",
      "votes": null
    },
    {
      "id": "3067135",
      "postDate": "12/09/2024 00:30:40",
      "content": "<p>In my view, there doesn’t seem to be any issue with the code itself. Regarding the strategy of incrementally increasing training data that you mentioned, there are two potential problems:</p>\n<ol>\n<li>The initial model might be affected by insufficient training data.</li>\n<li>As the data volume grows in the future, how can we control the training time to stay within the one-minute requirement?</li>\n</ol>\n<p>One approach I tried was maintaining a list of models. This list contained a total of six models, each trained on 30 days of data. This approach might allow the models to see more data, but the results were not very satisfactory, and I didn’t pursue it further.</p>\n<p>I transitioned to using NN models instead, but one idea that I believe could be explored further is to use the data from each day to generate a tree. When generating a new tree, you would simultaneously drop the earliest tree, thereby updating the information incrementally.</p>",
      "rawMarkdown": "In my view, there doesn’t seem to be any issue with the code itself. Regarding the strategy of incrementally increasing training data that you mentioned, there are two potential problems:\n\n1. The initial model might be affected by insufficient training data.\n2. As the data volume grows in the future, how can we control the training time to stay within the one-minute requirement?\n\nOne approach I tried was maintaining a list of models. This list contained a total of six models, each trained on 30 days of data. This approach might allow the models to see more data, but the results were not very satisfactory, and I didn’t pursue it further.\n\nI transitioned to using NN models instead, but one idea that I believe could be explored further is to use the data from each day to generate a tree. When generating a new tree, you would simultaneously drop the earliest tree, thereby updating the information incrementally.",
      "votes": null
    },
    {
      "id": "3067140",
      "postDate": "12/09/2024 00:47:43",
      "content": "<p>Thank you for the suggestion. Are you retraining on the fly the NN or are you just using it once you trained offline on the data provided in the competition?<br>\nThe 1 minute limitation is big and may be an issue with the approach provided.</p>",
      "rawMarkdown": "Thank you for the suggestion. Are you retraining on the fly the NN or are you just using it once you trained offline on the data provided in the competition?\nThe 1 minute limitation is big and may be an issue with the approach provided.",
      "votes": null
    },
    {
      "id": "3067147",
      "postDate": "12/09/2024 01:09:11",
      "content": "<p>I tried updating the model during the submission process, not by fully retraining it but by updating the parameters with a controlled learning rate. With this approach, the time constraint of one minute is manageable.</p>",
      "rawMarkdown": "I tried updating the model during the submission process, not by fully retraining it but by updating the parameters with a controlled learning rate. With this approach, the time constraint of one minute is manageable.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3067036,
      "author_name": "forecastingvibes",
      "author_url": "",
      "post_date": "12/08/2024 19:07:50",
      "content": "<p>How do you use the retrained model? </p>\n<p>I would presume that your 0.0044 lgbm is trained on a few years of data, while the \"new\" one is only trained on 30 days. If you simply replace the old with the new, your score should significantly suffer, even if your code creates all features correctly.</p>\n<p>Other than that, for debugging, I think it is quite helpful to create a small offline version of the API, so you can check whether you end up with the correct responder values. Debugging by tripple checking the code only gets you so far, tests work best. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3067042,
          "author_name": "simonedegasperis",
          "author_url": "",
          "post_date": "12/08/2024 19:29:05",
          "content": "<p>Thank you for your answer. This looks a good suggestion. Creating an offline version of the API can be really helpful, however I've tested that the code generates the expected features and labels as expected (or at least as I expect) in a playground notebook.</p>\n<p>For the initial model it has been trained for about the last 500 datetimes present in the train set by subsampling a fraction of data for each datetime (due to memory limitations).</p>\n<p>For the retrained model, I'm retraining on each 30 batches incrementally 30, 60, 90,… so you're right, first retrained models will suffer from that. This was a sort of quick experiment to see if I was going in the right direction considering that there's very small work or conversation on the subject of retraining in the competition.<br>\nIt is true that financial markets often follow cycles but very often past data is not significant for future data. This is also the reason why I tried a simple approach to see a possible outcome of retrain on new data even if small subset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3067135,
      "author_name": "xuliangxu",
      "author_url": "",
      "post_date": "12/09/2024 00:30:40",
      "content": "<p>In my view, there doesn’t seem to be any issue with the code itself. Regarding the strategy of incrementally increasing training data that you mentioned, there are two potential problems:</p>\n<ol>\n<li>The initial model might be affected by insufficient training data.</li>\n<li>As the data volume grows in the future, how can we control the training time to stay within the one-minute requirement?</li>\n</ol>\n<p>One approach I tried was maintaining a list of models. This list contained a total of six models, each trained on 30 days of data. This approach might allow the models to see more data, but the results were not very satisfactory, and I didn’t pursue it further.</p>\n<p>I transitioned to using NN models instead, but one idea that I believe could be explored further is to use the data from each day to generate a tree. When generating a new tree, you would simultaneously drop the earliest tree, thereby updating the information incrementally.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3067140,
          "author_name": "simonedegasperis",
          "author_url": "",
          "post_date": "12/09/2024 00:47:43",
          "content": "<p>Thank you for the suggestion. Are you retraining on the fly the NN or are you just using it once you trained offline on the data provided in the competition?<br>\nThe 1 minute limitation is big and may be an issue with the approach provided.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3067147,
              "author_name": "xuliangxu",
              "author_url": "",
              "post_date": "12/09/2024 01:09:11",
              "content": "<p>I tried updating the model during the submission process, not by fully retraining it but by updating the parameters with a controlled learning rate. With this approach, the time constraint of one minute is manageable.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3067009": "Hi, I was wondering if you have succesfullly retrained a model with the test data during inference.\nI gave a try with this notebook: https://www.kaggle.com/code/simonedegasperis/test-cache\n\nIn this notebook I've basically set up a cache where I concatenate each served batch at inference time.\nEvery 30 batches I'm retraining a simple lgbm model but the final score is worse, from 0.0044 (without re'training) to -0.014.\n\nIn order to build the ground truth responder_6, I'm shifting 1 datetime back the responder_6_lags because we know they are the values for each symbol at the end of the previous datetime.\n\nSee:\n\n```python\n    # re-train a model on the fly every 30 batches\n    if batch_count % 30 == 0:\n        print(\"Using cache data\")\n        labels = cache[['date_id', 'symbol_id', 'responder_6_lag_1']]\n        labels = labels.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).last()  # pick up last record of previous date\n        lag_cols_rename = {\"responder_6_lag_1\": \"responder_6\"}\n        labels = labels.rename(lag_cols_rename)\n        # I shift 1 day back because we know that responder_6_lag_1 correspond to the last recrd of the previous day\n        labels = labels.with_columns(\n            date_id = pl.col('date_id') - 1,  # lagged by 1 day\n        )\n        pl_train = cache.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).last()  # pick up last record of previous date\n        pl_train = pl_train.join(labels, on=[\"date_id\", \"symbol_id\"],  how=\"left\")\n        # after this process we will obtain the labels\n        X_train = pl_train[CONFIG.feature_cols].to_numpy()\n        y_train = pl_train.select(CONFIG.target_col).to_numpy().flatten()\n\n        train_data = lgb.Dataset(X_train, label=y_train)\n```\n\nI'm not sure if I'm doing some conceptual error.\n\nI think the big difficulty of this competition is to use the evaluation API to retrain the models or to add additional information to, in case, model the problem as a time series.\n\nI have seen that the solutions posted mainly use an ensemble of already trained models but this strategy is limited and will hardly have a high score.\n\nDo you have any suggestions or idea?\n\n\nUPDATE\n\nI've worked a little bit this idea and I've achieved a score of 0.006 on the new test set\n\nhttps://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700",
    "3067036": "How do you use the retrained model? \n\nI would presume that your 0.0044 lgbm is trained on a few years of data, while the \"new\" one is only trained on 30 days. If you simply replace the old with the new, your score should significantly suffer, even if your code creates all features correctly.\n\n\nOther than that, for debugging, I think it is quite helpful to create a small offline version of the API, so you can check whether you end up with the correct responder values. Debugging by tripple checking the code only gets you so far, tests work best.",
    "3067042": "Thank you for your answer. This looks a good suggestion. Creating an offline version of the API can be really helpful, however I've tested that the code generates the expected features and labels as expected (or at least as I expect) in a playground notebook.\n\nFor the initial model it has been trained for about the last 500 datetimes present in the train set by subsampling a fraction of data for each datetime (due to memory limitations).\n\nFor the retrained model, I'm retraining on each 30 batches incrementally 30, 60, 90,... so you're right, first retrained models will suffer from that. This was a sort of quick experiment to see if I was going in the right direction considering that there's very small work or conversation on the subject of retraining in the competition.\nIt is true that financial markets often follow cycles but very often past data is not significant for future data. This is also the reason why I tried a simple approach to see a possible outcome of retrain on new data even if small subset.",
    "3067135": "In my view, there doesn’t seem to be any issue with the code itself. Regarding the strategy of incrementally increasing training data that you mentioned, there are two potential problems:\n\n1. The initial model might be affected by insufficient training data.\n2. As the data volume grows in the future, how can we control the training time to stay within the one-minute requirement?\n\nOne approach I tried was maintaining a list of models. This list contained a total of six models, each trained on 30 days of data. This approach might allow the models to see more data, but the results were not very satisfactory, and I didn’t pursue it further.\n\nI transitioned to using NN models instead, but one idea that I believe could be explored further is to use the data from each day to generate a tree. When generating a new tree, you would simultaneously drop the earliest tree, thereby updating the information incrementally.",
    "3067140": "Thank you for the suggestion. Are you retraining on the fly the NN or are you just using it once you trained offline on the data provided in the competition?\nThe 1 minute limitation is big and may be an issue with the approach provided.",
    "3067147": "I tried updating the model during the submission process, not by fully retraining it but by updating the parameters with a controlled learning rate. With this approach, the time constraint of one minute is manageable."
  },
  "source": "meta"
}