{
  "id": 550985,
  "title": "Is retraining the key?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/550985",
  "author_name": "",
  "post_date": "2024-12-10T15:37:01.209810800Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I was wondering if retraining a model is really the key for this competition. </p>\n<p>I shared my online retraining solution here:<br>\n<a href=\"https://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700\" target=\"_blank\">https://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700</a> <br>\nwhere I scored 0.006 which is the highest notebook shared publicly so far.<br>\nI don't have a previous version of the notebook scored before the test set update.</p>\n<p>Maybe a simple model can outperform my solution.<br>\nWhich strategy are you using? I've described in the notebook the strategy I've adopted.</p>",
  "messages": [
    {
      "id": "3068722",
      "postDate": "12/10/2024 15:37:01",
      "content": "<p>I was wondering if retraining a model is really the key for this competition. </p>\n<p>I shared my online retraining solution here:<br>\n<a href=\"https://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700\" target=\"_blank\">https://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700</a> <br>\nwhere I scored 0.006 which is the highest notebook shared publicly so far.<br>\nI don't have a previous version of the notebook scored before the test set update.</p>\n<p>Maybe a simple model can outperform my solution.<br>\nWhich strategy are you using? I've described in the notebook the strategy I've adopted.</p>",
      "rawMarkdown": "I was wondering if retraining a model is really the key for this competition. \n\nI shared my online retraining solution here:\nhttps://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700 \nwhere I scored 0.006 which is the highest notebook shared publicly so far.\nI don't have a previous version of the notebook scored before the test set update.\n\nMaybe a simple model can outperform my solution.\nWhich strategy are you using? I've described in the notebook the strategy I've adopted.",
      "votes": null
    },
    {
      "id": "3068798",
      "postDate": "12/10/2024 17:03:17",
      "content": "<p>Thanks for sharing! So instead of continue training the pre-trained LGB booster, you are training completely new boosters using every 100 batch and then average ensemble with the pretrained ones. Am I getting there correctly?</p>\n<p>Maybe it is worth to compare the score with/-out having the continued part. It would be interesting to know how much increment the online training can contribute.</p>",
      "rawMarkdown": "Thanks for sharing! So instead of continue training the pre-trained LGB booster, you are training completely new boosters using every 100 batch and then average ensemble with the pretrained ones. Am I getting there correctly?\n\nMaybe it is worth to compare the score with/-out having the continued part. It would be interesting to know how much increment the online training can contribute.",
      "votes": null
    },
    {
      "id": "3068815",
      "postDate": "12/10/2024 17:26:21",
      "content": "<p>Yes, basically I've 2 models: the offline trained model and a new model that I'm training from scratch each 100 batches on all the online data that I hold in a cache. Additionally I've sampled a part of the original offline data that I'm adding to the data for the retrain. Finally I just did an average of the initial model and the model that I'm retraining each 100 batches.</p>\n<p>I don't think the final score is particularly high since I got 0.006 while I've another notebook where I've implemented another solution based on blending of offline trained models where I got 0.0073. Anyway it was a good attempt.</p>\n<p>Thanks for the suggestion.</p>",
      "rawMarkdown": "Yes, basically I've 2 models: the offline trained model and a new model that I'm training from scratch each 100 batches on all the online data that I hold in a cache. Additionally I've sampled a part of the original offline data that I'm adding to the data for the retrain. Finally I just did an average of the initial model and the model that I'm retraining each 100 batches.\n\nI don't think the final score is particularly high since I got 0.006 while I've another notebook where I've implemented another solution based on blending of offline trained models where I got 0.0073. Anyway it was a good attempt.\n\nThanks for the suggestion.",
      "votes": null
    },
    {
      "id": "3068919",
      "postDate": "12/10/2024 20:06:05",
      "content": "<p>I think it is due to the new test data. I have a very simple purely offline model that scores 0.0069 on LB, which is a lot higher than my previous similar models.</p>\n<p>I'm quite new at kaggle, so take what I say with a grain of salt, but I would NOT update so much on LB scores you got over the last 12 hours.</p>\n<p>If you want to draw conclusions from LB, you should retrain and resubmit old notebooks, and only compare with that. I see you don't have old notebooks, so I don't think any inference can be drawn from the example you posted.</p>",
      "rawMarkdown": "I think it is due to the new test data. I have a very simple purely offline model that scores 0.0069 on LB, which is a lot higher than my previous similar models.\n\nI'm quite new at kaggle, so take what I say with a grain of salt, but I would NOT update so much on LB scores you got over the last 12 hours.\n\nIf you want to draw conclusions from LB, you should retrain and resubmit old notebooks, and only compare with that. I see you don't have old notebooks, so I don't think any inference can be drawn from the example you posted.",
      "votes": null
    },
    {
      "id": "3069370",
      "postDate": "12/11/2024 12:08:01",
      "content": "<p>When you retrain the model, you can try to use some old train data instead of using all the new data. This method work for me.</p>",
      "rawMarkdown": "When you retrain the model, you can try to use some old train data instead of using all the new data. This method work for me.",
      "votes": null
    },
    {
      "id": "3069376",
      "postDate": "12/11/2024 12:13:42",
      "content": "<p>Thank you for sharing, in reality I'm mixing old data with the new data</p>",
      "rawMarkdown": "Thank you for sharing, in reality I'm mixing old data with the new data",
      "votes": null
    },
    {
      "id": "3069378",
      "postDate": "12/11/2024 12:14:35",
      "content": "<p>Thank you for your insights, this was also the point of the post, I think a simple model can outperform this solution</p>",
      "rawMarkdown": "Thank you for your insights, this was also the point of the post, I think a simple model can outperform this solution",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3068798,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "12/10/2024 17:03:17",
      "content": "<p>Thanks for sharing! So instead of continue training the pre-trained LGB booster, you are training completely new boosters using every 100 batch and then average ensemble with the pretrained ones. Am I getting there correctly?</p>\n<p>Maybe it is worth to compare the score with/-out having the continued part. It would be interesting to know how much increment the online training can contribute.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3068815,
          "author_name": "simonedegasperis",
          "author_url": "",
          "post_date": "12/10/2024 17:26:21",
          "content": "<p>Yes, basically I've 2 models: the offline trained model and a new model that I'm training from scratch each 100 batches on all the online data that I hold in a cache. Additionally I've sampled a part of the original offline data that I'm adding to the data for the retrain. Finally I just did an average of the initial model and the model that I'm retraining each 100 batches.</p>\n<p>I don't think the final score is particularly high since I got 0.006 while I've another notebook where I've implemented another solution based on blending of offline trained models where I got 0.0073. Anyway it was a good attempt.</p>\n<p>Thanks for the suggestion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3068919,
      "author_name": "reeeeeeeeeeeeeee",
      "author_url": "",
      "post_date": "12/10/2024 20:06:05",
      "content": "<p>I think it is due to the new test data. I have a very simple purely offline model that scores 0.0069 on LB, which is a lot higher than my previous similar models.</p>\n<p>I'm quite new at kaggle, so take what I say with a grain of salt, but I would NOT update so much on LB scores you got over the last 12 hours.</p>\n<p>If you want to draw conclusions from LB, you should retrain and resubmit old notebooks, and only compare with that. I see you don't have old notebooks, so I don't think any inference can be drawn from the example you posted.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3069378,
          "author_name": "simonedegasperis",
          "author_url": "",
          "post_date": "12/11/2024 12:14:35",
          "content": "<p>Thank you for your insights, this was also the point of the post, I think a simple model can outperform this solution</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3069370,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "12/11/2024 12:08:01",
      "content": "<p>When you retrain the model, you can try to use some old train data instead of using all the new data. This method work for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3069376,
          "author_name": "simonedegasperis",
          "author_url": "",
          "post_date": "12/11/2024 12:13:42",
          "content": "<p>Thank you for sharing, in reality I'm mixing old data with the new data</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3068722": "I was wondering if retraining a model is really the key for this competition. \n\nI shared my online retraining solution here:\nhttps://www.kaggle.com/code/simonedegasperis/online-retrain?scriptVersionId=212331700 \nwhere I scored 0.006 which is the highest notebook shared publicly so far.\nI don't have a previous version of the notebook scored before the test set update.\n\nMaybe a simple model can outperform my solution.\nWhich strategy are you using? I've described in the notebook the strategy I've adopted.",
    "3068798": "Thanks for sharing! So instead of continue training the pre-trained LGB booster, you are training completely new boosters using every 100 batch and then average ensemble with the pretrained ones. Am I getting there correctly?\n\nMaybe it is worth to compare the score with/-out having the continued part. It would be interesting to know how much increment the online training can contribute.",
    "3068815": "Yes, basically I've 2 models: the offline trained model and a new model that I'm training from scratch each 100 batches on all the online data that I hold in a cache. Additionally I've sampled a part of the original offline data that I'm adding to the data for the retrain. Finally I just did an average of the initial model and the model that I'm retraining each 100 batches.\n\nI don't think the final score is particularly high since I got 0.006 while I've another notebook where I've implemented another solution based on blending of offline trained models where I got 0.0073. Anyway it was a good attempt.\n\nThanks for the suggestion.",
    "3068919": "I think it is due to the new test data. I have a very simple purely offline model that scores 0.0069 on LB, which is a lot higher than my previous similar models.\n\nI'm quite new at kaggle, so take what I say with a grain of salt, but I would NOT update so much on LB scores you got over the last 12 hours.\n\nIf you want to draw conclusions from LB, you should retrain and resubmit old notebooks, and only compare with that. I see you don't have old notebooks, so I don't think any inference can be drawn from the example you posted.",
    "3069370": "When you retrain the model, you can try to use some old train data instead of using all the new data. This method work for me.",
    "3069376": "Thank you for sharing, in reality I'm mixing old data with the new data",
    "3069378": "Thank you for your insights, this was also the point of the post, I think a simple model can outperform this solution"
  },
  "source": "meta"
}