{
  "id": 551468,
  "title": "Why is there no improvement in LB when I use CatBoost for online learning?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/551468",
  "author_name": "Shiqiang Lee",
  "post_date": "2024-12-13T09:45:43.133000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I used 79 raw features, and during the testing phase, the model adds the previous day's test data to my cached training data(start date 1289), removes one day of old training data, and then retrains the model daily. However, with the same parameters and data volume, the LB only improved by 0.0001（0.0056 to 0.0057） in the end.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/shiqiangli/janestreetcbtonlineretrain</a></p>",
  "messages": [
    {
      "id": 3071070,
      "postDate": "2024-12-13T09:45:43.133Z",
      "content": "<p>I used 79 raw features, and during the testing phase, the model adds the previous day's test data to my cached training data(start date 1289), removes one day of old training data, and then retrains the model daily. However, with the same parameters and data volume, the LB only improved by 0.0001（0.0056 to 0.0057） in the end.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/shiqiangli/janestreetcbtonlineretrain</a></p>",
      "rawMarkdown": "I used 79 raw features, and during the testing phase, the model adds the previous day's test data to my cached training data(start date 1289), removes one day of old training data, and then retrains the model daily. However, with the same parameters and data volume, the LB only improved by 0.0001（0.0056 to 0.0057） in the end.\n[https://www.kaggle.com/code/shiqiangli/janestreetcbtonlineretrain](url)",
      "votes": 2
    },
    {
      "id": 3077050,
      "postDate": "2024-12-20T14:20:50.833Z",
      "content": "<p>it may be overfitting if retrain every day. accumulate more days of data see if its working.</p>",
      "rawMarkdown": "it may be overfitting if retrain every day. accumulate more days of data see if its working.",
      "replies": [
        {
          "id": 3077078,
          "postDate": "2024-12-20T14:44:37.897Z",
          "content": "<p>thanks bro, I will try it</p>",
          "rawMarkdown": "thanks bro, I will try it"
        },
        {
          "id": 3077081,
          "postDate": "2024-12-20T14:47:04.573Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3071122,
      "postDate": "2024-12-13T11:21:12.690Z",
      "content": "<p>same with NN when doing online learning the LB improved by 0.0001 only. we must be doing something wrong!</p>",
      "rawMarkdown": "same with NN when doing online learning the LB improved by 0.0001 only. we must be doing something wrong!\n\n",
      "replies": [
        {
          "id": 3071233,
          "postDate": "2024-12-13T14:09:29.387Z",
          "content": "<p>yes, I will also try mlp and some derivative models related to rnn and attention</p>",
          "rawMarkdown": "yes, I will also try mlp and some derivative models related to rnn and attention"
        },
        {
          "id": 3075579,
          "postDate": "2024-12-19T03:36:27.930Z",
          "content": "<p>I use a mlp and do the online learning. It also only improved LB 0.0002. And if I set the epoch number as a little bit larger, the results get worse. Have you figured out the reason behnd it? </p>",
          "rawMarkdown": "I use a mlp and do the online learning. It also only improved LB 0.0002. And if I set the epoch number as a little bit larger, the results get worse. Have you figured out the reason behnd it? ",
          "replies": [
            {
              "id": 3077079,
              "postDate": "2024-12-20T14:45:18.593Z",
              "content": "<p>not figure out</p>",
              "rawMarkdown": "not figure out"
            }
          ]
        }
      ]
    },
    {
      "id": 3075605,
      "postDate": "2024-12-19T04:17:22.777Z",
      "content": "<p>Good work! Thanks for sharing bro!</p>",
      "rawMarkdown": "Good work! Thanks for sharing bro!"
    }
  ],
  "comments": [
    {
      "id": 3077050,
      "author_name": "ZT",
      "author_url": "",
      "post_date": "2024-12-20T14:20:50.833000",
      "content": "<p>it may be overfitting if retrain every day. accumulate more days of data see if its working.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3077078,
          "author_name": "Shiqiang Lee",
          "author_url": "",
          "post_date": "2024-12-20T14:44:37.897000",
          "content": "<p>thanks bro, I will try it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3077081,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-12-20T14:47:04.573000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3071122,
      "author_name": "Abhi",
      "author_url": "",
      "post_date": "2024-12-13T11:21:12.690000",
      "content": "<p>same with NN when doing online learning the LB improved by 0.0001 only. we must be doing something wrong!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3071233,
          "author_name": "Shiqiang Lee",
          "author_url": "",
          "post_date": "2024-12-13T14:09:29.387000",
          "content": "<p>yes, I will also try mlp and some derivative models related to rnn and attention</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3075579,
          "author_name": "lele",
          "author_url": "",
          "post_date": "2024-12-19T03:36:27.930000",
          "content": "<p>I use a mlp and do the online learning. It also only improved LB 0.0002. And if I set the epoch number as a little bit larger, the results get worse. Have you figured out the reason behnd it? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3077079,
              "author_name": "Shiqiang Lee",
              "author_url": "",
              "post_date": "2024-12-20T14:45:18.593000",
              "content": "<p>not figure out</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3075605,
      "author_name": "MrSimple",
      "author_url": "",
      "post_date": "2024-12-19T04:17:22.777000",
      "content": "<p>Good work! Thanks for sharing bro!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3071070": "I used 79 raw features, and during the testing phase, the model adds the previous day's test data to my cached training data(start date 1289), removes one day of old training data, and then retrains the model daily. However, with the same parameters and data volume, the LB only improved by 0.0001（0.0056 to 0.0057） in the end.\n[https://www.kaggle.com/code/shiqiangli/janestreetcbtonlineretrain](url)",
    "3077050": "it may be overfitting if retrain every day. accumulate more days of data see if its working.",
    "3071122": "same with NN when doing online learning the LB improved by 0.0001 only. we must be doing something wrong!\n\n",
    "3075605": "Good work! Thanks for sharing bro!"
  }
}