{
  "id": 550197,
  "title": "Online Learning made the score NEGATIVE ???",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/550197",
  "author_name": "",
  "post_date": "2024-12-05T23:48:50.694586100Z",
  "votes": null,
  "comment_count": 14,
  "views": 0,
  "content": "<p>So we first did a base NN model, no feature engineering, the score is 0.0008.<br>\nand applied online learning, reduced LR to 1e-6, batch=32, and trained 2 epochs for each day.<br>\nthe score became <strong>-0.43</strong>…??</p>\n<p>We had no clue why is this happening.<br>\nAny suggestions?  </p>\n<hr>\n<p>Update</p>\n<p>Thank you all for commenting.<br>\nHere I will report my findings.<br>\nEliminate Batchnorm layers, it will reinitialize it self during training. if your online learning method does not learn enough information, batchnorm will cause catastrophic weight loss, which inevitibly result in the weights become random and score gets worse.</p>",
  "messages": [
    {
      "id": "3064721",
      "postDate": "12/05/2024 23:48:50",
      "content": "<p>So we first did a base NN model, no feature engineering, the score is 0.0008.<br>\nand applied online learning, reduced LR to 1e-6, batch=32, and trained 2 epochs for each day.<br>\nthe score became <strong>-0.43</strong>…??</p>\n<p>We had no clue why is this happening.<br>\nAny suggestions?  </p>\n<hr>\n<p>Update</p>\n<p>Thank you all for commenting.<br>\nHere I will report my findings.<br>\nEliminate Batchnorm layers, it will reinitialize it self during training. if your online learning method does not learn enough information, batchnorm will cause catastrophic weight loss, which inevitibly result in the weights become random and score gets worse.</p>",
      "rawMarkdown": "So we first did a base NN model, no feature engineering, the score is 0.0008.\nand applied online learning, reduced LR to 1e-6, batch=32, and trained 2 epochs for each day.\nthe score became **-0.43**...??\n\nWe had no clue why is this happening.\nAny suggestions?  \n\n\n-----------------------------------------\nUpdate\n\nThank you all for commenting.\nHere I will report my findings.\nEliminate Batchnorm layers, it will reinitialize it self during training. if your online learning method does not learn enough information, batchnorm will cause catastrophic weight loss, which inevitibly result in the weights become random and score gets worse.",
      "votes": null
    },
    {
      "id": "3064861",
      "postDate": "12/06/2024 04:00:20",
      "content": "<p>I'm sorry, this is not a suggestion, but a report.<br>\nI also tried adding a few layers to the base NN model and introducing online learning, but the evaluation metrics turned negative.<br>\nIt may be different from the online learning that the higher-ranking people are talking about.</p>",
      "rawMarkdown": "I'm sorry, this is not a suggestion, but a report.\nI also tried adding a few layers to the base NN model and introducing online learning, but the evaluation metrics turned negative.\nIt may be different from the online learning that the higher-ranking people are talking about.",
      "votes": null
    },
    {
      "id": "3064904",
      "postDate": "12/06/2024 05:34:09",
      "content": "<p>training with the data after submission. what else could it be …🤯</p>",
      "rawMarkdown": "training with the data after submission. what else could it be ...🤯",
      "votes": null
    },
    {
      "id": "3065439",
      "postDate": "12/06/2024 19:43:33",
      "content": "<p>Most probably you have a bug in your code or you are using wrong logic for the implementation. It took me while to realize it, but the most important step and the very first step in any ML project is to assure integrity and robustness of your code otherwise you'll will be deeming models and techniques to be useless while the actual problem is that you didn't implement them right.</p>",
      "rawMarkdown": "Most probably you have a bug in your code or you are using wrong logic for the implementation. It took me while to realize it, but the most important step and the very first step in any ML project is to assure integrity and robustness of your code otherwise you'll will be deeming models and techniques to be useless while the actual problem is that you didn't implement them right.",
      "votes": null
    },
    {
      "id": "3065511",
      "postDate": "12/06/2024 23:09:56",
      "content": "<p>OL doesn’t guarantee any boost. If not implemented properly, it can produce worse results. </p>\n<p>My assumption is that your model was fitted to the noise during the update rather than the signal. Ofc it could also be some bugs in the code :(</p>",
      "rawMarkdown": "OL doesn’t guarantee any boost. If not implemented properly, it can produce worse results. \n\nMy assumption is that your model was fitted to the noise during the update rather than the signal. Ofc it could also be some bugs in the code :(",
      "votes": null
    },
    {
      "id": "3066709",
      "postDate": "12/08/2024 12:04:06",
      "content": "<p>The same with me.I guess the simple NN Model is not suitable for online learning。</p>",
      "rawMarkdown": "The same with me.I guess the simple NN Model is not suitable for online learning。",
      "votes": null
    },
    {
      "id": "3067556",
      "postDate": "12/09/2024 12:04:23",
      "content": "<p>Out of curiosity what happens if you make your batch size 1024 or 2048? </p>",
      "rawMarkdown": "Out of curiosity what happens if you make your batch size 1024 or 2048?",
      "votes": null
    },
    {
      "id": "3077506",
      "postDate": "12/21/2024 04:22:07",
      "content": "<p>May I ask what function do you use for online learning of NN (if you use pytorch)? Is it official integrated function or you just write all the train process by hand? I have tested my data on a simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!</p>",
      "rawMarkdown": "May I ask what function do you use for online learning of NN (if you use pytorch)? Is it official integrated function or you just write all the train process by hand? I have tested my data on a simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!",
      "votes": null
    },
    {
      "id": "3077749",
      "postDate": "12/21/2024 11:04:56",
      "content": "<p>Online learning is simply to run the train loop again using the updated data. </p>",
      "rawMarkdown": "Online learning is simply to run the train loop again using the updated data.",
      "votes": null
    },
    {
      "id": "3077783",
      "postDate": "12/21/2024 11:56:19",
      "content": "<p>Hi! I agree, and online learning runs well on your simulator:<br>\n<a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test</a></p>\n<p>but fail on submission, have you encountered this problem and how you solve it?</p>",
      "rawMarkdown": "Hi! I agree, and online learning runs well on your simulator:\nhttps://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\n\nbut fail on submission, have you encountered this problem and how you solve it?",
      "votes": null
    },
    {
      "id": "3077813",
      "postDate": "12/21/2024 12:33:36",
      "content": "<p>Can you share your specific error message? Or better to share the code if you don’t mind (you can leave the main idea private if you care).</p>",
      "rawMarkdown": "Can you share your specific error message? Or better to share the code if you don’t mind (you can leave the main idea private if you care).",
      "votes": null
    },
    {
      "id": "3077827",
      "postDate": "12/21/2024 12:48:38",
      "content": "<p>Grateful for that if you would like to help!</p>\n<p>I use this one:</p>\n<p><a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test</a></p>\n<p>and it runs well, I use trainer to retrain, and my code is like this:</p>\n<pre><code>ds = NNDataset(X_nn_train, y_nn_train, w_nn_train)\ndl = DataLoader(ds, batch_size = 2048), =)\n\ntrainer = Trainer(\n    =CONFIG.nn_retrain_epochs,\n    =,\n    =,\n    =\n)\n\n model  nn_models:\n    model.train()\n    model.(device)\n    trainer.fit(model, dl)\n    model.eval()\n    model.(device)\n()\n</code></pre>\n<p>I use NNDataset to process data and DataLoader to feed data, may I know if I need to declare global or somthing else? Wish to have your insight :)</p>",
      "rawMarkdown": "Grateful for that if you would like to help!\n\nI use this one:\n\nhttps://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\n\nand it runs well, I use trainer to retrain, and my code is like this:\n\n    ds = NNDataset(X_nn_train, y_nn_train, w_nn_train)\n    dl = DataLoader(ds, batch_size = 2048), shuffle=False)\n\n    trainer = Trainer(\n        max_epochs=CONFIG.nn_retrain_epochs,\n        enable_progress_bar=True,\n        logger=False,\n        enable_checkpointing=False\n    )\n\n    for model in nn_models:\n        model.train()\n        model.to(device)\n        trainer.fit(model, dl)\n        model.eval()\n        model.to(device)\n    print(\"[Retrain] NN models fine-tuned.\")\n\nI use NNDataset to process data and DataLoader to feed data, may I know if I need to declare global or somthing else? Wish to have your insight :)",
      "votes": null
    },
    {
      "id": "3077828",
      "postDate": "12/21/2024 12:50:53",
      "content": "<p>Don’t see any obvious problems….. but what is the error msg?</p>",
      "rawMarkdown": "Don’t see any obvious problems….. but what is the error msg?",
      "votes": null
    },
    {
      "id": "3077867",
      "postDate": "12/21/2024 13:20:12",
      "content": "<p>Notebook Inference Server Error<br>\nYour submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed. See more debugging tips</p>",
      "rawMarkdown": "Notebook Inference Server Error\nYour submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed. See more debugging tips",
      "votes": null
    },
    {
      "id": "3077870",
      "postDate": "12/21/2024 13:23:46",
      "content": "<p>May I just ask do you use a integrated datamodule like DataModule of pytorchlightning in the retraining session or you write your own version?</p>",
      "rawMarkdown": "May I just ask do you use a integrated datamodule like DataModule of pytorchlightning in the retraining session or you write your own version?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3064861,
      "author_name": "thinkphil",
      "author_url": "",
      "post_date": "12/06/2024 04:00:20",
      "content": "<p>I'm sorry, this is not a suggestion, but a report.<br>\nI also tried adding a few layers to the base NN model and introducing online learning, but the evaluation metrics turned negative.<br>\nIt may be different from the online learning that the higher-ranking people are talking about.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3064904,
          "author_name": "zoutain",
          "author_url": "",
          "post_date": "12/06/2024 05:34:09",
          "content": "<p>training with the data after submission. what else could it be …🤯</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3065439,
      "author_name": "aymanallawi",
      "author_url": "",
      "post_date": "12/06/2024 19:43:33",
      "content": "<p>Most probably you have a bug in your code or you are using wrong logic for the implementation. It took me while to realize it, but the most important step and the very first step in any ML project is to assure integrity and robustness of your code otherwise you'll will be deeming models and techniques to be useless while the actual problem is that you didn't implement them right.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3065511,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "12/06/2024 23:09:56",
      "content": "<p>OL doesn’t guarantee any boost. If not implemented properly, it can produce worse results. </p>\n<p>My assumption is that your model was fitted to the noise during the update rather than the signal. Ofc it could also be some bugs in the code :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 3077506,
          "author_name": "larrylin666",
          "author_url": "",
          "post_date": "12/21/2024 04:22:07",
          "content": "<p>May I ask what function do you use for online learning of NN (if you use pytorch)? Is it official integrated function or you just write all the train process by hand? I have tested my data on a simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3077749,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "12/21/2024 11:04:56",
              "content": "<p>Online learning is simply to run the train loop again using the updated data. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3077783,
                  "author_name": "larrylin666",
                  "author_url": "",
                  "post_date": "12/21/2024 11:56:19",
                  "content": "<p>Hi! I agree, and online learning runs well on your simulator:<br>\n<a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test</a></p>\n<p>but fail on submission, have you encountered this problem and how you solve it?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3077813,
                      "author_name": "shiyili",
                      "author_url": "",
                      "post_date": "12/21/2024 12:33:36",
                      "content": "<p>Can you share your specific error message? Or better to share the code if you don’t mind (you can leave the main idea private if you care).</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3077827,
                          "author_name": "larrylin666",
                          "author_url": "",
                          "post_date": "12/21/2024 12:48:38",
                          "content": "<p>Grateful for that if you would like to help!</p>\n<p>I use this one:</p>\n<p><a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test</a></p>\n<p>and it runs well, I use trainer to retrain, and my code is like this:</p>\n<pre><code>ds = NNDataset(X_nn_train, y_nn_train, w_nn_train)\ndl = DataLoader(ds, batch_size = 2048), =)\n\ntrainer = Trainer(\n    =CONFIG.nn_retrain_epochs,\n    =,\n    =,\n    =\n)\n\n model  nn_models:\n    model.train()\n    model.(device)\n    trainer.fit(model, dl)\n    model.eval()\n    model.(device)\n()\n</code></pre>\n<p>I use NNDataset to process data and DataLoader to feed data, may I know if I need to declare global or somthing else? Wish to have your insight :)</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3077828,
                              "author_name": "shiyili",
                              "author_url": "",
                              "post_date": "12/21/2024 12:50:53",
                              "content": "<p>Don’t see any obvious problems….. but what is the error msg?</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3077867,
                                  "author_name": "larrylin666",
                                  "author_url": "",
                                  "post_date": "12/21/2024 13:20:12",
                                  "content": "<p>Notebook Inference Server Error<br>\nYour submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed. See more debugging tips</p>",
                                  "votes": null,
                                  "replies": []
                                },
                                {
                                  "id": 3077870,
                                  "author_name": "larrylin666",
                                  "author_url": "",
                                  "post_date": "12/21/2024 13:23:46",
                                  "content": "<p>May I just ask do you use a integrated datamodule like DataModule of pytorchlightning in the retraining session or you write your own version?</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3066709,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "12/08/2024 12:04:06",
      "content": "<p>The same with me.I guess the simple NN Model is not suitable for online learning。</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3067556,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "12/09/2024 12:04:23",
      "content": "<p>Out of curiosity what happens if you make your batch size 1024 or 2048? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3064721": "So we first did a base NN model, no feature engineering, the score is 0.0008.\nand applied online learning, reduced LR to 1e-6, batch=32, and trained 2 epochs for each day.\nthe score became **-0.43**...??\n\nWe had no clue why is this happening.\nAny suggestions?  \n\n\n-----------------------------------------\nUpdate\n\nThank you all for commenting.\nHere I will report my findings.\nEliminate Batchnorm layers, it will reinitialize it self during training. if your online learning method does not learn enough information, batchnorm will cause catastrophic weight loss, which inevitibly result in the weights become random and score gets worse.",
    "3064861": "I'm sorry, this is not a suggestion, but a report.\nI also tried adding a few layers to the base NN model and introducing online learning, but the evaluation metrics turned negative.\nIt may be different from the online learning that the higher-ranking people are talking about.",
    "3064904": "training with the data after submission. what else could it be ...🤯",
    "3065439": "Most probably you have a bug in your code or you are using wrong logic for the implementation. It took me while to realize it, but the most important step and the very first step in any ML project is to assure integrity and robustness of your code otherwise you'll will be deeming models and techniques to be useless while the actual problem is that you didn't implement them right.",
    "3065511": "OL doesn’t guarantee any boost. If not implemented properly, it can produce worse results. \n\nMy assumption is that your model was fitted to the noise during the update rather than the signal. Ofc it could also be some bugs in the code :(",
    "3066709": "The same with me.I guess the simple NN Model is not suitable for online learning。",
    "3067556": "Out of curiosity what happens if you make your batch size 1024 or 2048?",
    "3077506": "May I ask what function do you use for online learning of NN (if you use pytorch)? Is it official integrated function or you just write all the train process by hand? I have tested my data on a simulator, it works perfectly with a result but I failed immediately on the submission. Really appreciate your response!",
    "3077749": "Online learning is simply to run the train loop again using the updated data.",
    "3077783": "Hi! I agree, and online learning runs well on your simulator:\nhttps://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\n\nbut fail on submission, have you encountered this problem and how you solve it?",
    "3077813": "Can you share your specific error message? Or better to share the code if you don’t mind (you can leave the main idea private if you care).",
    "3077827": "Grateful for that if you would like to help!\n\nI use this one:\n\nhttps://www.kaggle.com/code/shiyili/js24-rmf-submission-api-debug-with-synthetic-test\n\nand it runs well, I use trainer to retrain, and my code is like this:\n\n    ds = NNDataset(X_nn_train, y_nn_train, w_nn_train)\n    dl = DataLoader(ds, batch_size = 2048), shuffle=False)\n\n    trainer = Trainer(\n        max_epochs=CONFIG.nn_retrain_epochs,\n        enable_progress_bar=True,\n        logger=False,\n        enable_checkpointing=False\n    )\n\n    for model in nn_models:\n        model.train()\n        model.to(device)\n        trainer.fit(model, dl)\n        model.eval()\n        model.to(device)\n    print(\"[Retrain] NN models fine-tuned.\")\n\nI use NNDataset to process data and DataLoader to feed data, may I know if I need to declare global or somthing else? Wish to have your insight :)",
    "3077828": "Don’t see any obvious problems….. but what is the error msg?",
    "3077867": "Notebook Inference Server Error\nYour submission notebook may not have started the inference server that is called to obtain predictions. This could mean you forgot to start it, or the notebook crashed. See more debugging tips",
    "3077870": "May I just ask do you use a integrated datamodule like DataModule of pytorchlightning in the retraining session or you write your own version?"
  },
  "source": "meta"
}