{
  "id": 555919,
  "title": "Online learning fail for batchnorm1d?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/555919",
  "author_name": "gezi",
  "post_date": "2025-01-10T05:42:11.860000",
  "votes": 6,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I tried online learning, tough not improve much it could improve LB 0.0001, but after my model with batchnorm1d layer, online learning will fail get local result -0.78.. like.  Very confused about this.</p>",
  "messages": [
    {
      "id": 3093156,
      "postDate": "2025-01-10T14:25:35.073Z",
      "content": "<p>From my experiments, all normalisation methods in models fail badly including batch normalisation, layer normalisation, instance normalisation etc. (some of them are ok, but fail badly in online learning). Guess it's because of the super non-stationary property of this dataset and how we train the model. So models without any normalisation in models are the best. </p>",
      "rawMarkdown": "From my experiments, all normalisation methods in models fail badly including batch normalisation, layer normalisation, instance normalisation etc. (some of them are ok, but fail badly in online learning). Guess it's because of the super non-stationary property of this dataset and how we train the model. So models without any normalisation in models are the best. ",
      "votes": 7,
      "replies": [
        {
          "id": 3093252,
          "postDate": "2025-01-10T17:03:24.520Z",
          "content": "<p>Thanks for sharing this insight! Do you mean without any normalisation in the entire model? Not just the normalising the input?</p>",
          "rawMarkdown": "Thanks for sharing this insight! Do you mean without any normalisation in the entire model? Not just the normalising the input?\n",
          "replies": [
            {
              "id": 3093272,
              "postDate": "2025-01-10T17:15:06.393Z",
              "content": "<p>Yes, using normalization in my option is one of the several most common traps resulting in so my teams couldn't build high performance nn models in this competition.</p>",
              "rawMarkdown": "Yes, using normalization in my option is one of the several most common traps resulting in so my teams couldn't build high performance nn models in this competition.",
              "votes": 1
            },
            {
              "id": 3093277,
              "postDate": "2025-01-10T17:17:47.023Z",
              "content": "<p>That's very generous advice, I'll give this a go</p>",
              "rawMarkdown": "That's very generous advice, I'll give this a go",
              "votes": 1
            },
            {
              "id": 3093278,
              "postDate": "2025-01-10T17:17:52.890Z",
              "content": "<p>I have spent so much time experimenting how to properly add norm modules to the model…tried all kinds of default norm methods and also designed some in-house norm modules, nothing had worked :(</p>",
              "rawMarkdown": "I have spent so much time experimenting how to properly add norm modules to the model...tried all kinds of default norm methods and also designed some in-house norm modules, nothing had worked :(",
              "votes": 1
            },
            {
              "id": 3093282,
              "postDate": "2025-01-10T17:25:05.313Z",
              "content": "<p>Yeah, one of the key learnings for me in this competition is knowing dropping normalization layers in some cases is so impactful. But considering what the data looks like and how normalization layers work, this isn't really surprising.</p>",
              "rawMarkdown": "Yeah, one of the key learnings for me in this competition is knowing dropping normalization layers in some cases is so impactful. But considering what the data looks like and how normalization layers work, this isn't really surprising."
            },
            {
              "id": 3093284,
              "postDate": "2025-01-10T17:30:06.617Z",
              "content": "<p>I tried many customised nomalization code as well. Instance norm was the best I found, but it won't give me too much performance boost. I was draining my head on improving nomalization before I read this post…</p>",
              "rawMarkdown": "I tried many customised nomalization code as well. Instance norm was the best I found, but it won't give me too much performance boost. I was draining my head on improving nomalization before I read this post..."
            }
          ]
        },
        {
          "id": 3093374,
          "postDate": "2025-01-10T19:58:11.040Z",
          "content": "<p>Having batch normalization works better for me. I dont know why it doesnt work for you, maybe you engineered a feature that is very sensitive to exact values.</p>",
          "rawMarkdown": "Having batch normalization works better for me. I dont know why it doesnt work for you, maybe you engineered a feature that is very sensitive to exact values.",
          "votes": 1,
          "replies": [
            {
              "id": 3093379,
              "postDate": "2025-01-10T20:05:21.907Z",
              "content": "<p>Interesting, but I didn't engineer any new features, just use the original 79 features. Probably it's related to model architecture? Since no matter where I added BN, the result will get much worse. Compared to BN, LN is actually OK, but much worse than no normalisation at all in online learning. </p>",
              "rawMarkdown": "Interesting, but I didn't engineer any new features, just use the original 79 features. Probably it's related to model architecture? Since no matter where I added BN, the result will get much worse. Compared to BN, LN is actually OK, but much worse than no normalisation at all in online learning. ",
              "votes": 2
            },
            {
              "id": 3093387,
              "postDate": "2025-01-10T20:25:50.070Z",
              "content": "<p>My validation score gets significantly worse if I switch to removing BNs and pre-normalizing the data. It depends on the architecture then (I guess)</p>",
              "rawMarkdown": "My validation score gets significantly worse if I switch to removing BNs and pre-normalizing the data. It depends on the architecture then (I guess)",
              "votes": 3
            },
            {
              "id": 3093395,
              "postDate": "2025-01-10T20:54:27.843Z",
              "content": "<p>I have the same experience as Ahmet, batch normalization is what works best for me. I am surprised that there are such different results depending on the architecture</p>",
              "rawMarkdown": "I have the same experience as Ahmet, batch normalization is what works best for me. I am surprised that there are such different results depending on the architecture",
              "votes": 1
            },
            {
              "id": 3093398,
              "postDate": "2025-01-10T21:04:10.347Z",
              "content": "<p>Yeah, this makes me rethink this problem… I will spend some time in the last 3 days to find this out (hopefully I can pass current bottleneck). </p>",
              "rawMarkdown": "Yeah, this makes me rethink this problem... I will spend some time in the last 3 days to find this out (hopefully I can pass current bottleneck). ",
              "votes": 2
            },
            {
              "id": 3093400,
              "postDate": "2025-01-10T21:10:40.447Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3093472,
              "postDate": "2025-01-11T01:07:34.737Z",
              "content": "<p>Was the validation score with BNs with or without online learning? I think the issue with batch norm occurs when there's online learning</p>",
              "rawMarkdown": "Was the validation score with BNs with or without online learning? I think the issue with batch norm occurs when there's online learning",
              "votes": 1
            },
            {
              "id": 3093484,
              "postDate": "2025-01-11T01:58:08.153Z",
              "content": "<p>What was your batch size? For batch size of 3 days and default parameters adding batch norm makes my validation score significantly worse.</p>",
              "rawMarkdown": "What was your batch size? For batch size of 3 days and default parameters adding batch norm makes my validation score significantly worse."
            }
          ]
        }
      ]
    },
    {
      "id": 3092830,
      "postDate": "2025-01-10T05:42:11.860Z",
      "content": "<p>I tried online learning, tough not improve much it could improve LB 0.0001, but after my model with batchnorm1d layer, online learning will fail get local result -0.78.. like.  Very confused about this.</p>",
      "rawMarkdown": "I tried online learning, tough not improve much it could improve LB 0.0001, but after my model with batchnorm1d layer, online learning will fail get local result -0.78.. like.  Very confused about this.",
      "votes": 6
    },
    {
      "id": 3093205,
      "postDate": "2025-01-10T15:55:48.380Z",
      "content": "<p>You can freeze the batchnorm layer. If you are using Pytorch, the default batchnorm momentum is 0.1, which is way too high.</p>",
      "rawMarkdown": "You can freeze the batchnorm layer. If you are using Pytorch, the default batchnorm momentum is 0.1, which is way too high.",
      "votes": 4,
      "replies": [
        {
          "id": 3093288,
          "postDate": "2025-01-10T17:35:39.267Z",
          "content": "<p>Thanks for your advice! I will give a try using tf default 0.01.</p>",
          "rawMarkdown": "Thanks for your advice! I will give a try using tf default 0.01."
        }
      ]
    },
    {
      "id": 3093186,
      "postDate": "2025-01-10T15:06:19.140Z",
      "content": "<p>For me, I just use mean-std normalization in data preprocessing instead of using normalization layer in the model.</p>",
      "rawMarkdown": "For me, I just use mean-std normalization in data preprocessing instead of using normalization layer in the model.",
      "votes": 4
    },
    {
      "id": 3093562,
      "postDate": "2025-01-11T05:39:21.060Z",
      "content": "<p>I only stopped running a RobustScaler on the features last week and it was the difference between a leaderboard score of 0.005 and 0.007. Up until that point I had found any normalisation layers in a NN had really negative impact (-R2).</p>\n<p>When i stopped scaling features before passing into NN I could add some minimal layer norm throughout the place without too much impact but I will try removing even more</p>",
      "rawMarkdown": "I only stopped running a RobustScaler on the features last week and it was the difference between a leaderboard score of 0.005 and 0.007. Up until that point I had found any normalisation layers in a NN had really negative impact (-R2).\n\nWhen i stopped scaling features before passing into NN I could add some minimal layer norm throughout the place without too much impact but I will try removing even more",
      "votes": 1,
      "replies": [
        {
          "id": 3094179,
          "postDate": "2025-01-11T17:55:46.500Z",
          "content": "<p>Interesting. My model has been hard stuck at 0.0055 with online learning, regardless of hyperparameters and I had a feeling batchnorm was the culprit. Hopefully I see similar results as you from removing it.</p>",
          "rawMarkdown": "Interesting. My model has been hard stuck at 0.0055 with online learning, regardless of hyperparameters and I had a feeling batchnorm was the culprit. Hopefully I see similar results as you from removing it."
        }
      ]
    },
    {
      "id": 3092901,
      "postDate": "2025-01-10T07:53:42.930Z",
      "content": "<p>Batch norm didn't work well on my regular training either. Layer norm is more stable.</p>",
      "rawMarkdown": "Batch norm didn't work well on my regular training either. Layer norm is more stable.",
      "votes": 1,
      "replies": [
        {
          "id": 3093091,
          "postDate": "2025-01-10T13:25:48.157Z",
          "content": "<p>Thanks! I will give layer norm a try, for my problem seem due to not shuffle the batch, but after fixing still some specific day data will make batchnorm1d work so bad, have not investigate the reason.</p>",
          "rawMarkdown": "Thanks! I will give layer norm a try, for my problem seem due to not shuffle the batch, but after fixing still some specific day data will make batchnorm1d work so bad, have not investigate the reason."
        },
        {
          "id": 3093149,
          "postDate": "2025-01-10T14:18:18.720Z",
          "content": "<p>For me, removing normalization layer also works well if I use StandardScaler when preprocessing </p>",
          "rawMarkdown": "For me, removing normalization layer also works well if I use StandardScaler when preprocessing ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3093156,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2025-01-10T14:25:35.073000",
      "content": "<p>From my experiments, all normalisation methods in models fail badly including batch normalisation, layer normalisation, instance normalisation etc. (some of them are ok, but fail badly in online learning). Guess it's because of the super non-stationary property of this dataset and how we train the model. So models without any normalisation in models are the best. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 3093252,
          "author_name": "herryxie",
          "author_url": "",
          "post_date": "2025-01-10T17:03:24.520000",
          "content": "<p>Thanks for sharing this insight! Do you mean without any normalisation in the entire model? Not just the normalising the input?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3093272,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-10T17:15:06.393000",
              "content": "<p>Yes, using normalization in my option is one of the several most common traps resulting in so my teams couldn't build high performance nn models in this competition.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3093277,
              "author_name": "herryxie",
              "author_url": "",
              "post_date": "2025-01-10T17:17:47.023000",
              "content": "<p>That's very generous advice, I'll give this a go</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3093278,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2025-01-10T17:17:52.890000",
              "content": "<p>I have spent so much time experimenting how to properly add norm modules to the model…tried all kinds of default norm methods and also designed some in-house norm modules, nothing had worked :(</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3093282,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-10T17:25:05.313000",
              "content": "<p>Yeah, one of the key learnings for me in this competition is knowing dropping normalization layers in some cases is so impactful. But considering what the data looks like and how normalization layers work, this isn't really surprising.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3093284,
              "author_name": "herryxie",
              "author_url": "",
              "post_date": "2025-01-10T17:30:06.617000",
              "content": "<p>I tried many customised nomalization code as well. Instance norm was the best I found, but it won't give me too much performance boost. I was draining my head on improving nomalization before I read this post…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3093374,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2025-01-10T19:58:11.040000",
          "content": "<p>Having batch normalization works better for me. I dont know why it doesnt work for you, maybe you engineered a feature that is very sensitive to exact values.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3093379,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-10T20:05:21.907000",
              "content": "<p>Interesting, but I didn't engineer any new features, just use the original 79 features. Probably it's related to model architecture? Since no matter where I added BN, the result will get much worse. Compared to BN, LN is actually OK, but much worse than no normalisation at all in online learning. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3093387,
              "author_name": "Ahmet Erdem",
              "author_url": "",
              "post_date": "2025-01-10T20:25:50.070000",
              "content": "<p>My validation score gets significantly worse if I switch to removing BNs and pre-normalizing the data. It depends on the architecture then (I guess)</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3093395,
              "author_name": "Fnoa",
              "author_url": "",
              "post_date": "2025-01-10T20:54:27.843000",
              "content": "<p>I have the same experience as Ahmet, batch normalization is what works best for me. I am surprised that there are such different results depending on the architecture</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3093398,
              "author_name": "HAO",
              "author_url": "",
              "post_date": "2025-01-10T21:04:10.347000",
              "content": "<p>Yeah, this makes me rethink this problem… I will spend some time in the last 3 days to find this out (hopefully I can pass current bottleneck). </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3093400,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-01-10T21:10:40.447000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3093472,
              "author_name": "rulan",
              "author_url": "",
              "post_date": "2025-01-11T01:07:34.737000",
              "content": "<p>Was the validation score with BNs with or without online learning? I think the issue with batch norm occurs when there's online learning</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3093484,
              "author_name": "Lu Bin Liu",
              "author_url": "",
              "post_date": "2025-01-11T01:58:08.153000",
              "content": "<p>What was your batch size? For batch size of 3 days and default parameters adding batch norm makes my validation score significantly worse.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3093205,
      "author_name": "xz",
      "author_url": "",
      "post_date": "2025-01-10T15:55:48.380000",
      "content": "<p>You can freeze the batchnorm layer. If you are using Pytorch, the default batchnorm momentum is 0.1, which is way too high.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3093288,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2025-01-10T17:35:39.267000",
          "content": "<p>Thanks for your advice! I will give a try using tf default 0.01.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3093186,
      "author_name": "I2nfinit3y",
      "author_url": "",
      "post_date": "2025-01-10T15:06:19.140000",
      "content": "<p>For me, I just use mean-std normalization in data preprocessing instead of using normalization layer in the model.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3093562,
      "author_name": "Michael Timbs",
      "author_url": "",
      "post_date": "2025-01-11T05:39:21.060000",
      "content": "<p>I only stopped running a RobustScaler on the features last week and it was the difference between a leaderboard score of 0.005 and 0.007. Up until that point I had found any normalisation layers in a NN had really negative impact (-R2).</p>\n<p>When i stopped scaling features before passing into NN I could add some minimal layer norm throughout the place without too much impact but I will try removing even more</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3094179,
          "author_name": "John Caresio",
          "author_url": "",
          "post_date": "2025-01-11T17:55:46.500000",
          "content": "<p>Interesting. My model has been hard stuck at 0.0055 with online learning, regardless of hyperparameters and I had a feeling batchnorm was the culprit. Hopefully I see similar results as you from removing it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3092901,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2025-01-10T07:53:42.930000",
      "content": "<p>Batch norm didn't work well on my regular training either. Layer norm is more stable.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3093091,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2025-01-10T13:25:48.157000",
          "content": "<p>Thanks! I will give layer norm a try, for my problem seem due to not shuffle the batch, but after fixing still some specific day data will make batchnorm1d work so bad, have not investigate the reason.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3093149,
          "author_name": "zhow",
          "author_url": "",
          "post_date": "2025-01-10T14:18:18.720000",
          "content": "<p>For me, removing normalization layer also works well if I use StandardScaler when preprocessing </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3093156": "From my experiments, all normalisation methods in models fail badly including batch normalisation, layer normalisation, instance normalisation etc. (some of them are ok, but fail badly in online learning). Guess it's because of the super non-stationary property of this dataset and how we train the model. So models without any normalisation in models are the best. ",
    "3092830": "I tried online learning, tough not improve much it could improve LB 0.0001, but after my model with batchnorm1d layer, online learning will fail get local result -0.78.. like.  Very confused about this.",
    "3093205": "You can freeze the batchnorm layer. If you are using Pytorch, the default batchnorm momentum is 0.1, which is way too high.",
    "3093186": "For me, I just use mean-std normalization in data preprocessing instead of using normalization layer in the model.",
    "3093562": "I only stopped running a RobustScaler on the features last week and it was the difference between a leaderboard score of 0.005 and 0.007. Up until that point I had found any normalisation layers in a NN had really negative impact (-R2).\n\nWhen i stopped scaling features before passing into NN I could add some minimal layer norm throughout the place without too much impact but I will try removing even more",
    "3092901": "Batch norm didn't work well on my regular training either. Layer norm is more stable."
  }
}