{
  "id": 548136,
  "title": "Normalization",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548136",
  "author_name": "Victor Shlepov",
  "post_date": "2024-11-25T10:54:30.709000",
  "votes": 24,
  "comment_count": 36,
  "views": 0,
  "content": "<p>Ok, let’s pick normalization as this week’s topic. After all, it’s one of the first steps in any AI/ML pipeline… Again, whatever I write here is just a collection of thoughts and hypotheses—not necessarily correct ones. There are still a couple of folks ahead of me by a solid margin on the LB, and the competition hasn’t even reached the halfway point. With that said, let’s get down to business.</p>\n<ol>\n<li>We’re all accustomed to applying some kind of normalization almost by default. This time, I’d suggest pausing for a moment—stop printing “BatchNormalization” or “LayerNormalization” (yes, I’m a TensorFlow addict)—and instead think about what we’re normalizing and why.</li>\n<li>The data isn’t “square.” Let’s take a single day of features with a shape of [steps, symbols, features]. I suppose most of you use some form of uniform inputs (except, perhaps, for the steps), so let’s say [None, 30, 79]. For any given day, you’ll have missing data points, either in the symbol dimension (non-traded symbols) or the feature dimension—or, quite likely, both. Whatever imputation strategy you choose (zero, forward fill, mean) - it will affect the normalization process in potentially strange ways.</li>\n<li>Global mean and variance. Well, normalizing features using global stats is nothing but leakage—the textbook classic. OK, that’s not to say I haven’t tried it; who am I to follow all the textbook advice, right? :) Still, it doesn’t yield any improvements either. A fundamentally flawed approach with no tangible results doesn’t sound like a promising combo, I guess…</li>\n<li>There might not be a true “norm” in financial markets—it’s a moving target, and the “norm” itself is constantly evolving. You might consider skipping normalization altogether or applying some kind of log normalization purely for numerical stability. I’ve tried dozens of approaches—even a trainable momentum. None of them have been good enough so far.</li>\n<li>It might be worth checking the dates when a massive number of new symbols or features appear and seeing how the model behaves around those days. That should provide some food for thought.</li>\n</ol>\n<p>As usual, your thoughts are highly welcome.</p>\n<p>[UPDATES]</p>\n<ol>\n<li>We have all sorts of feature distributions here - gaussian-like, log-like, whatsoever-like. And even 3 integer-encoded features. Further away - we know little to nothing about the features, it's not ever guaranteed that every feature has uniform mean, variance or whatever distribution parameters independent from the symbol. Maybe we deal with something in between 79 and 39x79 distributions, which makes almost any of the default normalization approaches a highly questionable so to say…</li>\n</ol>",
  "messages": [
    {
      "id": 3054939,
      "postDate": "2024-11-25T10:54:30.710Z",
      "content": "<p>Ok, let’s pick normalization as this week’s topic. After all, it’s one of the first steps in any AI/ML pipeline… Again, whatever I write here is just a collection of thoughts and hypotheses—not necessarily correct ones. There are still a couple of folks ahead of me by a solid margin on the LB, and the competition hasn’t even reached the halfway point. With that said, let’s get down to business.</p>\n<ol>\n<li>We’re all accustomed to applying some kind of normalization almost by default. This time, I’d suggest pausing for a moment—stop printing “BatchNormalization” or “LayerNormalization” (yes, I’m a TensorFlow addict)—and instead think about what we’re normalizing and why.</li>\n<li>The data isn’t “square.” Let’s take a single day of features with a shape of [steps, symbols, features]. I suppose most of you use some form of uniform inputs (except, perhaps, for the steps), so let’s say [None, 30, 79]. For any given day, you’ll have missing data points, either in the symbol dimension (non-traded symbols) or the feature dimension—or, quite likely, both. Whatever imputation strategy you choose (zero, forward fill, mean) - it will affect the normalization process in potentially strange ways.</li>\n<li>Global mean and variance. Well, normalizing features using global stats is nothing but leakage—the textbook classic. OK, that’s not to say I haven’t tried it; who am I to follow all the textbook advice, right? :) Still, it doesn’t yield any improvements either. A fundamentally flawed approach with no tangible results doesn’t sound like a promising combo, I guess…</li>\n<li>There might not be a true “norm” in financial markets—it’s a moving target, and the “norm” itself is constantly evolving. You might consider skipping normalization altogether or applying some kind of log normalization purely for numerical stability. I’ve tried dozens of approaches—even a trainable momentum. None of them have been good enough so far.</li>\n<li>It might be worth checking the dates when a massive number of new symbols or features appear and seeing how the model behaves around those days. That should provide some food for thought.</li>\n</ol>\n<p>As usual, your thoughts are highly welcome.</p>\n<p>[UPDATES]</p>\n<ol>\n<li>We have all sorts of feature distributions here - gaussian-like, log-like, whatsoever-like. And even 3 integer-encoded features. Further away - we know little to nothing about the features, it's not ever guaranteed that every feature has uniform mean, variance or whatever distribution parameters independent from the symbol. Maybe we deal with something in between 79 and 39x79 distributions, which makes almost any of the default normalization approaches a highly questionable so to say…</li>\n</ol>",
      "rawMarkdown": "Ok, let’s pick normalization as this week’s topic. After all, it’s one of the first steps in any AI/ML pipeline… Again, whatever I write here is just a collection of thoughts and hypotheses—not necessarily correct ones. There are still a couple of folks ahead of me by a solid margin on the LB, and the competition hasn’t even reached the halfway point. With that said, let’s get down to business.\n1. We’re all accustomed to applying some kind of normalization almost by default. This time, I’d suggest pausing for a moment—stop printing “BatchNormalization” or “LayerNormalization” (yes, I’m a TensorFlow addict)—and instead think about what we’re normalizing and why.\n2. The data isn’t “square.” Let’s take a single day of features with a shape of [steps, symbols, features]. I suppose most of you use some form of uniform inputs (except, perhaps, for the steps), so let’s say [None, 30, 79]. For any given day, you’ll have missing data points, either in the symbol dimension (non-traded symbols) or the feature dimension—or, quite likely, both. Whatever imputation strategy you choose (zero, forward fill, mean) - it will affect the normalization process in potentially strange ways.\n3. Global mean and variance. Well, normalizing features using global stats is nothing but leakage—the textbook classic. OK, that’s not to say I haven’t tried it; who am I to follow all the textbook advice, right? :) Still, it doesn’t yield any improvements either. A fundamentally flawed approach with no tangible results doesn’t sound like a promising combo, I guess…\n4. There might not be a true “norm” in financial markets—it’s a moving target, and the “norm” itself is constantly evolving. You might consider skipping normalization altogether or applying some kind of log normalization purely for numerical stability. I’ve tried dozens of approaches—even a trainable momentum. None of them have been good enough so far.\n5. It might be worth checking the dates when a massive number of new symbols or features appear and seeing how the model behaves around those days. That should provide some food for thought.\n\nAs usual, your thoughts are highly welcome.\n\n[UPDATES]\n\n6. We have all sorts of feature distributions here - gaussian-like, log-like, whatsoever-like. And even 3 integer-encoded features. Further away - we know little to nothing about the features, it's not ever guaranteed that every feature has uniform mean, variance or whatever distribution parameters independent from the symbol. Maybe we deal with something in between 79 and 39x79 distributions, which makes almost any of the default normalization approaches a highly questionable so to say...",
      "votes": 24
    },
    {
      "id": 3056935,
      "postDate": "2024-11-27T14:35:17.903Z",
      "content": "<p>Love the discussion. My thoughts on normalization are to look at why we perform it in the first place. There are two reasons that come to mind immediately. The first is due to the mechanics of how NNs work, especially with initialization of parameters. The weights and bias values are drawn from a distribution, the same distribution for each feature. But if we don’t normalize and a feature has a very different mean and/or std, and its weight/bias comes from the same distribution as that used for weight/bias for all other features (as is the case with any commonly used initialization scheme), then the impact of our feature within the network is biased. Theoretically, NNs can overcome this obstacle, but may get stuck in local optima along the way and would certainly have longer train times without normalization. So, normalization across features puts all features on a level playing field at the start. </p>\n<p>The second main reason for normalization, not across features as above, but within a feature, would be “align” the relationship between the feature and other features across samples/time; to deal with non-stationarity. Ultimately, any model (even a NN) is capturing the Mutual Information in the data. But MI is assuming stationary relationships. So, this second purpose for normalization is to make features more stationary across samples/time. </p>\n<p>The first reason would indicate that a global normalization should be enough. I’m not as worried about the data leakage here and there are methods to alleviate the leakage — calculate the feature mean and std in training only and apply the same values to your validation data — that will alleviate the leakage. </p>\n<p>The second reason for normalization, making the relationships more stable, is the black art. I like the “rolling window” approach on surface but worry that it is essentially filtering out lower frequency variations in the data; but perhaps we want that. I guess what I’m saying is that any rolling normalization is essentially creating a new feature that has more focus on the higher frequencies within that feature.  </p>\n<p>Another approach to making the data more stationary is to normalize using recent realized volatility. Volatility in the markets is known to cluster and thus once it spikes, it takes some time to return back to “normal” levels. Clearly, volatility spikes and clustering have an effect on stationarity and thus how the Mutual Information is captured. I’ve tried a few things here but nothing that’s stuck.  </p>\n<p>The above is my mental model around normalization. I share it in the hopes of making a contribution to this conversation.  Please challenge and augment the above as that is how we all learn. </p>\n<p>EDIT: I mentioned “global” normalization as being enough to satisfy the first reason. I want to clarify that this is global on a per feature basis, so column by column. It’s “global” in my head since it’s not time varying and not symbol specific and it’s not “batch” or “layer” normalization. IMHO, any batch or layer normalization is absolutely the wrong approach here as those were created to solve a different problem. </p>",
      "rawMarkdown": "Love the discussion. My thoughts on normalization are to look at why we perform it in the first place. There are two reasons that come to mind immediately. The first is due to the mechanics of how NNs work, especially with initialization of parameters. The weights and bias values are drawn from a distribution, the same distribution for each feature. But if we don’t normalize and a feature has a very different mean and/or std, and its weight/bias comes from the same distribution as that used for weight/bias for all other features (as is the case with any commonly used initialization scheme), then the impact of our feature within the network is biased. Theoretically, NNs can overcome this obstacle, but may get stuck in local optima along the way and would certainly have longer train times without normalization. So, normalization across features puts all features on a level playing field at the start. \n\nThe second main reason for normalization, not across features as above, but within a feature, would be “align” the relationship between the feature and other features across samples/time; to deal with non-stationarity. Ultimately, any model (even a NN) is capturing the Mutual Information in the data. But MI is assuming stationary relationships. So, this second purpose for normalization is to make features more stationary across samples/time. \n\nThe first reason would indicate that a global normalization should be enough. I’m not as worried about the data leakage here and there are methods to alleviate the leakage — calculate the feature mean and std in training only and apply the same values to your validation data — that will alleviate the leakage. \n\nThe second reason for normalization, making the relationships more stable, is the black art. I like the “rolling window” approach on surface but worry that it is essentially filtering out lower frequency variations in the data; but perhaps we want that. I guess what I’m saying is that any rolling normalization is essentially creating a new feature that has more focus on the higher frequencies within that feature.  \n\nAnother approach to making the data more stationary is to normalize using recent realized volatility. Volatility in the markets is known to cluster and thus once it spikes, it takes some time to return back to “normal” levels. Clearly, volatility spikes and clustering have an effect on stationarity and thus how the Mutual Information is captured. I’ve tried a few things here but nothing that’s stuck.  \n\nThe above is my mental model around normalization. I share it in the hopes of making a contribution to this conversation.  Please challenge and augment the above as that is how we all learn. \n\nEDIT: I mentioned “global” normalization as being enough to satisfy the first reason. I want to clarify that this is global on a per feature basis, so column by column. It’s “global” in my head since it’s not time varying and not symbol specific and it’s not “batch” or “layer” normalization. IMHO, any batch or layer normalization is absolutely the wrong approach here as those were created to solve a different problem. ",
      "votes": 14,
      "replies": [
        {
          "id": 3057069,
          "postDate": "2024-11-27T17:16:41.207Z",
          "content": "<p>Well, that's a reasonable approach, I personally don't see much to argue with :)</p>\n<blockquote>\n  <p>Another approach to making the data more stationary is to normalize using recent realized volatility.</p>\n</blockquote>\n<p>This might be worth trying out. I have not experimented with it yet - I don't like much the idea of a fixed moving windows, or any other unnecessary limitations to the model architecture, but it feels like there should be some interesting shortcuts…</p>",
          "rawMarkdown": "Well, that's a reasonable approach, I personally don't see much to argue with :)\n\n>Another approach to making the data more stationary is to normalize using recent realized volatility.\n\nThis might be worth trying out. I have not experimented with it yet - I don't like much the idea of a fixed moving windows, or any other unnecessary limitations to the model architecture, but it feels like there should be some interesting shortcuts..."
        },
        {
          "id": 3057722,
          "postDate": "2024-11-28T14:48:29.220Z",
          "content": "<blockquote>\n  <p>Another approach to making the data more stationary is to normalize using recent realized volatility. Volatility in the markets is known to cluster and thus once it spikes, it takes some time to return back to “normal” levels. Clearly, volatility spikes and clustering have an effect on stationarity and thus how the Mutual Information is captured. I’ve tried a few things here but nothing that’s stuck.</p>\n</blockquote>\n<p>How to do this?</p>",
          "rawMarkdown": ">Another approach to making the data more stationary is to normalize using recent realized volatility. Volatility in the markets is known to cluster and thus once it spikes, it takes some time to return back to “normal” levels. Clearly, volatility spikes and clustering have an effect on stationarity and thus how the Mutual Information is captured. I’ve tried a few things here but nothing that’s stuck.\n\nHow to do this?",
          "replies": [
            {
              "id": 3057813,
              "postDate": "2024-11-28T16:46:29.943Z",
              "content": "<p>Well, this is a bit challenging given that we don’t know what the features are. I have not been successfully at this with this data yet. </p>",
              "rawMarkdown": "Well, this is a bit challenging given that we don’t know what the features are. I have not been successfully at this with this data yet. "
            }
          ]
        }
      ]
    },
    {
      "id": 3057176,
      "postDate": "2024-11-27T20:01:21.380Z",
      "content": "<p>For NN models with sequence input, InstanceNorm is more intuitive for me. As said, each feature has its own distribution. Using either batch norm or layer norm will likely mix the distribution of different features. </p>",
      "rawMarkdown": "For NN models with sequence input, InstanceNorm is more intuitive for me. As said, each feature has its own distribution. Using either batch norm or layer norm will likely mix the distribution of different features. ",
      "votes": 3
    },
    {
      "id": 3057703,
      "postDate": "2024-11-28T14:10:48.640Z",
      "content": "<p>A lot of trading data is domain-knowledge scaled. E.g. a price extension in terms of the opening price range (9:30am to 10:30am UTC+5). I am assuming that is the case here and not focusing on scaling except for training stability. </p>",
      "rawMarkdown": "A lot of trading data is domain-knowledge scaled. E.g. a price extension in terms of the opening price range (9:30am to 10:30am UTC+5). I am assuming that is the case here and not focusing on scaling except for training stability. ",
      "votes": 1,
      "replies": [
        {
          "id": 3057750,
          "postDate": "2024-11-28T15:16:39.717Z",
          "content": "<blockquote>\n  <p>I am assuming that is the case here and not focusing on scaling except for training stability</p>\n</blockquote>\n<p>Well, it's all about training stability, but when you have a dozed of different feature distributions, random shifts over the time and ragged data - it's kind of a challenging task, isn't it? :)</p>",
          "rawMarkdown": ">I am assuming that is the case here and not focusing on scaling except for training stability\n\nWell, it's all about training stability, but when you have a dozed of different feature distributions, random shifts over the time and ragged data - it's kind of a challenging task, isn't it? :)",
          "replies": [
            {
              "id": 3057778,
              "postDate": "2024-11-28T15:49:34.337Z",
              "content": "<p>It is indeed, sounds like a prime job for an auto encoder ;)</p>",
              "rawMarkdown": "It is indeed, sounds like a prime job for an auto encoder ;)"
            }
          ]
        }
      ]
    },
    {
      "id": 3056275,
      "postDate": "2024-11-26T18:56:41.657Z",
      "content": "<p>These insights are wonderful!</p>",
      "rawMarkdown": "These insights are wonderful!",
      "votes": 1
    },
    {
      "id": 3056047,
      "postDate": "2024-11-26T13:50:01.707Z",
      "content": "<p>Also if we look at some features across time we can say for certain new value ranges are constantly popping up, and we can say for almost certain that also happens with test data. A feature that has been ranging from 0 and 1 up until day 1000, can go to, say, 3 on day 1001. No normalization technique will account for that. Based on pure experimentation, what has worked best for me so far is inputting 0 for missings, and standard scaling. I'm very sure, however, this is not the best approach.</p>",
      "rawMarkdown": "Also if we look at some features across time we can say for certain new value ranges are constantly popping up, and we can say for almost certain that also happens with test data. A feature that has been ranging from 0 and 1 up until day 1000, can go to, say, 3 on day 1001. No normalization technique will account for that. Based on pure experimentation, what has worked best for me so far is inputting 0 for missings, and standard scaling. I'm very sure, however, this is not the best approach.",
      "votes": 1,
      "replies": [
        {
          "id": 3056066,
          "postDate": "2024-11-26T14:22:47Z",
          "content": "<p>Yep. The major shifts is distributions is likely the reason my model feels like crazy starting from date_id 1000 or something…</p>\n<blockquote>\n  <p>No normalization technique will account for that</p>\n</blockquote>\n<p>Some technique would, I guess. It's just none of us discovered it yet… :)</p>\n<blockquote>\n  <p>Based on pure experimentation, what has worked best for me so far is inputting 0 for missings</p>\n</blockquote>\n<p>Worth to test zero- vs forward-fill imputation. Zero should likely distort true distributions heavily.</p>",
          "rawMarkdown": "Yep. The major shifts is distributions is likely the reason my model feels like crazy starting from date_id 1000 or something...\n\n>No normalization technique will account for that\n\nSome technique would, I guess. It's just none of us discovered it yet... :)\n\n>Based on pure experimentation, what has worked best for me so far is inputting 0 for missings\n\nWorth to test zero- vs forward-fill imputation. Zero should likely distort true distributions heavily.",
          "votes": 3,
          "replies": [
            {
              "id": 3056416,
              "postDate": "2024-11-27T00:35:34.683Z",
              "content": "<blockquote>\n  <p>Worth to test zero- vs forward-fill imputation. Zero should likely distort true distributions heavily.</p>\n</blockquote>\n<p>I agree. But again, zero worked best for me so far (I tested forward-fill)</p>",
              "rawMarkdown": "> Worth to test zero- vs forward-fill imputation. Zero should likely distort true distributions heavily.\n\nI agree. But again, zero worked best for me so far (I tested forward-fill)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3055156,
      "postDate": "2024-11-25T15:38:11.043Z",
      "content": "<p>Sliding window normalisation worth a try</p>",
      "rawMarkdown": "Sliding window normalisation worth a try",
      "votes": 1,
      "replies": [
        {
          "id": 3055168,
          "postDate": "2024-11-25T15:52:17.693Z",
          "content": "<p>I'd say it's almost the same as normalization with a trainable momentum, which is effectively a flexible learnable window… No major effect so far…</p>",
          "rawMarkdown": "I'd say it's almost the same as normalization with a trainable momentum, which is effectively a flexible learnable window... No major effect so far..."
        }
      ]
    },
    {
      "id": 3056249,
      "postDate": "2024-11-26T18:13:51.627Z",
      "content": "<p>Great thoughts (and the same to your another two posts, pure gold)! <br>\nI just switched from GBDTs to NN models, so haven't experimented too much yet. But from what I see, the proper normalisation method(maybe also some other preprocessing methods) could be the key to win in this competition. Since using global normalisation methods, no matter what models we choose (I have tried GRU model, 1d casaul cnn and transformer models, the local CV and LB are quite stable - for me now the scores are around 0.0060). Using moving window mean/std is slightly better than global normalisation (we know it's a bad choice).<br>\nNeed to dive deeper into the data/features to figure out a better way, just like what you do. Cheers!</p>",
      "rawMarkdown": "Great thoughts (and the same to your another two posts, pure gold)! \nI just switched from GBDTs to NN models, so haven't experimented too much yet. But from what I see, the proper normalisation method(maybe also some other preprocessing methods) could be the key to win in this competition. Since using global normalisation methods, no matter what models we choose (I have tried GRU model, 1d casaul cnn and transformer models, the local CV and LB are quite stable - for me now the scores are around 0.0060). Using moving window mean/std is slightly better than global normalisation (we know it's a bad choice).\nNeed to dive deeper into the data/features to figure out a better way, just like what you do. Cheers!",
      "votes": 2,
      "replies": [
        {
          "id": 3056278,
          "postDate": "2024-11-26T18:58:00.843Z",
          "content": "<p>The trouble with a moving window is that you effectively frame the focus of a model with some given period. Markets don't work this way. The window should not be a constant. It's not just about a normalization, it's about the whole approach. I'm still trying to come up with some structure which would have all the necessary degrees of freedom - no luck yet… :)</p>",
          "rawMarkdown": "The trouble with a moving window is that you effectively frame the focus of a model with some given period. Markets don't work this way. The window should not be a constant. It's not just about a normalization, it's about the whole approach. I'm still trying to come up with some structure which would have all the necessary degrees of freedom - no luck yet... :)"
        },
        {
          "id": 3056452,
          "postDate": "2024-11-27T01:30:17.283Z",
          "content": "<p>how could you use moving window mean/std in hidden test dataset?</p>",
          "rawMarkdown": "how could you use moving window mean/std in hidden test dataset?",
          "replies": [
            {
              "id": 3057400,
              "postDate": "2024-11-28T05:38:49.150Z",
              "content": "<p>Using cache or momentum, which is almost the same…</p>",
              "rawMarkdown": "Using cache or momentum, which is almost the same..."
            },
            {
              "id": 3061475,
              "postDate": "2024-12-02T16:55:18.107Z",
              "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> What kind of NN models have you found useful for this competition?</p>",
              "rawMarkdown": "@victorshlepov What kind of NN models have you found useful for this competition?"
            }
          ]
        }
      ]
    },
    {
      "id": 3055768,
      "postDate": "2024-11-26T05:41:59Z",
      "content": "<p>I feel responder is normalized within [-5,5], and features aren't having much outliers (if you look at distribution of each). I feel the purpose of this competition is more on predict methology rather than feature engineering.</p>\n<p>btw, looking to team up with someone struggling at 0.00043 as well, kaggle didn't allow me to contact if not kaggle contributer.</p>",
      "rawMarkdown": "I feel responder is normalized within [-5,5], and features aren't having much outliers (if you look at distribution of each). I feel the purpose of this competition is more on predict methology rather than feature engineering.\n\nbtw, looking to team up with someone struggling at 0.00043 as well, kaggle didn't allow me to contact if not kaggle contributer."
    },
    {
      "id": 3055805,
      "postDate": "2024-11-26T06:44:40.937Z",
      "content": "<p>Great insights on normalization! Rethinking default techniques like BatchNormalization and considering dynamic, context-aware methods is crucial, especially for evolving data like financial markets.<br>\n Avoiding data leakage with rolling stats and tracking structural changes are excellent suggestions.</p>",
      "rawMarkdown": "Great insights on normalization! Rethinking default techniques like BatchNormalization and considering dynamic, context-aware methods is crucial, especially for evolving data like financial markets.\n Avoiding data leakage with rolling stats and tracking structural changes are excellent suggestions.",
      "votes": -1
    },
    {
      "id": 3055764,
      "postDate": "2024-11-26T05:32:45.817Z",
      "content": "<p>I guess many standard methods fail when comes to such data (just consoling myself when I see all my struggles result in bad PB ranks)</p>\n<p>Normalization – Though tree models might be an option to avoid this…but still not tempting enough to switch over. Avoiding normalization in nn model sounds interesting but it cannot be as simple as it that. However thanks for your thoughts…at least some more experiments to try.</p>",
      "rawMarkdown": "I guess many standard methods fail when comes to such data (just consoling myself when I see all my struggles result in bad PB ranks)\n\nNormalization – Though tree models might be an option to avoid this...but still not tempting enough to switch over. Avoiding normalization in nn model sounds interesting but it cannot be as simple as it that. However thanks for your thoughts...at least some more experiments to try."
    },
    {
      "id": 3055340,
      "postDate": "2024-11-25T17:52:10.453Z",
      "content": "<blockquote>\n  <p>Well, normalizing features using global stats is nothing but leakage—the textbook classic. </p>\n</blockquote>\n<p>This is only true if the global mean / variance is calculated prior to splitting the data</p>",
      "rawMarkdown": "> Well, normalizing features using global stats is nothing but leakage—the textbook classic. \n\nThis is only true if the global mean / variance is calculated prior to splitting the data",
      "replies": [
        {
          "id": 3055357,
          "postDate": "2024-11-25T18:00:42.817Z",
          "content": "<p>It's always true. In real market settings you have no idea about the future mean/variance. This is a temporal leakage - straight from the book. Recall COVID19 - many people expected this kind of shifts in returns and volatility back in 2017-2018?<br>\nIf you or any of your friends know for sure the means and variances of quotes, returns or anything, say for the next week - give me a call, we'll make billions. Buffet will be serving us donuts for breakfast… :)</p>",
          "rawMarkdown": "It's always true. In real market settings you have no idea about the future mean/variance. This is a temporal leakage - straight from the book. Recall COVID19 - many people expected this kind of shifts in returns and volatility back in 2017-2018?\nIf you or any of your friends know for sure the means and variances of quotes, returns or anything, say for the next week - give me a call, we'll make billions. Buffet will be serving us donuts for breakfast... :)",
          "votes": 1,
          "replies": [
            {
              "id": 3055480,
              "postDate": "2024-11-25T19:48:07.120Z",
              "content": "<p>If you do a train / val / test split first and only use the mean and variance of the train to \"standardize\" the val and test, then there is no leakage</p>",
              "rawMarkdown": "If you do a train / val / test split first and only use the mean and variance of the train to \"standardize\" the val and test, then there is no leakage",
              "votes": 2
            },
            {
              "id": 3055488,
              "postDate": "2024-11-25T20:02:12.477Z",
              "content": "<p>There is leakage. </p>\n<p>df[(df['date_id'] == 0) &amp; (df['time_id'] == 0)] is being standardized using data from df[(df['date_id'] == 0) &amp; (df['time_id'] == 1)], etc. </p>",
              "rawMarkdown": "There is leakage. \n\ndf[(df['date_id'] == 0) & (df['time_id'] == 0)] is being standardized using data from df[(df['date_id'] == 0) & (df['time_id'] == 1)], etc. ",
              "votes": 1
            },
            {
              "id": 3055554,
              "postDate": "2024-11-25T21:27:19.803Z",
              "content": "<p>I was referring to leakage from val / test to the train, not within the training set itself (I'm assuming that df[(df['date_id'] == 0) &amp; (df['time_id'] == 0)]  and df[(df['date_id'] == 0) &amp; (df['time_id'] == 1)] are in the training set). </p>",
              "rawMarkdown": "I was referring to leakage from val / test to the train, not within the training set itself (I'm assuming that df[(df['date_id'] == 0) & (df['time_id'] == 0)]  and df[(df['date_id'] == 0) & (df['time_id'] == 1)] are in the training set). "
            }
          ]
        }
      ]
    },
    {
      "id": 3055275,
      "postDate": "2024-11-25T16:59:47.630Z",
      "content": "<p>Does the .diff not work?</p>",
      "rawMarkdown": "Does the .diff not work?",
      "replies": [
        {
          "id": 3055374,
          "postDate": "2024-11-25T18:22:08.457Z",
          "content": "<p>This does not fit the not-stationarity of data… Really, it's not about the python, pandas, or anything - it's about nature of the market: you never know when the next shift happens. This one by Eugene Fama is old, but still true: <a href=\"https://extranet.parisschoolofeconomics.eu/docs/ferriere-nathalie/fama1965.pdf\" target=\"_blank\">the past cannot be used to predict the future in any meaningful way</a>.</p>",
          "rawMarkdown": "This does not fit the not-stationarity of data... Really, it's not about the python, pandas, or anything - it's about nature of the market: you never know when the next shift happens. This one by Eugene Fama is old, but still true: [the past cannot be used to predict the future in any meaningful way](https://extranet.parisschoolofeconomics.eu/docs/ferriere-nathalie/fama1965.pdf).",
          "replies": [
            {
              "id": 3055381,
              "postDate": "2024-11-25T18:27:45.393Z",
              "content": "<p>Differentiation is one of the techniques used to transform a non-stationary series into a stationary one, isn't it?</p>\n<p>Or did you mean 'This does not align with the non-stationarity of the data' in terms of the model?</p>",
              "rawMarkdown": "Differentiation is one of the techniques used to transform a non-stationary series into a stationary one, isn't it?\n\nOr did you mean 'This does not align with the non-stationarity of the data' in terms of the model?"
            },
            {
              "id": 3055389,
              "postDate": "2024-11-25T18:33:36.207Z",
              "content": "<p>My bad, I thought of pandas.DataFrame().diff()</p>",
              "rawMarkdown": "My bad, I thought of pandas.DataFrame().diff()"
            },
            {
              "id": 3055446,
              "postDate": "2024-11-25T19:06:19.703Z",
              "content": "<p>Pandas .diff is the differentiation I mentioned; it's based on the same concept. In Marcos Lopes' book (Chapter 5), he suggests using fractional differentiation, which is almost the same thing.</p>",
              "rawMarkdown": "Pandas .diff is the differentiation I mentioned; it's based on the same concept. In Marcos Lopes' book (Chapter 5), he suggests using fractional differentiation, which is almost the same thing.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 3054969,
      "postDate": "2024-11-25T11:33:42.917Z",
      "content": "<p>Maxmin scaler may be more helpful than the standard scaler.</p>",
      "rawMarkdown": "Maxmin scaler may be more helpful than the standard scaler.",
      "replies": [
        {
          "id": 3055759,
          "postDate": "2024-11-26T05:24:34.317Z",
          "content": "<p>I tried maxmin of pred in the prediction function, the result seems way worse than simply doing a clip [-5,5]</p>",
          "rawMarkdown": "I tried maxmin of pred in the prediction function, the result seems way worse than simply doing a clip [-5,5]",
          "votes": 1,
          "replies": [
            {
              "id": 3055881,
              "postDate": "2024-11-26T08:30:50.113Z",
              "content": "<p>Thank you for your sharing!</p>",
              "rawMarkdown": "Thank you for your sharing!"
            }
          ]
        }
      ]
    },
    {
      "id": 3057290,
      "postDate": "2024-11-28T01:36:02.243Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3056935,
      "author_name": "Maciej Zawadzki",
      "author_url": "",
      "post_date": "2024-11-27T14:35:17.903000",
      "content": "<p>Love the discussion. My thoughts on normalization are to look at why we perform it in the first place. There are two reasons that come to mind immediately. The first is due to the mechanics of how NNs work, especially with initialization of parameters. The weights and bias values are drawn from a distribution, the same distribution for each feature. But if we don’t normalize and a feature has a very different mean and/or std, and its weight/bias comes from the same distribution as that used for weight/bias for all other features (as is the case with any commonly used initialization scheme), then the impact of our feature within the network is biased. Theoretically, NNs can overcome this obstacle, but may get stuck in local optima along the way and would certainly have longer train times without normalization. So, normalization across features puts all features on a level playing field at the start. </p>\n<p>The second main reason for normalization, not across features as above, but within a feature, would be “align” the relationship between the feature and other features across samples/time; to deal with non-stationarity. Ultimately, any model (even a NN) is capturing the Mutual Information in the data. But MI is assuming stationary relationships. So, this second purpose for normalization is to make features more stationary across samples/time. </p>\n<p>The first reason would indicate that a global normalization should be enough. I’m not as worried about the data leakage here and there are methods to alleviate the leakage — calculate the feature mean and std in training only and apply the same values to your validation data — that will alleviate the leakage. </p>\n<p>The second reason for normalization, making the relationships more stable, is the black art. I like the “rolling window” approach on surface but worry that it is essentially filtering out lower frequency variations in the data; but perhaps we want that. I guess what I’m saying is that any rolling normalization is essentially creating a new feature that has more focus on the higher frequencies within that feature.  </p>\n<p>Another approach to making the data more stationary is to normalize using recent realized volatility. Volatility in the markets is known to cluster and thus once it spikes, it takes some time to return back to “normal” levels. Clearly, volatility spikes and clustering have an effect on stationarity and thus how the Mutual Information is captured. I’ve tried a few things here but nothing that’s stuck.  </p>\n<p>The above is my mental model around normalization. I share it in the hopes of making a contribution to this conversation.  Please challenge and augment the above as that is how we all learn. </p>\n<p>EDIT: I mentioned “global” normalization as being enough to satisfy the first reason. I want to clarify that this is global on a per feature basis, so column by column. It’s “global” in my head since it’s not time varying and not symbol specific and it’s not “batch” or “layer” normalization. IMHO, any batch or layer normalization is absolutely the wrong approach here as those were created to solve a different problem. </p>",
      "votes": 14,
      "replies": [
        {
          "id": 3057069,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-27T17:16:41.207000",
          "content": "<p>Well, that's a reasonable approach, I personally don't see much to argue with :)</p>\n<blockquote>\n  <p>Another approach to making the data more stationary is to normalize using recent realized volatility.</p>\n</blockquote>\n<p>This might be worth trying out. I have not experimented with it yet - I don't like much the idea of a fixed moving windows, or any other unnecessary limitations to the model architecture, but it feels like there should be some interesting shortcuts…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3057722,
          "author_name": "Seqaeon",
          "author_url": "",
          "post_date": "2024-11-28T14:48:29.220000",
          "content": "<blockquote>\n  <p>Another approach to making the data more stationary is to normalize using recent realized volatility. Volatility in the markets is known to cluster and thus once it spikes, it takes some time to return back to “normal” levels. Clearly, volatility spikes and clustering have an effect on stationarity and thus how the Mutual Information is captured. I’ve tried a few things here but nothing that’s stuck.</p>\n</blockquote>\n<p>How to do this?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3057813,
              "author_name": "Maciej Zawadzki",
              "author_url": "",
              "post_date": "2024-11-28T16:46:29.943000",
              "content": "<p>Well, this is a bit challenging given that we don’t know what the features are. I have not been successfully at this with this data yet. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3057176,
      "author_name": "SLi",
      "author_url": "",
      "post_date": "2024-11-27T20:01:21.380000",
      "content": "<p>For NN models with sequence input, InstanceNorm is more intuitive for me. As said, each feature has its own distribution. Using either batch norm or layer norm will likely mix the distribution of different features. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3057703,
      "author_name": "Tucker Arrants",
      "author_url": "",
      "post_date": "2024-11-28T14:10:48.640000",
      "content": "<p>A lot of trading data is domain-knowledge scaled. E.g. a price extension in terms of the opening price range (9:30am to 10:30am UTC+5). I am assuming that is the case here and not focusing on scaling except for training stability. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3057750,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-28T15:16:39.717000",
          "content": "<blockquote>\n  <p>I am assuming that is the case here and not focusing on scaling except for training stability</p>\n</blockquote>\n<p>Well, it's all about training stability, but when you have a dozed of different feature distributions, random shifts over the time and ragged data - it's kind of a challenging task, isn't it? :)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3057778,
              "author_name": "Tucker Arrants",
              "author_url": "",
              "post_date": "2024-11-28T15:49:34.337000",
              "content": "<p>It is indeed, sounds like a prime job for an auto encoder ;)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3056275,
      "author_name": "Param2007",
      "author_url": "",
      "post_date": "2024-11-26T18:56:41.657000",
      "content": "<p>These insights are wonderful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3056047,
      "author_name": "Natan Labarrère",
      "author_url": "",
      "post_date": "2024-11-26T13:50:01.707000",
      "content": "<p>Also if we look at some features across time we can say for certain new value ranges are constantly popping up, and we can say for almost certain that also happens with test data. A feature that has been ranging from 0 and 1 up until day 1000, can go to, say, 3 on day 1001. No normalization technique will account for that. Based on pure experimentation, what has worked best for me so far is inputting 0 for missings, and standard scaling. I'm very sure, however, this is not the best approach.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3056066,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-26T14:22:47",
          "content": "<p>Yep. The major shifts is distributions is likely the reason my model feels like crazy starting from date_id 1000 or something…</p>\n<blockquote>\n  <p>No normalization technique will account for that</p>\n</blockquote>\n<p>Some technique would, I guess. It's just none of us discovered it yet… :)</p>\n<blockquote>\n  <p>Based on pure experimentation, what has worked best for me so far is inputting 0 for missings</p>\n</blockquote>\n<p>Worth to test zero- vs forward-fill imputation. Zero should likely distort true distributions heavily.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3056416,
              "author_name": "Natan Labarrère",
              "author_url": "",
              "post_date": "2024-11-27T00:35:34.683000",
              "content": "<blockquote>\n  <p>Worth to test zero- vs forward-fill imputation. Zero should likely distort true distributions heavily.</p>\n</blockquote>\n<p>I agree. But again, zero worked best for me so far (I tested forward-fill)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3055156,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2024-11-25T15:38:11.043000",
      "content": "<p>Sliding window normalisation worth a try</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3055168,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-25T15:52:17.693000",
          "content": "<p>I'd say it's almost the same as normalization with a trainable momentum, which is effectively a flexible learnable window… No major effect so far…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3056249,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2024-11-26T18:13:51.627000",
      "content": "<p>Great thoughts (and the same to your another two posts, pure gold)! <br>\nI just switched from GBDTs to NN models, so haven't experimented too much yet. But from what I see, the proper normalisation method(maybe also some other preprocessing methods) could be the key to win in this competition. Since using global normalisation methods, no matter what models we choose (I have tried GRU model, 1d casaul cnn and transformer models, the local CV and LB are quite stable - for me now the scores are around 0.0060). Using moving window mean/std is slightly better than global normalisation (we know it's a bad choice).<br>\nNeed to dive deeper into the data/features to figure out a better way, just like what you do. Cheers!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3056278,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-26T18:58:00.843000",
          "content": "<p>The trouble with a moving window is that you effectively frame the focus of a model with some given period. Markets don't work this way. The window should not be a constant. It's not just about a normalization, it's about the whole approach. I'm still trying to come up with some structure which would have all the necessary degrees of freedom - no luck yet… :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3056452,
          "author_name": "Yourui Wang",
          "author_url": "",
          "post_date": "2024-11-27T01:30:17.283000",
          "content": "<p>how could you use moving window mean/std in hidden test dataset?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3057400,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-28T05:38:49.150000",
              "content": "<p>Using cache or momentum, which is almost the same…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061475,
              "author_name": "ironrro",
              "author_url": "",
              "post_date": "2024-12-02T16:55:18.107000",
              "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> What kind of NN models have you found useful for this competition?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3055768,
      "author_name": "abdonson",
      "author_url": "",
      "post_date": "2024-11-26T05:41:59",
      "content": "<p>I feel responder is normalized within [-5,5], and features aren't having much outliers (if you look at distribution of each). I feel the purpose of this competition is more on predict methology rather than feature engineering.</p>\n<p>btw, looking to team up with someone struggling at 0.00043 as well, kaggle didn't allow me to contact if not kaggle contributer.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3055805,
      "author_name": "shruti",
      "author_url": "",
      "post_date": "2024-11-26T06:44:40.937000",
      "content": "<p>Great insights on normalization! Rethinking default techniques like BatchNormalization and considering dynamic, context-aware methods is crucial, especially for evolving data like financial markets.<br>\n Avoiding data leakage with rolling stats and tracking structural changes are excellent suggestions.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3055764,
      "author_name": "Viji",
      "author_url": "",
      "post_date": "2024-11-26T05:32:45.817000",
      "content": "<p>I guess many standard methods fail when comes to such data (just consoling myself when I see all my struggles result in bad PB ranks)</p>\n<p>Normalization – Though tree models might be an option to avoid this…but still not tempting enough to switch over. Avoiding normalization in nn model sounds interesting but it cannot be as simple as it that. However thanks for your thoughts…at least some more experiments to try.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3055340,
      "author_name": "Lu Bin Liu",
      "author_url": "",
      "post_date": "2024-11-25T17:52:10.453000",
      "content": "<blockquote>\n  <p>Well, normalizing features using global stats is nothing but leakage—the textbook classic. </p>\n</blockquote>\n<p>This is only true if the global mean / variance is calculated prior to splitting the data</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3055357,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-25T18:00:42.817000",
          "content": "<p>It's always true. In real market settings you have no idea about the future mean/variance. This is a temporal leakage - straight from the book. Recall COVID19 - many people expected this kind of shifts in returns and volatility back in 2017-2018?<br>\nIf you or any of your friends know for sure the means and variances of quotes, returns or anything, say for the next week - give me a call, we'll make billions. Buffet will be serving us donuts for breakfast… :)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3055480,
              "author_name": "Lu Bin Liu",
              "author_url": "",
              "post_date": "2024-11-25T19:48:07.120000",
              "content": "<p>If you do a train / val / test split first and only use the mean and variance of the train to \"standardize\" the val and test, then there is no leakage</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3055488,
              "author_name": "Jack",
              "author_url": "",
              "post_date": "2024-11-25T20:02:12.477000",
              "content": "<p>There is leakage. </p>\n<p>df[(df['date_id'] == 0) &amp; (df['time_id'] == 0)] is being standardized using data from df[(df['date_id'] == 0) &amp; (df['time_id'] == 1)], etc. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3055554,
              "author_name": "Lu Bin Liu",
              "author_url": "",
              "post_date": "2024-11-25T21:27:19.803000",
              "content": "<p>I was referring to leakage from val / test to the train, not within the training set itself (I'm assuming that df[(df['date_id'] == 0) &amp; (df['time_id'] == 0)]  and df[(df['date_id'] == 0) &amp; (df['time_id'] == 1)] are in the training set). </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3055275,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-11-25T16:59:47.630000",
      "content": "<p>Does the .diff not work?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3055374,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-25T18:22:08.457000",
          "content": "<p>This does not fit the not-stationarity of data… Really, it's not about the python, pandas, or anything - it's about nature of the market: you never know when the next shift happens. This one by Eugene Fama is old, but still true: <a href=\"https://extranet.parisschoolofeconomics.eu/docs/ferriere-nathalie/fama1965.pdf\" target=\"_blank\">the past cannot be used to predict the future in any meaningful way</a>.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3055381,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-11-25T18:27:45.393000",
              "content": "<p>Differentiation is one of the techniques used to transform a non-stationary series into a stationary one, isn't it?</p>\n<p>Or did you mean 'This does not align with the non-stationarity of the data' in terms of the model?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3055389,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-25T18:33:36.207000",
              "content": "<p>My bad, I thought of pandas.DataFrame().diff()</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3055446,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-11-25T19:06:19.703000",
              "content": "<p>Pandas .diff is the differentiation I mentioned; it's based on the same concept. In Marcos Lopes' book (Chapter 5), he suggests using fractional differentiation, which is almost the same thing.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3054969,
      "author_name": "#1BuBu",
      "author_url": "",
      "post_date": "2024-11-25T11:33:42.917000",
      "content": "<p>Maxmin scaler may be more helpful than the standard scaler.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3055759,
          "author_name": "abdonson",
          "author_url": "",
          "post_date": "2024-11-26T05:24:34.317000",
          "content": "<p>I tried maxmin of pred in the prediction function, the result seems way worse than simply doing a clip [-5,5]</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3055881,
              "author_name": "#1BuBu",
              "author_url": "",
              "post_date": "2024-11-26T08:30:50.113000",
              "content": "<p>Thank you for your sharing!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3057290,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-28T01:36:02.243000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3054939": "Ok, let’s pick normalization as this week’s topic. After all, it’s one of the first steps in any AI/ML pipeline… Again, whatever I write here is just a collection of thoughts and hypotheses—not necessarily correct ones. There are still a couple of folks ahead of me by a solid margin on the LB, and the competition hasn’t even reached the halfway point. With that said, let’s get down to business.\n1. We’re all accustomed to applying some kind of normalization almost by default. This time, I’d suggest pausing for a moment—stop printing “BatchNormalization” or “LayerNormalization” (yes, I’m a TensorFlow addict)—and instead think about what we’re normalizing and why.\n2. The data isn’t “square.” Let’s take a single day of features with a shape of [steps, symbols, features]. I suppose most of you use some form of uniform inputs (except, perhaps, for the steps), so let’s say [None, 30, 79]. For any given day, you’ll have missing data points, either in the symbol dimension (non-traded symbols) or the feature dimension—or, quite likely, both. Whatever imputation strategy you choose (zero, forward fill, mean) - it will affect the normalization process in potentially strange ways.\n3. Global mean and variance. Well, normalizing features using global stats is nothing but leakage—the textbook classic. OK, that’s not to say I haven’t tried it; who am I to follow all the textbook advice, right? :) Still, it doesn’t yield any improvements either. A fundamentally flawed approach with no tangible results doesn’t sound like a promising combo, I guess…\n4. There might not be a true “norm” in financial markets—it’s a moving target, and the “norm” itself is constantly evolving. You might consider skipping normalization altogether or applying some kind of log normalization purely for numerical stability. I’ve tried dozens of approaches—even a trainable momentum. None of them have been good enough so far.\n5. It might be worth checking the dates when a massive number of new symbols or features appear and seeing how the model behaves around those days. That should provide some food for thought.\n\nAs usual, your thoughts are highly welcome.\n\n[UPDATES]\n\n6. We have all sorts of feature distributions here - gaussian-like, log-like, whatsoever-like. And even 3 integer-encoded features. Further away - we know little to nothing about the features, it's not ever guaranteed that every feature has uniform mean, variance or whatever distribution parameters independent from the symbol. Maybe we deal with something in between 79 and 39x79 distributions, which makes almost any of the default normalization approaches a highly questionable so to say...",
    "3056935": "Love the discussion. My thoughts on normalization are to look at why we perform it in the first place. There are two reasons that come to mind immediately. The first is due to the mechanics of how NNs work, especially with initialization of parameters. The weights and bias values are drawn from a distribution, the same distribution for each feature. But if we don’t normalize and a feature has a very different mean and/or std, and its weight/bias comes from the same distribution as that used for weight/bias for all other features (as is the case with any commonly used initialization scheme), then the impact of our feature within the network is biased. Theoretically, NNs can overcome this obstacle, but may get stuck in local optima along the way and would certainly have longer train times without normalization. So, normalization across features puts all features on a level playing field at the start. \n\nThe second main reason for normalization, not across features as above, but within a feature, would be “align” the relationship between the feature and other features across samples/time; to deal with non-stationarity. Ultimately, any model (even a NN) is capturing the Mutual Information in the data. But MI is assuming stationary relationships. So, this second purpose for normalization is to make features more stationary across samples/time. \n\nThe first reason would indicate that a global normalization should be enough. I’m not as worried about the data leakage here and there are methods to alleviate the leakage — calculate the feature mean and std in training only and apply the same values to your validation data — that will alleviate the leakage. \n\nThe second reason for normalization, making the relationships more stable, is the black art. I like the “rolling window” approach on surface but worry that it is essentially filtering out lower frequency variations in the data; but perhaps we want that. I guess what I’m saying is that any rolling normalization is essentially creating a new feature that has more focus on the higher frequencies within that feature.  \n\nAnother approach to making the data more stationary is to normalize using recent realized volatility. Volatility in the markets is known to cluster and thus once it spikes, it takes some time to return back to “normal” levels. Clearly, volatility spikes and clustering have an effect on stationarity and thus how the Mutual Information is captured. I’ve tried a few things here but nothing that’s stuck.  \n\nThe above is my mental model around normalization. I share it in the hopes of making a contribution to this conversation.  Please challenge and augment the above as that is how we all learn. \n\nEDIT: I mentioned “global” normalization as being enough to satisfy the first reason. I want to clarify that this is global on a per feature basis, so column by column. It’s “global” in my head since it’s not time varying and not symbol specific and it’s not “batch” or “layer” normalization. IMHO, any batch or layer normalization is absolutely the wrong approach here as those were created to solve a different problem. ",
    "3057176": "For NN models with sequence input, InstanceNorm is more intuitive for me. As said, each feature has its own distribution. Using either batch norm or layer norm will likely mix the distribution of different features. ",
    "3057703": "A lot of trading data is domain-knowledge scaled. E.g. a price extension in terms of the opening price range (9:30am to 10:30am UTC+5). I am assuming that is the case here and not focusing on scaling except for training stability. ",
    "3056275": "These insights are wonderful!",
    "3056047": "Also if we look at some features across time we can say for certain new value ranges are constantly popping up, and we can say for almost certain that also happens with test data. A feature that has been ranging from 0 and 1 up until day 1000, can go to, say, 3 on day 1001. No normalization technique will account for that. Based on pure experimentation, what has worked best for me so far is inputting 0 for missings, and standard scaling. I'm very sure, however, this is not the best approach.",
    "3055156": "Sliding window normalisation worth a try",
    "3056249": "Great thoughts (and the same to your another two posts, pure gold)! \nI just switched from GBDTs to NN models, so haven't experimented too much yet. But from what I see, the proper normalisation method(maybe also some other preprocessing methods) could be the key to win in this competition. Since using global normalisation methods, no matter what models we choose (I have tried GRU model, 1d casaul cnn and transformer models, the local CV and LB are quite stable - for me now the scores are around 0.0060). Using moving window mean/std is slightly better than global normalisation (we know it's a bad choice).\nNeed to dive deeper into the data/features to figure out a better way, just like what you do. Cheers!",
    "3055768": "I feel responder is normalized within [-5,5], and features aren't having much outliers (if you look at distribution of each). I feel the purpose of this competition is more on predict methology rather than feature engineering.\n\nbtw, looking to team up with someone struggling at 0.00043 as well, kaggle didn't allow me to contact if not kaggle contributer.",
    "3055805": "Great insights on normalization! Rethinking default techniques like BatchNormalization and considering dynamic, context-aware methods is crucial, especially for evolving data like financial markets.\n Avoiding data leakage with rolling stats and tracking structural changes are excellent suggestions.",
    "3055764": "I guess many standard methods fail when comes to such data (just consoling myself when I see all my struggles result in bad PB ranks)\n\nNormalization – Though tree models might be an option to avoid this...but still not tempting enough to switch over. Avoiding normalization in nn model sounds interesting but it cannot be as simple as it that. However thanks for your thoughts...at least some more experiments to try.",
    "3055340": "> Well, normalizing features using global stats is nothing but leakage—the textbook classic. \n\nThis is only true if the global mean / variance is calculated prior to splitting the data",
    "3055275": "Does the .diff not work?",
    "3054969": "Maxmin scaler may be more helpful than the standard scaler.",
    "3057290": ""
  }
}