{
  "id": 576237,
  "title": "Question on Scaling and Offset from best public notebooks ",
  "url": "/competitions/waveform-inversion/discussion/576237",
  "author_name": "",
  "post_date": "2025-05-03T17:41:24.011527800Z",
  "votes": 9,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>A few things about the scaling and offset operation bothered me so I wanted to ask. If you played with the best public notebooks like <a href=\"https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22\" target=\"_blank\">this one</a>, you may have seen the final operation of the U-Net model. It's a simple scale and offset operation as below:<br>\n<code>output = logits * 1000.0 + 1500.0</code></p>\n<p>Honestly, this is new to me because I'd usually normalize targets in the dataset class, then train with normalized targets (e.g. possibly in the range of [-1, 1] or [0, 1]) so that the gradients won't be huge. With our dataset's range of (1500, 4500), I'd expect the training with this to be very unstable yet it seems pretty stable and scalable. </p>\n<p>Besides, if I change it and do a (Min-Max) normalization like so<br>\n<code>y = (y - 1500) / (4500 - 1500)</code><br>\nand add a Sigmoid to the model's forward function, I get worse converge performance. <br>\nWhat's the likely explanation here? Is there an intuitive or theoretical explanation or am I just doing something wrong?</p>\n<p>One other thing is where does 1000 come from? Isn't it supposed to be <code>pred * (max_value - min_value) + min_value</code> so <code>pred * 3000 + 1500</code>? Am I missing something?</p>",
  "messages": [
    {
      "id": "3193029",
      "postDate": "05/03/2025 17:41:24",
      "content": "<p>Hi all,</p>\n<p>A few things about the scaling and offset operation bothered me so I wanted to ask. If you played with the best public notebooks like <a href=\"https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22\" target=\"_blank\">this one</a>, you may have seen the final operation of the U-Net model. It's a simple scale and offset operation as below:<br>\n<code>output = logits * 1000.0 + 1500.0</code></p>\n<p>Honestly, this is new to me because I'd usually normalize targets in the dataset class, then train with normalized targets (e.g. possibly in the range of [-1, 1] or [0, 1]) so that the gradients won't be huge. With our dataset's range of (1500, 4500), I'd expect the training with this to be very unstable yet it seems pretty stable and scalable. </p>\n<p>Besides, if I change it and do a (Min-Max) normalization like so<br>\n<code>y = (y - 1500) / (4500 - 1500)</code><br>\nand add a Sigmoid to the model's forward function, I get worse converge performance. <br>\nWhat's the likely explanation here? Is there an intuitive or theoretical explanation or am I just doing something wrong?</p>\n<p>One other thing is where does 1000 come from? Isn't it supposed to be <code>pred * (max_value - min_value) + min_value</code> so <code>pred * 3000 + 1500</code>? Am I missing something?</p>",
      "rawMarkdown": "Hi all,\n\nA few things about the scaling and offset operation bothered me so I wanted to ask. If you played with the best public notebooks like [this one](https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22), you may have seen the final operation of the U-Net model. It's a simple scale and offset operation as below:\n`output = logits * 1000.0 + 1500.0`\n\nHonestly, this is new to me because I'd usually normalize targets in the dataset class, then train with normalized targets (e.g. possibly in the range of [-1, 1] or [0, 1]) so that the gradients won't be huge. With our dataset's range of (1500, 4500), I'd expect the training with this to be very unstable yet it seems pretty stable and scalable. \n\nBesides, if I change it and do a (Min-Max) normalization like so\n`y = (y - 1500) / (4500 - 1500)`\nand add a Sigmoid to the model's forward function, I get worse converge performance. \nWhat's the likely explanation here? Is there an intuitive or theoretical explanation or am I just doing something wrong?\n\nOne other thing is where does 1000 come from? Isn't it supposed to be `pred * (max_value - min_value) + min_value` so `pred * 3000 + 1500`? Am I missing something?",
      "votes": null
    },
    {
      "id": "3193042",
      "postDate": "05/03/2025 18:06:57",
      "content": "<p>If you train with a l1 loss, the scale of the loss does not matter, your gradients will be 1 no matter what. That being said, l1 is definitely the outlier among possible loss functions here and I personally would prefer just normalizing targets.</p>\n<blockquote>\n  <p>and add a Sigmoid to the model's forward function, I get worse converge performance.</p>\n</blockquote>\n<p>Not surprising. Sigmoid has small gradients towards the tails, but values are rather uniformly distributed in the 1500-4500 range, at least for the B class subsets.</p>",
      "rawMarkdown": "If you train with a l1 loss, the scale of the loss does not matter, your gradients will be 1 no matter what. That being said, l1 is definitely the outlier among possible loss functions here and I personally would prefer just normalizing targets.\n\n> and add a Sigmoid to the model's forward function, I get worse converge performance.\n\nNot surprising. Sigmoid has small gradients towards the tails, but values are rather uniformly distributed in the 1500-4500 range, at least for the B class subsets.",
      "votes": null
    },
    {
      "id": "3193059",
      "postDate": "05/03/2025 18:46:10",
      "content": "<p>Thanks for the L1 explanation, you're right, the gradients are the same but models still end up different in my experiences. And for what its worth, I tried HardTanh(0, 1) and Identity instead of Sigmoid, which wasn't any better. </p>",
      "rawMarkdown": "Thanks for the L1 explanation, you're right, the gradients are the same but models still end up different in my experiences. And for what its worth, I tried HardTanh(0, 1) and Identity instead of Sigmoid, which wasn't any better.",
      "votes": null
    },
    {
      "id": "3199354",
      "postDate": "05/10/2025 20:56:32",
      "content": "<p>Hey again, <br>\nafter a little going back and forth, I don't think this statement is true:</p>\n<blockquote>\n  <p>If you train with a l1 loss, the scale of the loss does not matter, your gradients will be 1 no matter what.</p>\n</blockquote>\n<p>d(k x Loss)/d(params) = k x d(Loss)/d(params)<br>\nThe gradients are always scaled whether it's L1, L2, or something else.</p>",
      "rawMarkdown": "Hey again, \nafter a little going back and forth, I don't think this statement is true:\n\n>If you train with a l1 loss, the scale of the loss does not matter, your gradients will be 1 no matter what.\n\nd(k x Loss)/d(params) = k x d(Loss)/d(params)\nThe gradients are always scaled whether it's L1, L2, or something else.",
      "votes": null
    },
    {
      "id": "3199637",
      "postDate": "05/11/2025 09:01:06",
      "content": "<p>What is k here? You are correct that scaling the loss itself is equivalent to scaling the learning rate. However, when you normalize/scale your targets, you are not scaling the loss directly but scaling the magnitude of the residuals.</p>\n<p><code>l1 = |r|</code>, whereas <code>l2 = r^2</code>, with <code>r = y - y'</code> being the residual / error. Hence, <code>d/dr l1 = sign(r)</code> and <code>d/dr l2 = 2r</code>.</p>\n<p>This makes l1 invariant to the scale of the targets / residuals, and l2 not. Let <code>y_s' = a y'</code> be a scaled version of the target. Then <code>r_s = y - a y'</code> and <code>d/dr_s l1 = sign(r_s), d/dy l1 = 1 * sign(r_s)</code>. Hence, <code>d/d theta l1</code> is scaled by 1 during back propagation.</p>",
      "rawMarkdown": "What is k here? You are correct that scaling the loss itself is equivalent to scaling the learning rate. However, when you normalize/scale your targets, you are not scaling the loss directly but scaling the magnitude of the residuals.\n\n`l1 = |r|`, whereas `l2 = r^2`, with `r = y - y'` being the residual / error. Hence, `d/dr l1 = sign(r)` and `d/dr l2 = 2r`.\n\nThis makes l1 invariant to the scale of the targets / residuals, and l2 not. Let `y_s' = a y'` be a scaled version of the target. Then `r_s = y - a y'` and `d/dr_s l1 = sign(r_s), d/dy l1 = 1 * sign(r_s)`. Hence, `d/d theta l1` is scaled by 1 during back propagation.",
      "votes": null
    },
    {
      "id": "3199646",
      "postDate": "05/11/2025 09:22:06",
      "content": "<p>Addendum: The absolute value of the targets DO however affect other things, even when using l1. In particular, how well the network will be initialized and how quickly it can be updated. For instance, imagine scaling targets by a factor of 1e3 and a single, one dimensional layer y = w1  x1. Given a fixed update with a learning rate of 1e-3 and assuming normalized inputs, it would take quite a while for the network to be updated. In that regard, I stand corrected. </p>",
      "rawMarkdown": "Addendum: The absolute value of the targets DO however affect other things, even when using l1. In particular, how well the network will be initialized and how quickly it can be updated. For instance, imagine scaling targets by a factor of 1e3 and a single, one dimensional layer y = w1  x1. Given a fixed update with a learning rate of 1e-3 and assuming normalized inputs, it would take quite a while for the network to be updated. In that regard, I stand corrected.",
      "votes": null
    },
    {
      "id": "3199706",
      "postDate": "05/11/2025 11:16:04",
      "content": "<p>k is 1/3000 with Min-Max normalization.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F615314%2Fc80ad24c46ac1ad524b06b2d5f4296fc%2Fl1_vs_l1-normed.png?generation=1746960779984207&amp;alt=media\" alt=\"\"></p>\n<p>You may have thought of a different normalization process but gradients being always 1 with L1 is not true. It depends on the function <code>f(x)</code>, or <code>y</code> in your definition. The residual is the error but we don't take the derivative with respect to the residual again as in your example but the parameters that make up <code>y</code> so instead of <code>d/dr l1</code> it should be <code>d/dx l1</code> where <code>y = f(x)</code>, which is not necessarily 1.</p>\n<blockquote>\n  <p>scaling the loss itself is equivalent to scaling the learning rate</p>\n</blockquote>\n<p>That's because gradients are scaled with loss.</p>",
      "rawMarkdown": "k is 1/3000 with Min-Max normalization.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F615314%2Fc80ad24c46ac1ad524b06b2d5f4296fc%2Fl1_vs_l1-normed.png?generation=1746960779984207&alt=media)\n\nYou may have thought of a different normalization process but gradients being always 1 with L1 is not true. It depends on the function `f(x)`, or `y` in your definition. The residual is the error but we don't take the derivative with respect to the residual again as in your example but the parameters that make up `y` so instead of `d/dr l1` it should be `d/dx l1` where `y = f(x)`, which is not necessarily 1.\n\n>scaling the loss itself is equivalent to scaling the learning rate\n\nThat's because gradients are scaled with loss.",
      "votes": null
    },
    {
      "id": "3199738",
      "postDate": "05/11/2025 12:13:13",
      "content": "<p>You might have misunderstood me then. I would solely normalize the targets, not the output of your network. I.e. leave out the stuff for f(x). Usually also as a separate preprocessing step which doesn't participate in backprop.</p>\n<p>What you are proposing is essentially a different output layer for your network. That affects gradients obviously. </p>",
      "rawMarkdown": "You might have misunderstood me then. I would solely normalize the targets, not the output of your network. I.e. leave out the stuff for f(x). Usually also as a separate preprocessing step which doesn't participate in backprop.\n\nWhat you are proposing is essentially a different output layer for your network. That affects gradients obviously.",
      "votes": null
    },
    {
      "id": "3200808",
      "postDate": "05/13/2025 04:54:39",
      "content": "<p>May I ask if you have changed the final mapping to pred * 3000+1500? Does this improve compared to pred * 1000+1500</p>",
      "rawMarkdown": "May I ask if you have changed the final mapping to pred * 3000+1500? Does this improve compared to pred * 1000+1500",
      "votes": null
    },
    {
      "id": "3200906",
      "postDate": "05/13/2025 08:00:03",
      "content": "<p>In my experience <code>pred * 3000+1500</code> wasn't better than <code>pred * 1000+1500</code> so no. I'm comparing the predictions with normalized target values.</p>",
      "rawMarkdown": "In my experience `pred * 3000+1500` wasn't better than `pred * 1000+1500` so no. I'm comparing the predictions with normalized target values.",
      "votes": null
    },
    {
      "id": "3200969",
      "postDate": "05/13/2025 09:22:00",
      "content": "<p>thanks for your sharing😁</p>",
      "rawMarkdown": "thanks for your sharing😁",
      "votes": null
    },
    {
      "id": "3202337",
      "postDate": "05/15/2025 08:19:47",
      "content": "<p>How is it working? </p>",
      "rawMarkdown": "How is it working?",
      "votes": null
    },
    {
      "id": "3203070",
      "postDate": "05/16/2025 09:11:20",
      "content": "<p>The median of the velocity values is roughly around 3000 (I observed this by calculating the median for each mini-batch of the train_samples data, though I didn't use the entire dataset). So I think setting b in the equation a * x + b to around 3000 makes sense, since the median minimizes the MAE.</p>\n<p>As for a, I felt that a value slightly larger than the standard deviation of the velocity would work well, so I chose 1000. After training for a few epochs, I found that starting with 1000 * x + 3000 clearly gave better initial loss, though after around 3 epochs, the difference became negligible.</p>",
      "rawMarkdown": "The median of the velocity values is roughly around 3000 (I observed this by calculating the median for each mini-batch of the train_samples data, though I didn't use the entire dataset). So I think setting b in the equation a * x + b to around 3000 makes sense, since the median minimizes the MAE.\n\nAs for a, I felt that a value slightly larger than the standard deviation of the velocity would work well, so I chose 1000. After training for a few epochs, I found that starting with 1000 * x + 3000 clearly gave better initial loss, though after around 3 epochs, the difference became negligible.",
      "votes": null
    },
    {
      "id": "3203122",
      "postDate": "05/16/2025 10:29:21",
      "content": "<p>Interesting, I always thought of 3000 as Max-Min difference. I guess your answer makes the most sense.</p>",
      "rawMarkdown": "Interesting, I always thought of 3000 as Max-Min difference. I guess your answer makes the most sense.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3193042,
      "author_name": "hoffmanns",
      "author_url": "",
      "post_date": "05/03/2025 18:06:57",
      "content": "<p>If you train with a l1 loss, the scale of the loss does not matter, your gradients will be 1 no matter what. That being said, l1 is definitely the outlier among possible loss functions here and I personally would prefer just normalizing targets.</p>\n<blockquote>\n  <p>and add a Sigmoid to the model's forward function, I get worse converge performance.</p>\n</blockquote>\n<p>Not surprising. Sigmoid has small gradients towards the tails, but values are rather uniformly distributed in the 1500-4500 range, at least for the B class subsets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3193059,
          "author_name": "mgoksu",
          "author_url": "",
          "post_date": "05/03/2025 18:46:10",
          "content": "<p>Thanks for the L1 explanation, you're right, the gradients are the same but models still end up different in my experiences. And for what its worth, I tried HardTanh(0, 1) and Identity instead of Sigmoid, which wasn't any better. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3199354,
          "author_name": "mgoksu",
          "author_url": "",
          "post_date": "05/10/2025 20:56:32",
          "content": "<p>Hey again, <br>\nafter a little going back and forth, I don't think this statement is true:</p>\n<blockquote>\n  <p>If you train with a l1 loss, the scale of the loss does not matter, your gradients will be 1 no matter what.</p>\n</blockquote>\n<p>d(k x Loss)/d(params) = k x d(Loss)/d(params)<br>\nThe gradients are always scaled whether it's L1, L2, or something else.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3199637,
              "author_name": "hoffmanns",
              "author_url": "",
              "post_date": "05/11/2025 09:01:06",
              "content": "<p>What is k here? You are correct that scaling the loss itself is equivalent to scaling the learning rate. However, when you normalize/scale your targets, you are not scaling the loss directly but scaling the magnitude of the residuals.</p>\n<p><code>l1 = |r|</code>, whereas <code>l2 = r^2</code>, with <code>r = y - y'</code> being the residual / error. Hence, <code>d/dr l1 = sign(r)</code> and <code>d/dr l2 = 2r</code>.</p>\n<p>This makes l1 invariant to the scale of the targets / residuals, and l2 not. Let <code>y_s' = a y'</code> be a scaled version of the target. Then <code>r_s = y - a y'</code> and <code>d/dr_s l1 = sign(r_s), d/dy l1 = 1 * sign(r_s)</code>. Hence, <code>d/d theta l1</code> is scaled by 1 during back propagation.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3199706,
                  "author_name": "mgoksu",
                  "author_url": "",
                  "post_date": "05/11/2025 11:16:04",
                  "content": "<p>k is 1/3000 with Min-Max normalization.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F615314%2Fc80ad24c46ac1ad524b06b2d5f4296fc%2Fl1_vs_l1-normed.png?generation=1746960779984207&amp;alt=media\" alt=\"\"></p>\n<p>You may have thought of a different normalization process but gradients being always 1 with L1 is not true. It depends on the function <code>f(x)</code>, or <code>y</code> in your definition. The residual is the error but we don't take the derivative with respect to the residual again as in your example but the parameters that make up <code>y</code> so instead of <code>d/dr l1</code> it should be <code>d/dx l1</code> where <code>y = f(x)</code>, which is not necessarily 1.</p>\n<blockquote>\n  <p>scaling the loss itself is equivalent to scaling the learning rate</p>\n</blockquote>\n<p>That's because gradients are scaled with loss.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3199738,
                      "author_name": "hoffmanns",
                      "author_url": "",
                      "post_date": "05/11/2025 12:13:13",
                      "content": "<p>You might have misunderstood me then. I would solely normalize the targets, not the output of your network. I.e. leave out the stuff for f(x). Usually also as a separate preprocessing step which doesn't participate in backprop.</p>\n<p>What you are proposing is essentially a different output layer for your network. That affects gradients obviously. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 3199646,
              "author_name": "hoffmanns",
              "author_url": "",
              "post_date": "05/11/2025 09:22:06",
              "content": "<p>Addendum: The absolute value of the targets DO however affect other things, even when using l1. In particular, how well the network will be initialized and how quickly it can be updated. For instance, imagine scaling targets by a factor of 1e3 and a single, one dimensional layer y = w1  x1. Given a fixed update with a learning rate of 1e-3 and assuming normalized inputs, it would take quite a while for the network to be updated. In that regard, I stand corrected. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3200808,
      "author_name": "peilwang",
      "author_url": "",
      "post_date": "05/13/2025 04:54:39",
      "content": "<p>May I ask if you have changed the final mapping to pred * 3000+1500? Does this improve compared to pred * 1000+1500</p>",
      "votes": null,
      "replies": [
        {
          "id": 3200906,
          "author_name": "mgoksu",
          "author_url": "",
          "post_date": "05/13/2025 08:00:03",
          "content": "<p>In my experience <code>pred * 3000+1500</code> wasn't better than <code>pred * 1000+1500</code> so no. I'm comparing the predictions with normalized target values.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3200969,
              "author_name": "peilwang",
              "author_url": "",
              "post_date": "05/13/2025 09:22:00",
              "content": "<p>thanks for your sharing😁</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3202337,
      "author_name": "rajiasultanagssc",
      "author_url": "",
      "post_date": "05/15/2025 08:19:47",
      "content": "<p>How is it working? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3203070,
      "author_name": "haruiig",
      "author_url": "",
      "post_date": "05/16/2025 09:11:20",
      "content": "<p>The median of the velocity values is roughly around 3000 (I observed this by calculating the median for each mini-batch of the train_samples data, though I didn't use the entire dataset). So I think setting b in the equation a * x + b to around 3000 makes sense, since the median minimizes the MAE.</p>\n<p>As for a, I felt that a value slightly larger than the standard deviation of the velocity would work well, so I chose 1000. After training for a few epochs, I found that starting with 1000 * x + 3000 clearly gave better initial loss, though after around 3 epochs, the difference became negligible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3203122,
          "author_name": "mgoksu",
          "author_url": "",
          "post_date": "05/16/2025 10:29:21",
          "content": "<p>Interesting, I always thought of 3000 as Max-Min difference. I guess your answer makes the most sense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3193029": "Hi all,\n\nA few things about the scaling and offset operation bothered me so I wanted to ask. If you played with the best public notebooks like [this one](https://www.kaggle.com/code/muhammadqasimshabbir/gwi-unet-with-float16-dataset22), you may have seen the final operation of the U-Net model. It's a simple scale and offset operation as below:\n`output = logits * 1000.0 + 1500.0`\n\nHonestly, this is new to me because I'd usually normalize targets in the dataset class, then train with normalized targets (e.g. possibly in the range of [-1, 1] or [0, 1]) so that the gradients won't be huge. With our dataset's range of (1500, 4500), I'd expect the training with this to be very unstable yet it seems pretty stable and scalable. \n\nBesides, if I change it and do a (Min-Max) normalization like so\n`y = (y - 1500) / (4500 - 1500)`\nand add a Sigmoid to the model's forward function, I get worse converge performance. \nWhat's the likely explanation here? Is there an intuitive or theoretical explanation or am I just doing something wrong?\n\nOne other thing is where does 1000 come from? Isn't it supposed to be `pred * (max_value - min_value) + min_value` so `pred * 3000 + 1500`? Am I missing something?",
    "3193042": "If you train with a l1 loss, the scale of the loss does not matter, your gradients will be 1 no matter what. That being said, l1 is definitely the outlier among possible loss functions here and I personally would prefer just normalizing targets.\n\n> and add a Sigmoid to the model's forward function, I get worse converge performance.\n\nNot surprising. Sigmoid has small gradients towards the tails, but values are rather uniformly distributed in the 1500-4500 range, at least for the B class subsets.",
    "3193059": "Thanks for the L1 explanation, you're right, the gradients are the same but models still end up different in my experiences. And for what its worth, I tried HardTanh(0, 1) and Identity instead of Sigmoid, which wasn't any better.",
    "3199354": "Hey again, \nafter a little going back and forth, I don't think this statement is true:\n\n>If you train with a l1 loss, the scale of the loss does not matter, your gradients will be 1 no matter what.\n\nd(k x Loss)/d(params) = k x d(Loss)/d(params)\nThe gradients are always scaled whether it's L1, L2, or something else.",
    "3199637": "What is k here? You are correct that scaling the loss itself is equivalent to scaling the learning rate. However, when you normalize/scale your targets, you are not scaling the loss directly but scaling the magnitude of the residuals.\n\n`l1 = |r|`, whereas `l2 = r^2`, with `r = y - y'` being the residual / error. Hence, `d/dr l1 = sign(r)` and `d/dr l2 = 2r`.\n\nThis makes l1 invariant to the scale of the targets / residuals, and l2 not. Let `y_s' = a y'` be a scaled version of the target. Then `r_s = y - a y'` and `d/dr_s l1 = sign(r_s), d/dy l1 = 1 * sign(r_s)`. Hence, `d/d theta l1` is scaled by 1 during back propagation.",
    "3199646": "Addendum: The absolute value of the targets DO however affect other things, even when using l1. In particular, how well the network will be initialized and how quickly it can be updated. For instance, imagine scaling targets by a factor of 1e3 and a single, one dimensional layer y = w1  x1. Given a fixed update with a learning rate of 1e-3 and assuming normalized inputs, it would take quite a while for the network to be updated. In that regard, I stand corrected.",
    "3199706": "k is 1/3000 with Min-Max normalization.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F615314%2Fc80ad24c46ac1ad524b06b2d5f4296fc%2Fl1_vs_l1-normed.png?generation=1746960779984207&alt=media)\n\nYou may have thought of a different normalization process but gradients being always 1 with L1 is not true. It depends on the function `f(x)`, or `y` in your definition. The residual is the error but we don't take the derivative with respect to the residual again as in your example but the parameters that make up `y` so instead of `d/dr l1` it should be `d/dx l1` where `y = f(x)`, which is not necessarily 1.\n\n>scaling the loss itself is equivalent to scaling the learning rate\n\nThat's because gradients are scaled with loss.",
    "3199738": "You might have misunderstood me then. I would solely normalize the targets, not the output of your network. I.e. leave out the stuff for f(x). Usually also as a separate preprocessing step which doesn't participate in backprop.\n\nWhat you are proposing is essentially a different output layer for your network. That affects gradients obviously.",
    "3200808": "May I ask if you have changed the final mapping to pred * 3000+1500? Does this improve compared to pred * 1000+1500",
    "3200906": "In my experience `pred * 3000+1500` wasn't better than `pred * 1000+1500` so no. I'm comparing the predictions with normalized target values.",
    "3200969": "thanks for your sharing😁",
    "3202337": "How is it working?",
    "3203070": "The median of the velocity values is roughly around 3000 (I observed this by calculating the median for each mini-batch of the train_samples data, though I didn't use the entire dataset). So I think setting b in the equation a * x + b to around 3000 makes sense, since the median minimizes the MAE.\n\nAs for a, I felt that a value slightly larger than the standard deviation of the velocity would work well, so I chose 1000. After training for a few epochs, I found that starting with 1000 * x + 3000 clearly gave better initial loss, though after around 3 epochs, the difference became negligible.",
    "3203122": "Interesting, I always thought of 3000 as Max-Min difference. I guess your answer makes the most sense."
  },
  "source": "meta"
}