{
  "id": 584580,
  "title": "How to break Bartley's strong base LB 28.8? ",
  "url": "/competitions/waveform-inversion/discussion/584580",
  "author_name": "",
  "post_date": "2025-06-14T12:30:07.324604500Z",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<blockquote>\n  <p>Tried different losses / models but not able to break the Bartley's strong base LB. Improved speed of training with 'training speed discussion post'<br>\n  Target normalize / target scale are learning rate too slow then without target normalization.</p>\n</blockquote>",
  "messages": [
    {
      "id": "3224121",
      "postDate": "06/14/2025 12:30:07",
      "content": "<blockquote>\n  <p>Tried different losses / models but not able to break the Bartley's strong base LB. Improved speed of training with 'training speed discussion post'<br>\n  Target normalize / target scale are learning rate too slow then without target normalization.</p>\n</blockquote>",
      "rawMarkdown": "> Tried different losses / models but not able to break the Bartley's strong base LB. Improved speed of training with 'training speed discussion post'\n> Target normalize / target scale are learning rate too slow then without target normalization.",
      "votes": null
    },
    {
      "id": "3224192",
      "postDate": "06/14/2025 14:23:45",
      "content": "<p>if using target normalization:</p>\n<pre><code> sklearn.preprocessing  StandardScaler\ny_scaler = StandardScaler()\ny_train_scaled = y_scaler.fit_transform(y_train.reshape(-, )).ravel()\n</code></pre>\n<p>make sure to:</p>\n<ul>\n<li>use inverse_transform after prediction.</li>\n<li>adjust learning rate: target scaling often needs lower LR (like <code>1e-3 or 5e-4</code> if youre using MSE loss).</li>\n<li>log-transform targets if the distribution is skewed. you can try:</li>\n</ul>\n<pre><code>y_train = np.log1p(y_train)\n</code></pre>",
      "rawMarkdown": "if using target normalization:\n\n```py\nfrom sklearn.preprocessing import StandardScaler\ny_scaler = StandardScaler()\ny_train_scaled = y_scaler.fit_transform(y_train.reshape(-1, 1)).ravel()\n```\n\nmake sure to:\n\n- use inverse_transform after prediction.\n- adjust learning rate: target scaling often needs lower LR (like `1e-3 or 5e-4` if youre using MSE loss).\n- log-transform targets if the distribution is skewed. you can try:\n\n```py\ny_train = np.log1p(y_train)\n```",
      "votes": null
    },
    {
      "id": "3224314",
      "postDate": "06/14/2025 16:56:18",
      "content": "<p>Is the 28.8 LB score the best we can do by predicting a float16 velocity map? Could someone confirm that the they are able to get a better score predicting in float16? I think I am wasting my time trying to fine-tune a model that predicts float16 velocity maps, and that I am not able to get better scores just because the precision is not enough. I highly doubt this is the problem, but would be nice to hear other experiences.</p>",
      "rawMarkdown": "Is the 28.8 LB score the best we can do by predicting a float16 velocity map? Could someone confirm that the they are able to get a better score predicting in float16? I think I am wasting my time trying to fine-tune a model that predicts float16 velocity maps, and that I am not able to get better scores just because the precision is not enough. I highly doubt this is the problem, but would be nice to hear other experiences.",
      "votes": null
    },
    {
      "id": "3224770",
      "postDate": "06/15/2025 13:07:35",
      "content": "<p>Even though the CV score of my improved model improved by 3 on float16, I ran into a huge problem where its predictions were all integers! I'm not sure if it's better to use float32</p>",
      "rawMarkdown": "Even though the CV score of my improved model improved by 3 on float16, I ran into a huge problem where its predictions were all integers! I'm not sure if it's better to use float32",
      "votes": null
    },
    {
      "id": "3224898",
      "postDate": "06/15/2025 16:11:51",
      "content": "<p>May I ask a question? Were your capabilities enhanced through architectural modifications to a pre-trained model, or primarily through fine-tuning? Thank you kindly.😭</p>",
      "rawMarkdown": "May I ask a question? Were your capabilities enhanced through architectural modifications to a pre-trained model, or primarily through fine-tuning? Thank you kindly.😭",
      "votes": null
    },
    {
      "id": "3224930",
      "postDate": "06/15/2025 17:11:41",
      "content": "<p>My feeling is that one should use network prediction as initial solution then refine using physics differential equation. That should solve the integer value problem</p>",
      "rawMarkdown": "My feeling is that one should use network prediction as initial solution then refine using physics differential equation. That should solve the integer value problem",
      "votes": null
    },
    {
      "id": "3224931",
      "postDate": "06/15/2025 17:12:53",
      "content": "<p>Think of how to use the test data in training! masked auto encoders … using pyshics for self/unsupervised, refining ….</p>",
      "rawMarkdown": "Think of how to use the test data in training! masked auto encoders ... using pyshics for self/unsupervised, refining ....",
      "votes": null
    },
    {
      "id": "3225083",
      "postDate": "06/16/2025 00:59:15",
      "content": "<p>note that this problem is not the same as regular regression problem. because you have the wave forward modeling eqn to check the correctness of your current prediction. so the final solution could be additive or sequential, i.e. you need not retrain and retrain to get best solution.</p>\n<p>a winning solution could be:</p>\n<ol>\n<li>we have current model m</li>\n<li>make prediction on test y0 = m0(x)</li>\n<li>measure fitness of y0, e.g. x0=wave(y0), err0 = error(x0,x)</li>\n<li>build a new model to improve: new model input = (x,err0)</li>\n<li>so next time, y1 = m1(x,err0)</li>\n<li>repeat</li>\n</ol>",
      "rawMarkdown": "note that this problem is not the same as regular regression problem. because you have the wave forward modeling eqn to check the correctness of your current prediction. so the final solution could be additive or sequential, i.e. you need not retrain and retrain to get best solution.\n\na winning solution could be:\n\n1. we have current model m\n2. make prediction on test y0 = m0(x)\n3. measure fitness of y0, e.g. x0=wave(y0), err0 = error(x0,x)\n4. build a new model to improve: new model input = (x,err0)\n5. so next time, y1 = m1(x,err0)\n6. repeat",
      "votes": null
    },
    {
      "id": "3225128",
      "postDate": "06/16/2025 03:20:06",
      "content": "<p>Thank you very much for your suggestion. I think it seems like a lot of work for me. I have only seen in machine learning before using the output of a model as input to another model, which has increased my knowledge. <br>\nI would like to ask, if err1=error (y1, y) and y1 and err1 are directly used as inputs for the new model, do you think this is different from the principle of using x1 and err0 as new inputs，and do you think this two-stage model improvement is certain?</p>",
      "rawMarkdown": "Thank you very much for your suggestion. I think it seems like a lot of work for me. I have only seen in machine learning before using the output of a model as input to another model, which has increased my knowledge. \nI would like to ask, if err1=error (y1, y) and y1 and err1 are directly used as inputs for the new model, do you think this is different from the principle of using x1 and err0 as new inputs，and do you think this two-stage model improvement is certain?",
      "votes": null
    },
    {
      "id": "3225133",
      "postDate": "06/16/2025 03:30:40",
      "content": "<p>You don't what is the true target y of test. So y cannot compute target error at test. You can only compute input reconstruction error.</p>",
      "rawMarkdown": "You don't what is the true target y of test. So y cannot compute target error at test. You can only compute input reconstruction error.",
      "votes": null
    },
    {
      "id": "3225253",
      "postDate": "06/16/2025 07:04:35",
      "content": "<p>Thank you for your suggestion.😘</p>",
      "rawMarkdown": "Thank you for your suggestion.😘",
      "votes": null
    },
    {
      "id": "3225318",
      "postDate": "06/16/2025 08:59:18",
      "content": "<p>the trick to learn iterative refinition is to have lots of data. if each \"new model input = (x,err0)\" is trained with new batch of data, results will be very good.</p>\n<p>actuall you just need one </p>\n<pre><code> &lt;-- train on (,err_t)\n</code></pre>\n<p>and then use it iteratively.</p>\n<p>you can think of this is as data-driven version of the update step whent you solve differentiable equation iteratively and mathemtically</p>",
      "rawMarkdown": "the trick to learn iterative refinition is to have lots of data. if each \"new model input = (x,err0)\" is trained with new batch of data, results will be very good.\n\nactuall you just need one \n```\nmodel_(t+1) <-- train on (x,err_t)\n```\n\nand then use it iteratively.\n\nyou can think of this is as data-driven version of the update step whent you solve differentiable equation iteratively and mathemtically",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3224192,
      "author_name": "dakshbhatnagar08",
      "author_url": "",
      "post_date": "06/14/2025 14:23:45",
      "content": "<p>if using target normalization:</p>\n<pre><code> sklearn.preprocessing  StandardScaler\ny_scaler = StandardScaler()\ny_train_scaled = y_scaler.fit_transform(y_train.reshape(-, )).ravel()\n</code></pre>\n<p>make sure to:</p>\n<ul>\n<li>use inverse_transform after prediction.</li>\n<li>adjust learning rate: target scaling often needs lower LR (like <code>1e-3 or 5e-4</code> if youre using MSE loss).</li>\n<li>log-transform targets if the distribution is skewed. you can try:</li>\n</ul>\n<pre><code>y_train = np.log1p(y_train)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3224314,
      "author_name": "fpeccia",
      "author_url": "",
      "post_date": "06/14/2025 16:56:18",
      "content": "<p>Is the 28.8 LB score the best we can do by predicting a float16 velocity map? Could someone confirm that the they are able to get a better score predicting in float16? I think I am wasting my time trying to fine-tune a model that predicts float16 velocity maps, and that I am not able to get better scores just because the precision is not enough. I highly doubt this is the problem, but would be nice to hear other experiences.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3224770,
          "author_name": "oliver34",
          "author_url": "",
          "post_date": "06/15/2025 13:07:35",
          "content": "<p>Even though the CV score of my improved model improved by 3 on float16, I ran into a huge problem where its predictions were all integers! I'm not sure if it's better to use float32</p>",
          "votes": null,
          "replies": [
            {
              "id": 3224898,
              "author_name": "z0x0zn",
              "author_url": "",
              "post_date": "06/15/2025 16:11:51",
              "content": "<p>May I ask a question? Were your capabilities enhanced through architectural modifications to a pre-trained model, or primarily through fine-tuning? Thank you kindly.😭</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3224930,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/15/2025 17:11:41",
              "content": "<p>My feeling is that one should use network prediction as initial solution then refine using physics differential equation. That should solve the integer value problem</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3225083,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "06/16/2025 00:59:15",
                  "content": "<p>note that this problem is not the same as regular regression problem. because you have the wave forward modeling eqn to check the correctness of your current prediction. so the final solution could be additive or sequential, i.e. you need not retrain and retrain to get best solution.</p>\n<p>a winning solution could be:</p>\n<ol>\n<li>we have current model m</li>\n<li>make prediction on test y0 = m0(x)</li>\n<li>measure fitness of y0, e.g. x0=wave(y0), err0 = error(x0,x)</li>\n<li>build a new model to improve: new model input = (x,err0)</li>\n<li>so next time, y1 = m1(x,err0)</li>\n<li>repeat</li>\n</ol>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3225128,
                      "author_name": "oliver34",
                      "author_url": "",
                      "post_date": "06/16/2025 03:20:06",
                      "content": "<p>Thank you very much for your suggestion. I think it seems like a lot of work for me. I have only seen in machine learning before using the output of a model as input to another model, which has increased my knowledge. <br>\nI would like to ask, if err1=error (y1, y) and y1 and err1 are directly used as inputs for the new model, do you think this is different from the principle of using x1 and err0 as new inputs，and do you think this two-stage model improvement is certain?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3225133,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "06/16/2025 03:30:40",
                          "content": "<p>You don't what is the true target y of test. So y cannot compute target error at test. You can only compute input reconstruction error.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    },
                    {
                      "id": 3225253,
                      "author_name": "z0x0zn",
                      "author_url": "",
                      "post_date": "06/16/2025 07:04:35",
                      "content": "<p>Thank you for your suggestion.😘</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3225318,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "06/16/2025 08:59:18",
                          "content": "<p>the trick to learn iterative refinition is to have lots of data. if each \"new model input = (x,err0)\" is trained with new batch of data, results will be very good.</p>\n<p>actuall you just need one </p>\n<pre><code> &lt;-- train on (,err_t)\n</code></pre>\n<p>and then use it iteratively.</p>\n<p>you can think of this is as data-driven version of the update step whent you solve differentiable equation iteratively and mathemtically</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3224931,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/15/2025 17:12:53",
      "content": "<p>Think of how to use the test data in training! masked auto encoders … using pyshics for self/unsupervised, refining ….</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3224121": "> Tried different losses / models but not able to break the Bartley's strong base LB. Improved speed of training with 'training speed discussion post'\n> Target normalize / target scale are learning rate too slow then without target normalization.",
    "3224192": "if using target normalization:\n\n```py\nfrom sklearn.preprocessing import StandardScaler\ny_scaler = StandardScaler()\ny_train_scaled = y_scaler.fit_transform(y_train.reshape(-1, 1)).ravel()\n```\n\nmake sure to:\n\n- use inverse_transform after prediction.\n- adjust learning rate: target scaling often needs lower LR (like `1e-3 or 5e-4` if youre using MSE loss).\n- log-transform targets if the distribution is skewed. you can try:\n\n```py\ny_train = np.log1p(y_train)\n```",
    "3224314": "Is the 28.8 LB score the best we can do by predicting a float16 velocity map? Could someone confirm that the they are able to get a better score predicting in float16? I think I am wasting my time trying to fine-tune a model that predicts float16 velocity maps, and that I am not able to get better scores just because the precision is not enough. I highly doubt this is the problem, but would be nice to hear other experiences.",
    "3224770": "Even though the CV score of my improved model improved by 3 on float16, I ran into a huge problem where its predictions were all integers! I'm not sure if it's better to use float32",
    "3224898": "May I ask a question? Were your capabilities enhanced through architectural modifications to a pre-trained model, or primarily through fine-tuning? Thank you kindly.😭",
    "3224930": "My feeling is that one should use network prediction as initial solution then refine using physics differential equation. That should solve the integer value problem",
    "3224931": "Think of how to use the test data in training! masked auto encoders ... using pyshics for self/unsupervised, refining ....",
    "3225083": "note that this problem is not the same as regular regression problem. because you have the wave forward modeling eqn to check the correctness of your current prediction. so the final solution could be additive or sequential, i.e. you need not retrain and retrain to get best solution.\n\na winning solution could be:\n\n1. we have current model m\n2. make prediction on test y0 = m0(x)\n3. measure fitness of y0, e.g. x0=wave(y0), err0 = error(x0,x)\n4. build a new model to improve: new model input = (x,err0)\n5. so next time, y1 = m1(x,err0)\n6. repeat",
    "3225128": "Thank you very much for your suggestion. I think it seems like a lot of work for me. I have only seen in machine learning before using the output of a model as input to another model, which has increased my knowledge. \nI would like to ask, if err1=error (y1, y) and y1 and err1 are directly used as inputs for the new model, do you think this is different from the principle of using x1 and err0 as new inputs，and do you think this two-stage model improvement is certain?",
    "3225133": "You don't what is the true target y of test. So y cannot compute target error at test. You can only compute input reconstruction error.",
    "3225253": "Thank you for your suggestion.😘",
    "3225318": "the trick to learn iterative refinition is to have lots of data. if each \"new model input = (x,err0)\" is trained with new batch of data, results will be very good.\n\nactuall you just need one \n```\nmodel_(t+1) <-- train on (x,err_t)\n```\n\nand then use it iteratively.\n\nyou can think of this is as data-driven version of the update step whent you solve differentiable equation iteratively and mathemtically"
  },
  "source": "meta"
}