{
  "id": 543268,
  "title": "Use weight as sample_weight when training?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543268",
  "author_name": "",
  "post_date": "2024-10-29T15:33:34.283369700Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Should we use weights as sample_weight when training? </p>\n<p>The distribution of weights in the test set must be different from that in the training set. If training with weights, the model learns the patterns in the training set but can't adapt to the test set.</p>\n<p>How to correctly understand weights?</p>",
  "messages": [
    {
      "id": "3031340",
      "postDate": "10/29/2024 15:33:34",
      "content": "<p>Should we use weights as sample_weight when training? </p>\n<p>The distribution of weights in the test set must be different from that in the training set. If training with weights, the model learns the patterns in the training set but can't adapt to the test set.</p>\n<p>How to correctly understand weights?</p>",
      "rawMarkdown": "Should we use weights as sample_weight when training? \n\nThe distribution of weights in the test set must be different from that in the training set. If training with weights, the model learns the patterns in the training set but can't adapt to the test set.\n\n How to correctly understand weights?",
      "votes": null
    },
    {
      "id": "3031357",
      "postDate": "10/29/2024 15:55:26",
      "content": "<p>I dont think the <code>weight</code> variable shoud be used as sample_weight as they represent different things. The <code>weight</code> variable is the weight of a symbol in a portforlio (my best guess), while the sample weight in lgbm is usually related to the distrubution of samples. </p>\n<p>However many work in the shared kernel used <code>weight</code> as sample weight and the score is fine. I would try to use <code>weight</code> as an input feature (as it is associated with the volatility of r6) and compare the result with using <code>weight</code> as sample_weight. </p>\n<p>I personally would calculate sample weights based on the target distribution. From EDA, it is not difficult to find that r6 follows a laplacian distribution, indicating an imbalance between high and low volatility samples. These ideas are just ideas and have not been implemented yet.</p>",
      "rawMarkdown": "I dont think the `weight` variable shoud be used as sample_weight as they represent different things. The `weight` variable is the weight of a symbol in a portforlio (my best guess), while the sample weight in lgbm is usually related to the distrubution of samples. \n\nHowever many work in the shared kernel used `weight` as sample weight and the score is fine. I would try to use `weight` as an input feature (as it is associated with the volatility of r6) and compare the result with using `weight` as sample_weight. \n\nI personally would calculate sample weights based on the target distribution. From EDA, it is not difficult to find that r6 follows a laplacian distribution, indicating an imbalance between high and low volatility samples. These ideas are just ideas and have not been implemented yet.",
      "votes": null
    },
    {
      "id": "3031594",
      "postDate": "10/29/2024 21:54:09",
      "content": "<p>on the data page it says: \"<code>weight</code> - The weighting used for calculating the scoring function.\" So i think we are meant to use it as sample_weight</p>",
      "rawMarkdown": "on the data page it says: \"`weight` - The weighting used for calculating the scoring function.\" So i think we are meant to use it as sample_weight",
      "votes": null
    },
    {
      "id": "3031610",
      "postDate": "10/29/2024 22:37:51",
      "content": "<p>Use it in the evaluation metric is different from using it as the training loss. </p>",
      "rawMarkdown": "Use it in the evaluation metric is different from using it as the training loss.",
      "votes": null
    },
    {
      "id": "3031616",
      "postDate": "10/29/2024 23:01:40",
      "content": "<p>Yes you're right I misunderstood</p>",
      "rawMarkdown": "Yes you're right I misunderstood",
      "votes": null
    },
    {
      "id": "3031753",
      "postDate": "10/30/2024 04:17:49",
      "content": "<p>Both of them adjusts the contribution of each sample to overall loss or final score. Why do you think they are different?</p>",
      "rawMarkdown": "Both of them adjusts the contribution of each sample to overall loss or final score. Why do you think they are different?",
      "votes": null
    },
    {
      "id": "3032156",
      "postDate": "10/30/2024 15:56:33",
      "content": "<p>A simple comparision:</p>\n<pre><code># OOF R2 Ensemble:   | :  |   sample_weight  train\n# OOF R2 Ensemble:  | :  |  sample_weight  train\n</code></pre>\n<p>Both are ensemble of lgb+xgb, kept the same hparams and cv scheme.</p>",
      "rawMarkdown": "A simple comparision:\n\n```\n# OOF R2 Ensemble: 0.00782361616536087  | LB: 0.0045 |  no sample_weight in train\n# OOF R2 Ensemble: 0.007549115052393751 | LB: 0.0041 | use sample_weight in train\n```\n\nBoth are ensemble of lgb+xgb, kept the same hparams and cv scheme.",
      "votes": null
    },
    {
      "id": "3032181",
      "postDate": "10/30/2024 16:36:53",
      "content": "<p>I had similar results.</p>",
      "rawMarkdown": "I had similar results.",
      "votes": null
    },
    {
      "id": "3032184",
      "postDate": "10/30/2024 16:40:01",
      "content": "<p>using weights as train loss will affect the gradient and heissen and cause unexpected bias in model training.</p>",
      "rawMarkdown": "using weights as train loss will affect the gradient and heissen and cause unexpected bias in model training.",
      "votes": null
    },
    {
      "id": "3032200",
      "postDate": "10/30/2024 16:54:31",
      "content": "<p>Yeah but they are essentially the same thing, sample-wise losses/errors are multiplied with sample-wise weights.</p>",
      "rawMarkdown": "Yeah but they are essentially the same thing, sample-wise losses/errors are multiplied with sample-wise weights.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3031357,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/29/2024 15:55:26",
      "content": "<p>I dont think the <code>weight</code> variable shoud be used as sample_weight as they represent different things. The <code>weight</code> variable is the weight of a symbol in a portforlio (my best guess), while the sample weight in lgbm is usually related to the distrubution of samples. </p>\n<p>However many work in the shared kernel used <code>weight</code> as sample weight and the score is fine. I would try to use <code>weight</code> as an input feature (as it is associated with the volatility of r6) and compare the result with using <code>weight</code> as sample_weight. </p>\n<p>I personally would calculate sample weights based on the target distribution. From EDA, it is not difficult to find that r6 follows a laplacian distribution, indicating an imbalance between high and low volatility samples. These ideas are just ideas and have not been implemented yet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3031594,
          "author_name": "snehalverma10",
          "author_url": "",
          "post_date": "10/29/2024 21:54:09",
          "content": "<p>on the data page it says: \"<code>weight</code> - The weighting used for calculating the scoring function.\" So i think we are meant to use it as sample_weight</p>",
          "votes": null,
          "replies": [
            {
              "id": 3031610,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "10/29/2024 22:37:51",
              "content": "<p>Use it in the evaluation metric is different from using it as the training loss. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3031616,
                  "author_name": "snehalverma10",
                  "author_url": "",
                  "post_date": "10/29/2024 23:01:40",
                  "content": "<p>Yes you're right I misunderstood</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3031753,
                  "author_name": "gunesevitan",
                  "author_url": "",
                  "post_date": "10/30/2024 04:17:49",
                  "content": "<p>Both of them adjusts the contribution of each sample to overall loss or final score. Why do you think they are different?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3032184,
                      "author_name": "shiyili",
                      "author_url": "",
                      "post_date": "10/30/2024 16:40:01",
                      "content": "<p>using weights as train loss will affect the gradient and heissen and cause unexpected bias in model training.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3032200,
                          "author_name": "gunesevitan",
                          "author_url": "",
                          "post_date": "10/30/2024 16:54:31",
                          "content": "<p>Yeah but they are essentially the same thing, sample-wise losses/errors are multiplied with sample-wise weights.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3032156,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/30/2024 15:56:33",
      "content": "<p>A simple comparision:</p>\n<pre><code># OOF R2 Ensemble:   | :  |   sample_weight  train\n# OOF R2 Ensemble:  | :  |  sample_weight  train\n</code></pre>\n<p>Both are ensemble of lgb+xgb, kept the same hparams and cv scheme.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3032181,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "10/30/2024 16:36:53",
          "content": "<p>I had similar results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3031340": "Should we use weights as sample_weight when training? \n\nThe distribution of weights in the test set must be different from that in the training set. If training with weights, the model learns the patterns in the training set but can't adapt to the test set.\n\n How to correctly understand weights?",
    "3031357": "I dont think the `weight` variable shoud be used as sample_weight as they represent different things. The `weight` variable is the weight of a symbol in a portforlio (my best guess), while the sample weight in lgbm is usually related to the distrubution of samples. \n\nHowever many work in the shared kernel used `weight` as sample weight and the score is fine. I would try to use `weight` as an input feature (as it is associated with the volatility of r6) and compare the result with using `weight` as sample_weight. \n\nI personally would calculate sample weights based on the target distribution. From EDA, it is not difficult to find that r6 follows a laplacian distribution, indicating an imbalance between high and low volatility samples. These ideas are just ideas and have not been implemented yet.",
    "3031594": "on the data page it says: \"`weight` - The weighting used for calculating the scoring function.\" So i think we are meant to use it as sample_weight",
    "3031610": "Use it in the evaluation metric is different from using it as the training loss.",
    "3031616": "Yes you're right I misunderstood",
    "3031753": "Both of them adjusts the contribution of each sample to overall loss or final score. Why do you think they are different?",
    "3032156": "A simple comparision:\n\n```\n# OOF R2 Ensemble: 0.00782361616536087  | LB: 0.0045 |  no sample_weight in train\n# OOF R2 Ensemble: 0.007549115052393751 | LB: 0.0041 | use sample_weight in train\n```\n\nBoth are ensemble of lgb+xgb, kept the same hparams and cv scheme.",
    "3032181": "I had similar results.",
    "3032184": "using weights as train loss will affect the gradient and heissen and cause unexpected bias in model training.",
    "3032200": "Yeah but they are essentially the same thing, sample-wise losses/errors are multiplied with sample-wise weights."
  },
  "source": "meta"
}