{
  "id": 551913,
  "title": "Why do we get noisy scores? Target(=sii) distribution may be the Key.",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551913",
  "author_name": "",
  "post_date": "2024-12-16T12:20:59.637520700Z",
  "votes": 15,
  "comment_count": 13,
  "views": 0,
  "content": "<p>We know from public notebook results that small differences in seed values(=42,2024,…) ​​can yield very noisy scores.<br>\nOne of the reasons for this may be insufficient training data, but in my opinion, <strong>target (=sii) distribution</strong> of the test data is different from the training data is also an important factor.</p>\n<p>Through several experiments, I found that the variance of QWK scores changes just by changing the distribution of sii (=0~3). This is due to the way QWK is calculated. In QWK calculations, the larger the difference between the predicted and correct labels, the more squared the penalty is given. (The figure below shows the penalties for each class.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10604629%2Fe806b9b43a4c8d4ce22673032f132e7e%2Fpicture_pc_ad440002b1e991af34e35e4785200043.webp?generation=1734348983215112&amp;alt=media\" alt=\"\"></p>\n<p>This experiments suggests that even in a completely random environment, the standard deviation of QWK increases when there are many labels such as 0 and 2. For example, the higher the ratio of sii=2, the higher the random QWK variation.</p>\n<table>\n<thead>\n<tr>\n<th>sii0/all=0.5, sii1/all=0.5-j, sii2/all=j</th>\n<th>j</th>\n<th>QWK's std</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td>0.0</td>\n<td>0.02226</td>\n</tr>\n<tr>\n<td></td>\n<td>0.1</td>\n<td>0.02727</td>\n</tr>\n<tr>\n<td></td>\n<td>0.2</td>\n<td>0.02945</td>\n</tr>\n<tr>\n<td></td>\n<td>0.3</td>\n<td>0.03056</td>\n</tr>\n<tr>\n<td></td>\n<td>0.4</td>\n<td>0.03099</td>\n</tr>\n<tr>\n<td></td>\n<td>0.5</td>\n<td>0.03120</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nIn other words, at least in the public test data, the dataset has many such labels(=0,2). For proof of this, just look at the 0.495 and 0.497 notebooks. The only difference between these two notebooks is that x0 (=initial value for optimization) in optim.minimize is changed from [0.5,1.5,2.5] to [0.5,1.49,2.5]. This means that only the threshold for allocation to each class has changed.</p>\n<pre><code>j = , thr = [  ]\n----&gt; || Optimized QWK SCORE ::  \nj = , thr = [   ]\n----&gt; || Optimized QWK SCORE ::  \n</code></pre>\n<p>The two values of thresholds are important. (You can ignore the last value, because the maximum predicted value in CV is about 2.00 at most, and there is no row that predicts sii=3.)</p>\n<p>The difference between these two thresholds shows that there are more rows that predict sii=0 than before.(because threshold[0] change from 0.588 to 0.600.) This means that the public test data contains more sii=0 than the predicted distribution in notebooks of 0.495.</p>\n<p>In other words, this means that the label distribution in test data is more biased toward 0 or 3 than in the training data, and when the above assumptions are combined, QWK with a larger variance can be obtained.</p>\n<p><strong>Note</strong>: However, even if we can predict the rough distribution of public test data, we can't predict the distribution of private test data. It can also be seen that simply changing threshold is simply overfitting public test data. But we are not sure it would give bad results on private test data.<br>\nIf it is close to the training data, it is advantageous to have a high CV, if it is close to the public test data, it is advantageous to have a high LB, and if it is neither, it is chaotic.</p>",
  "messages": [
    {
      "id": "3073374",
      "postDate": "12/16/2024 12:20:59",
      "content": "<p>We know from public notebook results that small differences in seed values(=42,2024,…) ​​can yield very noisy scores.<br>\nOne of the reasons for this may be insufficient training data, but in my opinion, <strong>target (=sii) distribution</strong> of the test data is different from the training data is also an important factor.</p>\n<p>Through several experiments, I found that the variance of QWK scores changes just by changing the distribution of sii (=0~3). This is due to the way QWK is calculated. In QWK calculations, the larger the difference between the predicted and correct labels, the more squared the penalty is given. (The figure below shows the penalties for each class.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10604629%2Fe806b9b43a4c8d4ce22673032f132e7e%2Fpicture_pc_ad440002b1e991af34e35e4785200043.webp?generation=1734348983215112&amp;alt=media\" alt=\"\"></p>\n<p>This experiments suggests that even in a completely random environment, the standard deviation of QWK increases when there are many labels such as 0 and 2. For example, the higher the ratio of sii=2, the higher the random QWK variation.</p>\n<table>\n<thead>\n<tr>\n<th>sii0/all=0.5, sii1/all=0.5-j, sii2/all=j</th>\n<th>j</th>\n<th>QWK's std</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td>0.0</td>\n<td>0.02226</td>\n</tr>\n<tr>\n<td></td>\n<td>0.1</td>\n<td>0.02727</td>\n</tr>\n<tr>\n<td></td>\n<td>0.2</td>\n<td>0.02945</td>\n</tr>\n<tr>\n<td></td>\n<td>0.3</td>\n<td>0.03056</td>\n</tr>\n<tr>\n<td></td>\n<td>0.4</td>\n<td>0.03099</td>\n</tr>\n<tr>\n<td></td>\n<td>0.5</td>\n<td>0.03120</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nIn other words, at least in the public test data, the dataset has many such labels(=0,2). For proof of this, just look at the 0.495 and 0.497 notebooks. The only difference between these two notebooks is that x0 (=initial value for optimization) in optim.minimize is changed from [0.5,1.5,2.5] to [0.5,1.49,2.5]. This means that only the threshold for allocation to each class has changed.</p>\n<pre><code>j = , thr = [  ]\n----&gt; || Optimized QWK SCORE ::  \nj = , thr = [   ]\n----&gt; || Optimized QWK SCORE ::  \n</code></pre>\n<p>The two values of thresholds are important. (You can ignore the last value, because the maximum predicted value in CV is about 2.00 at most, and there is no row that predicts sii=3.)</p>\n<p>The difference between these two thresholds shows that there are more rows that predict sii=0 than before.(because threshold[0] change from 0.588 to 0.600.) This means that the public test data contains more sii=0 than the predicted distribution in notebooks of 0.495.</p>\n<p>In other words, this means that the label distribution in test data is more biased toward 0 or 3 than in the training data, and when the above assumptions are combined, QWK with a larger variance can be obtained.</p>\n<p><strong>Note</strong>: However, even if we can predict the rough distribution of public test data, we can't predict the distribution of private test data. It can also be seen that simply changing threshold is simply overfitting public test data. But we are not sure it would give bad results on private test data.<br>\nIf it is close to the training data, it is advantageous to have a high CV, if it is close to the public test data, it is advantageous to have a high LB, and if it is neither, it is chaotic.</p>",
      "rawMarkdown": "We know from public notebook results that small differences in seed values(=42,2024,...) ​​can yield very noisy scores.\nOne of the reasons for this may be insufficient training data, but in my opinion, **target (=sii) distribution** of the test data is different from the training data is also an important factor.\n\nThrough several experiments, I found that the variance of QWK scores changes just by changing the distribution of sii (=0~3). This is due to the way QWK is calculated. In QWK calculations, the larger the difference between the predicted and correct labels, the more squared the penalty is given. (The figure below shows the penalties for each class.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10604629%2Fe806b9b43a4c8d4ce22673032f132e7e%2Fpicture_pc_ad440002b1e991af34e35e4785200043.webp?generation=1734348983215112&alt=media)\n\nThis experiments suggests that even in a completely random environment, the standard deviation of QWK increases when there are many labels such as 0 and 2. For example, the higher the ratio of sii=2, the higher the random QWK variation.\n| sii0/all=0.5, sii1/all=0.5-j, sii2/all=j | j | QWK's std |\n| --- | --- | --- |\n| | 0.0 | 0.02226 |\n| | 0.1 | 0.02727 |\n| | 0.2 | 0.02945 |\n| | 0.3 | 0.03056 |\n| | 0.4 | 0.03099 |\n| | 0.5 | 0.03120 |\n\n<br>\nIn other words, at least in the public test data, the dataset has many such labels(=0,2). For proof of this, just look at the 0.495 and 0.497 notebooks. The only difference between these two notebooks is that x0 (=initial value for optimization) in optim.minimize is changed from [0.5,1.5,2.5] to [0.5,1.49,2.5]. This means that only the threshold for allocation to each class has changed.\n\n```python\nj = 1.49, thr = [0.60088252 0.96162101 2.71690136]\n----> || Optimized QWK SCORE ::  0.456\nj = 1.5, thr = [0.5882359  0.96172902 2.67307225]\n----> || Optimized QWK SCORE ::  0.457\n```\n\nThe two values of thresholds are important. (You can ignore the last value, because the maximum predicted value in CV is about 2.00 at most, and there is no row that predicts sii=3.)\n\nThe difference between these two thresholds shows that there are more rows that predict sii=0 than before.(because threshold[0] change from 0.588 to 0.600.) This means that the public test data contains more sii=0 than the predicted distribution in notebooks of 0.495.\n\nIn other words, this means that the label distribution in test data is more biased toward 0 or 3 than in the training data, and when the above assumptions are combined, QWK with a larger variance can be obtained.\n\n**Note**: However, even if we can predict the rough distribution of public test data, we can't predict the distribution of private test data. It can also be seen that simply changing threshold is simply overfitting public test data. But we are not sure it would give bad results on private test data.\nIf it is close to the training data, it is advantageous to have a high CV, if it is close to the public test data, it is advantageous to have a high LB, and if it is neither, it is chaotic.",
      "votes": null
    },
    {
      "id": "3073394",
      "postDate": "12/16/2024 13:18:59",
      "content": "<p>Agreed, I did something similar in my initial experiments and have no major indicator of the distribution and if the private test data is close to the train/ public test set <a href=\"https://www.kaggle.com/minar44\" target=\"_blank\">@minar44</a> <br>\nI think the results will be chaotic given the nature of the data and the overall CV-LB score</p>",
      "rawMarkdown": "Agreed, I did something similar in my initial experiments and have no major indicator of the distribution and if the private test data is close to the train/ public test set @minar44 \nI think the results will be chaotic given the nature of the data and the overall CV-LB score",
      "votes": null
    },
    {
      "id": "3073422",
      "postDate": "12/16/2024 13:59:54",
      "content": "<p>In conclusion, get ready for the big lottery? 😍🤣</p>",
      "rawMarkdown": "In conclusion, get ready for the big lottery? 😍🤣",
      "votes": null
    },
    {
      "id": "3073425",
      "postDate": "12/16/2024 14:03:08",
      "content": "<p>Yes, this is a lottery <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <br>\nPeople winning golds here are one of the luckiest ones out there!</p>",
      "rawMarkdown": "Yes, this is a lottery @shlomoron \nPeople winning golds here are one of the luckiest ones out there!",
      "votes": null
    },
    {
      "id": "3073431",
      "postDate": "12/16/2024 14:04:38",
      "content": "<p>\"In other words, this means that the label distribution in test data is more biased toward 0 or 3 than in the training data\"<br>\nI don't think it necessarily has to mean that the distribution of sii=0 is higher. The thresholds depends on the data (not only the sii value). There could simply be a feature with a higher distribution than in the training df, which leads your model to a higher bias in general, so the thresholds move up aswell.<br>\nI haven't tested this, but I would assume that if you do the following:<br>\nSplit df into 3 parts with equal distribution of sii values.<br>\nTrain on one. Predict the other two. The optimal qwk for the predicted dfs will be different from each other, eventhough they have the same distribution of sii.<br>\nAnyway, we can't predict how the private LB looks anyway. We need alot of luck :)</p>",
      "rawMarkdown": "\"In other words, this means that the label distribution in test data is more biased toward 0 or 3 than in the training data\"\nI don't think it necessarily has to mean that the distribution of sii=0 is higher. The thresholds depends on the data (not only the sii value). There could simply be a feature with a higher distribution than in the training df, which leads your model to a higher bias in general, so the thresholds move up aswell.\nI haven't tested this, but I would assume that if you do the following:\nSplit df into 3 parts with equal distribution of sii values.\nTrain on one. Predict the other two. The optimal qwk for the predicted dfs will be different from each other, eventhough they have the same distribution of sii.\nAnyway, we can't predict how the private LB looks anyway. We need alot of luck :)",
      "votes": null
    },
    {
      "id": "3073449",
      "postDate": "12/16/2024 14:14:54",
      "content": "<p>Well, we can still maximize our chances of winning gold. If all are random, it's 0.0046, so with some robustness and avoiding pitfalls that exist in public NB, we can easily exceed the 0.01 chance for gold—and maybe much higher! For a lottery, it is really great odds.</p>",
      "rawMarkdown": "Well, we can still maximize our chances of winning gold. If all are random, it's 0.0046, so with some robustness and avoiding pitfalls that exist in public NB, we can easily exceed the 0.01 chance for gold—and maybe much higher! For a lottery, it is really great odds.",
      "votes": null
    },
    {
      "id": "3073453",
      "postDate": "12/16/2024 14:18:10",
      "content": "<p>I think there isn't much to do other than minimizing another metric on soft predictions and then make your bets.</p>",
      "rawMarkdown": "I think there isn't much to do other than minimizing another metric on soft predictions and then make your bets.",
      "votes": null
    },
    {
      "id": "3073535",
      "postDate": "12/16/2024 15:36:39",
      "content": "<blockquote>\n  <p>The two values of thresholds are important. (You can ignore the last value,</p>\n</blockquote>\n<p>The last value produces a <strong>local</strong> maximum, isn't it? The <strong>global</strong> maximum is obtained from a value around 1.5 if we look in the 0.495 notebook. What if we updated 0.497 notebook with these changes? In my experiments with NN, global maximums lead to a 0.01 improvement in CV and 0.005 in LB.&nbsp;But sometimes this improvement is even more if the first step optimization wasn't successful.</p>",
      "rawMarkdown": ">The two values of thresholds are important. (You can ignore the last value,\n\nThe last value produces a **local** maximum, isn't it? The **global** maximum is obtained from a value around 1.5 if we look in the 0.495 notebook. What if we updated 0.497 notebook with these changes? In my experiments with NN, global maximums lead to a 0.01 improvement in CV and 0.005 in LB. But sometimes this improvement is even more if the first step optimization wasn't successful.",
      "votes": null
    },
    {
      "id": "3074128",
      "postDate": "12/17/2024 10:54:50",
      "content": "<p>The two values and last value ​​here are thresholds. Thresholds are used to divide each class into 0, 1, 2, and 3.<br>\nFor example, the last value of the 0.497 notebook threshold is 2.71…, but the maximum predicted value of each model CV is as follows, and there is no value that is classified as 3.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>global maximum</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>model_1</td>\n<td>1.730926</td>\n</tr>\n<tr>\n<td>model_2</td>\n<td>2.003144</td>\n</tr>\n<tr>\n<td>model_3</td>\n<td>1.926503</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nAlso, I checked your notebook, it seems that the way the threshold is determined is different from that of the 0.497 notebook. Following that approach, the threshold would be [0.46, 0.91, 1.499]. I don't know if this will work as I haven't tried it.</p>",
      "rawMarkdown": "The two values and last value ​​here are thresholds. Thresholds are used to divide each class into 0, 1, 2, and 3.\nFor example, the last value of the 0.497 notebook threshold is 2.71..., but the maximum predicted value of each model CV is as follows, and there is no value that is classified as 3.\n| model | global maximum |\n| --- | --- |\n| model_1 | 1.730926 |\n| model_2 | 2.003144 |\n| model_3 | 1.926503 |\n\n<br>\nAlso, I checked your notebook, it seems that the way the threshold is determined is different from that of the 0.497 notebook. Following that approach, the threshold would be [0.46, 0.91, 1.499]. I don't know if this will work as I haven't tried it.",
      "votes": null
    },
    {
      "id": "3074146",
      "postDate": "12/17/2024 11:16:51",
      "content": "<p>Although it is higher than the notebook predicted sii=0 of 0.495, it cannot be said that bias is higher than training data. Thank you for pointing it out.</p>",
      "rawMarkdown": "Although it is higher than the notebook predicted sii=0 of 0.495, it cannot be said that bias is higher than training data. Thank you for pointing it out.",
      "votes": null
    },
    {
      "id": "3074294",
      "postDate": "12/17/2024 14:31:44",
      "content": "<p>I understand what you are pointing out about thresholds; it's okay. My question is why the author of the 0.495 notebook presents such great insight about global maximums and doesn't incorporate the calculation of the thresholds in that way in the further pipeline. The author of the 0.497 notebook is missing this discovery as well. So both best public notebooks can be even better, I think.</p>",
      "rawMarkdown": "I understand what you are pointing out about thresholds; it's okay. My question is why the author of the 0.495 notebook presents such great insight about global maximums and doesn't incorporate the calculation of the thresholds in that way in the further pipeline. The author of the 0.497 notebook is missing this discovery as well. So both best public notebooks can be even better, I think.",
      "votes": null
    },
    {
      "id": "3074303",
      "postDate": "12/17/2024 14:45:12",
      "content": "<p>I haven't tried it yet, so I can't say anything about it, but I think that using global maximum method may be overfits training data.<br>\nSure, that method allows us to find a combination of thresholds that optimizes CV score, but it doesn't mean it's the optimal threshold for test data too.</p>",
      "rawMarkdown": "I haven't tried it yet, so I can't say anything about it, but I think that using global maximum method may be overfits training data.\nSure, that method allows us to find a combination of thresholds that optimizes CV score, but it doesn't mean it's the optimal threshold for test data too.",
      "votes": null
    },
    {
      "id": "3074535",
      "postDate": "12/17/2024 19:13:30",
      "content": "<blockquote>\n  <p>using global maximum method may be overfits training data.  </p>\n</blockquote>\n<p>Is true also for local minimum.  <br>\nYou can avoid CV overfit by further splitting: train, train_treshold and validation.  <br>\nBut it's probably not worth it. Just submit and see if global minimum perform better than local minimum. Since it's only a small and very focused experiment you are not in danger of LB overfitting.  </p>",
      "rawMarkdown": ">using global maximum method may be overfits training data.  \n\nIs true also for local minimum.  \nYou can avoid CV overfit by further splitting: train, train_treshold and validation.  \nBut it's probably not worth it. Just submit and see if global minimum perform better than local minimum. Since it's only a small and very focused experiment you are not in danger of LB overfitting.",
      "votes": null
    },
    {
      "id": "3075680",
      "postDate": "12/19/2024 06:07:56",
      "content": "<p>The noisy Quadratic Weighted Kappa (QWK) scores can be attributed to a fundamental mismatch between training and test data distributions, particularly in how the target labels (sii values 0-3) are distributed. The test data contains a higher proportion of sii=0 cases than what was predicted from training data patterns, as evidenced by the threshold shift from 0.588 to 0.600 between optimization attempts.</p>\n<p>The nature of QWK calculation itself contributes to this volatility, as it employs a squared penalty system for prediction errors. This is clearly demonstrated in the penalty matrix (ω_i,j), where predictions that deviate further from the true class incur substantially larger penalties. For instance, confusing class 3 with class 1 results in a severe penalty of 4.</p>\n<p>The sensitivity of QWK scores becomes particularly apparent in how the standard deviation increases with parameter j, especially when there are numerous labels of 0 and 2. This sensitivity is so pronounced that even minor variations in seed values (like choosing between 42 and 2024) can lead to markedly different scores. This high variance in performance metrics stems from the combined effect of distributional differences between training and test sets, QWK's inherent quadratic penalty structure, and potentially insufficient training data, making the model's performance highly dependent on both initialization conditions and the underlying data distribution.</p>",
      "rawMarkdown": "The noisy Quadratic Weighted Kappa (QWK) scores can be attributed to a fundamental mismatch between training and test data distributions, particularly in how the target labels (sii values 0-3) are distributed. The test data contains a higher proportion of sii=0 cases than what was predicted from training data patterns, as evidenced by the threshold shift from 0.588 to 0.600 between optimization attempts.\n\nThe nature of QWK calculation itself contributes to this volatility, as it employs a squared penalty system for prediction errors. This is clearly demonstrated in the penalty matrix (ω_i,j), where predictions that deviate further from the true class incur substantially larger penalties. For instance, confusing class 3 with class 1 results in a severe penalty of 4.\n\nThe sensitivity of QWK scores becomes particularly apparent in how the standard deviation increases with parameter j, especially when there are numerous labels of 0 and 2. This sensitivity is so pronounced that even minor variations in seed values (like choosing between 42 and 2024) can lead to markedly different scores. This high variance in performance metrics stems from the combined effect of distributional differences between training and test sets, QWK's inherent quadratic penalty structure, and potentially insufficient training data, making the model's performance highly dependent on both initialization conditions and the underlying data distribution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3073394,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/16/2024 13:18:59",
      "content": "<p>Agreed, I did something similar in my initial experiments and have no major indicator of the distribution and if the private test data is close to the train/ public test set <a href=\"https://www.kaggle.com/minar44\" target=\"_blank\">@minar44</a> <br>\nI think the results will be chaotic given the nature of the data and the overall CV-LB score</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3073422,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "12/16/2024 13:59:54",
      "content": "<p>In conclusion, get ready for the big lottery? 😍🤣</p>",
      "votes": null,
      "replies": [
        {
          "id": 3073425,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/16/2024 14:03:08",
          "content": "<p>Yes, this is a lottery <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <br>\nPeople winning golds here are one of the luckiest ones out there!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3073449,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "12/16/2024 14:14:54",
              "content": "<p>Well, we can still maximize our chances of winning gold. If all are random, it's 0.0046, so with some robustness and avoiding pitfalls that exist in public NB, we can easily exceed the 0.01 chance for gold—and maybe much higher! For a lottery, it is really great odds.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3073453,
                  "author_name": "gunesevitan",
                  "author_url": "",
                  "post_date": "12/16/2024 14:18:10",
                  "content": "<p>I think there isn't much to do other than minimizing another metric on soft predictions and then make your bets.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3073431,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "12/16/2024 14:04:38",
      "content": "<p>\"In other words, this means that the label distribution in test data is more biased toward 0 or 3 than in the training data\"<br>\nI don't think it necessarily has to mean that the distribution of sii=0 is higher. The thresholds depends on the data (not only the sii value). There could simply be a feature with a higher distribution than in the training df, which leads your model to a higher bias in general, so the thresholds move up aswell.<br>\nI haven't tested this, but I would assume that if you do the following:<br>\nSplit df into 3 parts with equal distribution of sii values.<br>\nTrain on one. Predict the other two. The optimal qwk for the predicted dfs will be different from each other, eventhough they have the same distribution of sii.<br>\nAnyway, we can't predict how the private LB looks anyway. We need alot of luck :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3074146,
          "author_name": "minar44",
          "author_url": "",
          "post_date": "12/17/2024 11:16:51",
          "content": "<p>Although it is higher than the notebook predicted sii=0 of 0.495, it cannot be said that bias is higher than training data. Thank you for pointing it out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3073535,
      "author_name": "yekenot",
      "author_url": "",
      "post_date": "12/16/2024 15:36:39",
      "content": "<blockquote>\n  <p>The two values of thresholds are important. (You can ignore the last value,</p>\n</blockquote>\n<p>The last value produces a <strong>local</strong> maximum, isn't it? The <strong>global</strong> maximum is obtained from a value around 1.5 if we look in the 0.495 notebook. What if we updated 0.497 notebook with these changes? In my experiments with NN, global maximums lead to a 0.01 improvement in CV and 0.005 in LB.&nbsp;But sometimes this improvement is even more if the first step optimization wasn't successful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3074128,
          "author_name": "minar44",
          "author_url": "",
          "post_date": "12/17/2024 10:54:50",
          "content": "<p>The two values and last value ​​here are thresholds. Thresholds are used to divide each class into 0, 1, 2, and 3.<br>\nFor example, the last value of the 0.497 notebook threshold is 2.71…, but the maximum predicted value of each model CV is as follows, and there is no value that is classified as 3.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>global maximum</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>model_1</td>\n<td>1.730926</td>\n</tr>\n<tr>\n<td>model_2</td>\n<td>2.003144</td>\n</tr>\n<tr>\n<td>model_3</td>\n<td>1.926503</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nAlso, I checked your notebook, it seems that the way the threshold is determined is different from that of the 0.497 notebook. Following that approach, the threshold would be [0.46, 0.91, 1.499]. I don't know if this will work as I haven't tried it.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3074294,
              "author_name": "yekenot",
              "author_url": "",
              "post_date": "12/17/2024 14:31:44",
              "content": "<p>I understand what you are pointing out about thresholds; it's okay. My question is why the author of the 0.495 notebook presents such great insight about global maximums and doesn't incorporate the calculation of the thresholds in that way in the further pipeline. The author of the 0.497 notebook is missing this discovery as well. So both best public notebooks can be even better, I think.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3074303,
              "author_name": "minar44",
              "author_url": "",
              "post_date": "12/17/2024 14:45:12",
              "content": "<p>I haven't tried it yet, so I can't say anything about it, but I think that using global maximum method may be overfits training data.<br>\nSure, that method allows us to find a combination of thresholds that optimizes CV score, but it doesn't mean it's the optimal threshold for test data too.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3074535,
                  "author_name": "shlomoron",
                  "author_url": "",
                  "post_date": "12/17/2024 19:13:30",
                  "content": "<blockquote>\n  <p>using global maximum method may be overfits training data.  </p>\n</blockquote>\n<p>Is true also for local minimum.  <br>\nYou can avoid CV overfit by further splitting: train, train_treshold and validation.  <br>\nBut it's probably not worth it. Just submit and see if global minimum perform better than local minimum. Since it's only a small and very focused experiment you are not in danger of LB overfitting.  </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3075680,
      "author_name": "aadi611",
      "author_url": "",
      "post_date": "12/19/2024 06:07:56",
      "content": "<p>The noisy Quadratic Weighted Kappa (QWK) scores can be attributed to a fundamental mismatch between training and test data distributions, particularly in how the target labels (sii values 0-3) are distributed. The test data contains a higher proportion of sii=0 cases than what was predicted from training data patterns, as evidenced by the threshold shift from 0.588 to 0.600 between optimization attempts.</p>\n<p>The nature of QWK calculation itself contributes to this volatility, as it employs a squared penalty system for prediction errors. This is clearly demonstrated in the penalty matrix (ω_i,j), where predictions that deviate further from the true class incur substantially larger penalties. For instance, confusing class 3 with class 1 results in a severe penalty of 4.</p>\n<p>The sensitivity of QWK scores becomes particularly apparent in how the standard deviation increases with parameter j, especially when there are numerous labels of 0 and 2. This sensitivity is so pronounced that even minor variations in seed values (like choosing between 42 and 2024) can lead to markedly different scores. This high variance in performance metrics stems from the combined effect of distributional differences between training and test sets, QWK's inherent quadratic penalty structure, and potentially insufficient training data, making the model's performance highly dependent on both initialization conditions and the underlying data distribution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3073374": "We know from public notebook results that small differences in seed values(=42,2024,...) ​​can yield very noisy scores.\nOne of the reasons for this may be insufficient training data, but in my opinion, **target (=sii) distribution** of the test data is different from the training data is also an important factor.\n\nThrough several experiments, I found that the variance of QWK scores changes just by changing the distribution of sii (=0~3). This is due to the way QWK is calculated. In QWK calculations, the larger the difference between the predicted and correct labels, the more squared the penalty is given. (The figure below shows the penalties for each class.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10604629%2Fe806b9b43a4c8d4ce22673032f132e7e%2Fpicture_pc_ad440002b1e991af34e35e4785200043.webp?generation=1734348983215112&alt=media)\n\nThis experiments suggests that even in a completely random environment, the standard deviation of QWK increases when there are many labels such as 0 and 2. For example, the higher the ratio of sii=2, the higher the random QWK variation.\n| sii0/all=0.5, sii1/all=0.5-j, sii2/all=j | j | QWK's std |\n| --- | --- | --- |\n| | 0.0 | 0.02226 |\n| | 0.1 | 0.02727 |\n| | 0.2 | 0.02945 |\n| | 0.3 | 0.03056 |\n| | 0.4 | 0.03099 |\n| | 0.5 | 0.03120 |\n\n<br>\nIn other words, at least in the public test data, the dataset has many such labels(=0,2). For proof of this, just look at the 0.495 and 0.497 notebooks. The only difference between these two notebooks is that x0 (=initial value for optimization) in optim.minimize is changed from [0.5,1.5,2.5] to [0.5,1.49,2.5]. This means that only the threshold for allocation to each class has changed.\n\n```python\nj = 1.49, thr = [0.60088252 0.96162101 2.71690136]\n----> || Optimized QWK SCORE ::  0.456\nj = 1.5, thr = [0.5882359  0.96172902 2.67307225]\n----> || Optimized QWK SCORE ::  0.457\n```\n\nThe two values of thresholds are important. (You can ignore the last value, because the maximum predicted value in CV is about 2.00 at most, and there is no row that predicts sii=3.)\n\nThe difference between these two thresholds shows that there are more rows that predict sii=0 than before.(because threshold[0] change from 0.588 to 0.600.) This means that the public test data contains more sii=0 than the predicted distribution in notebooks of 0.495.\n\nIn other words, this means that the label distribution in test data is more biased toward 0 or 3 than in the training data, and when the above assumptions are combined, QWK with a larger variance can be obtained.\n\n**Note**: However, even if we can predict the rough distribution of public test data, we can't predict the distribution of private test data. It can also be seen that simply changing threshold is simply overfitting public test data. But we are not sure it would give bad results on private test data.\nIf it is close to the training data, it is advantageous to have a high CV, if it is close to the public test data, it is advantageous to have a high LB, and if it is neither, it is chaotic.",
    "3073394": "Agreed, I did something similar in my initial experiments and have no major indicator of the distribution and if the private test data is close to the train/ public test set @minar44 \nI think the results will be chaotic given the nature of the data and the overall CV-LB score",
    "3073422": "In conclusion, get ready for the big lottery? 😍🤣",
    "3073425": "Yes, this is a lottery @shlomoron \nPeople winning golds here are one of the luckiest ones out there!",
    "3073431": "\"In other words, this means that the label distribution in test data is more biased toward 0 or 3 than in the training data\"\nI don't think it necessarily has to mean that the distribution of sii=0 is higher. The thresholds depends on the data (not only the sii value). There could simply be a feature with a higher distribution than in the training df, which leads your model to a higher bias in general, so the thresholds move up aswell.\nI haven't tested this, but I would assume that if you do the following:\nSplit df into 3 parts with equal distribution of sii values.\nTrain on one. Predict the other two. The optimal qwk for the predicted dfs will be different from each other, eventhough they have the same distribution of sii.\nAnyway, we can't predict how the private LB looks anyway. We need alot of luck :)",
    "3073449": "Well, we can still maximize our chances of winning gold. If all are random, it's 0.0046, so with some robustness and avoiding pitfalls that exist in public NB, we can easily exceed the 0.01 chance for gold—and maybe much higher! For a lottery, it is really great odds.",
    "3073453": "I think there isn't much to do other than minimizing another metric on soft predictions and then make your bets.",
    "3073535": ">The two values of thresholds are important. (You can ignore the last value,\n\nThe last value produces a **local** maximum, isn't it? The **global** maximum is obtained from a value around 1.5 if we look in the 0.495 notebook. What if we updated 0.497 notebook with these changes? In my experiments with NN, global maximums lead to a 0.01 improvement in CV and 0.005 in LB. But sometimes this improvement is even more if the first step optimization wasn't successful.",
    "3074128": "The two values and last value ​​here are thresholds. Thresholds are used to divide each class into 0, 1, 2, and 3.\nFor example, the last value of the 0.497 notebook threshold is 2.71..., but the maximum predicted value of each model CV is as follows, and there is no value that is classified as 3.\n| model | global maximum |\n| --- | --- |\n| model_1 | 1.730926 |\n| model_2 | 2.003144 |\n| model_3 | 1.926503 |\n\n<br>\nAlso, I checked your notebook, it seems that the way the threshold is determined is different from that of the 0.497 notebook. Following that approach, the threshold would be [0.46, 0.91, 1.499]. I don't know if this will work as I haven't tried it.",
    "3074146": "Although it is higher than the notebook predicted sii=0 of 0.495, it cannot be said that bias is higher than training data. Thank you for pointing it out.",
    "3074294": "I understand what you are pointing out about thresholds; it's okay. My question is why the author of the 0.495 notebook presents such great insight about global maximums and doesn't incorporate the calculation of the thresholds in that way in the further pipeline. The author of the 0.497 notebook is missing this discovery as well. So both best public notebooks can be even better, I think.",
    "3074303": "I haven't tried it yet, so I can't say anything about it, but I think that using global maximum method may be overfits training data.\nSure, that method allows us to find a combination of thresholds that optimizes CV score, but it doesn't mean it's the optimal threshold for test data too.",
    "3074535": ">using global maximum method may be overfits training data.  \n\nIs true also for local minimum.  \nYou can avoid CV overfit by further splitting: train, train_treshold and validation.  \nBut it's probably not worth it. Just submit and see if global minimum perform better than local minimum. Since it's only a small and very focused experiment you are not in danger of LB overfitting.",
    "3075680": "The noisy Quadratic Weighted Kappa (QWK) scores can be attributed to a fundamental mismatch between training and test data distributions, particularly in how the target labels (sii values 0-3) are distributed. The test data contains a higher proportion of sii=0 cases than what was predicted from training data patterns, as evidenced by the threshold shift from 0.588 to 0.600 between optimization attempts.\n\nThe nature of QWK calculation itself contributes to this volatility, as it employs a squared penalty system for prediction errors. This is clearly demonstrated in the penalty matrix (ω_i,j), where predictions that deviate further from the true class incur substantially larger penalties. For instance, confusing class 3 with class 1 results in a severe penalty of 4.\n\nThe sensitivity of QWK scores becomes particularly apparent in how the standard deviation increases with parameter j, especially when there are numerous labels of 0 and 2. This sensitivity is so pronounced that even minor variations in seed values (like choosing between 42 and 2024) can lead to markedly different scores. This high variance in performance metrics stems from the combined effect of distributional differences between training and test sets, QWK's inherent quadratic penalty structure, and potentially insufficient training data, making the model's performance highly dependent on both initialization conditions and the underlying data distribution."
  },
  "source": "meta"
}