{
  "id": 552127,
  "title": "The optimized qwk dilemma",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552127",
  "author_name": "",
  "post_date": "2024-12-17T19:35:55.313804Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nafter spending way too much time on this competition, I wondered, why eventhough my notebook seemed to improve in CV (stable with different seeds), it wouldn't translate to the leaderboard. </p>\n<p>I finally stumbled upon the crux (at least for me): The unstable QWK thresholds.</p>\n<p>I tested this by splitting the train csv in 3 parts. I trained on two parts and then tested the third part. Here is one of the results, note that the resolution is not very high. The swings are also high if you zoom in on parts.</p>\n<p>oof predictions:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F5c2c2098b9ba42deed351bdb3a694151%2Flottery_train.PNG?generation=1734460508509910&amp;alt=media\" alt=\"Train\"></p>\n<p>test predictions:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fa33b22724fe1b87ef68519666a2f1339%2Flottery_val.PNG?generation=1734460543582369&amp;alt=media\" alt=\"Validation\"></p>\n<p>I predict PCIAT-Total in the screenshots, but it is the same for sii. As you can see the optimal qwk thresholds are not the same anymore. This is one of the more extreme outcomes, but it changes depending on the split seed, which model I use, etc.<br>\nEven if the maximums are not that far apart, it is still a huge qwk score swing and I can't find a way to stabilize it or work with it.</p>\n<p>Is there any way to fix this? I tried alot, but I assume it always depends on the target dataframe distribution?</p>",
  "messages": [
    {
      "id": "3074555",
      "postDate": "12/17/2024 19:35:55",
      "content": "<p>Hello everyone,<br>\nafter spending way too much time on this competition, I wondered, why eventhough my notebook seemed to improve in CV (stable with different seeds), it wouldn't translate to the leaderboard. </p>\n<p>I finally stumbled upon the crux (at least for me): The unstable QWK thresholds.</p>\n<p>I tested this by splitting the train csv in 3 parts. I trained on two parts and then tested the third part. Here is one of the results, note that the resolution is not very high. The swings are also high if you zoom in on parts.</p>\n<p>oof predictions:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F5c2c2098b9ba42deed351bdb3a694151%2Flottery_train.PNG?generation=1734460508509910&amp;alt=media\" alt=\"Train\"></p>\n<p>test predictions:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fa33b22724fe1b87ef68519666a2f1339%2Flottery_val.PNG?generation=1734460543582369&amp;alt=media\" alt=\"Validation\"></p>\n<p>I predict PCIAT-Total in the screenshots, but it is the same for sii. As you can see the optimal qwk thresholds are not the same anymore. This is one of the more extreme outcomes, but it changes depending on the split seed, which model I use, etc.<br>\nEven if the maximums are not that far apart, it is still a huge qwk score swing and I can't find a way to stabilize it or work with it.</p>\n<p>Is there any way to fix this? I tried alot, but I assume it always depends on the target dataframe distribution?</p>",
      "rawMarkdown": "Hello everyone,\nafter spending way too much time on this competition, I wondered, why eventhough my notebook seemed to improve in CV (stable with different seeds), it wouldn't translate to the leaderboard. \n\nI finally stumbled upon the crux (at least for me): The unstable QWK thresholds.\n\nI tested this by splitting the train csv in 3 parts. I trained on two parts and then tested the third part. Here is one of the results, note that the resolution is not very high. The swings are also high if you zoom in on parts.\n\noof predictions:\n![Train](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F5c2c2098b9ba42deed351bdb3a694151%2Flottery_train.PNG?generation=1734460508509910&alt=media)\n\ntest predictions:\n![Validation](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fa33b22724fe1b87ef68519666a2f1339%2Flottery_val.PNG?generation=1734460543582369&alt=media)\n\nI predict PCIAT-Total in the screenshots, but it is the same for sii. As you can see the optimal qwk thresholds are not the same anymore. This is one of the more extreme outcomes, but it changes depending on the split seed, which model I use, etc.\nEven if the maximums are not that far apart, it is still a huge qwk score swing and I can't find a way to stabilize it or work with it.\n\nIs there any way to fix this? I tried alot, but I assume it always depends on the target dataframe distribution?",
      "votes": null
    },
    {
      "id": "3074895",
      "postDate": "12/18/2024 07:18:44",
      "content": "<p>There are some ways of reducing variance like training on multiple seeds and taking the average. Very nice visualization btw, can you share the code?</p>",
      "rawMarkdown": "There are some ways of reducing variance like training on multiple seeds and taking the average. Very nice visualization btw, can you share the code?",
      "votes": null
    },
    {
      "id": "3074899",
      "postDate": "12/18/2024 07:28:45",
      "content": "<p>I'm taking the average of 15 seeds currently. I could probably take 100 seeds and it wouldnt change on the training data. But the test data will change, because it has its own distribution. I also made a post in your thread just now, because we probably have the same problem :)</p>\n<p>Here is the code:</p>\n<p>y_train_class are the 'sii' values and pred_light is the oof predictions <em>before</em> you use qwk thresholds to classify on them.<br>\nevaluate_predictions is the one everyone uses from the public notebook, it simply classifies the predictions with the given thresholds.</p>\n<p>`</p>\n<h1>Ranges for the threshold values 1 and 2</h1>\n<p>value_1_range = np.linspace(25, 33, 100) <br>\nvalue_2_range = np.linspace(35, 40, 100)  </p>\n<h1>Create a grid of all combinations</h1>\n<p>value_1_grid, value_2_grid = np.meshgrid(value_1_range, value_2_range)<br>\nresults = np.zeros_like(value_1_grid)</p>\n<h1>Go through every grid cell and check it's value</h1>\n<p>for i in range(value_1_grid.shape[0]):<br>\n    for j in range(value_1_grid.shape[1]):<br>\n        threshold = [value_1_grid[i, j], value_2_grid[i, j], 65]<br>\n        results[i, j] = evaluate_predictions(threshold, y_train_class, pred_light)</p>\n<h1>Heatmap plot</h1>\n<p>plt.figure(figsize=(14, 10))<br>\nplt.contourf(value_1_range, value_2_range, results, levels=30, cmap=\"viridis\")<br>\nplt.colorbar(label=\"Evaluation Metric\")<br>\nplt.title(\"Effect of Threshold on Predictions\")<br>\nplt.xlabel(\"Value 1 (Threshold 1)\")<br>\nplt.ylabel(\"Value 2 (Threshold 2)\")<br>\nplt.show()<br>\n`</p>\n<p>Edit: Weird, I'm posting it as code, but it doesn't take the inline spaces, which need to be there after every 'for'.</p>",
      "rawMarkdown": "I'm taking the average of 15 seeds currently. I could probably take 100 seeds and it wouldnt change on the training data. But the test data will change, because it has its own distribution. I also made a post in your thread just now, because we probably have the same problem :)\n\nHere is the code:\n\ny_train_class are the 'sii' values and pred_light is the oof predictions _before_ you use qwk thresholds to classify on them.\nevaluate_predictions is the one everyone uses from the public notebook, it simply classifies the predictions with the given thresholds.\n\n`\n# Ranges for the threshold values 1 and 2\nvalue_1_range = np.linspace(25, 33, 100) \nvalue_2_range = np.linspace(35, 40, 100)  \n\n# Create a grid of all combinations\nvalue_1_grid, value_2_grid = np.meshgrid(value_1_range, value_2_range)\nresults = np.zeros_like(value_1_grid)\n\n# Go through every grid cell and check it's value\nfor i in range(value_1_grid.shape[0]):\n    for j in range(value_1_grid.shape[1]):\n        threshold = [value_1_grid[i, j], value_2_grid[i, j], 65]\n        results[i, j] = evaluate_predictions(threshold, y_train_class, pred_light)\n      \n# Heatmap plot\nplt.figure(figsize=(14, 10))\nplt.contourf(value_1_range, value_2_range, results, levels=30, cmap=\"viridis\")\nplt.colorbar(label=\"Evaluation Metric\")\nplt.title(\"Effect of Threshold on Predictions\")\nplt.xlabel(\"Value 1 (Threshold 1)\")\nplt.ylabel(\"Value 2 (Threshold 2)\")\nplt.show()\n`\n\nEdit: Weird, I'm posting it as code, but it doesn't take the inline spaces, which need to be there after every 'for'.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3074895,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "12/18/2024 07:18:44",
      "content": "<p>There are some ways of reducing variance like training on multiple seeds and taking the average. Very nice visualization btw, can you share the code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3074899,
          "author_name": "mariusheuser",
          "author_url": "",
          "post_date": "12/18/2024 07:28:45",
          "content": "<p>I'm taking the average of 15 seeds currently. I could probably take 100 seeds and it wouldnt change on the training data. But the test data will change, because it has its own distribution. I also made a post in your thread just now, because we probably have the same problem :)</p>\n<p>Here is the code:</p>\n<p>y_train_class are the 'sii' values and pred_light is the oof predictions <em>before</em> you use qwk thresholds to classify on them.<br>\nevaluate_predictions is the one everyone uses from the public notebook, it simply classifies the predictions with the given thresholds.</p>\n<p>`</p>\n<h1>Ranges for the threshold values 1 and 2</h1>\n<p>value_1_range = np.linspace(25, 33, 100) <br>\nvalue_2_range = np.linspace(35, 40, 100)  </p>\n<h1>Create a grid of all combinations</h1>\n<p>value_1_grid, value_2_grid = np.meshgrid(value_1_range, value_2_range)<br>\nresults = np.zeros_like(value_1_grid)</p>\n<h1>Go through every grid cell and check it's value</h1>\n<p>for i in range(value_1_grid.shape[0]):<br>\n    for j in range(value_1_grid.shape[1]):<br>\n        threshold = [value_1_grid[i, j], value_2_grid[i, j], 65]<br>\n        results[i, j] = evaluate_predictions(threshold, y_train_class, pred_light)</p>\n<h1>Heatmap plot</h1>\n<p>plt.figure(figsize=(14, 10))<br>\nplt.contourf(value_1_range, value_2_range, results, levels=30, cmap=\"viridis\")<br>\nplt.colorbar(label=\"Evaluation Metric\")<br>\nplt.title(\"Effect of Threshold on Predictions\")<br>\nplt.xlabel(\"Value 1 (Threshold 1)\")<br>\nplt.ylabel(\"Value 2 (Threshold 2)\")<br>\nplt.show()<br>\n`</p>\n<p>Edit: Weird, I'm posting it as code, but it doesn't take the inline spaces, which need to be there after every 'for'.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3074555": "Hello everyone,\nafter spending way too much time on this competition, I wondered, why eventhough my notebook seemed to improve in CV (stable with different seeds), it wouldn't translate to the leaderboard. \n\nI finally stumbled upon the crux (at least for me): The unstable QWK thresholds.\n\nI tested this by splitting the train csv in 3 parts. I trained on two parts and then tested the third part. Here is one of the results, note that the resolution is not very high. The swings are also high if you zoom in on parts.\n\noof predictions:\n![Train](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F5c2c2098b9ba42deed351bdb3a694151%2Flottery_train.PNG?generation=1734460508509910&alt=media)\n\ntest predictions:\n![Validation](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fa33b22724fe1b87ef68519666a2f1339%2Flottery_val.PNG?generation=1734460543582369&alt=media)\n\nI predict PCIAT-Total in the screenshots, but it is the same for sii. As you can see the optimal qwk thresholds are not the same anymore. This is one of the more extreme outcomes, but it changes depending on the split seed, which model I use, etc.\nEven if the maximums are not that far apart, it is still a huge qwk score swing and I can't find a way to stabilize it or work with it.\n\nIs there any way to fix this? I tried alot, but I assume it always depends on the target dataframe distribution?",
    "3074895": "There are some ways of reducing variance like training on multiple seeds and taking the average. Very nice visualization btw, can you share the code?",
    "3074899": "I'm taking the average of 15 seeds currently. I could probably take 100 seeds and it wouldnt change on the training data. But the test data will change, because it has its own distribution. I also made a post in your thread just now, because we probably have the same problem :)\n\nHere is the code:\n\ny_train_class are the 'sii' values and pred_light is the oof predictions _before_ you use qwk thresholds to classify on them.\nevaluate_predictions is the one everyone uses from the public notebook, it simply classifies the predictions with the given thresholds.\n\n`\n# Ranges for the threshold values 1 and 2\nvalue_1_range = np.linspace(25, 33, 100) \nvalue_2_range = np.linspace(35, 40, 100)  \n\n# Create a grid of all combinations\nvalue_1_grid, value_2_grid = np.meshgrid(value_1_range, value_2_range)\nresults = np.zeros_like(value_1_grid)\n\n# Go through every grid cell and check it's value\nfor i in range(value_1_grid.shape[0]):\n    for j in range(value_1_grid.shape[1]):\n        threshold = [value_1_grid[i, j], value_2_grid[i, j], 65]\n        results[i, j] = evaluate_predictions(threshold, y_train_class, pred_light)\n      \n# Heatmap plot\nplt.figure(figsize=(14, 10))\nplt.contourf(value_1_range, value_2_range, results, levels=30, cmap=\"viridis\")\nplt.colorbar(label=\"Evaluation Metric\")\nplt.title(\"Effect of Threshold on Predictions\")\nplt.xlabel(\"Value 1 (Threshold 1)\")\nplt.ylabel(\"Value 2 (Threshold 2)\")\nplt.show()\n`\n\nEdit: Weird, I'm posting it as code, but it doesn't take the inline spaces, which need to be there after every 'for'."
  },
  "source": "meta"
}