{
  "id": 552183,
  "title": "Does blending work for QWK?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552183",
  "author_name": "",
  "post_date": "2024-12-18T04:53:01.671441700Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In my experiments, it never works for QWK even though the quality of soft predictions is improving in terms of mean squared error and log loss. Does it happen to anyone else?</p>",
  "messages": [
    {
      "id": "3074795",
      "postDate": "12/18/2024 04:53:01",
      "content": "<p>In my experiments, it never works for QWK even though the quality of soft predictions is improving in terms of mean squared error and log loss. Does it happen to anyone else?</p>",
      "rawMarkdown": "In my experiments, it never works for QWK even though the quality of soft predictions is improving in terms of mean squared error and log loss. Does it happen to anyone else?",
      "votes": null
    },
    {
      "id": "3074813",
      "postDate": "12/18/2024 05:22:16",
      "content": "<p>With this noisy dataset, ensemble technique can both improve the QWK score and reduce its variances. </p>\n<p>However, even if ensemble techniques are useful for smoothing out residual noise, there is more than just residual noise to smooth out in this dataset; there is strong background noise that can only be addressed by cleaning the data. By doing so, the variance in the score will be significantly reduced.</p>",
      "rawMarkdown": "With this noisy dataset, ensemble technique can both improve the QWK score and reduce its variances. \n\nHowever, even if ensemble techniques are useful for smoothing out residual noise, there is more than just residual noise to smooth out in this dataset; there is strong background noise that can only be addressed by cleaning the data. By doing so, the variance in the score will be significantly reduced.",
      "votes": null
    },
    {
      "id": "3074834",
      "postDate": "12/18/2024 05:43:47",
      "content": "<p>With this noisy dataset, ensemble technique can both improve the QWK score and reduce its variances.</p>",
      "rawMarkdown": "With this noisy dataset, ensemble technique can both improve the QWK score and reduce its variances.",
      "votes": null
    },
    {
      "id": "3074839",
      "postDate": "12/18/2024 05:48:19",
      "content": "<p>I used weighted hard voting and it gave me some sort of improvement and stabilization.</p>\n<pre><code> model v1: cv=., lb=.\n model v2: cv=., lb=.\n</code></pre>\n<p>You can see there is still a huge discrepancy between cv and lb. I probed the leaderboard several times and found that there are some features different a lot from train dataset to the test dataset. But some of them have high importance in all my tree-based models. I also have no idea why the public notebook could achieve such a high score with these features. Maybe the knnimputer made it hit the jackpot.</p>",
      "rawMarkdown": "I used weighted hard voting and it gave me some sort of improvement and stabilization.\n```\nvoting model v1: cv=0.4826, lb=0.462\nvoting model v2: cv=0.4854, lb=0.464\n```\nYou can see there is still a huge discrepancy between cv and lb. I probed the leaderboard several times and found that there are some features different a lot from train dataset to the test dataset. But some of them have high importance in all my tree-based models. I also have no idea why the public notebook could achieve such a high score with these features. Maybe the knnimputer made it hit the jackpot.",
      "votes": null
    },
    {
      "id": "3074850",
      "postDate": "12/18/2024 06:00:04",
      "content": "<p>Thanks, I only tried averaging soft predictions. I'll try voting now.</p>\n<p>Edit: I tried hard voting (taking mode of multiple class predictions) and it worked better than blending soft predictions.</p>",
      "rawMarkdown": "Thanks, I only tried averaging soft predictions. I'll try voting now.\n\nEdit: I tried hard voting (taking mode of multiple class predictions) and it worked better than blending soft predictions.",
      "votes": null
    },
    {
      "id": "3074896",
      "postDate": "12/18/2024 07:19:18",
      "content": "<p>I only tried blending so far. I assumed it wouldn't make a difference if I choose the qwk threshold before or after ensambling.<br>\nDoes it matter? What are the advantages and disadvantages?<br>\nI already made another post about this problem I run into: The optimal qwk thresholds heavily rely on the dataframe we are looking at.<br>\nDo you think choosing qwk for the models first and ensemble them afterwards is better? I feel like it is random anyway, but I dont have any experience with qwk.<br>\nI feel like the advantage of blending is, that the optimal qwk has more spikes (more close to optimal qwk) across the spectrum. But the highest point is probably lower? Didn't test hard voting yet.</p>\n<p>Optimal QWK after blending on train data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fce13bdfdc65cadf22ebe21ad4dd98e9a%2Ftrain_class.PNG?generation=1734506107639004&amp;alt=media\" alt=\"train\"></p>\n<p>Optimal QWK after blending on test data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F0a795cee1a8a808b10da8a9b91bc4189%2Fpredict_class.PNG?generation=1734506160174985&amp;alt=media\" alt=\"test\"></p>\n<p>I currently have lb highscore of 0.476 with a random threshold choice (cv is arround 0.48x for that) and I could probably improve it by randomly choosing another threshold, but this will change with private LB anyway. I'm lost :)</p>",
      "rawMarkdown": "I only tried blending so far. I assumed it wouldn't make a difference if I choose the qwk threshold before or after ensambling.\nDoes it matter? What are the advantages and disadvantages?\nI already made another post about this problem I run into: The optimal qwk thresholds heavily rely on the dataframe we are looking at.\nDo you think choosing qwk for the models first and ensemble them afterwards is better? I feel like it is random anyway, but I dont have any experience with qwk.\nI feel like the advantage of blending is, that the optimal qwk has more spikes (more close to optimal qwk) across the spectrum. But the highest point is probably lower? Didn't test hard voting yet.\n\nOptimal QWK after blending on train data:\n![train](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fce13bdfdc65cadf22ebe21ad4dd98e9a%2Ftrain_class.PNG?generation=1734506107639004&alt=media)\n\nOptimal QWK after blending on test data:\n![test](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F0a795cee1a8a808b10da8a9b91bc4189%2Fpredict_class.PNG?generation=1734506160174985&alt=media)\n\nI currently have lb highscore of 0.476 with a random threshold choice (cv is arround 0.48x for that) and I could probably improve it by randomly choosing another threshold, but this will change with private LB anyway. I'm lost :)",
      "votes": null
    },
    {
      "id": "3075356",
      "postDate": "12/18/2024 18:20:30",
      "content": "<p>I tried it also using soft predicitons and it didn't work for me also. </p>",
      "rawMarkdown": "I tried it also using soft predicitons and it didn't work for me also.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3074813,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "12/18/2024 05:22:16",
      "content": "<p>With this noisy dataset, ensemble technique can both improve the QWK score and reduce its variances. </p>\n<p>However, even if ensemble techniques are useful for smoothing out residual noise, there is more than just residual noise to smooth out in this dataset; there is strong background noise that can only be addressed by cleaning the data. By doing so, the variance in the score will be significantly reduced.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3074834,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "12/18/2024 05:43:47",
      "content": "<p>With this noisy dataset, ensemble technique can both improve the QWK score and reduce its variances.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3074839,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "12/18/2024 05:48:19",
      "content": "<p>I used weighted hard voting and it gave me some sort of improvement and stabilization.</p>\n<pre><code> model v1: cv=., lb=.\n model v2: cv=., lb=.\n</code></pre>\n<p>You can see there is still a huge discrepancy between cv and lb. I probed the leaderboard several times and found that there are some features different a lot from train dataset to the test dataset. But some of them have high importance in all my tree-based models. I also have no idea why the public notebook could achieve such a high score with these features. Maybe the knnimputer made it hit the jackpot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3074850,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/18/2024 06:00:04",
          "content": "<p>Thanks, I only tried averaging soft predictions. I'll try voting now.</p>\n<p>Edit: I tried hard voting (taking mode of multiple class predictions) and it worked better than blending soft predictions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3074896,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "12/18/2024 07:19:18",
      "content": "<p>I only tried blending so far. I assumed it wouldn't make a difference if I choose the qwk threshold before or after ensambling.<br>\nDoes it matter? What are the advantages and disadvantages?<br>\nI already made another post about this problem I run into: The optimal qwk thresholds heavily rely on the dataframe we are looking at.<br>\nDo you think choosing qwk for the models first and ensemble them afterwards is better? I feel like it is random anyway, but I dont have any experience with qwk.<br>\nI feel like the advantage of blending is, that the optimal qwk has more spikes (more close to optimal qwk) across the spectrum. But the highest point is probably lower? Didn't test hard voting yet.</p>\n<p>Optimal QWK after blending on train data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fce13bdfdc65cadf22ebe21ad4dd98e9a%2Ftrain_class.PNG?generation=1734506107639004&amp;alt=media\" alt=\"train\"></p>\n<p>Optimal QWK after blending on test data:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F0a795cee1a8a808b10da8a9b91bc4189%2Fpredict_class.PNG?generation=1734506160174985&amp;alt=media\" alt=\"test\"></p>\n<p>I currently have lb highscore of 0.476 with a random threshold choice (cv is arround 0.48x for that) and I could probably improve it by randomly choosing another threshold, but this will change with private LB anyway. I'm lost :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3075356,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "12/18/2024 18:20:30",
      "content": "<p>I tried it also using soft predicitons and it didn't work for me also. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3074795": "In my experiments, it never works for QWK even though the quality of soft predictions is improving in terms of mean squared error and log loss. Does it happen to anyone else?",
    "3074813": "With this noisy dataset, ensemble technique can both improve the QWK score and reduce its variances. \n\nHowever, even if ensemble techniques are useful for smoothing out residual noise, there is more than just residual noise to smooth out in this dataset; there is strong background noise that can only be addressed by cleaning the data. By doing so, the variance in the score will be significantly reduced.",
    "3074834": "With this noisy dataset, ensemble technique can both improve the QWK score and reduce its variances.",
    "3074839": "I used weighted hard voting and it gave me some sort of improvement and stabilization.\n```\nvoting model v1: cv=0.4826, lb=0.462\nvoting model v2: cv=0.4854, lb=0.464\n```\nYou can see there is still a huge discrepancy between cv and lb. I probed the leaderboard several times and found that there are some features different a lot from train dataset to the test dataset. But some of them have high importance in all my tree-based models. I also have no idea why the public notebook could achieve such a high score with these features. Maybe the knnimputer made it hit the jackpot.",
    "3074850": "Thanks, I only tried averaging soft predictions. I'll try voting now.\n\nEdit: I tried hard voting (taking mode of multiple class predictions) and it worked better than blending soft predictions.",
    "3074896": "I only tried blending so far. I assumed it wouldn't make a difference if I choose the qwk threshold before or after ensambling.\nDoes it matter? What are the advantages and disadvantages?\nI already made another post about this problem I run into: The optimal qwk thresholds heavily rely on the dataframe we are looking at.\nDo you think choosing qwk for the models first and ensemble them afterwards is better? I feel like it is random anyway, but I dont have any experience with qwk.\nI feel like the advantage of blending is, that the optimal qwk has more spikes (more close to optimal qwk) across the spectrum. But the highest point is probably lower? Didn't test hard voting yet.\n\nOptimal QWK after blending on train data:\n![train](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2Fce13bdfdc65cadf22ebe21ad4dd98e9a%2Ftrain_class.PNG?generation=1734506107639004&alt=media)\n\nOptimal QWK after blending on test data:\n![test](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20325352%2F0a795cee1a8a808b10da8a9b91bc4189%2Fpredict_class.PNG?generation=1734506160174985&alt=media)\n\nI currently have lb highscore of 0.476 with a random threshold choice (cv is arround 0.48x for that) and I could probably improve it by randomly choosing another threshold, but this will change with private LB anyway. I'm lost :)",
    "3075356": "I tried it also using soft predicitons and it didn't work for me also."
  },
  "source": "meta"
}