{
  "id": 543940,
  "title": "QWK vs LB with Automated Machine Learning",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/543940",
  "author_name": "",
  "post_date": "2024-11-02T11:09:12.672998300Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello </p>\n<p>Reference notebook: <a href=\"https://www.kaggle.com/code/taimour/automl-h2o-feature-engineering-piu\" target=\"_blank\">🌐 AutoML H2O - ⚙️ Feature Engineering - 📱 PIU</a></p>\n<table>\n<thead>\n<tr>\n<th>Version</th>\n<th>QWK</th>\n<th>Leader-board</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>32</td>\n<td>0.633</td>\n<td>0.464</td>\n</tr>\n<tr>\n<td>44</td>\n<td>0.848</td>\n<td>0.408</td>\n</tr>\n<tr>\n<td>45</td>\n<td>0.755</td>\n<td>0.423</td>\n</tr>\n<tr>\n<td>46</td>\n<td>0.800</td>\n<td>0.408</td>\n</tr>\n</tbody>\n</table>\n<p>I have observed in this notebook (in some others also) that when QWK starts to increase then score on the leader-board starts to decrease. Have you people also observed this in your notebooks? What are your views on it?</p>\n<p><em>Note: QWK has already been updated in this notebook to work with Automated machine learning algorithm H2O. There is no problem in the update of QWK, I am sure about it. In case you want to understand it, I have provided indepth details for it in the notebook.</em></p>",
  "messages": [
    {
      "id": "3034596",
      "postDate": "11/02/2024 11:09:12",
      "content": "<p>Hello </p>\n<p>Reference notebook: <a href=\"https://www.kaggle.com/code/taimour/automl-h2o-feature-engineering-piu\" target=\"_blank\">🌐 AutoML H2O - ⚙️ Feature Engineering - 📱 PIU</a></p>\n<table>\n<thead>\n<tr>\n<th>Version</th>\n<th>QWK</th>\n<th>Leader-board</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>32</td>\n<td>0.633</td>\n<td>0.464</td>\n</tr>\n<tr>\n<td>44</td>\n<td>0.848</td>\n<td>0.408</td>\n</tr>\n<tr>\n<td>45</td>\n<td>0.755</td>\n<td>0.423</td>\n</tr>\n<tr>\n<td>46</td>\n<td>0.800</td>\n<td>0.408</td>\n</tr>\n</tbody>\n</table>\n<p>I have observed in this notebook (in some others also) that when QWK starts to increase then score on the leader-board starts to decrease. Have you people also observed this in your notebooks? What are your views on it?</p>\n<p><em>Note: QWK has already been updated in this notebook to work with Automated machine learning algorithm H2O. There is no problem in the update of QWK, I am sure about it. In case you want to understand it, I have provided indepth details for it in the notebook.</em></p>",
      "rawMarkdown": "Hello \n\nReference notebook: [🌐 AutoML H2O - ⚙️ Feature Engineering - 📱 PIU](https://www.kaggle.com/code/taimour/automl-h2o-feature-engineering-piu)\n\n| Version | QWK | Leader-board |\n| --- | --- | --- |\n| 32 | 0.633 | 0.464 |\n| 44 | 0.848 | 0.408 |\n| 45 | 0.755 | 0.423 |\n| 46 | 0.800 | 0.408 |\n\nI have observed in this notebook (in some others also) that when QWK starts to increase then score on the leader-board starts to decrease. Have you people also observed this in your notebooks? What are your views on it?\n\n*Note: QWK has already been updated in this notebook to work with Automated machine learning algorithm H2O. There is no problem in the update of QWK, I am sure about it. In case you want to understand it, I have provided indepth details for it in the notebook.*",
      "votes": null
    },
    {
      "id": "3035817",
      "postDate": "11/03/2024 21:38:05",
      "content": "<p>Compared V32 and V44, V44 has used the KNN imputer.<br>\nI don't quite believe using in KNN imputer in this competition yet, may be that's why I keep staying in rank 900+.<br>\nIf you're filling missed sii in train set by KNN, I'm not surprised model can learn to predict within the \"new\" train set, which will derive a high QWK in cv, but this high score won't generalize to test set.<br>\nThis one is very tricky. Once we twist the train data, in particular filling the missed sii, the cv score can no longer be trusted, because we are filling the sii in the way we want.</p>",
      "rawMarkdown": "Compared V32 and V44, V44 has used the KNN imputer.\nI don't quite believe using in KNN imputer in this competition yet, may be that's why I keep staying in rank 900+.\nIf you're filling missed sii in train set by KNN, I'm not surprised model can learn to predict within the \"new\" train set, which will derive a high QWK in cv, but this high score won't generalize to test set.\nThis one is very tricky. Once we twist the train data, in particular filling the missed sii, the cv score can no longer be trusted, because we are filling the sii in the way we want.",
      "votes": null
    },
    {
      "id": "3036229",
      "postDate": "11/04/2024 11:34:13",
      "content": "<p>Filled nan sii can be used only in training dataset, because they are much easier to predict.<br>\nToo high scores may indicate a train-validation leak. </p>",
      "rawMarkdown": "Filled nan sii can be used only in training dataset, because they are much easier to predict.\nToo high scores may indicate a train-validation leak.",
      "votes": null
    },
    {
      "id": "3036518",
      "postDate": "11/04/2024 16:20:18",
      "content": "<p><a href=\"https://www.kaggle.com/taimour\" target=\"_blank\">@taimour</a> 'that when QWK starts to increase then score on the leader-board starts to decrease', that is means is means the your model is start to overfit the test data …</p>",
      "rawMarkdown": "taimour 'that when QWK starts to increase then score on the leader-board starts to decrease', that is means is means the your model is start to overfit the test data ...",
      "votes": null
    },
    {
      "id": "3037290",
      "postDate": "11/05/2024 15:10:24",
      "content": "<p>It's just very overfitted and the problem could be in a leakage, maybe you impute some data based on your targets and then the model just take these as a clear prediction signal.</p>",
      "rawMarkdown": "It's just very overfitted and the problem could be in a leakage, maybe you impute some data based on your targets and then the model just take these as a clear prediction signal.",
      "votes": null
    },
    {
      "id": "3038203",
      "postDate": "11/06/2024 17:45:07",
      "content": "<p>Yeah! actually I removed KNNImputer in new versions of this note book. I was just testing around with different techniques and trying to find what will be the result. Ideally KNNImputers shouldn't be used for target variables, I agree with that.</p>",
      "rawMarkdown": "Yeah! actually I removed KNNImputer in new versions of this note book. I was just testing around with different techniques and trying to find what will be the result. Ideally KNNImputers shouldn't be used for target variables, I agree with that.",
      "votes": null
    },
    {
      "id": "3038204",
      "postDate": "11/06/2024 17:45:26",
      "content": "<p>yes correct, I made changes in new versions.</p>",
      "rawMarkdown": "yes correct, I made changes in new versions.",
      "votes": null
    },
    {
      "id": "3038220",
      "postDate": "11/06/2024 18:07:17",
      "content": "<p>Yes, that's right! thank you. I have already made changes in new version to address this issue.</p>",
      "rawMarkdown": "Yes, that's right! thank you. I have already made changes in new version to address this issue.",
      "votes": null
    },
    {
      "id": "3038221",
      "postDate": "11/06/2024 18:07:47",
      "content": "<p>Yeah, right. I have made changes in new versions</p>",
      "rawMarkdown": "Yeah, right. I have made changes in new versions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3035817,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "11/03/2024 21:38:05",
      "content": "<p>Compared V32 and V44, V44 has used the KNN imputer.<br>\nI don't quite believe using in KNN imputer in this competition yet, may be that's why I keep staying in rank 900+.<br>\nIf you're filling missed sii in train set by KNN, I'm not surprised model can learn to predict within the \"new\" train set, which will derive a high QWK in cv, but this high score won't generalize to test set.<br>\nThis one is very tricky. Once we twist the train data, in particular filling the missed sii, the cv score can no longer be trusted, because we are filling the sii in the way we want.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3036229,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "11/04/2024 11:34:13",
          "content": "<p>Filled nan sii can be used only in training dataset, because they are much easier to predict.<br>\nToo high scores may indicate a train-validation leak. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3038204,
              "author_name": "taimour",
              "author_url": "",
              "post_date": "11/06/2024 17:45:26",
              "content": "<p>yes correct, I made changes in new versions.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3038203,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "11/06/2024 17:45:07",
          "content": "<p>Yeah! actually I removed KNNImputer in new versions of this note book. I was just testing around with different techniques and trying to find what will be the result. Ideally KNNImputers shouldn't be used for target variables, I agree with that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3036518,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "11/04/2024 16:20:18",
      "content": "<p><a href=\"https://www.kaggle.com/taimour\" target=\"_blank\">@taimour</a> 'that when QWK starts to increase then score on the leader-board starts to decrease', that is means is means the your model is start to overfit the test data …</p>",
      "votes": null,
      "replies": [
        {
          "id": 3038220,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "11/06/2024 18:07:17",
          "content": "<p>Yes, that's right! thank you. I have already made changes in new version to address this issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3037290,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "11/05/2024 15:10:24",
      "content": "<p>It's just very overfitted and the problem could be in a leakage, maybe you impute some data based on your targets and then the model just take these as a clear prediction signal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3038221,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "11/06/2024 18:07:47",
          "content": "<p>Yeah, right. I have made changes in new versions</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3034596": "Hello \n\nReference notebook: [🌐 AutoML H2O - ⚙️ Feature Engineering - 📱 PIU](https://www.kaggle.com/code/taimour/automl-h2o-feature-engineering-piu)\n\n| Version | QWK | Leader-board |\n| --- | --- | --- |\n| 32 | 0.633 | 0.464 |\n| 44 | 0.848 | 0.408 |\n| 45 | 0.755 | 0.423 |\n| 46 | 0.800 | 0.408 |\n\nI have observed in this notebook (in some others also) that when QWK starts to increase then score on the leader-board starts to decrease. Have you people also observed this in your notebooks? What are your views on it?\n\n*Note: QWK has already been updated in this notebook to work with Automated machine learning algorithm H2O. There is no problem in the update of QWK, I am sure about it. In case you want to understand it, I have provided indepth details for it in the notebook.*",
    "3035817": "Compared V32 and V44, V44 has used the KNN imputer.\nI don't quite believe using in KNN imputer in this competition yet, may be that's why I keep staying in rank 900+.\nIf you're filling missed sii in train set by KNN, I'm not surprised model can learn to predict within the \"new\" train set, which will derive a high QWK in cv, but this high score won't generalize to test set.\nThis one is very tricky. Once we twist the train data, in particular filling the missed sii, the cv score can no longer be trusted, because we are filling the sii in the way we want.",
    "3036229": "Filled nan sii can be used only in training dataset, because they are much easier to predict.\nToo high scores may indicate a train-validation leak.",
    "3036518": "taimour 'that when QWK starts to increase then score on the leader-board starts to decrease', that is means is means the your model is start to overfit the test data ...",
    "3037290": "It's just very overfitted and the problem could be in a leakage, maybe you impute some data based on your targets and then the model just take these as a clear prediction signal.",
    "3038203": "Yeah! actually I removed KNNImputer in new versions of this note book. I was just testing around with different techniques and trying to find what will be the result. Ideally KNNImputers shouldn't be used for target variables, I agree with that.",
    "3038204": "yes correct, I made changes in new versions.",
    "3038220": "Yes, that's right! thank you. I have already made changes in new version to address this issue.",
    "3038221": "Yeah, right. I have made changes in new versions"
  },
  "source": "meta"
}