{
  "id": 540738,
  "title": "Optimized QWK effect on the LB score",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/540738",
  "author_name": "",
  "post_date": "2024-10-15T21:32:11.660127900Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I was asking myself - the optimized QWK is not just an overfitting over the validation predictions? Usually a stable model have a CV score a little higher than the LB score, but people here have everything inversed with much higher LB score than the CV untuned score, so it smells of overfitting the LB based on the tuned thresholds, or do you have other experiences where this actually worked for the private LB too?</p>",
  "messages": [
    {
      "id": "3018590",
      "postDate": "10/15/2024 21:32:11",
      "content": "<p>I was asking myself - the optimized QWK is not just an overfitting over the validation predictions? Usually a stable model have a CV score a little higher than the LB score, but people here have everything inversed with much higher LB score than the CV untuned score, so it smells of overfitting the LB based on the tuned thresholds, or do you have other experiences where this actually worked for the private LB too?</p>",
      "rawMarkdown": "I was asking myself - the optimized QWK is not just an overfitting over the validation predictions? Usually a stable model have a CV score a little higher than the LB score, but people here have everything inversed with much higher LB score than the CV untuned score, so it smells of overfitting the LB based on the tuned thresholds, or do you have other experiences where this actually worked for the private LB too?",
      "votes": null
    },
    {
      "id": "3018790",
      "postDate": "10/16/2024 03:44:22",
      "content": "<p>well thats the big question.</p>",
      "rawMarkdown": "well thats the big question.",
      "votes": null
    },
    {
      "id": "3018899",
      "postDate": "10/16/2024 05:45:21",
      "content": "<p>Whether the threshold optimization works for private LB as well really depends on how the public and private test sets are split. From my experience, for competitions like this one, more often than not the threshold optimization doesn't work.</p>\n<p>In fact most people are overfitting with gradient boost ensembles. This reminds me of the ICR competition, and the shakeup was huge.</p>",
      "rawMarkdown": "Whether the threshold optimization works for private LB as well really depends on how the public and private test sets are split. From my experience, for competitions like this one, more often than not the threshold optimization doesn't work.\n\nIn fact most people are overfitting with gradient boost ensembles. This reminds me of the ICR competition, and the shakeup was huge.",
      "votes": null
    },
    {
      "id": "3018916",
      "postDate": "10/16/2024 06:08:18",
      "content": "<p>What actually worked in that competition?</p>",
      "rawMarkdown": "What actually worked in that competition?",
      "votes": null
    },
    {
      "id": "3018921",
      "postDate": "10/16/2024 06:21:23",
      "content": "<p>I agree with you, and there are people who on top of this tuning are applying magic coefficients other the results to squeeze every 0.001 of the point on the public LB and are angry on me when I'm telling that they are overfitting on the LB score, but everybody with it's own strategy :)</p>",
      "rawMarkdown": "I agree with you, and there are people who on top of this tuning are applying magic coefficients other the results to squeeze every 0.001 of the point on the public LB and are angry on me when I'm telling that they are overfitting on the LB score, but everybody with it's own strategy :)",
      "votes": null
    },
    {
      "id": "3018929",
      "postDate": "10/16/2024 06:34:13",
      "content": "<p>I largely agree and I only look at the raw CV-score - the optimized score is just good to know for me. I don't completely ignore the LB-score, as it's still the score of roughly 1400 samples my model wasn't trained on, but it's way too easy to overfit the public leaderboard with high variance models and some luck.</p>\n<p>(But it's also fun to exploit the randomness and overfit the leaderboard with some of the submissions.)</p>",
      "rawMarkdown": "I largely agree and I only look at the raw CV-score - the optimized score is just good to know for me. I don't completely ignore the LB-score, as it's still the score of roughly 1400 samples my model wasn't trained on, but it's way too easy to overfit the public leaderboard with high variance models and some luck.\n\n(But it's also fun to exploit the randomness and overfit the leaderboard with some of the submissions.)",
      "votes": null
    },
    {
      "id": "3018931",
      "postDate": "10/16/2024 06:38:08",
      "content": "<p>In ICR overfitting notebooks droped significally. I had a far-away score on the public LB, but was shoked with almost +3000 positions up in the private LB, my notebook is <a href=\"https://www.kaggle.com/code/eu1234/identifying-age-related-conditions\" target=\"_blank\">here </a> if of any interest </p>",
      "rawMarkdown": "In ICR overfitting notebooks droped significally. I had a far-away score on the public LB, but was shoked with almost +3000 positions up in the private LB, my notebook is [here ](https://www.kaggle.com/code/eu1234/identifying-age-related-conditions) if of any interest",
      "votes": null
    },
    {
      "id": "3018985",
      "postDate": "10/16/2024 07:30:16",
      "content": "<p>In ICR, a Catboost model trained with DEFAULT parameters was better than most contenders' solutions (LGBM &amp; XGB ensembles mostly). The winning solution was a DNN with variational feature selection, combined with a tricky CV scheme.</p>",
      "rawMarkdown": "In ICR, a Catboost model trained with DEFAULT parameters was better than most contenders' solutions (LGBM & XGB ensembles mostly). The winning solution was a DNN with variational feature selection, combined with a tricky CV scheme.",
      "votes": null
    },
    {
      "id": "3019107",
      "postDate": "10/16/2024 09:10:19",
      "content": "<p>When optimized QWK was generated, model didn't read the test set, it's optimized based on training set. So it can't overfit the LB?<br>\nMagic parameter is another issue. It's trying to overfit the LB.</p>\n<p>Here's my 2 best set (validation QWK -&gt; Optimized QWK -&gt; LB score)<br>\n0.37 --&gt; 0.43 --&gt; 0.46<br>\n0.38 --&gt; 0.45 --&gt; 0.46</p>\n<p>And here's my creative but bad model<br>\n0.55 --&gt; 0.6 --&gt; 0.38<br>\n0.93 --&gt; 0.93 --&gt; 0.33</p>\n<p>My question is how can we know if we're improving a model when we do feature engineering. validation score usually improves after feature engineering because we're making the train data look in the way we want.</p>",
      "rawMarkdown": "When optimized QWK was generated, model didn't read the test set, it's optimized based on training set. So it can't overfit the LB?\nMagic parameter is another issue. It's trying to overfit the LB.\n\nHere's my 2 best set (validation QWK -> Optimized QWK -> LB score)\n0.37 --> 0.43 --> 0.46\n0.38 --> 0.45 --> 0.46\n\nAnd here's my creative but bad model\n0.55 --> 0.6 --> 0.38\n0.93 --> 0.93 --> 0.33\n\nMy question is how can we know if we're improving a model when we do feature engineering. validation score usually improves after feature engineering because we're making the train data look in the way we want.",
      "votes": null
    },
    {
      "id": "3019130",
      "postDate": "10/16/2024 09:35:10",
      "content": "<p>If CV and LB scores rise hand-by-hand than the model is on the right track of improving, but a LB score much higher than CV just doesn't sound right for me</p>",
      "rawMarkdown": "If CV and LB scores rise hand-by-hand than the model is on the right track of improving, but a LB score much higher than CV just doesn't sound right for me",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3018790,
      "author_name": "chanpreetsingh07",
      "author_url": "",
      "post_date": "10/16/2024 03:44:22",
      "content": "<p>well thats the big question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3018899,
      "author_name": "passengerc07",
      "author_url": "",
      "post_date": "10/16/2024 05:45:21",
      "content": "<p>Whether the threshold optimization works for private LB as well really depends on how the public and private test sets are split. From my experience, for competitions like this one, more often than not the threshold optimization doesn't work.</p>\n<p>In fact most people are overfitting with gradient boost ensembles. This reminds me of the ICR competition, and the shakeup was huge.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3018916,
          "author_name": "chanpreetsingh07",
          "author_url": "",
          "post_date": "10/16/2024 06:08:18",
          "content": "<p>What actually worked in that competition?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3018931,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "10/16/2024 06:38:08",
              "content": "<p>In ICR overfitting notebooks droped significally. I had a far-away score on the public LB, but was shoked with almost +3000 positions up in the private LB, my notebook is <a href=\"https://www.kaggle.com/code/eu1234/identifying-age-related-conditions\" target=\"_blank\">here </a> if of any interest </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3018985,
              "author_name": "passengerc07",
              "author_url": "",
              "post_date": "10/16/2024 07:30:16",
              "content": "<p>In ICR, a Catboost model trained with DEFAULT parameters was better than most contenders' solutions (LGBM &amp; XGB ensembles mostly). The winning solution was a DNN with variational feature selection, combined with a tricky CV scheme.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3018921,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "10/16/2024 06:21:23",
          "content": "<p>I agree with you, and there are people who on top of this tuning are applying magic coefficients other the results to squeeze every 0.001 of the point on the public LB and are angry on me when I'm telling that they are overfitting on the LB score, but everybody with it's own strategy :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3018929,
      "author_name": "lennarthaupts",
      "author_url": "",
      "post_date": "10/16/2024 06:34:13",
      "content": "<p>I largely agree and I only look at the raw CV-score - the optimized score is just good to know for me. I don't completely ignore the LB-score, as it's still the score of roughly 1400 samples my model wasn't trained on, but it's way too easy to overfit the public leaderboard with high variance models and some luck.</p>\n<p>(But it's also fun to exploit the randomness and overfit the leaderboard with some of the submissions.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3019107,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "10/16/2024 09:10:19",
      "content": "<p>When optimized QWK was generated, model didn't read the test set, it's optimized based on training set. So it can't overfit the LB?<br>\nMagic parameter is another issue. It's trying to overfit the LB.</p>\n<p>Here's my 2 best set (validation QWK -&gt; Optimized QWK -&gt; LB score)<br>\n0.37 --&gt; 0.43 --&gt; 0.46<br>\n0.38 --&gt; 0.45 --&gt; 0.46</p>\n<p>And here's my creative but bad model<br>\n0.55 --&gt; 0.6 --&gt; 0.38<br>\n0.93 --&gt; 0.93 --&gt; 0.33</p>\n<p>My question is how can we know if we're improving a model when we do feature engineering. validation score usually improves after feature engineering because we're making the train data look in the way we want.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3019130,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "10/16/2024 09:35:10",
          "content": "<p>If CV and LB scores rise hand-by-hand than the model is on the right track of improving, but a LB score much higher than CV just doesn't sound right for me</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3018590": "I was asking myself - the optimized QWK is not just an overfitting over the validation predictions? Usually a stable model have a CV score a little higher than the LB score, but people here have everything inversed with much higher LB score than the CV untuned score, so it smells of overfitting the LB based on the tuned thresholds, or do you have other experiences where this actually worked for the private LB too?",
    "3018790": "well thats the big question.",
    "3018899": "Whether the threshold optimization works for private LB as well really depends on how the public and private test sets are split. From my experience, for competitions like this one, more often than not the threshold optimization doesn't work.\n\nIn fact most people are overfitting with gradient boost ensembles. This reminds me of the ICR competition, and the shakeup was huge.",
    "3018916": "What actually worked in that competition?",
    "3018921": "I agree with you, and there are people who on top of this tuning are applying magic coefficients other the results to squeeze every 0.001 of the point on the public LB and are angry on me when I'm telling that they are overfitting on the LB score, but everybody with it's own strategy :)",
    "3018929": "I largely agree and I only look at the raw CV-score - the optimized score is just good to know for me. I don't completely ignore the LB-score, as it's still the score of roughly 1400 samples my model wasn't trained on, but it's way too easy to overfit the public leaderboard with high variance models and some luck.\n\n(But it's also fun to exploit the randomness and overfit the leaderboard with some of the submissions.)",
    "3018931": "In ICR overfitting notebooks droped significally. I had a far-away score on the public LB, but was shoked with almost +3000 positions up in the private LB, my notebook is [here ](https://www.kaggle.com/code/eu1234/identifying-age-related-conditions) if of any interest",
    "3018985": "In ICR, a Catboost model trained with DEFAULT parameters was better than most contenders' solutions (LGBM & XGB ensembles mostly). The winning solution was a DNN with variational feature selection, combined with a tricky CV scheme.",
    "3019107": "When optimized QWK was generated, model didn't read the test set, it's optimized based on training set. So it can't overfit the LB?\nMagic parameter is another issue. It's trying to overfit the LB.\n\nHere's my 2 best set (validation QWK -> Optimized QWK -> LB score)\n0.37 --> 0.43 --> 0.46\n0.38 --> 0.45 --> 0.46\n\nAnd here's my creative but bad model\n0.55 --> 0.6 --> 0.38\n0.93 --> 0.93 --> 0.33\n\nMy question is how can we know if we're improving a model when we do feature engineering. validation score usually improves after feature engineering because we're making the train data look in the way we want.",
    "3019130": "If CV and LB scores rise hand-by-hand than the model is on the right track of improving, but a LB score much higher than CV just doesn't sound right for me"
  },
  "source": "meta"
}