{
  "id": 550035,
  "title": "A lottery game? The final score?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/550035",
  "author_name": "Reki",
  "post_date": "2024-12-05T06:04:42.199000",
  "votes": 2,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Run 1 - Seed 42 - QWK: 0.5186<br>\nRun 2 - Seed 52 - QWK: 0.5517<br>\nRun 3 - Seed 62 - QWK: 0.5165<br>\nRun 4 - Seed 72 - QWK: 0.5334<br>\nRun 5 - Seed 82 - QWK: 0.5298<br>\nRun 6 - Seed 92 - QWK: 0.5308<br>\nRun 7 - Seed 102 - QWK: 0.5432<br>\nRun 8 - Seed 112 - QWK: 0.5274<br>\nRun 9 - Seed 122 - QWK: 0.5276<br>\nRun 10 - Seed 132 - QWK: 0.5436</p>\n<p>----&gt; || Optimized QWK SCORE ::  0.607</p>\n<p>I feel that there is no overfitting in each of my twists, because the QWK score of my model on train is 0.63, and the average score of my verification set is 0.53, which can reach 0.607 after threshold adjustment, but the submission of this version only has a score of 0.378 on LB</p>\n<p>I am now stuck on how to choose to submit</p>",
  "messages": [
    {
      "id": 3071881,
      "postDate": "2024-12-14T12:42:39.757Z",
      "content": "<p>i work alot on this during first month of competition and after mang exp , The Good Cv Perform Worse on LB .. I did't know what the main issue and what to chose for final submision. <br>\nThis is Totally luck game. <br>\nThe Better Preprocessing the worse Lb result </p>",
      "rawMarkdown": " i work alot on this during first month of competition and after mang exp , The Good Cv Perform Worse on LB .. I did't know what the main issue and what to chose for final submision. \nThis is Totally luck game. \nThe Better Preprocessing the worse Lb result ",
      "votes": 1,
      "replies": [
        {
          "id": 3071929,
          "postDate": "2024-12-14T14:27:17.280Z",
          "content": "<p>I have a local CV0.59--LB0.415, then I went through some parameter optimization, let the score become CV0.49--LB0.441, I don't know if it is worth it?</p>",
          "rawMarkdown": "I have a local CV0.59--LB0.415, then I went through some parameter optimization, let the score become CV0.49--LB0.441, I don't know if it is worth it?",
          "votes": 1
        }
      ]
    },
    {
      "id": 3064018,
      "postDate": "2024-12-05T06:05:14.733Z",
      "content": "<p>I don't know if this is overfitting, any local CV results don't show that my model is overfitting</p>",
      "rawMarkdown": "I don't know if this is overfitting, any local CV results don't show that my model is overfitting",
      "votes": 1,
      "replies": [
        {
          "id": 3069949,
          "postDate": "2024-12-12T04:47:56.797Z",
          "content": "<p>Same here - I have CV-LB relation now after a lot of work. Hope it prevails on the private LB <a href=\"https://www.kaggle.com/ruichardliu\" target=\"_blank\">@ruichardliu</a> </p>",
          "rawMarkdown": "Same here - I have CV-LB relation now after a lot of work. Hope it prevails on the private LB @ruichardliu ",
          "votes": 2,
          "replies": [
            {
              "id": 3071883,
              "postDate": "2024-12-14T12:43:46.183Z",
              "content": "<p>I have also one Single model which has Optimized Score 0.452 and LB score of 0.450. <br>\nAfter many experiments. and I am Taking risk to chose it as final. </p>",
              "rawMarkdown": "I have also one Single model which has Optimized Score 0.452 and LB score of 0.450. \nAfter many experiments. and I am Taking risk to chose it as final. ",
              "votes": 1
            },
            {
              "id": 3071928,
              "postDate": "2024-12-14T14:25:45.850Z",
              "content": "<p>I will most likely choose submissions with similar scores, and when LB exceeds 0.46, the local CV score will start to decrease</p>",
              "rawMarkdown": "I will most likely choose submissions with similar scores, and when LB exceeds 0.46, the local CV score will start to decrease"
            }
          ]
        }
      ]
    },
    {
      "id": 3064017,
      "postDate": "2024-12-05T06:04:42.200Z",
      "content": "<p>Run 1 - Seed 42 - QWK: 0.5186<br>\nRun 2 - Seed 52 - QWK: 0.5517<br>\nRun 3 - Seed 62 - QWK: 0.5165<br>\nRun 4 - Seed 72 - QWK: 0.5334<br>\nRun 5 - Seed 82 - QWK: 0.5298<br>\nRun 6 - Seed 92 - QWK: 0.5308<br>\nRun 7 - Seed 102 - QWK: 0.5432<br>\nRun 8 - Seed 112 - QWK: 0.5274<br>\nRun 9 - Seed 122 - QWK: 0.5276<br>\nRun 10 - Seed 132 - QWK: 0.5436</p>\n<p>----&gt; || Optimized QWK SCORE ::  0.607</p>\n<p>I feel that there is no overfitting in each of my twists, because the QWK score of my model on train is 0.63, and the average score of my verification set is 0.53, which can reach 0.607 after threshold adjustment, but the submission of this version only has a score of 0.378 on LB</p>\n<p>I am now stuck on how to choose to submit</p>",
      "rawMarkdown": "Run 1 - Seed 42 - QWK: 0.5186\nRun 2 - Seed 52 - QWK: 0.5517\nRun 3 - Seed 62 - QWK: 0.5165\nRun 4 - Seed 72 - QWK: 0.5334\nRun 5 - Seed 82 - QWK: 0.5298\nRun 6 - Seed 92 - QWK: 0.5308\nRun 7 - Seed 102 - QWK: 0.5432\nRun 8 - Seed 112 - QWK: 0.5274\nRun 9 - Seed 122 - QWK: 0.5276\nRun 10 - Seed 132 - QWK: 0.5436\n\n----> || Optimized QWK SCORE ::  0.607\n\nI feel that there is no overfitting in each of my twists, because the QWK score of my model on train is 0.63, and the average score of my verification set is 0.53, which can reach 0.607 after threshold adjustment, but the submission of this version only has a score of 0.378 on LB\n\nI am now stuck on how to choose to submit",
      "votes": 2
    },
    {
      "id": 3065253,
      "postDate": "2024-12-06T15:27:22.423Z",
      "content": "<p>I believe that due to small number of samples compared to the number of features available, we do not have enough data to be able to generalize or characterize the distribution of the train. This leads to problems with generalizations in the actual test dataset.</p>",
      "rawMarkdown": "I believe that due to small number of samples compared to the number of features available, we do not have enough data to be able to generalize or characterize the distribution of the train. This leads to problems with generalizations in the actual test dataset."
    },
    {
      "id": 3064104,
      "postDate": "2024-12-05T08:31:46.420Z",
      "content": "<p>I agree, that this comp results will heavily rely on luck in the end, but there are probably still things you can do to improve your chances. At least that's what I'm trying to do aswell :)</p>\n<p>Your case seems weird to me. Did you use KNN Imputer on the target? Because my results haven't been this far apart (from CV to LB).<br>\nThen it is simply explained by data leakage</p>",
      "rawMarkdown": "I agree, that this comp results will heavily rely on luck in the end, but there are probably still things you can do to improve your chances. At least that's what I'm trying to do aswell :)\n\nYour case seems weird to me. Did you use KNN Imputer on the target? Because my results haven't been this far apart (from CV to LB).\nThen it is simply explained by data leakage",
      "replies": [
        {
          "id": 3064185,
          "postDate": "2024-12-05T10:40:04.663Z",
          "content": "<p>I did not use KNN-related filling, because I think the quality of this data set is not high, so the effect of KNN is not suitable for me. As for data leakage, I am checking this problem now, but I did not find out. I will share my notebook after the competition</p>",
          "rawMarkdown": "I did not use KNN-related filling, because I think the quality of this data set is not high, so the effect of KNN is not suitable for me. As for data leakage, I am checking this problem now, but I did not find out. I will share my notebook after the competition"
        },
        {
          "id": 3064262,
          "postDate": "2024-12-05T12:46:10.913Z",
          "content": "<p>I also used some clustering methods, CV,:5.1,LB:0.40, after a certain score, the value of CV and LB will become inversely proportional</p>",
          "rawMarkdown": "I also used some clustering methods, CV,:5.1,LB:0.40, after a certain score, the value of CV and LB will become inversely proportional"
        }
      ]
    },
    {
      "id": 3064019,
      "postDate": "2024-12-05T06:05:58.320Z",
      "content": "<p>I think it's safe to say that the data sets we're working with are vastly different from the ones on the leaderboard</p>",
      "rawMarkdown": "I think it's safe to say that the data sets we're working with are vastly different from the ones on the leaderboard",
      "replies": [
        {
          "id": 3064022,
          "postDate": "2024-12-05T06:10:30.253Z",
          "content": "<p>In such cases, it is not known which pretreatment is suitable, and any pretreatment may shake down under the final leaderboard</p>",
          "rawMarkdown": "In such cases, it is not known which pretreatment is suitable, and any pretreatment may shake down under the final leaderboard",
          "votes": 2,
          "replies": [
            {
              "id": 3071925,
              "postDate": "2024-12-14T14:22:23.827Z",
              "content": "<p><a href=\"https://www.kaggle.com/ruichardliu\" target=\"_blank\">@ruichardliu</a> agreed, this is a 100% luck game!</p>",
              "rawMarkdown": "@ruichardliu agreed, this is a 100% luck game!"
            }
          ]
        }
      ]
    },
    {
      "id": 3064402,
      "postDate": "2024-12-05T15:23:12.603Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 3064493,
          "postDate": "2024-12-05T16:56:21.680Z",
          "content": "<p>Yep，A lottery game </p>",
          "rawMarkdown": "Yep，A lottery game "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3071881,
      "author_name": "Sheikh Muhammad Abdullah",
      "author_url": "",
      "post_date": "2024-12-14T12:42:39.757000",
      "content": "<p>i work alot on this during first month of competition and after mang exp , The Good Cv Perform Worse on LB .. I did't know what the main issue and what to chose for final submision. <br>\nThis is Totally luck game. <br>\nThe Better Preprocessing the worse Lb result </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3071929,
          "author_name": "Reki",
          "author_url": "",
          "post_date": "2024-12-14T14:27:17.280000",
          "content": "<p>I have a local CV0.59--LB0.415, then I went through some parameter optimization, let the score become CV0.49--LB0.441, I don't know if it is worth it?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3064018,
      "author_name": "Reki",
      "author_url": "",
      "post_date": "2024-12-05T06:05:14.733000",
      "content": "<p>I don't know if this is overfitting, any local CV results don't show that my model is overfitting</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3069949,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-12-12T04:47:56.797000",
          "content": "<p>Same here - I have CV-LB relation now after a lot of work. Hope it prevails on the private LB <a href=\"https://www.kaggle.com/ruichardliu\" target=\"_blank\">@ruichardliu</a> </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3071883,
              "author_name": "Sheikh Muhammad Abdullah",
              "author_url": "",
              "post_date": "2024-12-14T12:43:46.183000",
              "content": "<p>I have also one Single model which has Optimized Score 0.452 and LB score of 0.450. <br>\nAfter many experiments. and I am Taking risk to chose it as final. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3071928,
              "author_name": "Reki",
              "author_url": "",
              "post_date": "2024-12-14T14:25:45.850000",
              "content": "<p>I will most likely choose submissions with similar scores, and when LB exceeds 0.46, the local CV score will start to decrease</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3065253,
      "author_name": "Alperen Duru",
      "author_url": "",
      "post_date": "2024-12-06T15:27:22.423000",
      "content": "<p>I believe that due to small number of samples compared to the number of features available, we do not have enough data to be able to generalize or characterize the distribution of the train. This leads to problems with generalizations in the actual test dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3064104,
      "author_name": "Marius Heuser",
      "author_url": "",
      "post_date": "2024-12-05T08:31:46.420000",
      "content": "<p>I agree, that this comp results will heavily rely on luck in the end, but there are probably still things you can do to improve your chances. At least that's what I'm trying to do aswell :)</p>\n<p>Your case seems weird to me. Did you use KNN Imputer on the target? Because my results haven't been this far apart (from CV to LB).<br>\nThen it is simply explained by data leakage</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3064185,
          "author_name": "Reki",
          "author_url": "",
          "post_date": "2024-12-05T10:40:04.663000",
          "content": "<p>I did not use KNN-related filling, because I think the quality of this data set is not high, so the effect of KNN is not suitable for me. As for data leakage, I am checking this problem now, but I did not find out. I will share my notebook after the competition</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3064262,
          "author_name": "Reki",
          "author_url": "",
          "post_date": "2024-12-05T12:46:10.913000",
          "content": "<p>I also used some clustering methods, CV,:5.1,LB:0.40, after a certain score, the value of CV and LB will become inversely proportional</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3064019,
      "author_name": "Reki",
      "author_url": "",
      "post_date": "2024-12-05T06:05:58.320000",
      "content": "<p>I think it's safe to say that the data sets we're working with are vastly different from the ones on the leaderboard</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3064022,
          "author_name": "Reki",
          "author_url": "",
          "post_date": "2024-12-05T06:10:30.253000",
          "content": "<p>In such cases, it is not known which pretreatment is suitable, and any pretreatment may shake down under the final leaderboard</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3071925,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-12-14T14:22:23.827000",
              "content": "<p><a href=\"https://www.kaggle.com/ruichardliu\" target=\"_blank\">@ruichardliu</a> agreed, this is a 100% luck game!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3064402,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-05T15:23:12.603000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 3064493,
          "author_name": "Reki",
          "author_url": "",
          "post_date": "2024-12-05T16:56:21.680000",
          "content": "<p>Yep，A lottery game </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3071881": " i work alot on this during first month of competition and after mang exp , The Good Cv Perform Worse on LB .. I did't know what the main issue and what to chose for final submision. \nThis is Totally luck game. \nThe Better Preprocessing the worse Lb result ",
    "3064018": "I don't know if this is overfitting, any local CV results don't show that my model is overfitting",
    "3064017": "Run 1 - Seed 42 - QWK: 0.5186\nRun 2 - Seed 52 - QWK: 0.5517\nRun 3 - Seed 62 - QWK: 0.5165\nRun 4 - Seed 72 - QWK: 0.5334\nRun 5 - Seed 82 - QWK: 0.5298\nRun 6 - Seed 92 - QWK: 0.5308\nRun 7 - Seed 102 - QWK: 0.5432\nRun 8 - Seed 112 - QWK: 0.5274\nRun 9 - Seed 122 - QWK: 0.5276\nRun 10 - Seed 132 - QWK: 0.5436\n\n----> || Optimized QWK SCORE ::  0.607\n\nI feel that there is no overfitting in each of my twists, because the QWK score of my model on train is 0.63, and the average score of my verification set is 0.53, which can reach 0.607 after threshold adjustment, but the submission of this version only has a score of 0.378 on LB\n\nI am now stuck on how to choose to submit",
    "3065253": "I believe that due to small number of samples compared to the number of features available, we do not have enough data to be able to generalize or characterize the distribution of the train. This leads to problems with generalizations in the actual test dataset.",
    "3064104": "I agree, that this comp results will heavily rely on luck in the end, but there are probably still things you can do to improve your chances. At least that's what I'm trying to do aswell :)\n\nYour case seems weird to me. Did you use KNN Imputer on the target? Because my results haven't been this far apart (from CV to LB).\nThen it is simply explained by data leakage",
    "3064019": "I think it's safe to say that the data sets we're working with are vastly different from the ones on the leaderboard",
    "3064402": ""
  }
}