{
  "id": 396970,
  "title": "LB 0.66 without training a model",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/396970",
  "author_name": "CPMP",
  "post_date": "2023-03-23T15:13:20.200000",
  "votes": 39,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I got 0.66 on the public LB by using constant values per question. You can see my full notebook here: <a href=\"https://www.kaggle.com/code/cpmpml/random-submission/\" target=\"_blank\">https://www.kaggle.com/code/cpmpml/random-submission/</a></p>\n<p>It sets the predictions as follows:</p>\n<pre><code>    df = sample_submission\n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] = \n    df.loc[df.question == , ] = \n    df.loc[df.question == , ] = \n    df.loc[df.question == , ] = \n</code></pre>\n<p>It means that models learn little, as they only improve LB by 0.04x compared to this baseline. We have some work to do!</p>\n<p><strong>Update:</strong> it looks like I rediscovered what others shared already, see <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> comment: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396970#2194152\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396970#2194152</a> <br>\nThe new piece of info is that CV is unchanged at 0.650 with the new data, but LB moved from 0.648 to 0.66 with the new data.</p>",
  "messages": [
    {
      "id": 2193886,
      "postDate": "2023-03-23T15:13:20.200Z",
      "content": "<p>I got 0.66 on the public LB by using constant values per question. You can see my full notebook here: <a href=\"https://www.kaggle.com/code/cpmpml/random-submission/\" target=\"_blank\">https://www.kaggle.com/code/cpmpml/random-submission/</a></p>\n<p>It sets the predictions as follows:</p>\n<pre><code>    df = sample_submission\n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] =  \n    df.loc[df.question == , ] = \n    df.loc[df.question == , ] = \n    df.loc[df.question == , ] = \n    df.loc[df.question == , ] = \n</code></pre>\n<p>It means that models learn little, as they only improve LB by 0.04x compared to this baseline. We have some work to do!</p>\n<p><strong>Update:</strong> it looks like I rediscovered what others shared already, see <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> comment: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396970#2194152\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396970#2194152</a> <br>\nThe new piece of info is that CV is unchanged at 0.650 with the new data, but LB moved from 0.648 to 0.66 with the new data.</p>",
      "rawMarkdown": "I got 0.66 on the public LB by using constant values per question. You can see my full notebook here: https://www.kaggle.com/code/cpmpml/random-submission/\n\nIt sets the predictions as follows:\n\n```python\n    df = sample_submission\n    df.loc[df.question == 1, 'correct'] = 1 \n    df.loc[df.question == 2, 'correct'] = 1 \n    df.loc[df.question == 3, 'correct'] = 1 \n    df.loc[df.question == 4, 'correct'] = 1 \n    df.loc[df.question == 5, 'correct'] = 0 \n    df.loc[df.question == 6, 'correct'] = 1 \n    df.loc[df.question == 7, 'correct'] = 1 \n    df.loc[df.question == 8, 'correct'] = 0 \n    df.loc[df.question == 9, 'correct'] = 1 \n    df.loc[df.question == 10, 'correct'] = 0 \n    df.loc[df.question == 11, 'correct'] = 1 \n    df.loc[df.question == 12, 'correct'] = 1 \n    df.loc[df.question == 13, 'correct'] = 0 \n    df.loc[df.question == 14, 'correct'] = 1 \n    df.loc[df.question == 15, 'correct'] = 0\n    df.loc[df.question == 16, 'correct'] = 1\n    df.loc[df.question == 17, 'correct'] = 1\n    df.loc[df.question == 18, 'correct'] = 1\n\n```\n\nIt means that models learn little, as they only improve LB by 0.04x compared to this baseline. We have some work to do!\n\n**Update:** it looks like I rediscovered what others shared already, see @cdeotte comment: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396970#2194152 \nThe new piece of info is that CV is unchanged at 0.650 with the new data, but LB moved from 0.648 to 0.66 with the new data.",
      "votes": 39
    },
    {
      "id": 2193944,
      "postDate": "2023-03-23T15:57:28.223Z",
      "content": "<p>This is interesting. The previous \"predict train mean\" baseline scored CV 0.650 and LB 0.648 (with original train and test data). And your new \"train mean\" baseline still has CV 0.650 (i verified locally) but has LB 0.660. This indicates that the new test data has a distribution shift relative from new train data. Whereas the old test and old train seemed more aligned.</p>",
      "rawMarkdown": "This is interesting. The previous \"predict train mean\" baseline scored CV 0.650 and LB 0.648 (with original train and test data). And your new \"train mean\" baseline still has CV 0.650 (i verified locally) but has LB 0.660. This indicates that the new test data has a distribution shift relative from new train data. Whereas the old test and old train seemed more aligned.",
      "votes": 8,
      "replies": [
        {
          "id": 2193977,
          "postDate": "2023-03-23T16:32:07.943Z",
          "content": "<p>Also it seems like the new test is relatively small (scoring time reduced by 300% for the same pipeline )</p>",
          "rawMarkdown": "Also it seems like the new test is relatively small (scoring time reduced by 300% for the same pipeline )",
          "votes": 5
        },
        {
          "id": 2194047,
          "postDate": "2023-03-23T17:36:18.643Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> right, I should have said that CV is 0.65 indeed. And yes, it means there is a small distribution shift now.</p>",
          "rawMarkdown": "@cdeotte right, I should have said that CV is 0.65 indeed. And yes, it means there is a small distribution shift now.",
          "votes": 1,
          "replies": [
            {
              "id": 2194102,
              "postDate": "2023-03-23T18:19:25.160Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> what is \"The previous \"predict train mean\" baseline\"?  Is it this one: <a href=\"https://www.kaggle.com/code/cdeotte/predict-train-mean-baseline-0-648\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/predict-train-mean-baseline-0-648</a>?</p>\n<p>I now see it does the same as my code, but it is not how I found it.</p>",
              "rawMarkdown": "@cdeotte what is \"The previous \"predict train mean\" baseline\"?  Is it this one: https://www.kaggle.com/code/cdeotte/predict-train-mean-baseline-0-648?\n\nI now see it does the same as my code, but it is not how I found it.",
              "votes": 2
            },
            {
              "id": 2194152,
              "postDate": "2023-03-23T18:53:29.230Z",
              "content": "<p>Yes that one. And Mayukh found it before me in his notebook version 2 <a href=\"https://www.kaggle.com/code/mayukh18/lgbm-regressor-ensemble-pipeline?scriptVersionId=118517533\" target=\"_blank\">here</a>. </p>\n<p>We didn't find this combination of 0 and 1 predictions at first either. I submitted version 2 of my \"train mean\" notebook for LB=0.573 then Mayukh submitted version 1 of his \"train mean\" notebook for LB=0.587. Then i submitted version 3 of my \"train mean\" notebook for LB=0.615. Then Mayukh found the optimal threshold of 0.62 and submitted version 2 of his \"train mean\" notebook for LB=0.648. Then i submitted version 4 of my \"train mean\" notebook with threshold plot. Afterward, everyone started finding and applying thresholds to all models.</p>",
              "rawMarkdown": "Yes that one. And Mayukh found it before me in his notebook version 2 [here][1]. \n\nWe didn't find this combination of 0 and 1 predictions at first either. I submitted version 2 of my \"train mean\" notebook for LB=0.573 then Mayukh submitted version 1 of his \"train mean\" notebook for LB=0.587. Then i submitted version 3 of my \"train mean\" notebook for LB=0.615. Then Mayukh found the optimal threshold of 0.62 and submitted version 2 of his \"train mean\" notebook for LB=0.648. Then i submitted version 4 of my \"train mean\" notebook with threshold plot. Afterward, everyone started finding and applying thresholds to all models.\n\n[1]: https://www.kaggle.com/code/mayukh18/lgbm-regressor-ensemble-pipeline?scriptVersionId=118517533",
              "votes": 3
            },
            {
              "id": 2195532,
              "postDate": "2023-03-24T17:43:20.233Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , I updated the post.</p>",
              "rawMarkdown": "Thanks @cdeotte , I updated the post.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2193901,
      "postDate": "2023-03-23T15:24:35.437Z",
      "content": "<p>Should be a strong contender for the efficiency prize. 😄</p>",
      "rawMarkdown": "Should be a strong contender for the efficiency prize. 😄",
      "votes": 3,
      "replies": [
        {
          "id": 2194076,
          "postDate": "2023-03-23T17:59:57.277Z",
          "content": "<p>Indeed! Let me select it to se ehow it would fare.</p>",
          "rawMarkdown": "Indeed! Let me select it to se ehow it would fare.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2308391,
      "postDate": "2023-06-18T22:15:50.153Z",
      "content": "<p>This is great and interesting post. When I read the last part of assign values for 18 questions, I got confused at the beginning. So I modified the last part to avoid manual assign value and could know why some are assigned as \"0\". Hope it is useful for others like me. </p>\n<p><a href=\"https://www.kaggle.com/chennanli/random-submission-modify-value-assignment\" target=\"_blank\">https://www.kaggle.com/chennanli/random-submission-modify-value-assignment</a></p>",
      "rawMarkdown": "This is great and interesting post. When I read the last part of assign values for 18 questions, I got confused at the beginning. So I modified the last part to avoid manual assign value and could know why some are assigned as \"0\". Hope it is useful for others like me. \n\nhttps://www.kaggle.com/chennanli/random-submission-modify-value-assignment"
    },
    {
      "id": 2197137,
      "postDate": "2023-03-25T23:26:41.690Z",
      "content": "<p>Thanks for attaining the interesting achievement👍. Your work has a great inspiration for us all.</p>",
      "rawMarkdown": "Thanks for attaining the interesting achievement👍. Your work has a great inspiration for us all."
    },
    {
      "id": 2308388,
      "postDate": "2023-06-18T22:11:33.830Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2308389,
          "postDate": "2023-06-18T22:13:05.573Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2193944,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-03-23T15:57:28.223000",
      "content": "<p>This is interesting. The previous \"predict train mean\" baseline scored CV 0.650 and LB 0.648 (with original train and test data). And your new \"train mean\" baseline still has CV 0.650 (i verified locally) but has LB 0.660. This indicates that the new test data has a distribution shift relative from new train data. Whereas the old test and old train seemed more aligned.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2193977,
          "author_name": "Reacher",
          "author_url": "",
          "post_date": "2023-03-23T16:32:07.943000",
          "content": "<p>Also it seems like the new test is relatively small (scoring time reduced by 300% for the same pipeline )</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2194047,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2023-03-23T17:36:18.643000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> right, I should have said that CV is 0.65 indeed. And yes, it means there is a small distribution shift now.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2194102,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-23T18:19:25.160000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> what is \"The previous \"predict train mean\" baseline\"?  Is it this one: <a href=\"https://www.kaggle.com/code/cdeotte/predict-train-mean-baseline-0-648\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/predict-train-mean-baseline-0-648</a>?</p>\n<p>I now see it does the same as my code, but it is not how I found it.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2194152,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-03-23T18:53:29.230000",
              "content": "<p>Yes that one. And Mayukh found it before me in his notebook version 2 <a href=\"https://www.kaggle.com/code/mayukh18/lgbm-regressor-ensemble-pipeline?scriptVersionId=118517533\" target=\"_blank\">here</a>. </p>\n<p>We didn't find this combination of 0 and 1 predictions at first either. I submitted version 2 of my \"train mean\" notebook for LB=0.573 then Mayukh submitted version 1 of his \"train mean\" notebook for LB=0.587. Then i submitted version 3 of my \"train mean\" notebook for LB=0.615. Then Mayukh found the optimal threshold of 0.62 and submitted version 2 of his \"train mean\" notebook for LB=0.648. Then i submitted version 4 of my \"train mean\" notebook with threshold plot. Afterward, everyone started finding and applying thresholds to all models.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2195532,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2023-03-24T17:43:20.233000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , I updated the post.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2193901,
      "author_name": "Bill Cruise",
      "author_url": "",
      "post_date": "2023-03-23T15:24:35.437000",
      "content": "<p>Should be a strong contender for the efficiency prize. 😄</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2194076,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2023-03-23T17:59:57.277000",
          "content": "<p>Indeed! Let me select it to se ehow it would fare.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2308391,
      "author_name": "cli3",
      "author_url": "",
      "post_date": "2023-06-18T22:15:50.153000",
      "content": "<p>This is great and interesting post. When I read the last part of assign values for 18 questions, I got confused at the beginning. So I modified the last part to avoid manual assign value and could know why some are assigned as \"0\". Hope it is useful for others like me. </p>\n<p><a href=\"https://www.kaggle.com/chennanli/random-submission-modify-value-assignment\" target=\"_blank\">https://www.kaggle.com/chennanli/random-submission-modify-value-assignment</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2197137,
      "author_name": "Tariq Mahmood",
      "author_url": "",
      "post_date": "2023-03-25T23:26:41.690000",
      "content": "<p>Thanks for attaining the interesting achievement👍. Your work has a great inspiration for us all.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2308388,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-18T22:11:33.830000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2308389,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-18T22:13:05.573000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2193886": "I got 0.66 on the public LB by using constant values per question. You can see my full notebook here: https://www.kaggle.com/code/cpmpml/random-submission/\n\nIt sets the predictions as follows:\n\n```python\n    df = sample_submission\n    df.loc[df.question == 1, 'correct'] = 1 \n    df.loc[df.question == 2, 'correct'] = 1 \n    df.loc[df.question == 3, 'correct'] = 1 \n    df.loc[df.question == 4, 'correct'] = 1 \n    df.loc[df.question == 5, 'correct'] = 0 \n    df.loc[df.question == 6, 'correct'] = 1 \n    df.loc[df.question == 7, 'correct'] = 1 \n    df.loc[df.question == 8, 'correct'] = 0 \n    df.loc[df.question == 9, 'correct'] = 1 \n    df.loc[df.question == 10, 'correct'] = 0 \n    df.loc[df.question == 11, 'correct'] = 1 \n    df.loc[df.question == 12, 'correct'] = 1 \n    df.loc[df.question == 13, 'correct'] = 0 \n    df.loc[df.question == 14, 'correct'] = 1 \n    df.loc[df.question == 15, 'correct'] = 0\n    df.loc[df.question == 16, 'correct'] = 1\n    df.loc[df.question == 17, 'correct'] = 1\n    df.loc[df.question == 18, 'correct'] = 1\n\n```\n\nIt means that models learn little, as they only improve LB by 0.04x compared to this baseline. We have some work to do!\n\n**Update:** it looks like I rediscovered what others shared already, see @cdeotte comment: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396970#2194152 \nThe new piece of info is that CV is unchanged at 0.650 with the new data, but LB moved from 0.648 to 0.66 with the new data.",
    "2193944": "This is interesting. The previous \"predict train mean\" baseline scored CV 0.650 and LB 0.648 (with original train and test data). And your new \"train mean\" baseline still has CV 0.650 (i verified locally) but has LB 0.660. This indicates that the new test data has a distribution shift relative from new train data. Whereas the old test and old train seemed more aligned.",
    "2193901": "Should be a strong contender for the efficiency prize. 😄",
    "2308391": "This is great and interesting post. When I read the last part of assign values for 18 questions, I got confused at the beginning. So I modified the last part to avoid manual assign value and could know why some are assigned as \"0\". Hope it is useful for others like me. \n\nhttps://www.kaggle.com/chennanli/random-submission-modify-value-assignment",
    "2197137": "Thanks for attaining the interesting achievement👍. Your work has a great inspiration for us all.",
    "2308388": ""
  }
}