{
  "id": 400819,
  "title": "Confusing CV/LB scores",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/400819",
  "author_name": "",
  "post_date": "2023-04-10T12:20:30.253018Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I've been training several models which have gotten a F1-score of around 0.84 when predicting on a validation set (80/20 split). I quickly realized that these scores wasn't realistic, and decided to calculate what the f1-score would be predicting only 1's. This resulted in (for the test data, hence the small \"numbers\"):</p>\n<p>58804 1's<br>\n23294 0's</p>\n<p>Which would result in:<br>\nTP = 58804<br>\nFP = 0<br>\nFN = 23294</p>\n<p>And a f1-score of:<br>\n58804<em>2 / (54804</em>2 + 0 + 23294) = 0.8346794</p>\n<p>I then looked and ran the notebook:<br>\n<a href=\"https://www.kaggle.com/code/cpmpml/random-submission\" target=\"_blank\">https://www.kaggle.com/code/cpmpml/random-submission</a></p>\n<p>Where CPMP basically just submits predictions with predictions based on the test datas average outcome. (if majority is 1, then 1, if majority 0, then 0).</p>\n<p>However, the LB score for this notebook is just 0.66. Reasonable, I thought, since the f1 score is quite punishing regarding predicting to many 0's so I changed it up and had it predict 1's only (as when calculating it manually, it should result in a better f1-score). How ever, the LB-score was lowered drastically to around 0.55. </p>\n<p>Whenever I look at people submissions and what CV's/LB's they are presenting I can't really understand why the numbers are so \"low\". It almost feels like everyone is using aucurracy as a test metric instead of the f1-score. Or some kind of weighted f1-score?</p>\n<p>Is there something crucial that I am missing?</p>",
  "messages": [
    {
      "id": "2216868",
      "postDate": "04/10/2023 12:20:30",
      "content": "<p>I've been training several models which have gotten a F1-score of around 0.84 when predicting on a validation set (80/20 split). I quickly realized that these scores wasn't realistic, and decided to calculate what the f1-score would be predicting only 1's. This resulted in (for the test data, hence the small \"numbers\"):</p>\n<p>58804 1's<br>\n23294 0's</p>\n<p>Which would result in:<br>\nTP = 58804<br>\nFP = 0<br>\nFN = 23294</p>\n<p>And a f1-score of:<br>\n58804<em>2 / (54804</em>2 + 0 + 23294) = 0.8346794</p>\n<p>I then looked and ran the notebook:<br>\n<a href=\"https://www.kaggle.com/code/cpmpml/random-submission\" target=\"_blank\">https://www.kaggle.com/code/cpmpml/random-submission</a></p>\n<p>Where CPMP basically just submits predictions with predictions based on the test datas average outcome. (if majority is 1, then 1, if majority 0, then 0).</p>\n<p>However, the LB score for this notebook is just 0.66. Reasonable, I thought, since the f1 score is quite punishing regarding predicting to many 0's so I changed it up and had it predict 1's only (as when calculating it manually, it should result in a better f1-score). How ever, the LB-score was lowered drastically to around 0.55. </p>\n<p>Whenever I look at people submissions and what CV's/LB's they are presenting I can't really understand why the numbers are so \"low\". It almost feels like everyone is using aucurracy as a test metric instead of the f1-score. Or some kind of weighted f1-score?</p>\n<p>Is there something crucial that I am missing?</p>",
      "rawMarkdown": "I've been training several models which have gotten a F1-score of around 0.84 when predicting on a validation set (80/20 split). I quickly realized that these scores wasn't realistic, and decided to calculate what the f1-score would be predicting only 1's. This resulted in (for the test data, hence the small \"numbers\"):\n\n58804 1's\n23294 0's\n\nWhich would result in:\nTP = 58804\nFP = 0\nFN = 23294\n\nAnd a f1-score of:\n58804*2 / (54804*2 + 0 + 23294) = 0.8346794\n\nI then looked and ran the notebook:\nhttps://www.kaggle.com/code/cpmpml/random-submission\n\nWhere CPMP basically just submits predictions with predictions based on the test datas average outcome. (if majority is 1, then 1, if majority 0, then 0).\n\nHowever, the LB score for this notebook is just 0.66. Reasonable, I thought, since the f1 score is quite punishing regarding predicting to many 0's so I changed it up and had it predict 1's only (as when calculating it manually, it should result in a better f1-score). How ever, the LB-score was lowered drastically to around 0.55. \n\nWhenever I look at people submissions and what CV's/LB's they are presenting I can't really understand why the numbers are so \"low\". It almost feels like everyone is using aucurracy as a test metric instead of the f1-score. Or some kind of weighted f1-score?\n\nIs there something crucial that I am missing?",
      "votes": null
    },
    {
      "id": "2216963",
      "postDate": "04/10/2023 13:43:06",
      "content": "<p>You need to <em>macro average</em> the f1-score.  In this case, you divide into two classes, the <strong>1</strong> and the <strong>0</strong> classes and calculate the true positive and false negative rates on each, then calculate the f1-score on each and average them.  A good explanation is found here:</p>\n<p><a href=\"https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f\" target=\"_blank\">https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f</a></p>",
      "rawMarkdown": "You need to *macro average* the f1-score.  In this case, you divide into two classes, the **1** and the **0** classes and calculate the true positive and false negative rates on each, then calculate the f1-score on each and average them.  A good explanation is found here:\n\n[https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f](https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f)",
      "votes": null
    },
    {
      "id": "2217036",
      "postDate": "04/10/2023 14:47:57",
      "content": "<p>Thanks a lot! This was exactly what I was looking for.</p>",
      "rawMarkdown": "Thanks a lot! This was exactly what I was looking for.",
      "votes": null
    },
    {
      "id": "2217546",
      "postDate": "04/11/2023 02:16:51",
      "content": "<p>I think the reason why your LB score is lower than your manual calculation is because the test data is not balanced. According to the data description, the test data has 50% of students who passed and 50% who failed, while the train data has 71.7% of students who passed and 28.3% who failed. This means that predicting all 1's on the test data will result in a lot of false positives and lower your precision and f1-score.</p>\n<p>The notebook you mentioned by CPMP uses the mean outcome of the train data (0.717) as a threshold to predict 1 or 0 on the test data. This is a simple way to account for the imbalance in the train data, but it may not be optimal for the test data. A better way would be to use cross-validation or grid search to find the best threshold that maximizes your f1-score on the validation set.</p>",
      "rawMarkdown": "I think the reason why your LB score is lower than your manual calculation is because the test data is not balanced. According to the data description, the test data has 50% of students who passed and 50% who failed, while the train data has 71.7% of students who passed and 28.3% who failed. This means that predicting all 1's on the test data will result in a lot of false positives and lower your precision and f1-score.\n\nThe notebook you mentioned by CPMP uses the mean outcome of the train data (0.717) as a threshold to predict 1 or 0 on the test data. This is a simple way to account for the imbalance in the train data, but it may not be optimal for the test data. A better way would be to use cross-validation or grid search to find the best threshold that maximizes your f1-score on the validation set.",
      "votes": null
    },
    {
      "id": "2271127",
      "postDate": "05/23/2023 16:33:39",
      "content": "<p>\"According to the data description, the test data has 50% of students who passed and 50% who failed\"</p>\n<p>Where is it written?</p>",
      "rawMarkdown": "\"According to the data description, the test data has 50% of students who passed and 50% who failed\"\n\nWhere is it written?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2216963,
      "author_name": "danielphalen",
      "author_url": "",
      "post_date": "04/10/2023 13:43:06",
      "content": "<p>You need to <em>macro average</em> the f1-score.  In this case, you divide into two classes, the <strong>1</strong> and the <strong>0</strong> classes and calculate the true positive and false negative rates on each, then calculate the f1-score on each and average them.  A good explanation is found here:</p>\n<p><a href=\"https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f\" target=\"_blank\">https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2217036,
          "author_name": "eliasforsblom",
          "author_url": "",
          "post_date": "04/10/2023 14:47:57",
          "content": "<p>Thanks a lot! This was exactly what I was looking for.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2217546,
      "author_name": "yus002",
      "author_url": "",
      "post_date": "04/11/2023 02:16:51",
      "content": "<p>I think the reason why your LB score is lower than your manual calculation is because the test data is not balanced. According to the data description, the test data has 50% of students who passed and 50% who failed, while the train data has 71.7% of students who passed and 28.3% who failed. This means that predicting all 1's on the test data will result in a lot of false positives and lower your precision and f1-score.</p>\n<p>The notebook you mentioned by CPMP uses the mean outcome of the train data (0.717) as a threshold to predict 1 or 0 on the test data. This is a simple way to account for the imbalance in the train data, but it may not be optimal for the test data. A better way would be to use cross-validation or grid search to find the best threshold that maximizes your f1-score on the validation set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2271127,
          "author_name": "miloszm",
          "author_url": "",
          "post_date": "05/23/2023 16:33:39",
          "content": "<p>\"According to the data description, the test data has 50% of students who passed and 50% who failed\"</p>\n<p>Where is it written?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2216868": "I've been training several models which have gotten a F1-score of around 0.84 when predicting on a validation set (80/20 split). I quickly realized that these scores wasn't realistic, and decided to calculate what the f1-score would be predicting only 1's. This resulted in (for the test data, hence the small \"numbers\"):\n\n58804 1's\n23294 0's\n\nWhich would result in:\nTP = 58804\nFP = 0\nFN = 23294\n\nAnd a f1-score of:\n58804*2 / (54804*2 + 0 + 23294) = 0.8346794\n\nI then looked and ran the notebook:\nhttps://www.kaggle.com/code/cpmpml/random-submission\n\nWhere CPMP basically just submits predictions with predictions based on the test datas average outcome. (if majority is 1, then 1, if majority 0, then 0).\n\nHowever, the LB score for this notebook is just 0.66. Reasonable, I thought, since the f1 score is quite punishing regarding predicting to many 0's so I changed it up and had it predict 1's only (as when calculating it manually, it should result in a better f1-score). How ever, the LB-score was lowered drastically to around 0.55. \n\nWhenever I look at people submissions and what CV's/LB's they are presenting I can't really understand why the numbers are so \"low\". It almost feels like everyone is using aucurracy as a test metric instead of the f1-score. Or some kind of weighted f1-score?\n\nIs there something crucial that I am missing?",
    "2216963": "You need to *macro average* the f1-score.  In this case, you divide into two classes, the **1** and the **0** classes and calculate the true positive and false negative rates on each, then calculate the f1-score on each and average them.  A good explanation is found here:\n\n[https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f](https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f)",
    "2217036": "Thanks a lot! This was exactly what I was looking for.",
    "2217546": "I think the reason why your LB score is lower than your manual calculation is because the test data is not balanced. According to the data description, the test data has 50% of students who passed and 50% who failed, while the train data has 71.7% of students who passed and 28.3% who failed. This means that predicting all 1's on the test data will result in a lot of false positives and lower your precision and f1-score.\n\nThe notebook you mentioned by CPMP uses the mean outcome of the train data (0.717) as a threshold to predict 1 or 0 on the test data. This is a simple way to account for the imbalance in the train data, but it may not be optimal for the test data. A better way would be to use cross-validation or grid search to find the best threshold that maximizes your f1-score on the validation set.",
    "2271127": "\"According to the data description, the test data has 50% of students who passed and 50% who failed\"\n\nWhere is it written?"
  },
  "source": "meta"
}