{
  "id": 406289,
  "title": "\"Random submission\" with unrealistic LB-scores?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/406289",
  "author_name": "",
  "post_date": "2023-05-01T21:31:21.914580Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I've been playing around some with the random-submissions notebooks, for example the one CPMP made, <a href=\"https://www.kaggle.com/code/cpmpml/random-submission\" target=\"_blank\">https://www.kaggle.com/code/cpmpml/random-submission</a>, which scores an LB of 0.66. Notebooks that simply predict 1's or 0's for each questions.</p>\n<p>However, isn't that LB-score very unrealistic, considering that the highest achievable macro average f1-score for a model where you just predict one out of two labels is when you have a perfectly balanced label (which rarely is the case for the different questions in this competition) and results in a macro average f1-score of 0.6667. Any other balance of the labels would result in a lower macro-average f1-score (unless one label is 100% dominant, where the macro-average f1-score of course would be 1.0). </p>\n<p>I am assuming that I am making some simple error, but I can't really find out where my thought process is flawed. Can it have something to do with how the composite macro average f1-score is calculated for all the questions combined which has some effect on the LB-score?</p>",
  "messages": [
    {
      "id": "2241848",
      "postDate": "05/01/2023 21:31:21",
      "content": "<p>I've been playing around some with the random-submissions notebooks, for example the one CPMP made, <a href=\"https://www.kaggle.com/code/cpmpml/random-submission\" target=\"_blank\">https://www.kaggle.com/code/cpmpml/random-submission</a>, which scores an LB of 0.66. Notebooks that simply predict 1's or 0's for each questions.</p>\n<p>However, isn't that LB-score very unrealistic, considering that the highest achievable macro average f1-score for a model where you just predict one out of two labels is when you have a perfectly balanced label (which rarely is the case for the different questions in this competition) and results in a macro average f1-score of 0.6667. Any other balance of the labels would result in a lower macro-average f1-score (unless one label is 100% dominant, where the macro-average f1-score of course would be 1.0). </p>\n<p>I am assuming that I am making some simple error, but I can't really find out where my thought process is flawed. Can it have something to do with how the composite macro average f1-score is calculated for all the questions combined which has some effect on the LB-score?</p>",
      "rawMarkdown": "I've been playing around some with the random-submissions notebooks, for example the one CPMP made, https://www.kaggle.com/code/cpmpml/random-submission, which scores an LB of 0.66. Notebooks that simply predict 1's or 0's for each questions.\n\nHowever, isn't that LB-score very unrealistic, considering that the highest achievable macro average f1-score for a model where you just predict one out of two labels is when you have a perfectly balanced label (which rarely is the case for the different questions in this competition) and results in a macro average f1-score of 0.6667. Any other balance of the labels would result in a lower macro-average f1-score (unless one label is 100% dominant, where the macro-average f1-score of course would be 1.0). \n\nI am assuming that I am making some simple error, but I can't really find out where my thought process is flawed. Can it have something to do with how the composite macro average f1-score is calculated for all the questions combined which has some effect on the LB-score?",
      "votes": null
    },
    {
      "id": "2242968",
      "postDate": "05/02/2023 15:55:22",
      "content": "<p>The competition metric (which is <strong>macro F1</strong>) ignores question number and just concatenates all the predictions together. Therefore the <strong>single list</strong> of predictions will have both 0's and 1's. (Even though each question individually will have all 0's or all 1's predictions)</p>",
      "rawMarkdown": "The competition metric (which is **macro F1**) ignores question number and just concatenates all the predictions together. Therefore the **single list** of predictions will have both 0's and 1's. (Even though each question individually will have all 0's or all 1's predictions)",
      "votes": null
    },
    {
      "id": "2243146",
      "postDate": "05/02/2023 17:49:36",
      "content": "<p>Thank you a lot for the answer, that clears things up for me!</p>",
      "rawMarkdown": "Thank you a lot for the answer, that clears things up for me!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2242968,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/02/2023 15:55:22",
      "content": "<p>The competition metric (which is <strong>macro F1</strong>) ignores question number and just concatenates all the predictions together. Therefore the <strong>single list</strong> of predictions will have both 0's and 1's. (Even though each question individually will have all 0's or all 1's predictions)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2243146,
          "author_name": "eliasforsblom",
          "author_url": "",
          "post_date": "05/02/2023 17:49:36",
          "content": "<p>Thank you a lot for the answer, that clears things up for me!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2241848": "I've been playing around some with the random-submissions notebooks, for example the one CPMP made, https://www.kaggle.com/code/cpmpml/random-submission, which scores an LB of 0.66. Notebooks that simply predict 1's or 0's for each questions.\n\nHowever, isn't that LB-score very unrealistic, considering that the highest achievable macro average f1-score for a model where you just predict one out of two labels is when you have a perfectly balanced label (which rarely is the case for the different questions in this competition) and results in a macro average f1-score of 0.6667. Any other balance of the labels would result in a lower macro-average f1-score (unless one label is 100% dominant, where the macro-average f1-score of course would be 1.0). \n\nI am assuming that I am making some simple error, but I can't really find out where my thought process is flawed. Can it have something to do with how the composite macro average f1-score is calculated for all the questions combined which has some effect on the LB-score?",
    "2242968": "The competition metric (which is **macro F1**) ignores question number and just concatenates all the predictions together. Therefore the **single list** of predictions will have both 0's and 1's. (Even though each question individually will have all 0's or all 1's predictions)",
    "2243146": "Thank you a lot for the answer, that clears things up for me!"
  },
  "source": "meta"
}