{
  "id": 409497,
  "title": "Improvement in Individual Question f1_scores, But Overall f1_score Decreased",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/409497",
  "author_name": "",
  "post_date": "2023-05-11T09:55:02.508556500Z",
  "votes": 7,
  "comment_count": 12,
  "views": 0,
  "content": "<p>The f1_score detail shows that while the majority of questions saw an improvement in their f1_score, there was a significant decrease in the overall f1_score. Can anyone explain the reason to me?</p>\n<pre><code>Overall F1 = 0.69957               Overall F1 = 0.68136\nQ1: F1 = 0.6738612983942933        Q1: F1 = 0.6687508371066515\nQ2: F1 = 0.5241114949036766        Q2: F1 = 0.5428246778296058      up\nQ3: F1 = 0.5281256850641168        Q3: F1 = 0.5976987601426149      up\nQ4: F1 = 0.6847349996014611        Q4: F1 = 0.691227857140134       up\nQ5: F1 = 0.6386584171090336        Q5: F1 = 0.6610093954862739      up\nQ6: F1 = 0.6483030909060505        Q6: F1 = 0.6551849554407907      up\nQ7: F1 = 0.6331689410426152        Q7: F1 = 0.6378148957541732      up\nQ8: F1 = 0.572312054606764         Q8: F1 = 0.5742240529668264      up\nQ9: F1 = 0.6356655516338143        Q9: F1 = 0.6382906478040109      up\nQ10: F1 = 0.5857512326742619       Q10: F1 = 0.6385312015191333     up\nQ11: F1 = 0.613713893572035        Q11: F1 = 0.61114392829887       \nQ12: F1 = 0.523497421681268        Q12: F1 = 0.5761879984295377     up\nQ13: F1 = 0.48744590669612414      Q13: F1 = 0.624941543011704      up\nQ14: F1 = 0.6407764212193775       Q14: F1 = 0.6429682077343519     up\nQ15: F1 = 0.6092252876576862       Q15: F1 = 0.6619349085654576     up\nQ16: F1 = 0.5108702335029486       Q16: F1 = 0.5459830163285537     up\nQ17: F1 = 0.5671329910293095       Q17: F1 = 0.568876045449799      up\nQ18: F1 = 0.5011066899244714       Q18: F1 = 0.5814679922796635     up\n</code></pre>",
  "messages": [
    {
      "id": "2254862",
      "postDate": "05/11/2023 09:55:02",
      "content": "<p>The f1_score detail shows that while the majority of questions saw an improvement in their f1_score, there was a significant decrease in the overall f1_score. Can anyone explain the reason to me?</p>\n<pre><code>Overall F1 = 0.69957               Overall F1 = 0.68136\nQ1: F1 = 0.6738612983942933        Q1: F1 = 0.6687508371066515\nQ2: F1 = 0.5241114949036766        Q2: F1 = 0.5428246778296058      up\nQ3: F1 = 0.5281256850641168        Q3: F1 = 0.5976987601426149      up\nQ4: F1 = 0.6847349996014611        Q4: F1 = 0.691227857140134       up\nQ5: F1 = 0.6386584171090336        Q5: F1 = 0.6610093954862739      up\nQ6: F1 = 0.6483030909060505        Q6: F1 = 0.6551849554407907      up\nQ7: F1 = 0.6331689410426152        Q7: F1 = 0.6378148957541732      up\nQ8: F1 = 0.572312054606764         Q8: F1 = 0.5742240529668264      up\nQ9: F1 = 0.6356655516338143        Q9: F1 = 0.6382906478040109      up\nQ10: F1 = 0.5857512326742619       Q10: F1 = 0.6385312015191333     up\nQ11: F1 = 0.613713893572035        Q11: F1 = 0.61114392829887       \nQ12: F1 = 0.523497421681268        Q12: F1 = 0.5761879984295377     up\nQ13: F1 = 0.48744590669612414      Q13: F1 = 0.624941543011704      up\nQ14: F1 = 0.6407764212193775       Q14: F1 = 0.6429682077343519     up\nQ15: F1 = 0.6092252876576862       Q15: F1 = 0.6619349085654576     up\nQ16: F1 = 0.5108702335029486       Q16: F1 = 0.5459830163285537     up\nQ17: F1 = 0.5671329910293095       Q17: F1 = 0.568876045449799      up\nQ18: F1 = 0.5011066899244714       Q18: F1 = 0.5814679922796635     up\n</code></pre>",
      "rawMarkdown": "The f1_score detail shows that while the majority of questions saw an improvement in their f1_score, there was a significant decrease in the overall f1_score. Can anyone explain the reason to me?\n```bash\nOverall F1 = 0.69957               Overall F1 = 0.68136\nQ1: F1 = 0.6738612983942933        Q1: F1 = 0.6687508371066515\nQ2: F1 = 0.5241114949036766        Q2: F1 = 0.5428246778296058      up\nQ3: F1 = 0.5281256850641168        Q3: F1 = 0.5976987601426149      up\nQ4: F1 = 0.6847349996014611        Q4: F1 = 0.691227857140134       up\nQ5: F1 = 0.6386584171090336        Q5: F1 = 0.6610093954862739      up\nQ6: F1 = 0.6483030909060505        Q6: F1 = 0.6551849554407907      up\nQ7: F1 = 0.6331689410426152        Q7: F1 = 0.6378148957541732      up\nQ8: F1 = 0.572312054606764         Q8: F1 = 0.5742240529668264      up\nQ9: F1 = 0.6356655516338143        Q9: F1 = 0.6382906478040109      up\nQ10: F1 = 0.5857512326742619       Q10: F1 = 0.6385312015191333     up\nQ11: F1 = 0.613713893572035        Q11: F1 = 0.61114392829887       \nQ12: F1 = 0.523497421681268        Q12: F1 = 0.5761879984295377     up\nQ13: F1 = 0.48744590669612414      Q13: F1 = 0.624941543011704      up\nQ14: F1 = 0.6407764212193775       Q14: F1 = 0.6429682077343519     up\nQ15: F1 = 0.6092252876576862       Q15: F1 = 0.6619349085654576     up\nQ16: F1 = 0.5108702335029486       Q16: F1 = 0.5459830163285537     up\nQ17: F1 = 0.5671329910293095       Q17: F1 = 0.568876045449799      up\nQ18: F1 = 0.5011066899244714       Q18: F1 = 0.5814679922796635     up\n```",
      "votes": null
    },
    {
      "id": "2254886",
      "postDate": "05/11/2023 10:20:31",
      "content": "<p>The positive and negative ratios are different for different questions, hence optimizing F1 scores (or accuracy) for a single question does not guarantee an increase in macro F1, at least that's how I understand it.</p>",
      "rawMarkdown": "The positive and negative ratios are different for different questions, hence optimizing F1 scores (or accuracy) for a single question does not guarantee an increase in macro F1, at least that's how I understand it.",
      "votes": null
    },
    {
      "id": "2255596",
      "postDate": "05/11/2023 20:45:40",
      "content": "<p>Is this also how the overall F1 can be greater than any of the individual F1s?</p>",
      "rawMarkdown": "Is this also how the overall F1 can be greater than any of the individual F1s?",
      "votes": null
    },
    {
      "id": "2255612",
      "postDate": "05/11/2023 21:14:27",
      "content": "<p>That's a good question, and I have absolutely no idea. </p>",
      "rawMarkdown": "That's a good question, and I have absolutely no idea.",
      "votes": null
    },
    {
      "id": "2255970",
      "postDate": "05/12/2023 06:54:20",
      "content": "<p>Yup, it's a interesting porints. Is training 18 models for each individual question a vible approach? If we are unable to optimize the model for a single question, what steps can we take to enhance the overall model? </p>",
      "rawMarkdown": "Yup, it's a interesting porints. Is training 18 models for each individual question a vible approach? If we are unable to optimize the model for a single question, what steps can we take to enhance the overall model?",
      "votes": null
    },
    {
      "id": "2256024",
      "postDate": "05/12/2023 08:04:39",
      "content": "<p>I've trying to find a way to optimize base on macro F1 in the past weeks, so far no success 😞</p>",
      "rawMarkdown": "I've trying to find a way to optimize base on macro F1 in the past weeks, so far no success 😞",
      "votes": null
    },
    {
      "id": "2256102",
      "postDate": "05/12/2023 09:02:33",
      "content": "<p>Macro f1_score is the unweighted mean of f1_score calculated for each label (that is 0 and 1 for this competition)</p>\n<p>I made a toy model with only 2 questions and 100 answers for each question:<br>\nQ1 is unbalanced: ground truth is  (1:90, 0:10), while Q2 is balanced: ground truth is (1:50, 0:50)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F14168fdf8ea538f374e16ffe9313638c%2Fmacro_f1_score.png?generation=1683882087511015&amp;alt=media\" alt=\"\"></p>\n<p>As you can see macro f1_score for each questions is lower than macro f1_score for Q1+Q2,<br>\nbecause what matters for macro f1_score is global TP, FP e FN for each label (0 and 1)</p>\n<p>If you add others questions you can easly build models that are better at almost every question but worse at global level. </p>",
      "rawMarkdown": "Macro f1_score is the unweighted mean of f1_score calculated for each label (that is 0 and 1 for this competition)\n\n\nI made a toy model with only 2 questions and 100 answers for each question:\nQ1 is unbalanced: ground truth is  (1:90, 0:10), while Q2 is balanced: ground truth is (1:50, 0:50)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F14168fdf8ea538f374e16ffe9313638c%2Fmacro_f1_score.png?generation=1683882087511015&alt=media)\n\n\nAs you can see macro f1_score for each questions is lower than macro f1_score for Q1+Q2,\nbecause what matters for macro f1_score is global TP, FP e FN for each label (0 and 1)\n\nIf you add others questions you can easly build models that are better at almost every question but worse at global level.",
      "votes": null
    },
    {
      "id": "2256143",
      "postDate": "05/12/2023 09:30:18",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> ! Is having balanced data crucial? I can achieve this by either upsampling or downsampling. The improvement in macro F1_score for each question is a result of data balancing. Unfortunately, balancing the data resulted in a poor overall macro F1_score.</p>",
      "rawMarkdown": "Thanks @steubk ! Is having balanced data crucial? I can achieve this by either upsampling or downsampling. The improvement in macro F1_score for each question is a result of data balancing. Unfortunately, balancing the data resulted in a poor overall macro F1_score.",
      "votes": null
    },
    {
      "id": "2256155",
      "postDate": "05/12/2023 09:37:03",
      "content": "<p>I am still in the process of optimizing the model. While I have added enough features and am paying attention to parameter tuning and the training phase, unfortunately, I have not made any progress yet.</p>",
      "rawMarkdown": "I am still in the process of optimizing the model. While I have added enough features and am paying attention to parameter tuning and the training phase, unfortunately, I have not made any progress yet.",
      "votes": null
    },
    {
      "id": "2256218",
      "postDate": "05/12/2023 10:36:46",
      "content": "<p>It's really an annoying thing.I think unbalanced 0-1 in different problems cause this. When I tried to create balanced dataset,but the score will drop. Then I tried to create 3 different threholds for 3 stages,but still didn't work. There must be something wrong, i just cant find way out. (QAQ)</p>",
      "rawMarkdown": "It's really an annoying thing.I think unbalanced 0-1 in different problems cause this. When I tried to create balanced dataset,but the score will drop. Then I tried to create 3 different threholds for 3 stages,but still didn't work. There must be something wrong, i just cant find way out. (QAQ)",
      "votes": null
    },
    {
      "id": "2262403",
      "postDate": "05/16/2023 22:44:59",
      "content": "<p>I am also struggling with this problem, but have yet to come up with a good solution. I have tried to combine the models into one with the question number as a categorical variable, but that didn't work either.</p>",
      "rawMarkdown": "I am also struggling with this problem, but have yet to come up with a good solution. I have tried to combine the models into one with the question number as a categorical variable, but that didn't work either.",
      "votes": null
    },
    {
      "id": "2262740",
      "postDate": "05/17/2023 06:00:39",
      "content": "<p>Combining multiple models into a single one is a brilliant concept. However, I am of the opinion that combining models into one may result in the loss of crucial information as different questions pertaining to various stages of the game may be overlooked.</p>",
      "rawMarkdown": "Combining multiple models into a single one is a brilliant concept. However, I am of the opinion that combining models into one may result in the loss of crucial information as different questions pertaining to various stages of the game may be overlooked.",
      "votes": null
    },
    {
      "id": "2267954",
      "postDate": "05/21/2023 11:00:22",
      "content": "<p>I see. That is certainly a possibility. It would be a good solution if we could build such a model while preventing missing information.<br>\nI wanted to discuss this idea further, so I started a new thread <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/411867\" target=\"_blank\">here</a>. Please feel free to discuss it with me here as well!</p>",
      "rawMarkdown": "I see. That is certainly a possibility. It would be a good solution if we could build such a model while preventing missing information.\nI wanted to discuss this idea further, so I started a new thread [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/411867). Please feel free to discuss it with me here as well!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2254886,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "05/11/2023 10:20:31",
      "content": "<p>The positive and negative ratios are different for different questions, hence optimizing F1 scores (or accuracy) for a single question does not guarantee an increase in macro F1, at least that's how I understand it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2255596,
          "author_name": "kevinmcgoldrick12",
          "author_url": "",
          "post_date": "05/11/2023 20:45:40",
          "content": "<p>Is this also how the overall F1 can be greater than any of the individual F1s?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2255612,
              "author_name": "woprime",
              "author_url": "",
              "post_date": "05/11/2023 21:14:27",
              "content": "<p>That's a good question, and I have absolutely no idea. </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2255970,
          "author_name": "mengvision",
          "author_url": "",
          "post_date": "05/12/2023 06:54:20",
          "content": "<p>Yup, it's a interesting porints. Is training 18 models for each individual question a vible approach? If we are unable to optimize the model for a single question, what steps can we take to enhance the overall model? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2256024,
              "author_name": "woprime",
              "author_url": "",
              "post_date": "05/12/2023 08:04:39",
              "content": "<p>I've trying to find a way to optimize base on macro F1 in the past weeks, so far no success 😞</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2256155,
                  "author_name": "mengvision",
                  "author_url": "",
                  "post_date": "05/12/2023 09:37:03",
                  "content": "<p>I am still in the process of optimizing the model. While I have added enough features and am paying attention to parameter tuning and the training phase, unfortunately, I have not made any progress yet.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2262403,
                      "author_name": "jinmiyashita",
                      "author_url": "",
                      "post_date": "05/16/2023 22:44:59",
                      "content": "<p>I am also struggling with this problem, but have yet to come up with a good solution. I have tried to combine the models into one with the question number as a categorical variable, but that didn't work either.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2262740,
                          "author_name": "mengvision",
                          "author_url": "",
                          "post_date": "05/17/2023 06:00:39",
                          "content": "<p>Combining multiple models into a single one is a brilliant concept. However, I am of the opinion that combining models into one may result in the loss of crucial information as different questions pertaining to various stages of the game may be overlooked.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2267954,
                              "author_name": "jinmiyashita",
                              "author_url": "",
                              "post_date": "05/21/2023 11:00:22",
                              "content": "<p>I see. That is certainly a possibility. It would be a good solution if we could build such a model while preventing missing information.<br>\nI wanted to discuss this idea further, so I started a new thread <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/411867\" target=\"_blank\">here</a>. Please feel free to discuss it with me here as well!</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2256102,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "05/12/2023 09:02:33",
      "content": "<p>Macro f1_score is the unweighted mean of f1_score calculated for each label (that is 0 and 1 for this competition)</p>\n<p>I made a toy model with only 2 questions and 100 answers for each question:<br>\nQ1 is unbalanced: ground truth is  (1:90, 0:10), while Q2 is balanced: ground truth is (1:50, 0:50)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F14168fdf8ea538f374e16ffe9313638c%2Fmacro_f1_score.png?generation=1683882087511015&amp;alt=media\" alt=\"\"></p>\n<p>As you can see macro f1_score for each questions is lower than macro f1_score for Q1+Q2,<br>\nbecause what matters for macro f1_score is global TP, FP e FN for each label (0 and 1)</p>\n<p>If you add others questions you can easly build models that are better at almost every question but worse at global level. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2256143,
          "author_name": "mengvision",
          "author_url": "",
          "post_date": "05/12/2023 09:30:18",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> ! Is having balanced data crucial? I can achieve this by either upsampling or downsampling. The improvement in macro F1_score for each question is a result of data balancing. Unfortunately, balancing the data resulted in a poor overall macro F1_score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2256218,
      "author_name": "rkxuan",
      "author_url": "",
      "post_date": "05/12/2023 10:36:46",
      "content": "<p>It's really an annoying thing.I think unbalanced 0-1 in different problems cause this. When I tried to create balanced dataset,but the score will drop. Then I tried to create 3 different threholds for 3 stages,but still didn't work. There must be something wrong, i just cant find way out. (QAQ)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2254862": "The f1_score detail shows that while the majority of questions saw an improvement in their f1_score, there was a significant decrease in the overall f1_score. Can anyone explain the reason to me?\n```bash\nOverall F1 = 0.69957               Overall F1 = 0.68136\nQ1: F1 = 0.6738612983942933        Q1: F1 = 0.6687508371066515\nQ2: F1 = 0.5241114949036766        Q2: F1 = 0.5428246778296058      up\nQ3: F1 = 0.5281256850641168        Q3: F1 = 0.5976987601426149      up\nQ4: F1 = 0.6847349996014611        Q4: F1 = 0.691227857140134       up\nQ5: F1 = 0.6386584171090336        Q5: F1 = 0.6610093954862739      up\nQ6: F1 = 0.6483030909060505        Q6: F1 = 0.6551849554407907      up\nQ7: F1 = 0.6331689410426152        Q7: F1 = 0.6378148957541732      up\nQ8: F1 = 0.572312054606764         Q8: F1 = 0.5742240529668264      up\nQ9: F1 = 0.6356655516338143        Q9: F1 = 0.6382906478040109      up\nQ10: F1 = 0.5857512326742619       Q10: F1 = 0.6385312015191333     up\nQ11: F1 = 0.613713893572035        Q11: F1 = 0.61114392829887       \nQ12: F1 = 0.523497421681268        Q12: F1 = 0.5761879984295377     up\nQ13: F1 = 0.48744590669612414      Q13: F1 = 0.624941543011704      up\nQ14: F1 = 0.6407764212193775       Q14: F1 = 0.6429682077343519     up\nQ15: F1 = 0.6092252876576862       Q15: F1 = 0.6619349085654576     up\nQ16: F1 = 0.5108702335029486       Q16: F1 = 0.5459830163285537     up\nQ17: F1 = 0.5671329910293095       Q17: F1 = 0.568876045449799      up\nQ18: F1 = 0.5011066899244714       Q18: F1 = 0.5814679922796635     up\n```",
    "2254886": "The positive and negative ratios are different for different questions, hence optimizing F1 scores (or accuracy) for a single question does not guarantee an increase in macro F1, at least that's how I understand it.",
    "2255596": "Is this also how the overall F1 can be greater than any of the individual F1s?",
    "2255612": "That's a good question, and I have absolutely no idea.",
    "2255970": "Yup, it's a interesting porints. Is training 18 models for each individual question a vible approach? If we are unable to optimize the model for a single question, what steps can we take to enhance the overall model?",
    "2256024": "I've trying to find a way to optimize base on macro F1 in the past weeks, so far no success 😞",
    "2256102": "Macro f1_score is the unweighted mean of f1_score calculated for each label (that is 0 and 1 for this competition)\n\n\nI made a toy model with only 2 questions and 100 answers for each question:\nQ1 is unbalanced: ground truth is  (1:90, 0:10), while Q2 is balanced: ground truth is (1:50, 0:50)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F14168fdf8ea538f374e16ffe9313638c%2Fmacro_f1_score.png?generation=1683882087511015&alt=media)\n\n\nAs you can see macro f1_score for each questions is lower than macro f1_score for Q1+Q2,\nbecause what matters for macro f1_score is global TP, FP e FN for each label (0 and 1)\n\nIf you add others questions you can easly build models that are better at almost every question but worse at global level.",
    "2256143": "Thanks @steubk ! Is having balanced data crucial? I can achieve this by either upsampling or downsampling. The improvement in macro F1_score for each question is a result of data balancing. Unfortunately, balancing the data resulted in a poor overall macro F1_score.",
    "2256155": "I am still in the process of optimizing the model. While I have added enough features and am paying attention to parameter tuning and the training phase, unfortunately, I have not made any progress yet.",
    "2256218": "It's really an annoying thing.I think unbalanced 0-1 in different problems cause this. When I tried to create balanced dataset,but the score will drop. Then I tried to create 3 different threholds for 3 stages,but still didn't work. There must be something wrong, i just cant find way out. (QAQ)",
    "2262403": "I am also struggling with this problem, but have yet to come up with a good solution. I have tried to combine the models into one with the question number as a categorical variable, but that didn't work either.",
    "2262740": "Combining multiple models into a single one is a brilliant concept. However, I am of the opinion that combining models into one may result in the loss of crucial information as different questions pertaining to various stages of the game may be overlooked.",
    "2267954": "I see. That is certainly a possibility. It would be a good solution if we could build such a model while preventing missing information.\nI wanted to discuss this idea further, so I started a new thread [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/411867). Please feel free to discuss it with me here as well!"
  },
  "source": "meta"
}