{
  "id": 417359,
  "title": "Why optimizing each question independently doesn't necessarily give good results ",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/417359",
  "author_name": "",
  "post_date": "2023-06-15T11:44:11.653146500Z",
  "votes": 18,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Have you tried to optimize the F1 score at each question independently and then put it all together?<br>\n<strong>For example, did you try any of the following:</strong></p>\n<ul>\n<li>Select the best Model for each question</li>\n<li>Select the best Hyperparameters for each question</li>\n<li>Select the best Threshold for each question</li>\n<li>Select the best Features for each questions</li>\n</ul>\n<p>I tried all of these and most of the times the final LB score was worse!!<br>\nMy intuitive explanation was: all these techniques were overfitting the train data.<br>\nBut then I had an even more bizarre result. I was selecting, for each question, the model that was giving the best F1-score at the question level.  The result was very surprising, the overall CV F1 score was worse than with only one single good model for all questions!! so even before submitting to the LB, at the CV level, it was not working. Overfitting could not be the explanation. <br>\n<strong>There was no choice but to dig into the scoring mechanism.</strong><br>\nBelow is a graph showing how it works.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F7c388f97ad6f1d4faccd3d058662966d%2FF1Score.png?generation=1686826873533216&amp;alt=media\" alt=\"\"></p>\n<p>The key is the <strong>harmonic mean</strong> . It is very different from a simple mean, it is greatly penalized by the worst. <strong>If the recall is very bad and the precision is excellent, the F1 will be very bad.</strong><br>\nIt is exactly what was happening. One model was giving a very bad recall, but an excellent precision for one question. Its F1 was low. So I was discarding it, but actually, <strong>when diluted with all the other questions,</strong> the bad recall didn't have such a negative impact in regards with the beneficial impact on the overall precision.</p>\n<p><strong>Conclusion:</strong><br>\nThe score in this competition is the F1 score macro on the overall questions taken altogether. It is very different from the average F1 scores for each question (otherwise it would simply be 18 separate competitions taken together). All the questions are put together in a same bag and then the F1 score is calculated.</p>\n<p>The name <strong>harmonic mean</strong> has really been well chosen. You get a good score only if you have a model that harmoniously treats the precision and the recall. It has to find the right balance. BUT, a perfect harmony at each individual question doesn't necessarily mean the best harmony for the 18 questions taken altogether.</p>",
  "messages": [
    {
      "id": "2303623",
      "postDate": "06/15/2023 11:44:11",
      "content": "<p>Have you tried to optimize the F1 score at each question independently and then put it all together?<br>\n<strong>For example, did you try any of the following:</strong></p>\n<ul>\n<li>Select the best Model for each question</li>\n<li>Select the best Hyperparameters for each question</li>\n<li>Select the best Threshold for each question</li>\n<li>Select the best Features for each questions</li>\n</ul>\n<p>I tried all of these and most of the times the final LB score was worse!!<br>\nMy intuitive explanation was: all these techniques were overfitting the train data.<br>\nBut then I had an even more bizarre result. I was selecting, for each question, the model that was giving the best F1-score at the question level.  The result was very surprising, the overall CV F1 score was worse than with only one single good model for all questions!! so even before submitting to the LB, at the CV level, it was not working. Overfitting could not be the explanation. <br>\n<strong>There was no choice but to dig into the scoring mechanism.</strong><br>\nBelow is a graph showing how it works.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F7c388f97ad6f1d4faccd3d058662966d%2FF1Score.png?generation=1686826873533216&amp;alt=media\" alt=\"\"></p>\n<p>The key is the <strong>harmonic mean</strong> . It is very different from a simple mean, it is greatly penalized by the worst. <strong>If the recall is very bad and the precision is excellent, the F1 will be very bad.</strong><br>\nIt is exactly what was happening. One model was giving a very bad recall, but an excellent precision for one question. Its F1 was low. So I was discarding it, but actually, <strong>when diluted with all the other questions,</strong> the bad recall didn't have such a negative impact in regards with the beneficial impact on the overall precision.</p>\n<p><strong>Conclusion:</strong><br>\nThe score in this competition is the F1 score macro on the overall questions taken altogether. It is very different from the average F1 scores for each question (otherwise it would simply be 18 separate competitions taken together). All the questions are put together in a same bag and then the F1 score is calculated.</p>\n<p>The name <strong>harmonic mean</strong> has really been well chosen. You get a good score only if you have a model that harmoniously treats the precision and the recall. It has to find the right balance. BUT, a perfect harmony at each individual question doesn't necessarily mean the best harmony for the 18 questions taken altogether.</p>",
      "rawMarkdown": "Have you tried to optimize the F1 score at each question independently and then put it all together?\n**For example, did you try any of the following:**\n\n- Select the best Model for each question\n- Select the best Hyperparameters for each question\n- Select the best Threshold for each question\n- Select the best Features for each questions\n\nI tried all of these and most of the times the final LB score was worse!!\nMy intuitive explanation was: all these techniques were overfitting the train data.\nBut then I had an even more bizarre result. I was selecting, for each question, the model that was giving the best F1-score at the question level.  The result was very surprising, the overall CV F1 score was worse than with only one single good model for all questions!! so even before submitting to the LB, at the CV level, it was not working. Overfitting could not be the explanation. \n**There was no choice but to dig into the scoring mechanism.**\nBelow is a graph showing how it works.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F7c388f97ad6f1d4faccd3d058662966d%2FF1Score.png?generation=1686826873533216&alt=media)\n\nThe key is the **harmonic mean** . It is very different from a simple mean, it is greatly penalized by the worst. **If the recall is very bad and the precision is excellent, the F1 will be very bad.**\nIt is exactly what was happening. One model was giving a very bad recall, but an excellent precision for one question. Its F1 was low. So I was discarding it, but actually, **when diluted with all the other questions,** the bad recall didn't have such a negative impact in regards with the beneficial impact on the overall precision.\n\n**Conclusion:**\nThe score in this competition is the F1 score macro on the overall questions taken altogether. It is very different from the average F1 scores for each question (otherwise it would simply be 18 separate competitions taken together). All the questions are put together in a same bag and then the F1 score is calculated.\n\nThe name **harmonic mean** has really been well chosen. You get a good score only if you have a model that harmoniously treats the precision and the recall. It has to find the right balance. BUT, a perfect harmony at each individual question doesn't necessarily mean the best harmony for the 18 questions taken altogether.",
      "votes": null
    },
    {
      "id": "2303903",
      "postDate": "06/15/2023 14:39:24",
      "content": "<p>Thank you for your very clear clarification! </p>\n<p>I have a small question related to your statement: \"The result was very surprising, the overall CV F1 score was worse than with only one single good model for all questions!!\" Were you trying to say one model for all 18 questions got better score than one model per question set-up?</p>",
      "rawMarkdown": "Thank you for your very clear clarification! \n\nI have a small question related to your statement: \"The result was very surprising, the overall CV F1 score was worse than with only one single good model for all questions!!\" Were you trying to say one model for all 18 questions got better score than one model per question set-up?",
      "votes": null
    },
    {
      "id": "2304118",
      "postDate": "06/15/2023 17:34:04",
      "content": "<p>Yes exactly <a href=\"https://www.kaggle.com/xiaosufrankhu\" target=\"_blank\">@xiaosufrankhu</a> . Let's say there are two models M1 and M2. M1 has a better overall F1 score. <br>\nF1(overall,M1) &gt; F1(overall,M2)<br>\nbut on a question basis let's say M2 has a better F1 score than M1 for question 2,5,7.<br>\nF1(q,M2) &gt; F1(q,M1)  for q in [2,5,7]</p>\n<p>Then you may think that a solution where you use M2 for questions [2,5,7] and M1 for all the other questions will give you a better score than M1 alone for all questions. But in my tests It was not the case, EVEN at the CV level. So it was not about overfitting but about the way the scoring works. </p>",
      "rawMarkdown": "Yes exactly @xiaosufrankhu . Let's say there are two models M1 and M2. M1 has a better overall F1 score. \nF1(overall,M1) > F1(overall,M2)\nbut on a question basis let's say M2 has a better F1 score than M1 for question 2,5,7.\nF1(q,M2) > F1(q,M1)  for q in [2,5,7]\n\nThen you may think that a solution where you use M2 for questions [2,5,7] and M1 for all the other questions will give you a better score than M1 alone for all questions. But in my tests It was not the case, EVEN at the CV level. So it was not about overfitting but about the way the scoring works.",
      "votes": null
    },
    {
      "id": "2304158",
      "postDate": "06/15/2023 18:10:18",
      "content": "<p>I think the scoring function is not the main cause of this problem. </p>\n<p>I've tried to optimize for validation macro F1 by first training a model for initial prediction, then updating predictions for each question by selecting the model which results in most macro F1 gain. I think by doing this I raised CV score by ~0.002 - 0.003, but this did not generalize to LB at all. I've also tested my suspicion on a held-out test set, and not to my surprise, the test also hasn't benefited from this CV score gain.</p>\n<p>So I guess the lesson here is, don't trust your CV, trust your CV's CV. 🙃</p>",
      "rawMarkdown": "I think the scoring function is not the main cause of this problem. \n\nI've tried to optimize for validation macro F1 by first training a model for initial prediction, then updating predictions for each question by selecting the model which results in most macro F1 gain. I think by doing this I raised CV score by ~0.002 - 0.003, but this did not generalize to LB at all. I've also tested my suspicion on a held-out test set, and not to my surprise, the test also hasn't benefited from this CV score gain.\n\nSo I guess the lesson here is, don't trust your CV, trust your CV's CV. 🙃",
      "votes": null
    },
    {
      "id": "2304345",
      "postDate": "06/16/2023 00:08:32",
      "content": "<p>Thanks for the insightful post. If you don't mind, I have posted a <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478\" target=\"_blank\">sequel</a> with some math behind it :)</p>",
      "rawMarkdown": "Thanks for the insightful post. If you don't mind, I have posted a [sequel](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478) with some math behind it :)",
      "votes": null
    },
    {
      "id": "2304412",
      "postDate": "06/16/2023 01:38:34",
      "content": "<p>I don't think optimizing models based on each question has anything to do with metric</p>\n<p>Let us say the ground truth  and predictions are as follows (consider 3 questions, 4 samples) - using a single model - hypothetical case where question 3 alone is predicted completely wrongly:</p>\n<p>y_true = np.array([ [1, 1, 1],<br>\n                               [1, 0, 1],<br>\n                               [1, 0, 0],<br>\n                               [1, 0, 0] ]) </p>\n<p>y_pred1 =  np.array([[1, 1, 0],<br>\n                                  [1, 0, 0],<br>\n                                  [1, 0, 1],<br>\n                                  [1, 0, 1]])</p>\n<p>Now, suppose you use separate model for each question and arrive at predictions as below (where third question is also predicted perfectly)</p>\n<p>y_pred2 =  np.array([[1, 1, 1],<br>\n                                   [1, 0, 1],<br>\n                                   [1, 0, 0],<br>\n                                   [1, 0, 0]])</p>\n<p>Why would the LB give lower score?      </p>\n<p>(Note: CV based on current competition metric will always be lower than CV of metric that averages individual f1 score of each question - given metric is conservative)</p>",
      "rawMarkdown": "I don't think optimizing models based on each question has anything to do with metric\n\nLet us say the ground truth  and predictions are as follows (consider 3 questions, 4 samples) - using a single model - hypothetical case where question 3 alone is predicted completely wrongly:\n\ny_true = np.array([ [1, 1, 1],\n                               [1, 0, 1],\n                               [1, 0, 0],\n                               [1, 0, 0] ]) \n\ny_pred1 =  np.array([[1, 1, 0],\n                                  [1, 0, 0],\n                                  [1, 0, 1],\n                                  [1, 0, 1]])\n\nNow, suppose you use separate model for each question and arrive at predictions as below (where third question is also predicted perfectly)\n\ny_pred2 =  np.array([[1, 1, 1],\n                                   [1, 0, 1],\n                                   [1, 0, 0],\n                                   [1, 0, 0]])\n\nWhy would the LB give lower score?      \n              \n(Note: CV based on current competition metric will always be lower than CV of metric that averages individual f1 score of each question - given metric is conservative)",
      "votes": null
    },
    {
      "id": "2306769",
      "postDate": "06/17/2023 14:58:56",
      "content": "<p>Thank you for your clarification!</p>",
      "rawMarkdown": "Thank you for your clarification!",
      "votes": null
    },
    {
      "id": "2307415",
      "postDate": "06/18/2023 07:07:54",
      "content": "<p>I see what you mean <a href=\"https://www.kaggle.com/woprime\" target=\"_blank\">@woprime</a> 🙂. It is definitely easy to overfit in this competition. As soon as our CV is in contact with the Labels, there are unsuspected possibilities of overfitting.</p>",
      "rawMarkdown": "I see what you mean @woprime 🙂. It is definitely easy to overfit in this competition. As soon as our CV is in contact with the Labels, there are unsuspected possibilities of overfitting.",
      "votes": null
    },
    {
      "id": "2307421",
      "postDate": "06/18/2023 07:19:40",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a>!</p>",
      "rawMarkdown": "Thank you very much @kononenko!",
      "votes": null
    },
    {
      "id": "2309272",
      "postDate": "06/19/2023 13:52:04",
      "content": "<p>Can this be something to do with best_threshold? Do you find a new threshold after mixing both the models or use separate thresholds to convert probabilities?</p>",
      "rawMarkdown": "Can this be something to do with best_threshold? Do you find a new threshold after mixing both the models or use separate thresholds to convert probabilities?",
      "votes": null
    },
    {
      "id": "2310787",
      "postDate": "06/20/2023 15:51:43",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/murugesann\" target=\"_blank\">@murugesann</a>. Regarding the threshold specifically, I tried to select the best threshold for each question to optimize the  F1 score of each question looking at the CV but the results in the LB was awful. Here, I suspect it is more than overfitting, it is because the best \"harmony\" between the precision and the recall for the 18 questions altogether cannot be obtained by adjusting the best harmony for each question independently. In other words, the global perfect balance between precision and recall cannot be reached by getting independently all the 18 local perfect balances.</p>\n<p>Then I tried to find the best threshold for each question to optimize the global CV F1 score for the 18 questions altogether. The LB score was not awful, but it didn't really improve it (here it is probably simply overfitting, because by doing this we put the training data in contact with the labels).</p>",
      "rawMarkdown": "Hi @murugesann. Regarding the threshold specifically, I tried to select the best threshold for each question to optimize the  F1 score of each question looking at the CV but the results in the LB was awful. Here, I suspect it is more than overfitting, it is because the best \"harmony\" between the precision and the recall for the 18 questions altogether cannot be obtained by adjusting the best harmony for each question independently. In other words, the global perfect balance between precision and recall cannot be reached by getting independently all the 18 local perfect balances.\n\nThen I tried to find the best threshold for each question to optimize the global CV F1 score for the 18 questions altogether. The LB score was not awful, but it didn't really improve it (here it is probably simply overfitting, because by doing this we put the training data in contact with the labels).",
      "votes": null
    },
    {
      "id": "2310805",
      "postDate": "06/20/2023 16:01:49",
      "content": "<p>I am combining a transformer model with catboost model, as using only transformer model times out in submission after 9 hours. I  find the overall LB score less than the score obtained by catboost model alone!</p>\n<p>It is not clear whether the transformer models are performing badly (they are giving above 0.8 for all questions in validation dataset) or whether it is due to the phenomenon mentioned above. As of now, I still believe it is not because of mixing and optimizing, but transformer models themselves are performing lower, perhaps due to overfitting, to bring down the combined score.</p>\n<p>For example, the best public notebook with catboost has given a score 0.7 in LB but when combined with transformers I get 0.692 :-(  </p>\n<p>I  had used a threshold of 0.5 for transformer models, and 0.62 for catboost which could also be the reason…( a sanity check with different threshold for transformer models did not give better score so I use 0.5)</p>\n<p>One experiment - calculate overall threshold after using the both models predictions  for traininig data - which I am trying now - have you tried this?</p>",
      "rawMarkdown": "I am combining a transformer model with catboost model, as using only transformer model times out in submission after 9 hours. I  find the overall LB score less than the score obtained by catboost model alone!\n\nIt is not clear whether the transformer models are performing badly (they are giving above 0.8 for all questions in validation dataset) or whether it is due to the phenomenon mentioned above. As of now, I still believe it is not because of mixing and optimizing, but transformer models themselves are performing lower, perhaps due to overfitting, to bring down the combined score.\n\nFor example, the best public notebook with catboost has given a score 0.7 in LB but when combined with transformers I get 0.692 :-(  \n\nI  had used a threshold of 0.5 for transformer models, and 0.62 for catboost which could also be the reason...( a sanity check with different threshold for transformer models did not give better score so I use 0.5)\n\nOne experiment - calculate overall threshold after using the both models predictions  for traininig data - which I am trying now - have you tried this?",
      "votes": null
    },
    {
      "id": "2312177",
      "postDate": "06/21/2023 17:52:05",
      "content": "<p>It looks like your Transformer models are overfitting. 0.8 is way above anything achieved. Are you using F1-score with average='macro'?</p>",
      "rawMarkdown": "It looks like your Transformer models are overfitting. 0.8 is way above anything achieved. Are you using F1-score with average='macro'?",
      "votes": null
    },
    {
      "id": "2313176",
      "postDate": "06/22/2023 14:03:12",
      "content": "<p>yes, I got mired with overcoming the memory and inference time problems for transformer models - all initial models timed out - my objective was to get some workable model done (I joined this competition 10 days or so ago) and if anything goes wrong, the LB score will reveal -  that I did not notice that I was referring to F1 score of the majority class (though I have done a complete analysis of macro score for binary classification)!! So, stupid of me…happens frequently in data science for me now-a-days!! </p>\n<p>My transformer models are not giving f1 macro score beyond 0.62 or so…though I have not done any hyper-tuning or parameter changes for the basic model..</p>",
      "rawMarkdown": "yes, I got mired with overcoming the memory and inference time problems for transformer models - all initial models timed out - my objective was to get some workable model done (I joined this competition 10 days or so ago) and if anything goes wrong, the LB score will reveal -  that I did not notice that I was referring to F1 score of the majority class (though I have done a complete analysis of macro score for binary classification)!! So, stupid of me...happens frequently in data science for me now-a-days!! \n\nMy transformer models are not giving f1 macro score beyond 0.62 or so...though I have not done any hyper-tuning or parameter changes for the basic model..",
      "votes": null
    },
    {
      "id": "2313184",
      "postDate": "06/22/2023 14:09:19",
      "content": "<p>Also, it looks like 'best_threshold' concept does not impact much on LB score.  I changed the threshold from 0.5 to 0.63 (which is the threshold used for rest of the models) for four transformer models, the LB score came out to be exactly same. So, may be those who were finding optimizing not working may investigate other aspects of their scoring methodology…</p>",
      "rawMarkdown": "Also, it looks like 'best_threshold' concept does not impact much on LB score.  I changed the threshold from 0.5 to 0.63 (which is the threshold used for rest of the models) for four transformer models, the LB score came out to be exactly same. So, may be those who were finding optimizing not working may investigate other aspects of their scoring methodology...",
      "votes": null
    },
    {
      "id": "2313531",
      "postDate": "06/22/2023 18:18:56",
      "content": "<p>10 days is not much for this competition which is very time consuming. Good luck!</p>",
      "rawMarkdown": "10 days is not much for this competition which is very time consuming. Good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2303903,
      "author_name": "xiaosufrankhu",
      "author_url": "",
      "post_date": "06/15/2023 14:39:24",
      "content": "<p>Thank you for your very clear clarification! </p>\n<p>I have a small question related to your statement: \"The result was very surprising, the overall CV F1 score was worse than with only one single good model for all questions!!\" Were you trying to say one model for all 18 questions got better score than one model per question set-up?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2304118,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "06/15/2023 17:34:04",
          "content": "<p>Yes exactly <a href=\"https://www.kaggle.com/xiaosufrankhu\" target=\"_blank\">@xiaosufrankhu</a> . Let's say there are two models M1 and M2. M1 has a better overall F1 score. <br>\nF1(overall,M1) &gt; F1(overall,M2)<br>\nbut on a question basis let's say M2 has a better F1 score than M1 for question 2,5,7.<br>\nF1(q,M2) &gt; F1(q,M1)  for q in [2,5,7]</p>\n<p>Then you may think that a solution where you use M2 for questions [2,5,7] and M1 for all the other questions will give you a better score than M1 alone for all questions. But in my tests It was not the case, EVEN at the CV level. So it was not about overfitting but about the way the scoring works. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2306769,
              "author_name": "xiaosufrankhu",
              "author_url": "",
              "post_date": "06/17/2023 14:58:56",
              "content": "<p>Thank you for your clarification!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2309272,
              "author_name": "murugesann",
              "author_url": "",
              "post_date": "06/19/2023 13:52:04",
              "content": "<p>Can this be something to do with best_threshold? Do you find a new threshold after mixing both the models or use separate thresholds to convert probabilities?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2310787,
                  "author_name": "gehallak",
                  "author_url": "",
                  "post_date": "06/20/2023 15:51:43",
                  "content": "<p>Hi <a href=\"https://www.kaggle.com/murugesann\" target=\"_blank\">@murugesann</a>. Regarding the threshold specifically, I tried to select the best threshold for each question to optimize the  F1 score of each question looking at the CV but the results in the LB was awful. Here, I suspect it is more than overfitting, it is because the best \"harmony\" between the precision and the recall for the 18 questions altogether cannot be obtained by adjusting the best harmony for each question independently. In other words, the global perfect balance between precision and recall cannot be reached by getting independently all the 18 local perfect balances.</p>\n<p>Then I tried to find the best threshold for each question to optimize the global CV F1 score for the 18 questions altogether. The LB score was not awful, but it didn't really improve it (here it is probably simply overfitting, because by doing this we put the training data in contact with the labels).</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2310805,
                      "author_name": "murugesann",
                      "author_url": "",
                      "post_date": "06/20/2023 16:01:49",
                      "content": "<p>I am combining a transformer model with catboost model, as using only transformer model times out in submission after 9 hours. I  find the overall LB score less than the score obtained by catboost model alone!</p>\n<p>It is not clear whether the transformer models are performing badly (they are giving above 0.8 for all questions in validation dataset) or whether it is due to the phenomenon mentioned above. As of now, I still believe it is not because of mixing and optimizing, but transformer models themselves are performing lower, perhaps due to overfitting, to bring down the combined score.</p>\n<p>For example, the best public notebook with catboost has given a score 0.7 in LB but when combined with transformers I get 0.692 :-(  </p>\n<p>I  had used a threshold of 0.5 for transformer models, and 0.62 for catboost which could also be the reason…( a sanity check with different threshold for transformer models did not give better score so I use 0.5)</p>\n<p>One experiment - calculate overall threshold after using the both models predictions  for traininig data - which I am trying now - have you tried this?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2312177,
                          "author_name": "gehallak",
                          "author_url": "",
                          "post_date": "06/21/2023 17:52:05",
                          "content": "<p>It looks like your Transformer models are overfitting. 0.8 is way above anything achieved. Are you using F1-score with average='macro'?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2313176,
                              "author_name": "murugesann",
                              "author_url": "",
                              "post_date": "06/22/2023 14:03:12",
                              "content": "<p>yes, I got mired with overcoming the memory and inference time problems for transformer models - all initial models timed out - my objective was to get some workable model done (I joined this competition 10 days or so ago) and if anything goes wrong, the LB score will reveal -  that I did not notice that I was referring to F1 score of the majority class (though I have done a complete analysis of macro score for binary classification)!! So, stupid of me…happens frequently in data science for me now-a-days!! </p>\n<p>My transformer models are not giving f1 macro score beyond 0.62 or so…though I have not done any hyper-tuning or parameter changes for the basic model..</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2313184,
                                  "author_name": "murugesann",
                                  "author_url": "",
                                  "post_date": "06/22/2023 14:09:19",
                                  "content": "<p>Also, it looks like 'best_threshold' concept does not impact much on LB score.  I changed the threshold from 0.5 to 0.63 (which is the threshold used for rest of the models) for four transformer models, the LB score came out to be exactly same. So, may be those who were finding optimizing not working may investigate other aspects of their scoring methodology…</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2313531,
                                      "author_name": "gehallak",
                                      "author_url": "",
                                      "post_date": "06/22/2023 18:18:56",
                                      "content": "<p>10 days is not much for this competition which is very time consuming. Good luck!</p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2304158,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "06/15/2023 18:10:18",
      "content": "<p>I think the scoring function is not the main cause of this problem. </p>\n<p>I've tried to optimize for validation macro F1 by first training a model for initial prediction, then updating predictions for each question by selecting the model which results in most macro F1 gain. I think by doing this I raised CV score by ~0.002 - 0.003, but this did not generalize to LB at all. I've also tested my suspicion on a held-out test set, and not to my surprise, the test also hasn't benefited from this CV score gain.</p>\n<p>So I guess the lesson here is, don't trust your CV, trust your CV's CV. 🙃</p>",
      "votes": null,
      "replies": [
        {
          "id": 2307415,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "06/18/2023 07:07:54",
          "content": "<p>I see what you mean <a href=\"https://www.kaggle.com/woprime\" target=\"_blank\">@woprime</a> 🙂. It is definitely easy to overfit in this competition. As soon as our CV is in contact with the Labels, there are unsuspected possibilities of overfitting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2304345,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "06/16/2023 00:08:32",
      "content": "<p>Thanks for the insightful post. If you don't mind, I have posted a <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478\" target=\"_blank\">sequel</a> with some math behind it :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2307421,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "06/18/2023 07:19:40",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2304412,
      "author_name": "murugesann",
      "author_url": "",
      "post_date": "06/16/2023 01:38:34",
      "content": "<p>I don't think optimizing models based on each question has anything to do with metric</p>\n<p>Let us say the ground truth  and predictions are as follows (consider 3 questions, 4 samples) - using a single model - hypothetical case where question 3 alone is predicted completely wrongly:</p>\n<p>y_true = np.array([ [1, 1, 1],<br>\n                               [1, 0, 1],<br>\n                               [1, 0, 0],<br>\n                               [1, 0, 0] ]) </p>\n<p>y_pred1 =  np.array([[1, 1, 0],<br>\n                                  [1, 0, 0],<br>\n                                  [1, 0, 1],<br>\n                                  [1, 0, 1]])</p>\n<p>Now, suppose you use separate model for each question and arrive at predictions as below (where third question is also predicted perfectly)</p>\n<p>y_pred2 =  np.array([[1, 1, 1],<br>\n                                   [1, 0, 1],<br>\n                                   [1, 0, 0],<br>\n                                   [1, 0, 0]])</p>\n<p>Why would the LB give lower score?      </p>\n<p>(Note: CV based on current competition metric will always be lower than CV of metric that averages individual f1 score of each question - given metric is conservative)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2303623": "Have you tried to optimize the F1 score at each question independently and then put it all together?\n**For example, did you try any of the following:**\n\n- Select the best Model for each question\n- Select the best Hyperparameters for each question\n- Select the best Threshold for each question\n- Select the best Features for each questions\n\nI tried all of these and most of the times the final LB score was worse!!\nMy intuitive explanation was: all these techniques were overfitting the train data.\nBut then I had an even more bizarre result. I was selecting, for each question, the model that was giving the best F1-score at the question level.  The result was very surprising, the overall CV F1 score was worse than with only one single good model for all questions!! so even before submitting to the LB, at the CV level, it was not working. Overfitting could not be the explanation. \n**There was no choice but to dig into the scoring mechanism.**\nBelow is a graph showing how it works.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F7c388f97ad6f1d4faccd3d058662966d%2FF1Score.png?generation=1686826873533216&alt=media)\n\nThe key is the **harmonic mean** . It is very different from a simple mean, it is greatly penalized by the worst. **If the recall is very bad and the precision is excellent, the F1 will be very bad.**\nIt is exactly what was happening. One model was giving a very bad recall, but an excellent precision for one question. Its F1 was low. So I was discarding it, but actually, **when diluted with all the other questions,** the bad recall didn't have such a negative impact in regards with the beneficial impact on the overall precision.\n\n**Conclusion:**\nThe score in this competition is the F1 score macro on the overall questions taken altogether. It is very different from the average F1 scores for each question (otherwise it would simply be 18 separate competitions taken together). All the questions are put together in a same bag and then the F1 score is calculated.\n\nThe name **harmonic mean** has really been well chosen. You get a good score only if you have a model that harmoniously treats the precision and the recall. It has to find the right balance. BUT, a perfect harmony at each individual question doesn't necessarily mean the best harmony for the 18 questions taken altogether.",
    "2303903": "Thank you for your very clear clarification! \n\nI have a small question related to your statement: \"The result was very surprising, the overall CV F1 score was worse than with only one single good model for all questions!!\" Were you trying to say one model for all 18 questions got better score than one model per question set-up?",
    "2304118": "Yes exactly @xiaosufrankhu . Let's say there are two models M1 and M2. M1 has a better overall F1 score. \nF1(overall,M1) > F1(overall,M2)\nbut on a question basis let's say M2 has a better F1 score than M1 for question 2,5,7.\nF1(q,M2) > F1(q,M1)  for q in [2,5,7]\n\nThen you may think that a solution where you use M2 for questions [2,5,7] and M1 for all the other questions will give you a better score than M1 alone for all questions. But in my tests It was not the case, EVEN at the CV level. So it was not about overfitting but about the way the scoring works.",
    "2304158": "I think the scoring function is not the main cause of this problem. \n\nI've tried to optimize for validation macro F1 by first training a model for initial prediction, then updating predictions for each question by selecting the model which results in most macro F1 gain. I think by doing this I raised CV score by ~0.002 - 0.003, but this did not generalize to LB at all. I've also tested my suspicion on a held-out test set, and not to my surprise, the test also hasn't benefited from this CV score gain.\n\nSo I guess the lesson here is, don't trust your CV, trust your CV's CV. 🙃",
    "2304345": "Thanks for the insightful post. If you don't mind, I have posted a [sequel](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417478) with some math behind it :)",
    "2304412": "I don't think optimizing models based on each question has anything to do with metric\n\nLet us say the ground truth  and predictions are as follows (consider 3 questions, 4 samples) - using a single model - hypothetical case where question 3 alone is predicted completely wrongly:\n\ny_true = np.array([ [1, 1, 1],\n                               [1, 0, 1],\n                               [1, 0, 0],\n                               [1, 0, 0] ]) \n\ny_pred1 =  np.array([[1, 1, 0],\n                                  [1, 0, 0],\n                                  [1, 0, 1],\n                                  [1, 0, 1]])\n\nNow, suppose you use separate model for each question and arrive at predictions as below (where third question is also predicted perfectly)\n\ny_pred2 =  np.array([[1, 1, 1],\n                                   [1, 0, 1],\n                                   [1, 0, 0],\n                                   [1, 0, 0]])\n\nWhy would the LB give lower score?      \n              \n(Note: CV based on current competition metric will always be lower than CV of metric that averages individual f1 score of each question - given metric is conservative)",
    "2306769": "Thank you for your clarification!",
    "2307415": "I see what you mean @woprime 🙂. It is definitely easy to overfit in this competition. As soon as our CV is in contact with the Labels, there are unsuspected possibilities of overfitting.",
    "2307421": "Thank you very much @kononenko!",
    "2309272": "Can this be something to do with best_threshold? Do you find a new threshold after mixing both the models or use separate thresholds to convert probabilities?",
    "2310787": "Hi @murugesann. Regarding the threshold specifically, I tried to select the best threshold for each question to optimize the  F1 score of each question looking at the CV but the results in the LB was awful. Here, I suspect it is more than overfitting, it is because the best \"harmony\" between the precision and the recall for the 18 questions altogether cannot be obtained by adjusting the best harmony for each question independently. In other words, the global perfect balance between precision and recall cannot be reached by getting independently all the 18 local perfect balances.\n\nThen I tried to find the best threshold for each question to optimize the global CV F1 score for the 18 questions altogether. The LB score was not awful, but it didn't really improve it (here it is probably simply overfitting, because by doing this we put the training data in contact with the labels).",
    "2310805": "I am combining a transformer model with catboost model, as using only transformer model times out in submission after 9 hours. I  find the overall LB score less than the score obtained by catboost model alone!\n\nIt is not clear whether the transformer models are performing badly (they are giving above 0.8 for all questions in validation dataset) or whether it is due to the phenomenon mentioned above. As of now, I still believe it is not because of mixing and optimizing, but transformer models themselves are performing lower, perhaps due to overfitting, to bring down the combined score.\n\nFor example, the best public notebook with catboost has given a score 0.7 in LB but when combined with transformers I get 0.692 :-(  \n\nI  had used a threshold of 0.5 for transformer models, and 0.62 for catboost which could also be the reason...( a sanity check with different threshold for transformer models did not give better score so I use 0.5)\n\nOne experiment - calculate overall threshold after using the both models predictions  for traininig data - which I am trying now - have you tried this?",
    "2312177": "It looks like your Transformer models are overfitting. 0.8 is way above anything achieved. Are you using F1-score with average='macro'?",
    "2313176": "yes, I got mired with overcoming the memory and inference time problems for transformer models - all initial models timed out - my objective was to get some workable model done (I joined this competition 10 days or so ago) and if anything goes wrong, the LB score will reveal -  that I did not notice that I was referring to F1 score of the majority class (though I have done a complete analysis of macro score for binary classification)!! So, stupid of me...happens frequently in data science for me now-a-days!! \n\nMy transformer models are not giving f1 macro score beyond 0.62 or so...though I have not done any hyper-tuning or parameter changes for the basic model..",
    "2313184": "Also, it looks like 'best_threshold' concept does not impact much on LB score.  I changed the threshold from 0.5 to 0.63 (which is the threshold used for rest of the models) for four transformer models, the LB score came out to be exactly same. So, may be those who were finding optimizing not working may investigate other aspects of their scoring methodology...",
    "2313531": "10 days is not much for this competition which is very time consuming. Good luck!"
  },
  "source": "meta"
}