{
  "id": 391062,
  "title": "What happen with F1-score for all question?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/391062",
  "author_name": "",
  "post_date": "2023-02-28T09:22:36.815863700Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Reference <a href=\"https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678\" target=\"_blank\">https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678</a><br>\nThe F1 score for each question like that<br>\n`When using optimal threshold…</p>\n<p>Q0: F1 = 0.592719804577749<br>\nQ1: F1 = 0.4946370344945941<br>\nQ2: F1 = 0.48868007326992496<br>\nQ3: F1 = 0.6120129283127926<br>\nQ4: F1 = 0.5600008987023111<br>\nQ5: F1 = 0.6067685167749097<br>\nQ6: F1 = 0.5913329448001666<br>\nQ7: F1 = 0.5221774461804971<br>\nQ8: F1 = 0.6022998262503276<br>\nQ9: F1 = 0.48323984605671066<br>\nQ10: F1 = 0.5959341670818572<br>\nQ11: F1 = 0.4918731493595805<br>\nQ12: F1 = 0.4330470569811089<br>\nQ13: F1 = 0.6079573795339904<br>\nQ14: F1 = 0.4876872946307056<br>\nQ15: F1 = 0.4553931721004565<br>\nQ16: F1 = 0.5412756417017907<br>\nQ17: F1 = 0.48889449728140544<br>\n==&gt; Overall F1 = 0.6785645580477335`<br>\nBut the average|max F1 - score is less than the Overall F1, I don't know why? <br>\nCan Everyone explain it to me? Thank you</p>",
  "messages": [
    {
      "id": "2162537",
      "postDate": "02/28/2023 09:22:36",
      "content": "<p>Reference <a href=\"https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678\" target=\"_blank\">https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678</a><br>\nThe F1 score for each question like that<br>\n`When using optimal threshold…</p>\n<p>Q0: F1 = 0.592719804577749<br>\nQ1: F1 = 0.4946370344945941<br>\nQ2: F1 = 0.48868007326992496<br>\nQ3: F1 = 0.6120129283127926<br>\nQ4: F1 = 0.5600008987023111<br>\nQ5: F1 = 0.6067685167749097<br>\nQ6: F1 = 0.5913329448001666<br>\nQ7: F1 = 0.5221774461804971<br>\nQ8: F1 = 0.6022998262503276<br>\nQ9: F1 = 0.48323984605671066<br>\nQ10: F1 = 0.5959341670818572<br>\nQ11: F1 = 0.4918731493595805<br>\nQ12: F1 = 0.4330470569811089<br>\nQ13: F1 = 0.6079573795339904<br>\nQ14: F1 = 0.4876872946307056<br>\nQ15: F1 = 0.4553931721004565<br>\nQ16: F1 = 0.5412756417017907<br>\nQ17: F1 = 0.48889449728140544<br>\n==&gt; Overall F1 = 0.6785645580477335`<br>\nBut the average|max F1 - score is less than the Overall F1, I don't know why? <br>\nCan Everyone explain it to me? Thank you</p>",
      "rawMarkdown": "Reference https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678\nThe F1 score for each question like that\n`When using optimal threshold...\n\nQ0: F1 = 0.592719804577749\nQ1: F1 = 0.4946370344945941\nQ2: F1 = 0.48868007326992496\nQ3: F1 = 0.6120129283127926\nQ4: F1 = 0.5600008987023111\nQ5: F1 = 0.6067685167749097\nQ6: F1 = 0.5913329448001666\nQ7: F1 = 0.5221774461804971\nQ8: F1 = 0.6022998262503276\nQ9: F1 = 0.48323984605671066\nQ10: F1 = 0.5959341670818572\nQ11: F1 = 0.4918731493595805\nQ12: F1 = 0.4330470569811089\nQ13: F1 = 0.6079573795339904\nQ14: F1 = 0.4876872946307056\nQ15: F1 = 0.4553931721004565\nQ16: F1 = 0.5412756417017907\nQ17: F1 = 0.48889449728140544\n==> Overall F1 = 0.6785645580477335`\nBut the average|max F1 - score is less than the Overall F1, I don't know why? \nCan Everyone explain it to me? Thank you",
      "votes": null
    },
    {
      "id": "2164025",
      "postDate": "03/01/2023 08:53:16",
      "content": "<p><code>F1</code> is calculated based on <code>precision</code> and <code>recall</code> (<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\" target=\"_blank\">source</a>). For each different <code>threshold</code>, we have different <code>precision</code> and <code>recall</code>, and thus different <code>F1</code>.</p>\n<p>When we choose a <code>best_overall_threshold</code> to have maximum <code>overall_F1</code>, we are balancing out between <code>overall_precision</code> and <code>overall_recall</code>. However, such <code>best_overall_threshold</code>, when applied to each question, doesn't guarantee the same balance, and therefore lowers the individual questions' <code>F1</code>s. That's why you saw <code>overall_F1</code> being higher than all individual <code>F1</code>s</p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217\" target=\"_blank\">People have tried using different thresholds for different questions</a>, but so far it doesn't seem as optimal as using a single overall threshold.</p>\n<p>Hope this helps!</p>",
      "rawMarkdown": "`F1` is calculated based on `precision` and `recall` ([source](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html)). For each different `threshold`, we have different `precision` and `recall`, and thus different `F1`.\n\nWhen we choose a `best_overall_threshold` to have maximum `overall_F1`, we are balancing out between `overall_precision` and `overall_recall`. However, such `best_overall_threshold`, when applied to each question, doesn't guarantee the same balance, and therefore lowers the individual questions' `F1`s. That's why you saw `overall_F1` being higher than all individual `F1`s\n\n[People have tried using different thresholds for different questions](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217), but so far it doesn't seem as optimal as using a single overall threshold.\n\nHope this helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2164025,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "03/01/2023 08:53:16",
      "content": "<p><code>F1</code> is calculated based on <code>precision</code> and <code>recall</code> (<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\" target=\"_blank\">source</a>). For each different <code>threshold</code>, we have different <code>precision</code> and <code>recall</code>, and thus different <code>F1</code>.</p>\n<p>When we choose a <code>best_overall_threshold</code> to have maximum <code>overall_F1</code>, we are balancing out between <code>overall_precision</code> and <code>overall_recall</code>. However, such <code>best_overall_threshold</code>, when applied to each question, doesn't guarantee the same balance, and therefore lowers the individual questions' <code>F1</code>s. That's why you saw <code>overall_F1</code> being higher than all individual <code>F1</code>s</p>\n<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217\" target=\"_blank\">People have tried using different thresholds for different questions</a>, but so far it doesn't seem as optimal as using a single overall threshold.</p>\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2162537": "Reference https://www.kaggle.com/code/starfalllover/xgboost-baseline-0-678\nThe F1 score for each question like that\n`When using optimal threshold...\n\nQ0: F1 = 0.592719804577749\nQ1: F1 = 0.4946370344945941\nQ2: F1 = 0.48868007326992496\nQ3: F1 = 0.6120129283127926\nQ4: F1 = 0.5600008987023111\nQ5: F1 = 0.6067685167749097\nQ6: F1 = 0.5913329448001666\nQ7: F1 = 0.5221774461804971\nQ8: F1 = 0.6022998262503276\nQ9: F1 = 0.48323984605671066\nQ10: F1 = 0.5959341670818572\nQ11: F1 = 0.4918731493595805\nQ12: F1 = 0.4330470569811089\nQ13: F1 = 0.6079573795339904\nQ14: F1 = 0.4876872946307056\nQ15: F1 = 0.4553931721004565\nQ16: F1 = 0.5412756417017907\nQ17: F1 = 0.48889449728140544\n==> Overall F1 = 0.6785645580477335`\nBut the average|max F1 - score is less than the Overall F1, I don't know why? \nCan Everyone explain it to me? Thank you",
    "2164025": "`F1` is calculated based on `precision` and `recall` ([source](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html)). For each different `threshold`, we have different `precision` and `recall`, and thus different `F1`.\n\nWhen we choose a `best_overall_threshold` to have maximum `overall_F1`, we are balancing out between `overall_precision` and `overall_recall`. However, such `best_overall_threshold`, when applied to each question, doesn't guarantee the same balance, and therefore lowers the individual questions' `F1`s. That's why you saw `overall_F1` being higher than all individual `F1`s\n\n[People have tried using different thresholds for different questions](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217), but so far it doesn't seem as optimal as using a single overall threshold.\n\nHope this helps!"
  },
  "source": "meta"
}