{
  "id": 391837,
  "title": "Different hyper-parameters for each question's model",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/391837",
  "author_name": "",
  "post_date": "2023-03-02T18:53:06.033669900Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have tried to <strong>optimise each question separately</strong> to improve overall F1 score:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F5176185e2ed37b2170c29398b031bfcd%2FScreenshot_7.jpg?generation=1677782997268274&amp;alt=media\" alt=\"\"><br>\n<em>(I didn't use the best_threshold after seeing <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217\" target=\"_blank\">this</a> notebook)</em><br>\nBut my global F1 dropped significantly, any idea why is that ? souldn't the model be more accurate globally ?<br>\n/<br>\n<strong>Note: I am using a XGB model.</strong></p>",
  "messages": [
    {
      "id": "2166381",
      "postDate": "03/02/2023 18:53:06",
      "content": "<p>I have tried to <strong>optimise each question separately</strong> to improve overall F1 score:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F5176185e2ed37b2170c29398b031bfcd%2FScreenshot_7.jpg?generation=1677782997268274&amp;alt=media\" alt=\"\"><br>\n<em>(I didn't use the best_threshold after seeing <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217\" target=\"_blank\">this</a> notebook)</em><br>\nBut my global F1 dropped significantly, any idea why is that ? souldn't the model be more accurate globally ?<br>\n/<br>\n<strong>Note: I am using a XGB model.</strong></p>",
      "rawMarkdown": "I have tried to **optimise each question separately** to improve overall F1 score:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F5176185e2ed37b2170c29398b031bfcd%2FScreenshot_7.jpg?generation=1677782997268274&alt=media)\n*(I didn't use the best_threshold after seeing [this](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217) notebook)*\nBut my global F1 dropped significantly, any idea why is that ? souldn't the model be more accurate globally ?\n/\n**Note: I am using a XGB model.**",
      "votes": null
    },
    {
      "id": "2166455",
      "postDate": "03/02/2023 20:06:48",
      "content": "<p>I think you may optimize each model, but you have to do it at the same time: not one model at a time. </p>",
      "rawMarkdown": "I think you may optimize each model, but you have to do it at the same time: not one model at a time.",
      "votes": null
    },
    {
      "id": "2166615",
      "postDate": "03/02/2023 21:48:59",
      "content": "<p>Thanks. Can you pleae guide us with a way to do it?</p>",
      "rawMarkdown": "Thanks. Can you pleae guide us with a way to do it?",
      "votes": null
    },
    {
      "id": "2168941",
      "postDate": "03/04/2023 16:46:02",
      "content": "<p>How do you \"optimize\" lr? The lower you set it - the better. It's all about the computational budget. 0.3 and 0.1 do look too high. Set lr at something like 0.02 and use the validation set to do the early stopping.</p>",
      "rawMarkdown": "How do you \"optimize\" lr? The lower you set it - the better. It's all about the computational budget. 0.3 and 0.1 do look too high. Set lr at something like 0.02 and use the validation set to do the early stopping.",
      "votes": null
    },
    {
      "id": "2170853",
      "postDate": "03/06/2023 10:44:48",
      "content": "<p>actually, in theory that it how it should work<br>\nbut as you can see, lr is not the same for each question, and even the global f1 isn't obtained when the lower lr is set because of overfitting.</p>",
      "rawMarkdown": "actually, in theory that it how it should work\nbut as you can see, lr is not the same for each question, and even the global f1 isn't obtained when the lower lr is set because of overfitting.",
      "votes": null
    },
    {
      "id": "2202166",
      "postDate": "03/29/2023 19:42:10",
      "content": "<p>Not sure what you are doing, but the optimization of 'best_threshold' might be partially to blame.  If you are doing something like this:</p>\n<ol>\n<li>Create my features</li>\n<li>FOREACH question:<ul>\n<li>Fit a model forecasting each question</li>\n<li>Find the best threshold for marking as 0 or 1 to optimize the F1 score on this question</li></ul></li>\n<li>Calculate the total F1 score with best parameters</li>\n</ol>\n<p>This will fail because the F1 score is macro averaged.  That is, the true values of 1 F1 score is calculated, then the true values of 0 F1 score is calculated, then those two F1 scores are averaged.</p>\n<p>So for some questions, like #2, which most everyone gets right, it is much more important to forecast the people who get it wrong when you are fitting the threshold of the individual question.  However, in the overall scheme when the number of wrong answers for #2 is lower, this will not matter as much.</p>\n<p>Try looking at the output of <code>classification_report</code> in the <code>sklearn.metrics</code> module.</p>\n<p>If you want to individually optimize thresholds, code like this might be better</p>\n<p>First get a dataframe <code>predictions</code> with your predictions and the true values with columns like: <br>\n<code>[index, session_id, correct, question#, prediction_prob]</code></p>\n<p>Then something like this will work (Excuse the poor formatting):<br>\n<code>import scipy.optimize as spo</code><br>\n<code>def fun(x):</code><br>\n<code>\\t for q in range(1,19):</code><br>\n<code>\\t\\t        thresh = x[q-1]</code><br>\n<code>\\t\\t        mask = predictions['question'] == q</code><br>\n<code>\\t\\t        predictions.loc[mask, 'temp'] = predictions.loc[mask, 'pred'].map(lambda x: 1 if x &gt; thresh else 0)</code><br>\n<code>\\t    m = f1_score(predictions['correct'].values, predictions['temp'].values, average='macro')</code><br>\n<code>\\t    return -m</code></p>\n<p><code>x0 = np.ones(18)*0.5 + np.random.normal(scale=0.01, size=18)</code></p>\n<p><code>soln = spo.minimize(fun=fun, x0=x0, method='Nelder-Mead', options={'disp': True, 'fatol': 1.e-5})</code></p>",
      "rawMarkdown": "Not sure what you are doing, but the optimization of 'best_threshold' might be partially to blame.  If you are doing something like this:\n\n1. Create my features\n2. FOREACH question:\n    - Fit a model forecasting each question\n    - Find the best threshold for marking as 0 or 1 to optimize the F1 score on this question\n3. Calculate the total F1 score with best parameters\n\nThis will fail because the F1 score is macro averaged.  That is, the true values of 1 F1 score is calculated, then the true values of 0 F1 score is calculated, then those two F1 scores are averaged.\n\nSo for some questions, like #2, which most everyone gets right, it is much more important to forecast the people who get it wrong when you are fitting the threshold of the individual question.  However, in the overall scheme when the number of wrong answers for #2 is lower, this will not matter as much.\n\nTry looking at the output of `classification_report` in the `sklearn.metrics` module.\n\nIf you want to individually optimize thresholds, code like this might be better\n\nFirst get a dataframe `predictions` with your predictions and the true values with columns like: \n`[index, session_id, correct, question#, prediction_prob]`\n\nThen something like this will work (Excuse the poor formatting):\n`import scipy.optimize as spo`\n`def fun(x):`\n`\\t for q in range(1,19):`\n`\\t\\t        thresh = x[q-1]`\n`\\t\\t        mask = predictions['question'] == q`\n`\\t\\t        predictions.loc[mask, 'temp'] = predictions.loc[mask, 'pred'].map(lambda x: 1 if x > thresh else 0)`\n`\\t    m = f1_score(predictions['correct'].values, predictions['temp'].values, average='macro')`\n`\\t    return -m`\n\n`x0 = np.ones(18)*0.5 + np.random.normal(scale=0.01, size=18)`\n\n`soln = spo.minimize(fun=fun, x0=x0, method='Nelder-Mead', options={'disp': True, 'fatol': 1.e-5})`",
      "votes": null
    },
    {
      "id": "2202775",
      "postDate": "03/30/2023 10:13:58",
      "content": "<p>I've explored this idea, but I optimized for accuracy rather than F1 score for each model, in many cases I can get a better CV. I haven't checked on the math, but my intuition tells me that accuracy correlates more with macro F1 score in different subsets (questions) with varying positive-negative ratios.</p>",
      "rawMarkdown": "I've explored this idea, but I optimized for accuracy rather than F1 score for each model, in many cases I can get a better CV. I haven't checked on the math, but my intuition tells me that accuracy correlates more with macro F1 score in different subsets (questions) with varying positive-negative ratios.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2166455,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "03/02/2023 20:06:48",
      "content": "<p>I think you may optimize each model, but you have to do it at the same time: not one model at a time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2166615,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "03/02/2023 21:48:59",
          "content": "<p>Thanks. Can you pleae guide us with a way to do it?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2168941,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "03/04/2023 16:46:02",
      "content": "<p>How do you \"optimize\" lr? The lower you set it - the better. It's all about the computational budget. 0.3 and 0.1 do look too high. Set lr at something like 0.02 and use the validation set to do the early stopping.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2170853,
          "author_name": "janmpia",
          "author_url": "",
          "post_date": "03/06/2023 10:44:48",
          "content": "<p>actually, in theory that it how it should work<br>\nbut as you can see, lr is not the same for each question, and even the global f1 isn't obtained when the lower lr is set because of overfitting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2202166,
      "author_name": "danielphalen",
      "author_url": "",
      "post_date": "03/29/2023 19:42:10",
      "content": "<p>Not sure what you are doing, but the optimization of 'best_threshold' might be partially to blame.  If you are doing something like this:</p>\n<ol>\n<li>Create my features</li>\n<li>FOREACH question:<ul>\n<li>Fit a model forecasting each question</li>\n<li>Find the best threshold for marking as 0 or 1 to optimize the F1 score on this question</li></ul></li>\n<li>Calculate the total F1 score with best parameters</li>\n</ol>\n<p>This will fail because the F1 score is macro averaged.  That is, the true values of 1 F1 score is calculated, then the true values of 0 F1 score is calculated, then those two F1 scores are averaged.</p>\n<p>So for some questions, like #2, which most everyone gets right, it is much more important to forecast the people who get it wrong when you are fitting the threshold of the individual question.  However, in the overall scheme when the number of wrong answers for #2 is lower, this will not matter as much.</p>\n<p>Try looking at the output of <code>classification_report</code> in the <code>sklearn.metrics</code> module.</p>\n<p>If you want to individually optimize thresholds, code like this might be better</p>\n<p>First get a dataframe <code>predictions</code> with your predictions and the true values with columns like: <br>\n<code>[index, session_id, correct, question#, prediction_prob]</code></p>\n<p>Then something like this will work (Excuse the poor formatting):<br>\n<code>import scipy.optimize as spo</code><br>\n<code>def fun(x):</code><br>\n<code>\\t for q in range(1,19):</code><br>\n<code>\\t\\t        thresh = x[q-1]</code><br>\n<code>\\t\\t        mask = predictions['question'] == q</code><br>\n<code>\\t\\t        predictions.loc[mask, 'temp'] = predictions.loc[mask, 'pred'].map(lambda x: 1 if x &gt; thresh else 0)</code><br>\n<code>\\t    m = f1_score(predictions['correct'].values, predictions['temp'].values, average='macro')</code><br>\n<code>\\t    return -m</code></p>\n<p><code>x0 = np.ones(18)*0.5 + np.random.normal(scale=0.01, size=18)</code></p>\n<p><code>soln = spo.minimize(fun=fun, x0=x0, method='Nelder-Mead', options={'disp': True, 'fatol': 1.e-5})</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2202775,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "03/30/2023 10:13:58",
      "content": "<p>I've explored this idea, but I optimized for accuracy rather than F1 score for each model, in many cases I can get a better CV. I haven't checked on the math, but my intuition tells me that accuracy correlates more with macro F1 score in different subsets (questions) with varying positive-negative ratios.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2166381": "I have tried to **optimise each question separately** to improve overall F1 score:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F5176185e2ed37b2170c29398b031bfcd%2FScreenshot_7.jpg?generation=1677782997268274&alt=media)\n*(I didn't use the best_threshold after seeing [this](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/389217) notebook)*\nBut my global F1 dropped significantly, any idea why is that ? souldn't the model be more accurate globally ?\n/\n**Note: I am using a XGB model.**",
    "2166455": "I think you may optimize each model, but you have to do it at the same time: not one model at a time.",
    "2166615": "Thanks. Can you pleae guide us with a way to do it?",
    "2168941": "How do you \"optimize\" lr? The lower you set it - the better. It's all about the computational budget. 0.3 and 0.1 do look too high. Set lr at something like 0.02 and use the validation set to do the early stopping.",
    "2170853": "actually, in theory that it how it should work\nbut as you can see, lr is not the same for each question, and even the global f1 isn't obtained when the lower lr is set because of overfitting.",
    "2202166": "Not sure what you are doing, but the optimization of 'best_threshold' might be partially to blame.  If you are doing something like this:\n\n1. Create my features\n2. FOREACH question:\n    - Fit a model forecasting each question\n    - Find the best threshold for marking as 0 or 1 to optimize the F1 score on this question\n3. Calculate the total F1 score with best parameters\n\nThis will fail because the F1 score is macro averaged.  That is, the true values of 1 F1 score is calculated, then the true values of 0 F1 score is calculated, then those two F1 scores are averaged.\n\nSo for some questions, like #2, which most everyone gets right, it is much more important to forecast the people who get it wrong when you are fitting the threshold of the individual question.  However, in the overall scheme when the number of wrong answers for #2 is lower, this will not matter as much.\n\nTry looking at the output of `classification_report` in the `sklearn.metrics` module.\n\nIf you want to individually optimize thresholds, code like this might be better\n\nFirst get a dataframe `predictions` with your predictions and the true values with columns like: \n`[index, session_id, correct, question#, prediction_prob]`\n\nThen something like this will work (Excuse the poor formatting):\n`import scipy.optimize as spo`\n`def fun(x):`\n`\\t for q in range(1,19):`\n`\\t\\t        thresh = x[q-1]`\n`\\t\\t        mask = predictions['question'] == q`\n`\\t\\t        predictions.loc[mask, 'temp'] = predictions.loc[mask, 'pred'].map(lambda x: 1 if x > thresh else 0)`\n`\\t    m = f1_score(predictions['correct'].values, predictions['temp'].values, average='macro')`\n`\\t    return -m`\n\n`x0 = np.ones(18)*0.5 + np.random.normal(scale=0.01, size=18)`\n\n`soln = spo.minimize(fun=fun, x0=x0, method='Nelder-Mead', options={'disp': True, 'fatol': 1.e-5})`",
    "2202775": "I've explored this idea, but I optimized for accuracy rather than F1 score for each model, in many cases I can get a better CV. I haven't checked on the math, but my intuition tells me that accuracy correlates more with macro F1 score in different subsets (questions) with varying positive-negative ratios."
  },
  "source": "meta"
}