{
  "id": 73744,
  "title": "How to train to maximize F1 score instead of minimizing binary cross entropy",
  "url": "/competitions/quora-insincere-questions-classification/discussion/73744",
  "author_name": "",
  "post_date": "2018-12-05T09:52:05.475836300Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Since the submission is graded based on F1 score instead of accuracy, I was thinking if we should try changing the loss function somehow so that the F1 score gets maximized.\nOne candidate I can think of inverse of F1 score. \nBut I am not sure about the differentiability for the backprop.\nAlso it would be good if the model gave me the threshold value to threshold upon to classify as 0 vs 1</p>\n\n<p>My existing solution (and everything I found online) minimizes the binary cross-entropy loss function itself, but after predicting the prob values, based on the eval set, I decide the threshold to classify as 0 vs 1 so that the F1 score gets maximized. Then for the actual test set, I use the same threshold. The assumption here is that it will follow the exact same ratio of 0s and 1s (which may not be correct).</p>\n\n<p>All comments/suggestions/criticisms most welcome.</p>",
  "messages": [
    {
      "id": "433655",
      "postDate": "12/05/2018 09:52:05",
      "content": "<p>Since the submission is graded based on F1 score instead of accuracy, I was thinking if we should try changing the loss function somehow so that the F1 score gets maximized.\nOne candidate I can think of inverse of F1 score. \nBut I am not sure about the differentiability for the backprop.\nAlso it would be good if the model gave me the threshold value to threshold upon to classify as 0 vs 1</p>\n\n<p>My existing solution (and everything I found online) minimizes the binary cross-entropy loss function itself, but after predicting the prob values, based on the eval set, I decide the threshold to classify as 0 vs 1 so that the F1 score gets maximized. Then for the actual test set, I use the same threshold. The assumption here is that it will follow the exact same ratio of 0s and 1s (which may not be correct).</p>\n\n<p>All comments/suggestions/criticisms most welcome.</p>",
      "rawMarkdown": "Since the submission is graded based on F1 score instead of accuracy, I was thinking if we should try changing the loss function somehow so that the F1 score gets maximized.\nOne candidate I can think of inverse of F1 score. \nBut I am not sure about the differentiability for the backprop.\nAlso it would be good if the model gave me the threshold value to threshold upon to classify as 0 vs 1\n\nMy existing solution (and everything I found online) minimizes the binary cross-entropy loss function itself, but after predicting the prob values, based on the eval set, I decide the threshold to classify as 0 vs 1 so that the F1 score gets maximized. Then for the actual test set, I use the same threshold. The assumption here is that it will follow the exact same ratio of 0s and 1s (which may not be correct).\n\nAll comments/suggestions/criticisms most welcome.",
      "votes": null
    },
    {
      "id": "433851",
      "postDate": "12/05/2018 15:10:50",
      "content": "<p>I thought this too but when I tried it didn't work out well Here is the loss function I found on another kaggle kernel.</p>\n\n<p>def f1_loss(y_true, y_pred):</p>\n\n<pre><code>tp = K.sum(K.cast(y_true*y_pred, 'float'), axis=0)\n\nfp = K.sum(K.cast((1-y_true)*y_pred, 'float'), axis=0)\nfn = K.sum(K.cast(y_true*(1-y_pred), 'float'), axis=0)\n\np = tp / (tp + fp + K.epsilon())\nr = tp / (tp + fn + K.epsilon())\n\nf1 = 2*p*r / (p+r+K.epsilon())\nf1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n\nreturn 1 - K.mean(f1)\n</code></pre>",
      "rawMarkdown": "I thought this too but when I tried it didn't work out well Here is the loss function I found on another kaggle kernel.\n\ndef f1_loss(y_true, y_pred):\n    \n    tp = K.sum(K.cast(y_true*y_pred, 'float'), axis=0)\n    \n    fp = K.sum(K.cast((1-y_true)*y_pred, 'float'), axis=0)\n    fn = K.sum(K.cast(y_true*(1-y_pred), 'float'), axis=0)\n\n    p = tp / (tp + fp + K.epsilon())\n    r = tp / (tp + fn + K.epsilon())\n\n    f1 = 2*p*r / (p+r+K.epsilon())\n    f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n    \n    return 1 - K.mean(f1)",
      "votes": null
    },
    {
      "id": "434922",
      "postDate": "12/07/2018 06:53:50",
      "content": "<p>Did you try to change the loss functions? Did it work?</p>",
      "rawMarkdown": "Did you try to change the loss functions? Did it work?",
      "votes": null
    },
    {
      "id": "435063",
      "postDate": "12/07/2018 12:50:20",
      "content": "<p>Not yet. Havn't got time to. Will try in the weekend.</p>",
      "rawMarkdown": "Not yet. Havn't got time to. Will try in the weekend.",
      "votes": null
    },
    {
      "id": "435171",
      "postDate": "12/07/2018 16:02:53",
      "content": "<p>I think this is a very hard problem since you can't backpropagate the gradients through non-differentiable functions.</p>",
      "rawMarkdown": "I think this is a very hard problem since you can't backpropagate the gradients through non-differentiable functions.",
      "votes": null
    },
    {
      "id": "436080",
      "postDate": "12/09/2018 14:38:17",
      "content": "<p>You can tune threshold automatically, using idea from my kernel: <a href=\"https://www.kaggle.com/sashamn/fast-alogrithm-for-optimal-threshold-in-f1-metric\">https://www.kaggle.com/sashamn/fast-alogrithm-for-optimal-threshold-in-f1-metric</a></p>",
      "rawMarkdown": "You can tune threshold automatically, using idea from my kernel: https://www.kaggle.com/sashamn/fast-alogrithm-for-optimal-threshold-in-f1-metric",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 433851,
      "author_name": "joeytaj",
      "author_url": "",
      "post_date": "12/05/2018 15:10:50",
      "content": "<p>I thought this too but when I tried it didn't work out well Here is the loss function I found on another kaggle kernel.</p>\n\n<p>def f1_loss(y_true, y_pred):</p>\n\n<pre><code>tp = K.sum(K.cast(y_true*y_pred, 'float'), axis=0)\n\nfp = K.sum(K.cast((1-y_true)*y_pred, 'float'), axis=0)\nfn = K.sum(K.cast(y_true*(1-y_pred), 'float'), axis=0)\n\np = tp / (tp + fp + K.epsilon())\nr = tp / (tp + fn + K.epsilon())\n\nf1 = 2*p*r / (p+r+K.epsilon())\nf1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n\nreturn 1 - K.mean(f1)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 436080,
          "author_name": "sashamn",
          "author_url": "",
          "post_date": "12/09/2018 14:38:17",
          "content": "<p>You can tune threshold automatically, using idea from my kernel: <a href=\"https://www.kaggle.com/sashamn/fast-alogrithm-for-optimal-threshold-in-f1-metric\">https://www.kaggle.com/sashamn/fast-alogrithm-for-optimal-threshold-in-f1-metric</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434922,
      "author_name": "maying",
      "author_url": "",
      "post_date": "12/07/2018 06:53:50",
      "content": "<p>Did you try to change the loss functions? Did it work?</p>",
      "votes": null,
      "replies": [
        {
          "id": 435063,
          "author_name": "piyushkgp",
          "author_url": "",
          "post_date": "12/07/2018 12:50:20",
          "content": "<p>Not yet. Havn't got time to. Will try in the weekend.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 435171,
      "author_name": "sashamn",
      "author_url": "",
      "post_date": "12/07/2018 16:02:53",
      "content": "<p>I think this is a very hard problem since you can't backpropagate the gradients through non-differentiable functions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "433655": "Since the submission is graded based on F1 score instead of accuracy, I was thinking if we should try changing the loss function somehow so that the F1 score gets maximized.\nOne candidate I can think of inverse of F1 score. \nBut I am not sure about the differentiability for the backprop.\nAlso it would be good if the model gave me the threshold value to threshold upon to classify as 0 vs 1\n\nMy existing solution (and everything I found online) minimizes the binary cross-entropy loss function itself, but after predicting the prob values, based on the eval set, I decide the threshold to classify as 0 vs 1 so that the F1 score gets maximized. Then for the actual test set, I use the same threshold. The assumption here is that it will follow the exact same ratio of 0s and 1s (which may not be correct).\n\nAll comments/suggestions/criticisms most welcome.",
    "433851": "I thought this too but when I tried it didn't work out well Here is the loss function I found on another kaggle kernel.\n\ndef f1_loss(y_true, y_pred):\n    \n    tp = K.sum(K.cast(y_true*y_pred, 'float'), axis=0)\n    \n    fp = K.sum(K.cast((1-y_true)*y_pred, 'float'), axis=0)\n    fn = K.sum(K.cast(y_true*(1-y_pred), 'float'), axis=0)\n\n    p = tp / (tp + fp + K.epsilon())\n    r = tp / (tp + fn + K.epsilon())\n\n    f1 = 2*p*r / (p+r+K.epsilon())\n    f1 = tf.where(tf.is_nan(f1), tf.zeros_like(f1), f1)\n    \n    return 1 - K.mean(f1)",
    "434922": "Did you try to change the loss functions? Did it work?",
    "435063": "Not yet. Havn't got time to. Will try in the weekend.",
    "435171": "I think this is a very hard problem since you can't backpropagate the gradients through non-differentiable functions.",
    "436080": "You can tune threshold automatically, using idea from my kernel: https://www.kaggle.com/sashamn/fast-alogrithm-for-optimal-threshold-in-f1-metric"
  },
  "source": "meta"
}