{
  "id": 76391,
  "title": "Optimal threshold leads to score overestimation",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76391",
  "author_name": "",
  "post_date": "2019-01-02T09:17:42.803871200Z",
  "votes": 54,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey.</p>\n\n<p>The golden rule of machine learning is to test your models on the unseen data. In the opposite case, you would get overestimated scores. In this competition, many kagglers use validation dataset to estimate the optimal threshold and use the same dataset to calculate the validation score. This is incorrect. Instead, you may use this code to get 'real' score of your model.</p>\n\n<pre><code>def scoring(y_true, y_proba, verbose=True):\n    from sklearn.metrics import roc_curve, precision_recall_curve, f1_score\n    from sklearn.model_selection import RepeatedStratifiedKFold\n\n    def threshold_search(y_true, y_proba):\n        precision , recall, thresholds = precision_recall_curve(y_true, y_proba)\n        thresholds = np.append(thresholds, 1.001) \n        F = 2 / (1/precision + 1/recall)\n        best_score = np.max(F)\n        best_th = thresholds[np.argmax(F)]\n        return best_th \n\n\n    rkf = RepeatedStratifiedKFold(n_splits=5, n_repeats=10)\n\n    scores = []\n    ths = []\n    for train_index, test_index in rkf.split(y_true, y_true):\n        y_prob_train, y_prob_test = y_proba[train_index], y_proba[test_index]\n        y_true_train, y_true_test = y_true[train_index], y_true[test_index]\n\n        # determine best threshold on 'train' part \n        best_threshold = threshold_search(y_true_train, y_prob_train)\n\n        # use this threshold on 'test' part for score \n        sc = f1_score(y_true_test, (y_prob_test &amp;gt;= best_threshold).astype(int))\n        scores.append(sc)\n        ths.append(best_threshold)\n\n    best_th = np.mean(ths)\n    score = np.mean(scores)\n\n    if verbose: print(f'Best threshold: {np.round(best_th, 4)}, Score: {np.round(score,5)}')\n\n    return best_th, score\n</code></pre>\n\n<p>I tested it in my pipeline: I splitted data in train and validation (80/20%). Train was used for model training and then I made predictions for validation set. Then I splitted validation set into <code>val1</code> and <code>val2</code>. Val1 was used for picking best threshold and <code>val2</code> was used to test this approach. I also calculated scores for <code>val1</code>. In the end, I got following results: </p>\n\n<ol>\n<li>Best threshold search on the whole <code>val1</code> always gave better results than proposed approach. The differences were about 0.001 F1 score (this incicates overfit). </li>\n<li>Proposed approach increased the score on <code>val2</code> by 0.0001 comparing to the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75735#445122\">previous version</a> of optimal threshold finding. This calculation took about 4 sec to run and personally I will take this trade any day. </li>\n</ol>\n\n<p>However, in this competition the amount of data is quite big, so a difference between 'real' score and 'overfitted' score is not too dramatic. But you should be more accurate during your actual work.</p>\n\n<p>PS. Just for curiosity, you may try to compare old and new approaches on smaller dataset. I bet I would be surprised 😯</p>",
  "messages": [
    {
      "id": "448866",
      "postDate": "01/02/2019 09:17:42",
      "content": "<p>Hey.</p>\n\n<p>The golden rule of machine learning is to test your models on the unseen data. In the opposite case, you would get overestimated scores. In this competition, many kagglers use validation dataset to estimate the optimal threshold and use the same dataset to calculate the validation score. This is incorrect. Instead, you may use this code to get 'real' score of your model.</p>\n\n<pre><code>def scoring(y_true, y_proba, verbose=True):\n    from sklearn.metrics import roc_curve, precision_recall_curve, f1_score\n    from sklearn.model_selection import RepeatedStratifiedKFold\n\n    def threshold_search(y_true, y_proba):\n        precision , recall, thresholds = precision_recall_curve(y_true, y_proba)\n        thresholds = np.append(thresholds, 1.001) \n        F = 2 / (1/precision + 1/recall)\n        best_score = np.max(F)\n        best_th = thresholds[np.argmax(F)]\n        return best_th \n\n\n    rkf = RepeatedStratifiedKFold(n_splits=5, n_repeats=10)\n\n    scores = []\n    ths = []\n    for train_index, test_index in rkf.split(y_true, y_true):\n        y_prob_train, y_prob_test = y_proba[train_index], y_proba[test_index]\n        y_true_train, y_true_test = y_true[train_index], y_true[test_index]\n\n        # determine best threshold on 'train' part \n        best_threshold = threshold_search(y_true_train, y_prob_train)\n\n        # use this threshold on 'test' part for score \n        sc = f1_score(y_true_test, (y_prob_test &amp;gt;= best_threshold).astype(int))\n        scores.append(sc)\n        ths.append(best_threshold)\n\n    best_th = np.mean(ths)\n    score = np.mean(scores)\n\n    if verbose: print(f'Best threshold: {np.round(best_th, 4)}, Score: {np.round(score,5)}')\n\n    return best_th, score\n</code></pre>\n\n<p>I tested it in my pipeline: I splitted data in train and validation (80/20%). Train was used for model training and then I made predictions for validation set. Then I splitted validation set into <code>val1</code> and <code>val2</code>. Val1 was used for picking best threshold and <code>val2</code> was used to test this approach. I also calculated scores for <code>val1</code>. In the end, I got following results: </p>\n\n<ol>\n<li>Best threshold search on the whole <code>val1</code> always gave better results than proposed approach. The differences were about 0.001 F1 score (this incicates overfit). </li>\n<li>Proposed approach increased the score on <code>val2</code> by 0.0001 comparing to the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75735#445122\">previous version</a> of optimal threshold finding. This calculation took about 4 sec to run and personally I will take this trade any day. </li>\n</ol>\n\n<p>However, in this competition the amount of data is quite big, so a difference between 'real' score and 'overfitted' score is not too dramatic. But you should be more accurate during your actual work.</p>\n\n<p>PS. Just for curiosity, you may try to compare old and new approaches on smaller dataset. I bet I would be surprised 😯</p>",
      "rawMarkdown": "Hey.\n\nThe golden rule of machine learning is to test your models on the unseen data. In the opposite case, you would get overestimated scores. In this competition, many kagglers use validation dataset to estimate the optimal threshold and use the same dataset to calculate the validation score. This is incorrect. Instead, you may use this code to get 'real' score of your model.\n\n    def scoring(y_true, y_proba, verbose=True):\n        from sklearn.metrics import roc_curve, precision_recall_curve, f1_score\n        from sklearn.model_selection import RepeatedStratifiedKFold\n        \n        def threshold_search(y_true, y_proba):\n            precision , recall, thresholds = precision_recall_curve(y_true, y_proba)\n            thresholds = np.append(thresholds, 1.001) \n            F = 2 / (1/precision + 1/recall)\n            best_score = np.max(F)\n            best_th = thresholds[np.argmax(F)]\n            return best_th \n        \n        \n        rkf = RepeatedStratifiedKFold(n_splits=5, n_repeats=10)\n    \n        scores = []\n        ths = []\n        for train_index, test_index in rkf.split(y_true, y_true):\n            y_prob_train, y_prob_test = y_proba[train_index], y_proba[test_index]\n            y_true_train, y_true_test = y_true[train_index], y_true[test_index]\n        \n            # determine best threshold on 'train' part \n            best_threshold = threshold_search(y_true_train, y_prob_train)\n            \n            # use this threshold on 'test' part for score \n            sc = f1_score(y_true_test, (y_prob_test &gt;= best_threshold).astype(int))\n            scores.append(sc)\n            ths.append(best_threshold)\n         \n        best_th = np.mean(ths)\n        score = np.mean(scores)\n        \n        if verbose: print(f'Best threshold: {np.round(best_th, 4)}, Score: {np.round(score,5)}')\n            \n        return best_th, score\n\n\nI tested it in my pipeline: I splitted data in train and validation (80/20%). Train was used for model training and then I made predictions for validation set. Then I splitted validation set into `val1` and `val2`. Val1 was used for picking best threshold and `val2` was used to test this approach. I also calculated scores for `val1`. In the end, I got following results: \n\n1. Best threshold search on the whole `val1` always gave better results than proposed approach. The differences were about 0.001 F1 score (this incicates overfit). \n2. Proposed approach increased the score on `val2` by 0.0001 comparing to the [previous version][1] of optimal threshold finding. This calculation took about 4 sec to run and personally I will take this trade any day. \n\nHowever, in this competition the amount of data is quite big, so a difference between 'real' score and 'overfitted' score is not too dramatic. But you should be more accurate during your actual work.\n\nPS. Just for curiosity, you may try to compare old and new approaches on smaller dataset. I bet I would be surprised 😯\n\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75735#445122",
      "votes": null
    },
    {
      "id": "448876",
      "postDate": "01/02/2019 09:42:33",
      "content": "<p>I have the same idea with you to use two validation dataset,  but it seems that when the model contain ensemble, more splits are required to estimate the optimal threshold on different stages, which can be kind of waste. \nDo you think it's a good idea to use the same dataset to estimate each models and the ensembled one?</p>",
      "rawMarkdown": "I have the same idea with you to use two validation dataset,  but it seems that when the model contain ensemble, more splits are required to estimate the optimal threshold on different stages, which can be kind of waste. \nDo you think it's a good idea to use the same dataset to estimate each models and the ensembled one?",
      "votes": null
    },
    {
      "id": "448882",
      "postDate": "01/02/2019 09:53:49",
      "content": "<p>I think you could select the optimal parameters of models (number of epochs) based on your validation (KFold or several train-val splits) on stage 1 of the competition and then use it on stage 2 without any change. In this case, validation set would be used for ensembling only. </p>",
      "rawMarkdown": "I think you could select the optimal parameters of models (number of epochs) based on your validation (KFold or several train-val splits) on stage 1 of the competition and then use it on stage 2 without any change. In this case, validation set would be used for ensembling only.",
      "votes": null
    },
    {
      "id": "448958",
      "postDate": "01/02/2019 12:40:23",
      "content": "<p>Very clever ideal. Thanks for sharing!</p>",
      "rawMarkdown": "Very clever ideal. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "449030",
      "postDate": "01/02/2019 15:03:23",
      "content": "<p>I calculate the best threshold by saving the validation predictions from each fold, giving me predictions for the entire 1.2M row dataset and then using the ground truth labels to find the best threshold.</p>\n\n<p>This might be more robust since the best threshold on the entire dataset should generalize better to the stage 2 test set than a threshold found on a smaller part of the dataset.</p>",
      "rawMarkdown": "I calculate the best threshold by saving the validation predictions from each fold, giving me predictions for the entire 1.2M row dataset and then using the ground truth labels to find the best threshold.\n\nThis might be more robust since the best threshold on the entire dataset should generalize better to the stage 2 test set than a threshold found on a smaller part of the dataset.",
      "votes": null
    },
    {
      "id": "449089",
      "postDate": "01/02/2019 16:16:48",
      "content": "<p>Based on my experience in previous competition with F1 score, voting of several folds with the best threshold in each gives better results. But it might not be the case in this one. Will see 😊</p>",
      "rawMarkdown": "Based on my experience in previous competition with F1 score, voting of several folds with the best threshold in each gives better results. But it might not be the case in this one. Will see 😊",
      "votes": null
    },
    {
      "id": "450157",
      "postDate": "01/04/2019 12:07:40",
      "content": "<p>Votin is much more unstable based on my observations.</p>",
      "rawMarkdown": "Votin is much more unstable based on my observations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 448876,
      "author_name": "peining",
      "author_url": "",
      "post_date": "01/02/2019 09:42:33",
      "content": "<p>I have the same idea with you to use two validation dataset,  but it seems that when the model contain ensemble, more splits are required to estimate the optimal threshold on different stages, which can be kind of waste. \nDo you think it's a good idea to use the same dataset to estimate each models and the ensembled one?</p>",
      "votes": null,
      "replies": [
        {
          "id": 448882,
          "author_name": "",
          "author_url": "",
          "post_date": "01/02/2019 09:53:49",
          "content": "<p>I think you could select the optimal parameters of models (number of epochs) based on your validation (KFold or several train-val splits) on stage 1 of the competition and then use it on stage 2 without any change. In this case, validation set would be used for ensembling only. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 448958,
      "author_name": "sunnymarkliu",
      "author_url": "",
      "post_date": "01/02/2019 12:40:23",
      "content": "<p>Very clever ideal. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449030,
      "author_name": "bkkaggle",
      "author_url": "",
      "post_date": "01/02/2019 15:03:23",
      "content": "<p>I calculate the best threshold by saving the validation predictions from each fold, giving me predictions for the entire 1.2M row dataset and then using the ground truth labels to find the best threshold.</p>\n\n<p>This might be more robust since the best threshold on the entire dataset should generalize better to the stage 2 test set than a threshold found on a smaller part of the dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 449089,
          "author_name": "",
          "author_url": "",
          "post_date": "01/02/2019 16:16:48",
          "content": "<p>Based on my experience in previous competition with F1 score, voting of several folds with the best threshold in each gives better results. But it might not be the case in this one. Will see 😊</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 450157,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "01/04/2019 12:07:40",
          "content": "<p>Votin is much more unstable based on my observations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "448866": "Hey.\n\nThe golden rule of machine learning is to test your models on the unseen data. In the opposite case, you would get overestimated scores. In this competition, many kagglers use validation dataset to estimate the optimal threshold and use the same dataset to calculate the validation score. This is incorrect. Instead, you may use this code to get 'real' score of your model.\n\n    def scoring(y_true, y_proba, verbose=True):\n        from sklearn.metrics import roc_curve, precision_recall_curve, f1_score\n        from sklearn.model_selection import RepeatedStratifiedKFold\n        \n        def threshold_search(y_true, y_proba):\n            precision , recall, thresholds = precision_recall_curve(y_true, y_proba)\n            thresholds = np.append(thresholds, 1.001) \n            F = 2 / (1/precision + 1/recall)\n            best_score = np.max(F)\n            best_th = thresholds[np.argmax(F)]\n            return best_th \n        \n        \n        rkf = RepeatedStratifiedKFold(n_splits=5, n_repeats=10)\n    \n        scores = []\n        ths = []\n        for train_index, test_index in rkf.split(y_true, y_true):\n            y_prob_train, y_prob_test = y_proba[train_index], y_proba[test_index]\n            y_true_train, y_true_test = y_true[train_index], y_true[test_index]\n        \n            # determine best threshold on 'train' part \n            best_threshold = threshold_search(y_true_train, y_prob_train)\n            \n            # use this threshold on 'test' part for score \n            sc = f1_score(y_true_test, (y_prob_test &gt;= best_threshold).astype(int))\n            scores.append(sc)\n            ths.append(best_threshold)\n         \n        best_th = np.mean(ths)\n        score = np.mean(scores)\n        \n        if verbose: print(f'Best threshold: {np.round(best_th, 4)}, Score: {np.round(score,5)}')\n            \n        return best_th, score\n\n\nI tested it in my pipeline: I splitted data in train and validation (80/20%). Train was used for model training and then I made predictions for validation set. Then I splitted validation set into `val1` and `val2`. Val1 was used for picking best threshold and `val2` was used to test this approach. I also calculated scores for `val1`. In the end, I got following results: \n\n1. Best threshold search on the whole `val1` always gave better results than proposed approach. The differences were about 0.001 F1 score (this incicates overfit). \n2. Proposed approach increased the score on `val2` by 0.0001 comparing to the [previous version][1] of optimal threshold finding. This calculation took about 4 sec to run and personally I will take this trade any day. \n\nHowever, in this competition the amount of data is quite big, so a difference between 'real' score and 'overfitted' score is not too dramatic. But you should be more accurate during your actual work.\n\nPS. Just for curiosity, you may try to compare old and new approaches on smaller dataset. I bet I would be surprised 😯\n\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75735#445122",
    "448876": "I have the same idea with you to use two validation dataset,  but it seems that when the model contain ensemble, more splits are required to estimate the optimal threshold on different stages, which can be kind of waste. \nDo you think it's a good idea to use the same dataset to estimate each models and the ensembled one?",
    "448882": "I think you could select the optimal parameters of models (number of epochs) based on your validation (KFold or several train-val splits) on stage 1 of the competition and then use it on stage 2 without any change. In this case, validation set would be used for ensembling only.",
    "448958": "Very clever ideal. Thanks for sharing!",
    "449030": "I calculate the best threshold by saving the validation predictions from each fold, giving me predictions for the entire 1.2M row dataset and then using the ground truth labels to find the best threshold.\n\nThis might be more robust since the best threshold on the entire dataset should generalize better to the stage 2 test set than a threshold found on a smaller part of the dataset.",
    "449089": "Based on my experience in previous competition with F1 score, voting of several folds with the best threshold in each gives better results. But it might not be the case in this one. Will see 😊",
    "450157": "Votin is much more unstable based on my observations."
  },
  "source": "meta"
}