{
  "id": 380379,
  "title": "I can't use early stopping for LGBMranker.",
  "url": "/competitions/otto-recommender-system/discussion/380379",
  "author_name": "",
  "post_date": "2023-01-23T03:07:27.061678700Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I can't use early stopping for LGBMranker because ndcg@20 and recall@20 of validation data are not proportional.<br>\nWhat should I do?</p>",
  "messages": [
    {
      "id": "2111612",
      "postDate": "01/23/2023 03:07:27",
      "content": "<p>I can't use early stopping for LGBMranker because ndcg@20 and recall@20 of validation data are not proportional.<br>\nWhat should I do?</p>",
      "rawMarkdown": "I can't use early stopping for LGBMranker because ndcg@20 and recall@20 of validation data are not proportional.\nWhat should I do?",
      "votes": null
    },
    {
      "id": "2112137",
      "postDate": "01/23/2023 12:30:27",
      "content": "<p>You're right, why doesn't you select the best iteration directly using the <code>bst.best_iteration</code>? </p>\n<pre><code>    lgbm = lgb.train(params, lgb_train, num_boost_round=, valid_sets=lgb_eval, early_stopping_rounds=)    \n    y_pred = bst.predict(X_test, num_iteration=bst.best_iteration)\n</code></pre>\n<p>During the kaggle days, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  recommands using regularization and most importantly learning rate scheduler ( that you can use via callback ) rather than using  early stopping.<br>\nHope it helps !</p>",
      "rawMarkdown": "You're right, why doesn't you select the best iteration directly using the ``bst.best_iteration ``? \n```python\n    lgbm = lgb.train(params, lgb_train, num_boost_round=800, valid_sets=lgb_eval, early_stopping_rounds=20)    \n    y_pred = bst.predict(X_test, num_iteration=bst.best_iteration)\n```\nDuring the kaggle days, @philippsinger  recommands using regularization and most importantly learning rate scheduler ( that you can use via callback ) rather than using  early stopping.\nHope it helps !",
      "votes": null
    },
    {
      "id": "2112144",
      "postDate": "01/23/2023 12:33:55",
      "content": "<p>MAP@K metric is somewhat correlated with recall. It makes more sense to use it instead of ndcg because if I understood correctly, ndcg also considers the exact order of predictions. </p>",
      "rawMarkdown": "MAP@K metric is somewhat correlated with recall. It makes more sense to use it instead of ndcg because if I understood correctly, ndcg also considers the exact order of predictions.",
      "votes": null
    },
    {
      "id": "2112282",
      "postDate": "01/23/2023 14:10:59",
      "content": "<p>Thank you very much for your useful information!!</p>",
      "rawMarkdown": "Thank you very much for your useful information!!",
      "votes": null
    },
    {
      "id": "2112305",
      "postDate": "01/23/2023 14:26:45",
      "content": "<p>Thank you for your comment.<br>\nIsn't bst.best_iteration also the number of iterations for which metric(ndcg) is highest, and not the number of iterations for which recall is best?</p>",
      "rawMarkdown": "Thank you for your comment.\nIsn't bst.best_iteration also the number of iterations for which metric(ndcg) is highest, and not the number of iterations for which recall is best?",
      "votes": null
    },
    {
      "id": "2112454",
      "postDate": "01/23/2023 16:24:31",
      "content": "<p>I made a notebook with example of custom metric, please take a look <a href=\"https://www.kaggle.com/greenwolf/lightgbm-fast-recall-20\" target=\"_blank\">https://www.kaggle.com/greenwolf/lightgbm-fast-recall-20</a></p>",
      "rawMarkdown": "I made a notebook with example of custom metric, please take a look https://www.kaggle.com/greenwolf/lightgbm-fast-recall-20",
      "votes": null
    },
    {
      "id": "2113864",
      "postDate": "01/24/2023 14:51:42",
      "content": "<p>Thank you very much for your great notebook!!</p>",
      "rawMarkdown": "Thank you very much for your great notebook!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2112137,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "01/23/2023 12:30:27",
      "content": "<p>You're right, why doesn't you select the best iteration directly using the <code>bst.best_iteration</code>? </p>\n<pre><code>    lgbm = lgb.train(params, lgb_train, num_boost_round=, valid_sets=lgb_eval, early_stopping_rounds=)    \n    y_pred = bst.predict(X_test, num_iteration=bst.best_iteration)\n</code></pre>\n<p>During the kaggle days, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  recommands using regularization and most importantly learning rate scheduler ( that you can use via callback ) rather than using  early stopping.<br>\nHope it helps !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2112305,
          "author_name": "yurimaeda",
          "author_url": "",
          "post_date": "01/23/2023 14:26:45",
          "content": "<p>Thank you for your comment.<br>\nIsn't bst.best_iteration also the number of iterations for which metric(ndcg) is highest, and not the number of iterations for which recall is best?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2112144,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/23/2023 12:33:55",
      "content": "<p>MAP@K metric is somewhat correlated with recall. It makes more sense to use it instead of ndcg because if I understood correctly, ndcg also considers the exact order of predictions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2112282,
          "author_name": "yurimaeda",
          "author_url": "",
          "post_date": "01/23/2023 14:10:59",
          "content": "<p>Thank you very much for your useful information!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2112454,
      "author_name": "greenwolf",
      "author_url": "",
      "post_date": "01/23/2023 16:24:31",
      "content": "<p>I made a notebook with example of custom metric, please take a look <a href=\"https://www.kaggle.com/greenwolf/lightgbm-fast-recall-20\" target=\"_blank\">https://www.kaggle.com/greenwolf/lightgbm-fast-recall-20</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2113864,
          "author_name": "yurimaeda",
          "author_url": "",
          "post_date": "01/24/2023 14:51:42",
          "content": "<p>Thank you very much for your great notebook!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2111612": "I can't use early stopping for LGBMranker because ndcg@20 and recall@20 of validation data are not proportional.\nWhat should I do?",
    "2112137": "You're right, why doesn't you select the best iteration directly using the ``bst.best_iteration ``? \n```python\n    lgbm = lgb.train(params, lgb_train, num_boost_round=800, valid_sets=lgb_eval, early_stopping_rounds=20)    \n    y_pred = bst.predict(X_test, num_iteration=bst.best_iteration)\n```\nDuring the kaggle days, @philippsinger  recommands using regularization and most importantly learning rate scheduler ( that you can use via callback ) rather than using  early stopping.\nHope it helps !",
    "2112144": "MAP@K metric is somewhat correlated with recall. It makes more sense to use it instead of ndcg because if I understood correctly, ndcg also considers the exact order of predictions.",
    "2112282": "Thank you very much for your useful information!!",
    "2112305": "Thank you for your comment.\nIsn't bst.best_iteration also the number of iterations for which metric(ndcg) is highest, and not the number of iterations for which recall is best?",
    "2112454": "I made a notebook with example of custom metric, please take a look https://www.kaggle.com/greenwolf/lightgbm-fast-recall-20",
    "2113864": "Thank you very much for your great notebook!!"
  },
  "source": "meta"
}