{
  "id": 361813,
  "title": "LogLoss explanation",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/361813",
  "author_name": "",
  "post_date": "2022-10-23T22:19:43.718903700Z",
  "votes": 17,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi colleagues,</p>\n<p>I have created the new kernel <a href=\"https://www.kaggle.com/code/alexryzhkov/how-to-kill-all-your-efforts\" target=\"_blank\">\"How to kill all your efforts?\"</a> to show why you should be careful with target mean of your datasets if you are working with <code>LogLoss</code> metric.</p>\n<p>To be clear - the target mean for the team A is around 6% and if you work carefully with the data splitting, you will receive the <strong>0.20583</strong> and <strong>0.20236</strong> LogLoss metrics for validation and test sets respectively using LightGBM model. The predictions histograms will look like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fc322431ccac008b1209641fcdb02f78d%2F2022-10-24%20%2001.13.03.png?generation=1666563246475132&amp;alt=media\" alt=\"\"></p>\n<p>But if you sample the train and validation data to have 33% target rate (less negative with all the positive objects), you will receive the dramatically worse scores - <strong>0.38444</strong> and <strong>0.38042</strong> LogLoss metrics for validation and test sets respectively using the same LightGBM model. The predictions histograms will now look like:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2F2e781446f4d82ac35475dbcb03959b86%2F2022-10-24%20%2001.13.38.png?generation=1666563435040349&amp;alt=media\" alt=\"\"></p>\n<p>As you can see they move to the right and it kills all of your work - it doesn't mind that great features you have in your dataset. So please <strong>be careful with the data preparation if you want to get a good score</strong>.</p>\n<p>Hope this helps you!</p>\n<p>Alex</p>",
  "messages": [
    {
      "id": "2001286",
      "postDate": "10/23/2022 22:19:43",
      "content": "<p>Hi colleagues,</p>\n<p>I have created the new kernel <a href=\"https://www.kaggle.com/code/alexryzhkov/how-to-kill-all-your-efforts\" target=\"_blank\">\"How to kill all your efforts?\"</a> to show why you should be careful with target mean of your datasets if you are working with <code>LogLoss</code> metric.</p>\n<p>To be clear - the target mean for the team A is around 6% and if you work carefully with the data splitting, you will receive the <strong>0.20583</strong> and <strong>0.20236</strong> LogLoss metrics for validation and test sets respectively using LightGBM model. The predictions histograms will look like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fc322431ccac008b1209641fcdb02f78d%2F2022-10-24%20%2001.13.03.png?generation=1666563246475132&amp;alt=media\" alt=\"\"></p>\n<p>But if you sample the train and validation data to have 33% target rate (less negative with all the positive objects), you will receive the dramatically worse scores - <strong>0.38444</strong> and <strong>0.38042</strong> LogLoss metrics for validation and test sets respectively using the same LightGBM model. The predictions histograms will now look like:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2F2e781446f4d82ac35475dbcb03959b86%2F2022-10-24%20%2001.13.38.png?generation=1666563435040349&amp;alt=media\" alt=\"\"></p>\n<p>As you can see they move to the right and it kills all of your work - it doesn't mind that great features you have in your dataset. So please <strong>be careful with the data preparation if you want to get a good score</strong>.</p>\n<p>Hope this helps you!</p>\n<p>Alex</p>",
      "rawMarkdown": "Hi colleagues,\n\nI have created the new kernel [\"How to kill all your efforts?\"](https://www.kaggle.com/code/alexryzhkov/how-to-kill-all-your-efforts) to show why you should be careful with target mean of your datasets if you are working with `LogLoss` metric.\n\nTo be clear - the target mean for the team A is around 6% and if you work carefully with the data splitting, you will receive the **0.20583** and **0.20236** LogLoss metrics for validation and test sets respectively using LightGBM model. The predictions histograms will look like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fc322431ccac008b1209641fcdb02f78d%2F2022-10-24%20%2001.13.03.png?generation=1666563246475132&alt=media)\n\nBut if you sample the train and validation data to have 33% target rate (less negative with all the positive objects), you will receive the dramatically worse scores - **0.38444** and **0.38042** LogLoss metrics for validation and test sets respectively using the same LightGBM model. The predictions histograms will now look like:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2F2e781446f4d82ac35475dbcb03959b86%2F2022-10-24%20%2001.13.38.png?generation=1666563435040349&alt=media)\n\nAs you can see they move to the right and it kills all of your work - it doesn't mind that great features you have in your dataset. So please **be careful with the data preparation if you want to get a good score**.\n\nHope this helps you!\n\nAlex",
      "votes": null
    },
    {
      "id": "2001428",
      "postDate": "10/24/2022 03:29:27",
      "content": "<p>Great explanation, so we should use stratified sample or always test the behavior of your split? I just read today about adversarial validation, but I am not sure if that could work in this case. Thanks.</p>",
      "rawMarkdown": "Great explanation, so we should use stratified sample or always test the behavior of your split? I just read today about adversarial validation, but I am not sure if that could work in this case. Thanks.",
      "votes": null
    },
    {
      "id": "2013429",
      "postDate": "11/01/2022 22:34:19",
      "content": "<p>Thank you for the explanation. It was very helpful to read.</p>",
      "rawMarkdown": "Thank you for the explanation. It was very helpful to read.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2001428,
      "author_name": "pastorsoto",
      "author_url": "",
      "post_date": "10/24/2022 03:29:27",
      "content": "<p>Great explanation, so we should use stratified sample or always test the behavior of your split? I just read today about adversarial validation, but I am not sure if that could work in this case. Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2013429,
      "author_name": "rafkhat",
      "author_url": "",
      "post_date": "11/01/2022 22:34:19",
      "content": "<p>Thank you for the explanation. It was very helpful to read.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2001286": "Hi colleagues,\n\nI have created the new kernel [\"How to kill all your efforts?\"](https://www.kaggle.com/code/alexryzhkov/how-to-kill-all-your-efforts) to show why you should be careful with target mean of your datasets if you are working with `LogLoss` metric.\n\nTo be clear - the target mean for the team A is around 6% and if you work carefully with the data splitting, you will receive the **0.20583** and **0.20236** LogLoss metrics for validation and test sets respectively using LightGBM model. The predictions histograms will look like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fc322431ccac008b1209641fcdb02f78d%2F2022-10-24%20%2001.13.03.png?generation=1666563246475132&alt=media)\n\nBut if you sample the train and validation data to have 33% target rate (less negative with all the positive objects), you will receive the dramatically worse scores - **0.38444** and **0.38042** LogLoss metrics for validation and test sets respectively using the same LightGBM model. The predictions histograms will now look like:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2F2e781446f4d82ac35475dbcb03959b86%2F2022-10-24%20%2001.13.38.png?generation=1666563435040349&alt=media)\n\nAs you can see they move to the right and it kills all of your work - it doesn't mind that great features you have in your dataset. So please **be careful with the data preparation if you want to get a good score**.\n\nHope this helps you!\n\nAlex",
    "2001428": "Great explanation, so we should use stratified sample or always test the behavior of your split? I just read today about adversarial validation, but I am not sure if that could work in this case. Thanks.",
    "2013429": "Thank you for the explanation. It was very helpful to read."
  },
  "source": "meta"
}