{
  "id": 539509,
  "title": "Error Analysis of CMI | Best Single Model",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/539509",
  "author_name": "",
  "post_date": "2024-10-09T09:04:06.711134Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have done an error analysis of the model <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/abdmental01/cmi-best-single-model</a> provided by <a href=\"url\" target=\"_blank\">https://www.kaggle.com/abdmental01</a>. What I notice is that <code>sii=3</code> is never predicted correctly.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Fb95dfc08c2f841267905631ce8a4712f%2Fval_error.png?generation=1728463788684094&amp;alt=media\" alt=\"\"></p>\n<p>So, I thought why not remove all the rows containing <code>sii=3</code> and then apply the same model. What I notice is that the model performance has gone down.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Ff630d4227cb8a30256754e4c489bbad6%2Fval_error_without_sii3.png?generation=1728464284444499&amp;alt=media\" alt=\"\"></p>\n<p>I think a possible explanation is that I should use other parameter set. Other than that I would not no why performance decreases. Does anybody has any other idea why the performance has decreased?</p>",
  "messages": [
    {
      "id": "3012685",
      "postDate": "10/09/2024 09:04:06",
      "content": "<p>I have done an error analysis of the model <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/abdmental01/cmi-best-single-model</a> provided by <a href=\"url\" target=\"_blank\">https://www.kaggle.com/abdmental01</a>. What I notice is that <code>sii=3</code> is never predicted correctly.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Fb95dfc08c2f841267905631ce8a4712f%2Fval_error.png?generation=1728463788684094&amp;alt=media\" alt=\"\"></p>\n<p>So, I thought why not remove all the rows containing <code>sii=3</code> and then apply the same model. What I notice is that the model performance has gone down.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Ff630d4227cb8a30256754e4c489bbad6%2Fval_error_without_sii3.png?generation=1728464284444499&amp;alt=media\" alt=\"\"></p>\n<p>I think a possible explanation is that I should use other parameter set. Other than that I would not no why performance decreases. Does anybody has any other idea why the performance has decreased?</p>",
      "rawMarkdown": "I have done an error analysis of the model [https://www.kaggle.com/code/abdmental01/cmi-best-single-model](url) provided by [https://www.kaggle.com/abdmental01](url). What I notice is that `sii=3` is never predicted correctly.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Fb95dfc08c2f841267905631ce8a4712f%2Fval_error.png?generation=1728463788684094&alt=media)\n\nSo, I thought why not remove all the rows containing `sii=3` and then apply the same model. What I notice is that the model performance has gone down.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Ff630d4227cb8a30256754e4c489bbad6%2Fval_error_without_sii3.png?generation=1728464284444499&alt=media)\n\nI think a possible explanation is that I should use other parameter set. Other than that I would not no why performance decreases. Does anybody has any other idea why the performance has decreased?",
      "votes": null
    },
    {
      "id": "3013048",
      "postDate": "10/09/2024 15:58:54",
      "content": "<p>Two reasons I can think of</p>\n<ol>\n<li>Without sii=3, the sii=2 become outliers and model finds it difficult to predict a 2.</li>\n<li>Without a 3, I guess all hyperparameters must be retuned?</li>\n</ol>\n<p>btw, is this error analysis based on train set or what? I had a few confusion matrix from cv, and 3 can be predicted, of course not very good. Here's a random one I picked.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22136130%2Ffd251315f6b75ec57b5ad7e24dc9b03f%2Fconfusion.png?generation=1728489442165539&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Two reasons I can think of\n1. Without sii=3, the sii=2 become outliers and model finds it difficult to predict a 2.\n2. Without a 3, I guess all hyperparameters must be retuned?\n\nbtw, is this error analysis based on train set or what? I had a few confusion matrix from cv, and 3 can be predicted, of course not very good. Here's a random one I picked.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22136130%2Ffd251315f6b75ec57b5ad7e24dc9b03f%2Fconfusion.png?generation=1728489442165539&alt=media)",
      "votes": null
    },
    {
      "id": "3013164",
      "postDate": "10/09/2024 17:37:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a>, thanks for your response. Regarding your points:</p>\n<ol>\n<li>This I did not think of before (so thanks for the tip). So I did a little research and found that one can also use the <code>LightGBM's</code> built-in parameters to control the effect of outliers. I will try this later. Although I am still curious why removing the <code>outliers</code> does not help the model to perform beter?😫</li>\n<li>I did retune the parameters with Optuna but the results got worse.</li>\n</ol>\n<p>The error analysis that I posted is based on the results of the validation set.</p>",
      "rawMarkdown": "Hi @tomyuen, thanks for your response. Regarding your points:\n\n1. This I did not think of before (so thanks for the tip). So I did a little research and found that one can also use the `LightGBM's` built-in parameters to control the effect of outliers. I will try this later. Although I am still curious why removing the `outliers` does not help the model to perform beter?😫\n2. I did retune the parameters with Optuna but the results got worse.\n\nThe error analysis that I posted is based on the results of the validation set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3013048,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "10/09/2024 15:58:54",
      "content": "<p>Two reasons I can think of</p>\n<ol>\n<li>Without sii=3, the sii=2 become outliers and model finds it difficult to predict a 2.</li>\n<li>Without a 3, I guess all hyperparameters must be retuned?</li>\n</ol>\n<p>btw, is this error analysis based on train set or what? I had a few confusion matrix from cv, and 3 can be predicted, of course not very good. Here's a random one I picked.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22136130%2Ffd251315f6b75ec57b5ad7e24dc9b03f%2Fconfusion.png?generation=1728489442165539&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3013164,
          "author_name": "wti200",
          "author_url": "",
          "post_date": "10/09/2024 17:37:46",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a>, thanks for your response. Regarding your points:</p>\n<ol>\n<li>This I did not think of before (so thanks for the tip). So I did a little research and found that one can also use the <code>LightGBM's</code> built-in parameters to control the effect of outliers. I will try this later. Although I am still curious why removing the <code>outliers</code> does not help the model to perform beter?😫</li>\n<li>I did retune the parameters with Optuna but the results got worse.</li>\n</ol>\n<p>The error analysis that I posted is based on the results of the validation set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3012685": "I have done an error analysis of the model [https://www.kaggle.com/code/abdmental01/cmi-best-single-model](url) provided by [https://www.kaggle.com/abdmental01](url). What I notice is that `sii=3` is never predicted correctly.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Fb95dfc08c2f841267905631ce8a4712f%2Fval_error.png?generation=1728463788684094&alt=media)\n\nSo, I thought why not remove all the rows containing `sii=3` and then apply the same model. What I notice is that the model performance has gone down.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Ff630d4227cb8a30256754e4c489bbad6%2Fval_error_without_sii3.png?generation=1728464284444499&alt=media)\n\nI think a possible explanation is that I should use other parameter set. Other than that I would not no why performance decreases. Does anybody has any other idea why the performance has decreased?",
    "3013048": "Two reasons I can think of\n1. Without sii=3, the sii=2 become outliers and model finds it difficult to predict a 2.\n2. Without a 3, I guess all hyperparameters must be retuned?\n\nbtw, is this error analysis based on train set or what? I had a few confusion matrix from cv, and 3 can be predicted, of course not very good. Here's a random one I picked.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22136130%2Ffd251315f6b75ec57b5ad7e24dc9b03f%2Fconfusion.png?generation=1728489442165539&alt=media)",
    "3013164": "Hi @tomyuen, thanks for your response. Regarding your points:\n\n1. This I did not think of before (so thanks for the tip). So I did a little research and found that one can also use the `LightGBM's` built-in parameters to control the effect of outliers. I will try this later. Although I am still curious why removing the `outliers` does not help the model to perform beter?😫\n2. I did retune the parameters with Optuna but the results got worse.\n\nThe error analysis that I posted is based on the results of the validation set."
  },
  "source": "meta"
}