{
  "id": 75824,
  "title": "only predict '0' target values",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/75824",
  "author_name": "",
  "post_date": "2018-12-27T00:29:08.636633400Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am primarily using CatBoost.  It seems that no matter how much I tune the parameters, I am never able to predict a target value of 1, only 0's on the submission.</p>\n\n<p>Has anyone else encountered this?  Does anyone else have any advice?\nThanks!</p>",
  "messages": [
    {
      "id": "445698",
      "postDate": "12/27/2018 00:29:08",
      "content": "<p>I am primarily using CatBoost.  It seems that no matter how much I tune the parameters, I am never able to predict a target value of 1, only 0's on the submission.</p>\n\n<p>Has anyone else encountered this?  Does anyone else have any advice?\nThanks!</p>",
      "rawMarkdown": "I am primarily using CatBoost.  It seems that no matter how much I tune the parameters, I am never able to predict a target value of 1, only 0's on the submission.\n\nHas anyone else encountered this?  Does anyone else have any advice?\nThanks!",
      "votes": null
    },
    {
      "id": "445893",
      "postDate": "12/27/2018 07:19:52",
      "content": "<p>Have you looked at a few of the kernels that are doing better than you? My first suggestion would be to post your own kernel to get more specific advice about your signal processing, training methods, cross-validation methods, etc.</p>\n\n<p>It definitely sounds like you are overfitting the imbalanced dataset. When you have only 6% positive samples, predicting all '0' values would give you a validation accuracy of about 94% despite being a terrible model. One way to work with an imbalanced dataset is to rebalance it so that the classes are equally likely. This may involve removing some '0's or adding some '1's. </p>\n\n<p>As far as I know, CatBoost doesn't support time series data so I assume you are summarizing the signals somehow? Maybe the features you are creating don't have any information value? Regarding the model, you might need to reduce the complexity of the model and increase the regularization.</p>\n\n<p>I don't mean to sound like an a**, but your question is essentially \"My model sucks, help me\" and without more information on exactly what you are working with and what you have tried to do, it's almost impossible to give advice without randomly guessing.</p>",
      "rawMarkdown": "Have you looked at a few of the kernels that are doing better than you? My first suggestion would be to post your own kernel to get more specific advice about your signal processing, training methods, cross-validation methods, etc.\n\nIt definitely sounds like you are overfitting the imbalanced dataset. When you have only 6% positive samples, predicting all '0' values would give you a validation accuracy of about 94% despite being a terrible model. One way to work with an imbalanced dataset is to rebalance it so that the classes are equally likely. This may involve removing some '0's or adding some '1's. \n\nAs far as I know, CatBoost doesn't support time series data so I assume you are summarizing the signals somehow? Maybe the features you are creating don't have any information value? Regarding the model, you might need to reduce the complexity of the model and increase the regularization.\n\nI don't mean to sound like an a**, but your question is essentially \"My model sucks, help me\" and without more information on exactly what you are working with and what you have tried to do, it's almost impossible to give advice without randomly guessing.",
      "votes": null
    },
    {
      "id": "445983",
      "postDate": "12/27/2018 09:54:59",
      "content": "<p>What loss function are you optimising during the training?</p>",
      "rawMarkdown": "What loss function are you optimising during the training?",
      "votes": null
    },
    {
      "id": "446037",
      "postDate": "12/27/2018 11:32:40",
      "content": "<p>Yes, it's true.  I'm a noob, my model sucks, and i didn't know what question to ask.   ABE_, thanks for your input.  I was not utilizing cross validation, I will try that next.  Catboost has a version of cross validation built in as cv.  Also, I now know my model was overfitting. <br>\nMy noob mind was suggesting that my model was 'underfitting' somehow.  I was imagining a line with a curve that barely moved.   As soon as you said 'overfitting' the obviousness dawned on me.   The overfitting on the training data that tested positive or '1' simply didn't match any combination in the test data therefore emitting only '0's'.</p>\n\n<p>This gives me a direction to work in.  Adams:  thanks for your comment too.  Before hand I was working on several loss functions, trying different parameters in a desperate attempt to predict at least one '1' in my submission.  So that answer was 'several'.</p>\n\n<p>Thanks guys, I see that the Kaggle community is very supportive.  This only motivates me to keep trying and dig deeper!</p>",
      "rawMarkdown": "Yes, it's true.  I'm a noob, my model sucks, and i didn't know what question to ask.   ABE_, thanks for your input.  I was not utilizing cross validation, I will try that next.  Catboost has a version of cross validation built in as cv.  Also, I now know my model was overfitting.  \nMy noob mind was suggesting that my model was 'underfitting' somehow.  I was imagining a line with a curve that barely moved.   As soon as you said 'overfitting' the obviousness dawned on me.   The overfitting on the training data that tested positive or '1' simply didn't match any combination in the test data therefore emitting only '0's'.\n\nThis gives me a direction to work in.  Adams:  thanks for your comment too.  Before hand I was working on several loss functions, trying different parameters in a desperate attempt to predict at least one '1' in my submission.  So that answer was 'several'.\n\nThanks guys, I see that the Kaggle community is very supportive.  This only motivates me to keep trying and dig deeper!",
      "votes": null
    },
    {
      "id": "446499",
      "postDate": "12/28/2018 07:34:44",
      "content": "<p>Another thing to think about is the choice of a threshold. For example, in some instances the probabilities could look like [0.0001, 0.001, 0.01, 0.02, 0.05, 0.00005] with correct labels of [0, 0, 1, 1, 1, 0]. Choosing 0.5 as a threshold would lead to all the predictions being 0 whereas choosing 0.005 would lead to all correct predictions.</p>",
      "rawMarkdown": "Another thing to think about is the choice of a threshold. For example, in some instances the probabilities could look like [0.0001, 0.001, 0.01, 0.02, 0.05, 0.00005] with correct labels of [0, 0, 1, 1, 1, 0]. Choosing 0.5 as a threshold would lead to all the predictions being 0 whereas choosing 0.005 would lead to all correct predictions.",
      "votes": null
    },
    {
      "id": "450439",
      "postDate": "01/05/2019 00:32:38",
      "content": "<p>As the training data only have 6% in 1s, maybe some up-sampling in the 1s would help.</p>",
      "rawMarkdown": "As the training data only have 6% in 1s, maybe some up-sampling in the 1s would help.",
      "votes": null
    },
    {
      "id": "452727",
      "postDate": "01/09/2019 04:56:54",
      "content": "<p>Couples of things you can try:\nAs @ABE_ mentioned , don't used catboost.predict for calculating predictions as it will be using 0.5 threshold by default. Use Predict_proba as mentioned over here :\n<a href=\"https://tech.yandex.com/catboost/doc/dg/concepts/python-reference_catboostclassifier_predict_proba-docpage/\">https://tech.yandex.com/catboost/doc/dg/concepts/python-reference_catboostclassifier_predict_proba-docpage/</a>\nOptimal threshold you can calculate from the auc: find the threshold which will maximize the difference of true positive rate and false positive rate.</p>\n\n<p>Also, For better results , try using scale_pos_weight features of catboost.</p>",
      "rawMarkdown": "Couples of things you can try:\nAs @ABE_ mentioned , don't used catboost.predict for calculating predictions as it will be using 0.5 threshold by default. Use Predict_proba as mentioned over here :\nhttps://tech.yandex.com/catboost/doc/dg/concepts/python-reference_catboostclassifier_predict_proba-docpage/\nOptimal threshold you can calculate from the auc: find the threshold which will maximize the difference of true positive rate and false positive rate.\n\nAlso, For better results , try using scale_pos_weight features of catboost.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 445893,
      "author_name": "jfr311",
      "author_url": "",
      "post_date": "12/27/2018 07:19:52",
      "content": "<p>Have you looked at a few of the kernels that are doing better than you? My first suggestion would be to post your own kernel to get more specific advice about your signal processing, training methods, cross-validation methods, etc.</p>\n\n<p>It definitely sounds like you are overfitting the imbalanced dataset. When you have only 6% positive samples, predicting all '0' values would give you a validation accuracy of about 94% despite being a terrible model. One way to work with an imbalanced dataset is to rebalance it so that the classes are equally likely. This may involve removing some '0's or adding some '1's. </p>\n\n<p>As far as I know, CatBoost doesn't support time series data so I assume you are summarizing the signals somehow? Maybe the features you are creating don't have any information value? Regarding the model, you might need to reduce the complexity of the model and increase the regularization.</p>\n\n<p>I don't mean to sound like an a**, but your question is essentially \"My model sucks, help me\" and without more information on exactly what you are working with and what you have tried to do, it's almost impossible to give advice without randomly guessing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445983,
      "author_name": "adamsfei",
      "author_url": "",
      "post_date": "12/27/2018 09:54:59",
      "content": "<p>What loss function are you optimising during the training?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 446037,
      "author_name": "oneday2one",
      "author_url": "",
      "post_date": "12/27/2018 11:32:40",
      "content": "<p>Yes, it's true.  I'm a noob, my model sucks, and i didn't know what question to ask.   ABE_, thanks for your input.  I was not utilizing cross validation, I will try that next.  Catboost has a version of cross validation built in as cv.  Also, I now know my model was overfitting. <br>\nMy noob mind was suggesting that my model was 'underfitting' somehow.  I was imagining a line with a curve that barely moved.   As soon as you said 'overfitting' the obviousness dawned on me.   The overfitting on the training data that tested positive or '1' simply didn't match any combination in the test data therefore emitting only '0's'.</p>\n\n<p>This gives me a direction to work in.  Adams:  thanks for your comment too.  Before hand I was working on several loss functions, trying different parameters in a desperate attempt to predict at least one '1' in my submission.  So that answer was 'several'.</p>\n\n<p>Thanks guys, I see that the Kaggle community is very supportive.  This only motivates me to keep trying and dig deeper!</p>",
      "votes": null,
      "replies": [
        {
          "id": 446499,
          "author_name": "jfr311",
          "author_url": "",
          "post_date": "12/28/2018 07:34:44",
          "content": "<p>Another thing to think about is the choice of a threshold. For example, in some instances the probabilities could look like [0.0001, 0.001, 0.01, 0.02, 0.05, 0.00005] with correct labels of [0, 0, 1, 1, 1, 0]. Choosing 0.5 as a threshold would lead to all the predictions being 0 whereas choosing 0.005 would lead to all correct predictions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 450439,
      "author_name": "cchen8888",
      "author_url": "",
      "post_date": "01/05/2019 00:32:38",
      "content": "<p>As the training data only have 6% in 1s, maybe some up-sampling in the 1s would help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 452727,
      "author_name": "harshit92",
      "author_url": "",
      "post_date": "01/09/2019 04:56:54",
      "content": "<p>Couples of things you can try:\nAs @ABE_ mentioned , don't used catboost.predict for calculating predictions as it will be using 0.5 threshold by default. Use Predict_proba as mentioned over here :\n<a href=\"https://tech.yandex.com/catboost/doc/dg/concepts/python-reference_catboostclassifier_predict_proba-docpage/\">https://tech.yandex.com/catboost/doc/dg/concepts/python-reference_catboostclassifier_predict_proba-docpage/</a>\nOptimal threshold you can calculate from the auc: find the threshold which will maximize the difference of true positive rate and false positive rate.</p>\n\n<p>Also, For better results , try using scale_pos_weight features of catboost.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "445698": "I am primarily using CatBoost.  It seems that no matter how much I tune the parameters, I am never able to predict a target value of 1, only 0's on the submission.\n\nHas anyone else encountered this?  Does anyone else have any advice?\nThanks!",
    "445893": "Have you looked at a few of the kernels that are doing better than you? My first suggestion would be to post your own kernel to get more specific advice about your signal processing, training methods, cross-validation methods, etc.\n\nIt definitely sounds like you are overfitting the imbalanced dataset. When you have only 6% positive samples, predicting all '0' values would give you a validation accuracy of about 94% despite being a terrible model. One way to work with an imbalanced dataset is to rebalance it so that the classes are equally likely. This may involve removing some '0's or adding some '1's. \n\nAs far as I know, CatBoost doesn't support time series data so I assume you are summarizing the signals somehow? Maybe the features you are creating don't have any information value? Regarding the model, you might need to reduce the complexity of the model and increase the regularization.\n\nI don't mean to sound like an a**, but your question is essentially \"My model sucks, help me\" and without more information on exactly what you are working with and what you have tried to do, it's almost impossible to give advice without randomly guessing.",
    "445983": "What loss function are you optimising during the training?",
    "446037": "Yes, it's true.  I'm a noob, my model sucks, and i didn't know what question to ask.   ABE_, thanks for your input.  I was not utilizing cross validation, I will try that next.  Catboost has a version of cross validation built in as cv.  Also, I now know my model was overfitting.  \nMy noob mind was suggesting that my model was 'underfitting' somehow.  I was imagining a line with a curve that barely moved.   As soon as you said 'overfitting' the obviousness dawned on me.   The overfitting on the training data that tested positive or '1' simply didn't match any combination in the test data therefore emitting only '0's'.\n\nThis gives me a direction to work in.  Adams:  thanks for your comment too.  Before hand I was working on several loss functions, trying different parameters in a desperate attempt to predict at least one '1' in my submission.  So that answer was 'several'.\n\nThanks guys, I see that the Kaggle community is very supportive.  This only motivates me to keep trying and dig deeper!",
    "446499": "Another thing to think about is the choice of a threshold. For example, in some instances the probabilities could look like [0.0001, 0.001, 0.01, 0.02, 0.05, 0.00005] with correct labels of [0, 0, 1, 1, 1, 0]. Choosing 0.5 as a threshold would lead to all the predictions being 0 whereas choosing 0.005 would lead to all correct predictions.",
    "450439": "As the training data only have 6% in 1s, maybe some up-sampling in the 1s would help.",
    "452727": "Couples of things you can try:\nAs @ABE_ mentioned , don't used catboost.predict for calculating predictions as it will be using 0.5 threshold by default. Use Predict_proba as mentioned over here :\nhttps://tech.yandex.com/catboost/doc/dg/concepts/python-reference_catboostclassifier_predict_proba-docpage/\nOptimal threshold you can calculate from the auc: find the threshold which will maximize the difference of true positive rate and false positive rate.\n\nAlso, For better results , try using scale_pos_weight features of catboost."
  },
  "source": "meta"
}