{
  "id": 344825,
  "title": "One solution to the Out of Memory message",
  "url": "/competitions/amex-default-prediction/discussion/344825",
  "author_name": "Juan Smith Perera",
  "post_date": "2022-08-16T18:56:44.145000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>For the last 10 days I have not been able to run my notebook without getting the Out of Memory message. It all started when I tried to increase the number of features beyond 500s.</p>\n<p>After trying a CUDA versión (works but it gave be lower CV) and methods to reduce the size of the train set, I came out with a simple solution:</p>\n<ol>\n<li>I divided <a href=\"https://www.kaggle.com/code/jsmithperera/amex-lightgbm-v3\" target=\"_blank\">my NB</a> into two parts: training and submission. </li>\n</ol>\n<p>2.- The output of my<a href=\"https://www.kaggle.com/code/jsmithperera/lgbm-v3-model\" target=\"_blank\"> training NB</a> is the model, saved as a Booster from LightGBM: model.booster_.save_model(\"./amex-model.txt\")</p>\n<p>3.- The <a href=\"https://www.kaggle.com/code/jsmithperera/lgbm-v3-submission\" target=\"_blank\">submission NB</a> reads the model and the test set, and makes the submission. </p>\n<p>The trick here is that the Booster model does not have the prediction_proba attribute but, as it turns out, this instruction </p>\n<p>model.predict(test[features],raw_score=True)</p>\n<p>produces the log-odds of the predictions: ln(p/(1-p)), which has a one-to-one relationship with p, the probability of default. Better yet, the amex-metric is invariant to the log-odds function. So, log-odds are okey for submission. </p>\n<p>If you want to know how to pass a model from one ND to another, <a href=\"https://www.kaggle.com/discussions/getting-started/333636\" target=\"_blank\">read this</a>, which I prepared for the JPX Competition. </p>\n<p>I hope this helps.</p>",
  "messages": [
    {
      "id": 1901571,
      "postDate": "2022-08-16T18:56:44.147Z",
      "content": "<p>For the last 10 days I have not been able to run my notebook without getting the Out of Memory message. It all started when I tried to increase the number of features beyond 500s.</p>\n<p>After trying a CUDA versión (works but it gave be lower CV) and methods to reduce the size of the train set, I came out with a simple solution:</p>\n<ol>\n<li>I divided <a href=\"https://www.kaggle.com/code/jsmithperera/amex-lightgbm-v3\" target=\"_blank\">my NB</a> into two parts: training and submission. </li>\n</ol>\n<p>2.- The output of my<a href=\"https://www.kaggle.com/code/jsmithperera/lgbm-v3-model\" target=\"_blank\"> training NB</a> is the model, saved as a Booster from LightGBM: model.booster_.save_model(\"./amex-model.txt\")</p>\n<p>3.- The <a href=\"https://www.kaggle.com/code/jsmithperera/lgbm-v3-submission\" target=\"_blank\">submission NB</a> reads the model and the test set, and makes the submission. </p>\n<p>The trick here is that the Booster model does not have the prediction_proba attribute but, as it turns out, this instruction </p>\n<p>model.predict(test[features],raw_score=True)</p>\n<p>produces the log-odds of the predictions: ln(p/(1-p)), which has a one-to-one relationship with p, the probability of default. Better yet, the amex-metric is invariant to the log-odds function. So, log-odds are okey for submission. </p>\n<p>If you want to know how to pass a model from one ND to another, <a href=\"https://www.kaggle.com/discussions/getting-started/333636\" target=\"_blank\">read this</a>, which I prepared for the JPX Competition. </p>\n<p>I hope this helps.</p>",
      "rawMarkdown": "For the last 10 days I have not been able to run my notebook without getting the Out of Memory message. It all started when I tried to increase the number of features beyond 500s.\n\nAfter trying a CUDA versión (works but it gave be lower CV) and methods to reduce the size of the train set, I came out with a simple solution:\n\n1. I divided [my NB](https://www.kaggle.com/code/jsmithperera/amex-lightgbm-v3) into two parts: training and submission. \n\n2.- The output of my[ training NB](https://www.kaggle.com/code/jsmithperera/lgbm-v3-model) is the model, saved as a Booster from LightGBM: model.booster_.save_model(\"./amex-model.txt\")\n\n3.- The [submission NB](https://www.kaggle.com/code/jsmithperera/lgbm-v3-submission) reads the model and the test set, and makes the submission. \n\nThe trick here is that the Booster model does not have the prediction_proba attribute but, as it turns out, this instruction \n\nmodel.predict(test[features],raw_score=True)\n\nproduces the log-odds of the predictions: ln(p/(1-p)), which has a one-to-one relationship with p, the probability of default. Better yet, the amex-metric is invariant to the log-odds function. So, log-odds are okey for submission. \n\nIf you want to know how to pass a model from one ND to another, [read this](https://www.kaggle.com/discussions/getting-started/333636), which I prepared for the JPX Competition. \n\nI hope this helps.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1901571": "For the last 10 days I have not been able to run my notebook without getting the Out of Memory message. It all started when I tried to increase the number of features beyond 500s.\n\nAfter trying a CUDA versión (works but it gave be lower CV) and methods to reduce the size of the train set, I came out with a simple solution:\n\n1. I divided [my NB](https://www.kaggle.com/code/jsmithperera/amex-lightgbm-v3) into two parts: training and submission. \n\n2.- The output of my[ training NB](https://www.kaggle.com/code/jsmithperera/lgbm-v3-model) is the model, saved as a Booster from LightGBM: model.booster_.save_model(\"./amex-model.txt\")\n\n3.- The [submission NB](https://www.kaggle.com/code/jsmithperera/lgbm-v3-submission) reads the model and the test set, and makes the submission. \n\nThe trick here is that the Booster model does not have the prediction_proba attribute but, as it turns out, this instruction \n\nmodel.predict(test[features],raw_score=True)\n\nproduces the log-odds of the predictions: ln(p/(1-p)), which has a one-to-one relationship with p, the probability of default. Better yet, the amex-metric is invariant to the log-odds function. So, log-odds are okey for submission. \n\nIf you want to know how to pass a model from one ND to another, [read this](https://www.kaggle.com/discussions/getting-started/333636), which I prepared for the JPX Competition. \n\nI hope this helps."
  }
}