{
  "id": 71466,
  "title": "Why OOF in Logistic Regression ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71466",
  "author_name": "",
  "post_date": "2018-11-13T23:33:17.810793Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi guys,\nCan someone please explain what is OOF that is being used in most of the logistic regression codes in the kernels. I am trying to implement logistic regression without it but the kernel is crashing everytime. Is it because there are very large number of variables? \nAnkur</p>",
  "messages": [
    {
      "id": "420637",
      "postDate": "11/13/2018 23:33:17",
      "content": "<p>Hi guys,\nCan someone please explain what is OOF that is being used in most of the logistic regression codes in the kernels. I am trying to implement logistic regression without it but the kernel is crashing everytime. Is it because there are very large number of variables? \nAnkur</p>",
      "rawMarkdown": "Hi guys,\nCan someone please explain what is OOF that is being used in most of the logistic regression codes in the kernels. I am trying to implement logistic regression without it but the kernel is crashing everytime. Is it because there are very large number of variables? \nAnkur",
      "votes": null
    },
    {
      "id": "420645",
      "postDate": "11/13/2018 23:51:53",
      "content": "<p>The amount of data in this kernel adds up to 6GB which is lower than the amount of memory the kernels are supposed to provide, so that shouldn't be an issue. I'm not sure what OOF is that you're referring to. It's not present in the parameters or documentation for any of SciKit-Learn's implementations of logistic regressions. </p>\n\n<p>Outside of the context of logistic regressions, OOF refers to \"out-of-fold\", which is when you evaluate the aggregated predictions of each testing process conducted when conducting cross validation. One would use OOF to get a more robust measure of the performance of their model (relative to a simple test-train split). Depending on the implementation, it might be coded in such a way as to save a copy of every OOF prediction for every single data point in the dataset. Other ways could just save the overall performance of teach cross-validation pass and then average them at the end. Overall, there's no reason to believe that anything involving OOF would cause a crash for this project, even with an inefficient implementation. If that were the case, simply switching to a traditional test-train split should solve it. Overall I've been hearing of people having a lot of serious issues with Kaggle kernels recently and I myself have noticed very strange behavior from them, so that may well be the issue. </p>",
      "rawMarkdown": "The amount of data in this kernel adds up to 6GB which is lower than the amount of memory the kernels are supposed to provide, so that shouldn't be an issue. I'm not sure what OOF is that you're referring to. It's not present in the parameters or documentation for any of SciKit-Learn's implementations of logistic regressions. \n\nOutside of the context of logistic regressions, OOF refers to \"out-of-fold\", which is when you evaluate the aggregated predictions of each testing process conducted when conducting cross validation. One would use OOF to get a more robust measure of the performance of their model (relative to a simple test-train split). Depending on the implementation, it might be coded in such a way as to save a copy of every OOF prediction for every single data point in the dataset. Other ways could just save the overall performance of teach cross-validation pass and then average them at the end. Overall, there's no reason to believe that anything involving OOF would cause a crash for this project, even with an inefficient implementation. If that were the case, simply switching to a traditional test-train split should solve it. Overall I've been hearing of people having a lot of serious issues with Kaggle kernels recently and I myself have noticed very strange behavior from them, so that may well be the issue.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 420645,
      "author_name": "alecthekulak",
      "author_url": "",
      "post_date": "11/13/2018 23:51:53",
      "content": "<p>The amount of data in this kernel adds up to 6GB which is lower than the amount of memory the kernels are supposed to provide, so that shouldn't be an issue. I'm not sure what OOF is that you're referring to. It's not present in the parameters or documentation for any of SciKit-Learn's implementations of logistic regressions. </p>\n\n<p>Outside of the context of logistic regressions, OOF refers to \"out-of-fold\", which is when you evaluate the aggregated predictions of each testing process conducted when conducting cross validation. One would use OOF to get a more robust measure of the performance of their model (relative to a simple test-train split). Depending on the implementation, it might be coded in such a way as to save a copy of every OOF prediction for every single data point in the dataset. Other ways could just save the overall performance of teach cross-validation pass and then average them at the end. Overall, there's no reason to believe that anything involving OOF would cause a crash for this project, even with an inefficient implementation. If that were the case, simply switching to a traditional test-train split should solve it. Overall I've been hearing of people having a lot of serious issues with Kaggle kernels recently and I myself have noticed very strange behavior from them, so that may well be the issue. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "420637": "Hi guys,\nCan someone please explain what is OOF that is being used in most of the logistic regression codes in the kernels. I am trying to implement logistic regression without it but the kernel is crashing everytime. Is it because there are very large number of variables? \nAnkur",
    "420645": "The amount of data in this kernel adds up to 6GB which is lower than the amount of memory the kernels are supposed to provide, so that shouldn't be an issue. I'm not sure what OOF is that you're referring to. It's not present in the parameters or documentation for any of SciKit-Learn's implementations of logistic regressions. \n\nOutside of the context of logistic regressions, OOF refers to \"out-of-fold\", which is when you evaluate the aggregated predictions of each testing process conducted when conducting cross validation. One would use OOF to get a more robust measure of the performance of their model (relative to a simple test-train split). Depending on the implementation, it might be coded in such a way as to save a copy of every OOF prediction for every single data point in the dataset. Other ways could just save the overall performance of teach cross-validation pass and then average them at the end. Overall, there's no reason to believe that anything involving OOF would cause a crash for this project, even with an inefficient implementation. If that were the case, simply switching to a traditional test-train split should solve it. Overall I've been hearing of people having a lot of serious issues with Kaggle kernels recently and I myself have noticed very strange behavior from them, so that may well be the issue."
  },
  "source": "meta"
}