{
  "id": 54539,
  "title": "How do you choose a logistic regularization parameter?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54539",
  "author_name": "",
  "post_date": "2018-04-14T15:23:45.874655900Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>So I'm trying to do blending the respectable way, by fitting a meta-learner on out-of-sample predictions. But I've just realized (well re-realized: it's not the first time I've had this realization, but I guess it's not sinking in) that there's a huge elephant sitting in the middle of the room where I'm trying to blend.  The elephant's name is C.  (Actually his name is One Over Lambda, but Scikit-Learn is a close personal friend of his and always refers to him by his nickname, C.)  </p>\n\n<p>By default, he equals 1.  But that default is entirely arbitrary, one might even say meaningless, unless the regressors (base model predictions, in this case) are scaled.  If you express the inputs as raw ranks (in a range going from zero to several million), the effect of C=1 will provide only very modest regularization, whereas if you divide by N to get a range from 0 to 1, the effect of C=1 will give so much regularization that the results are often complete nonsense.  If you grasp the elephant's leg, you can say with great confidence that he's a tree, but if you grasp his trunk, you will be certain that he's a snake.</p>\n\n<p>Even if the value C had some scale-independent meaning, it's not clear what would be a reasonable default.  (Conversely, it's not clear what would be the appropriate way to scale your regressors so that the Scikit-Learn default of 1 becomes a reasonable one.)  As best I can tell, we're back to making intuitive judgments (looking at the weights implied by the regression coefficients and adjusting C until we get a set of weights that we like) or using trial and error (adjusting C until we get a LB score the we like).  </p>\n\n<p>Maybe the best approach (feasible in this case because there data are so large) is to divide the validation data in half, choose values of C for the first half, and optimize for the AUC value that results when the coefficients are applied to the second half.  In a simple world without a problematic time component we could just use k-fold CV, and it would be a straightforward case of optimizing stacking parameters with reusable training data.</p>",
  "messages": [
    {
      "id": "314088",
      "postDate": "04/14/2018 15:23:45",
      "content": "<p>So I'm trying to do blending the respectable way, by fitting a meta-learner on out-of-sample predictions. But I've just realized (well re-realized: it's not the first time I've had this realization, but I guess it's not sinking in) that there's a huge elephant sitting in the middle of the room where I'm trying to blend.  The elephant's name is C.  (Actually his name is One Over Lambda, but Scikit-Learn is a close personal friend of his and always refers to him by his nickname, C.)  </p>\n\n<p>By default, he equals 1.  But that default is entirely arbitrary, one might even say meaningless, unless the regressors (base model predictions, in this case) are scaled.  If you express the inputs as raw ranks (in a range going from zero to several million), the effect of C=1 will provide only very modest regularization, whereas if you divide by N to get a range from 0 to 1, the effect of C=1 will give so much regularization that the results are often complete nonsense.  If you grasp the elephant's leg, you can say with great confidence that he's a tree, but if you grasp his trunk, you will be certain that he's a snake.</p>\n\n<p>Even if the value C had some scale-independent meaning, it's not clear what would be a reasonable default.  (Conversely, it's not clear what would be the appropriate way to scale your regressors so that the Scikit-Learn default of 1 becomes a reasonable one.)  As best I can tell, we're back to making intuitive judgments (looking at the weights implied by the regression coefficients and adjusting C until we get a set of weights that we like) or using trial and error (adjusting C until we get a LB score the we like).  </p>\n\n<p>Maybe the best approach (feasible in this case because there data are so large) is to divide the validation data in half, choose values of C for the first half, and optimize for the AUC value that results when the coefficients are applied to the second half.  In a simple world without a problematic time component we could just use k-fold CV, and it would be a straightforward case of optimizing stacking parameters with reusable training data.</p>",
      "rawMarkdown": "So I'm trying to do blending the respectable way, by fitting a meta-learner on out-of-sample predictions. But I've just realized (well re-realized: it's not the first time I've had this realization, but I guess it's not sinking in) that there's a huge elephant sitting in the middle of the room where I'm trying to blend.  The elephant's name is C.  (Actually his name is One Over Lambda, but Scikit-Learn is a close personal friend of his and always refers to him by his nickname, C.)  \n\nBy default, he equals 1.  But that default is entirely arbitrary, one might even say meaningless, unless the regressors (base model predictions, in this case) are scaled.  If you express the inputs as raw ranks (in a range going from zero to several million), the effect of C=1 will provide only very modest regularization, whereas if you divide by N to get a range from 0 to 1, the effect of C=1 will give so much regularization that the results are often complete nonsense.  If you grasp the elephant's leg, you can say with great confidence that he's a tree, but if you grasp his trunk, you will be certain that he's a snake.\n\nEven if the value C had some scale-independent meaning, it's not clear what would be a reasonable default.  (Conversely, it's not clear what would be the appropriate way to scale your regressors so that the Scikit-Learn default of 1 becomes a reasonable one.)  As best I can tell, we're back to making intuitive judgments (looking at the weights implied by the regression coefficients and adjusting C until we get a set of weights that we like) or using trial and error (adjusting C until we get a LB score the we like).  \n\nMaybe the best approach (feasible in this case because there data are so large) is to divide the validation data in half, choose values of C for the first half, and optimize for the AUC value that results when the coefficients are applied to the second half.  In a simple world without a problematic time component we could just use k-fold CV, and it would be a straightforward case of optimizing stacking parameters with reusable training data.",
      "votes": null
    },
    {
      "id": "314170",
      "postDate": "04/14/2018 19:15:39",
      "content": "<blockquote>\n  <p>Even if the value C had some scale-independent meaning, it's not clear what would be a reasonable default.</p>\n</blockquote>\n\n<p>The answer to your question is the same as to many other questions on Kaggle: <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegressionCV.html\"><strong>cross-validation</strong></a>.</p>\n\n<p>By the way, I like that I am not the only one ranting about these issues O_O</p>",
      "rawMarkdown": "&gt; Even if the value C had some scale-independent meaning, it's not clear what would be a reasonable default.\n\nThe answer to your question is the same as to many other questions on Kaggle: [__cross-validation__](http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegressionCV.html).\n\nBy the way, I like that I am not the only one ranting about these issues O_O",
      "votes": null
    },
    {
      "id": "314219",
      "postDate": "04/14/2018 22:19:13",
      "content": "<p>My intuition would be that it may not be as bad as you think. In my experience, if you standard scale your feature inputs default C=1 actually tends to work pretty well and isn't overkill for that scale. I can't find any documentation on it, but my guess for why sklearn chooses that default is empirical success with it.</p>\n\n<p>In addition, validating/CV-ing a choice of C shouldn't need to involve a wide parameter search space. Starting from C=1, you can check if your results improve or decline when you increase or decrease regularization (say C=0.1 vs. C=10), and then have a strong sense of the appropriate direction to go in for regularization strength - you just need to understand whether you are under or overfitting. From there, you can check different orders of magnitude in a binary search style then narrow down to a fine-tuning range if you want to get it really really precise. Realistically I think you can find very good parameters with only a small handful of runs this way, especially because there are only a few orders of magnitude around C=1 that are actually relevant. Even if you want to grid search across say 8 orders of magnitude (I pretty much would never recommend grid searching like this on non-trivially sized datasets), that's still not <em>that</em> many runs, especially if you're using a small set of model predictions as features.</p>",
      "rawMarkdown": "My intuition would be that it may not be as bad as you think. In my experience, if you standard scale your feature inputs default C=1 actually tends to work pretty well and isn't overkill for that scale. I can't find any documentation on it, but my guess for why sklearn chooses that default is empirical success with it.\n\nIn addition, validating/CV-ing a choice of C shouldn't need to involve a wide parameter search space. Starting from C=1, you can check if your results improve or decline when you increase or decrease regularization (say C=0.1 vs. C=10), and then have a strong sense of the appropriate direction to go in for regularization strength - you just need to understand whether you are under or overfitting. From there, you can check different orders of magnitude in a binary search style then narrow down to a fine-tuning range if you want to get it really really precise. Realistically I think you can find very good parameters with only a small handful of runs this way, especially because there are only a few orders of magnitude around C=1 that are actually relevant. Even if you want to grid search across say 8 orders of magnitude (I pretty much would never recommend grid searching like this on non-trivially sized datasets), that's still not *that* many runs, especially if you're using a small set of model predictions as features.",
      "votes": null
    },
    {
      "id": "314225",
      "postDate": "04/14/2018 22:43:55",
      "content": "<blockquote>\n  <p>Starting from C=1, you can check if your results improve or decline when you increase or decrease regularization (say C=0.1 vs. C=10), and then have a strong sense of the appropriate direction to go in for regularization strength - you just need to understand whether you are under or overfitting.</p>\n</blockquote>\n\n<p>That's one way of doing it. Alternatively, you can use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegressionCV.html\"><strong>LogisticRegressionCV</strong></a>, select how many different Cs you want tested, and sklearn will do everything else for you. Including cross-validation.</p>",
      "rawMarkdown": "&gt; Starting from C=1, you can check if your results improve or decline when you increase or decrease regularization (say C=0.1 vs. C=10), and then have a strong sense of the appropriate direction to go in for regularization strength - you just need to understand whether you are under or overfitting.\n\nThat's one way of doing it. Alternatively, you can use [__LogisticRegressionCV__](http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegressionCV.html), select how many different Cs you want tested, and sklearn will do everything else for you. Including cross-validation.",
      "votes": null
    },
    {
      "id": "314227",
      "postDate": "04/14/2018 22:48:07",
      "content": "<blockquote>\n  <p>In my experience, if you standard scale your feature inputs default C=1 actually tends to work pretty well and isn't overkill for that scale.</p>\n</blockquote>\n\n<p>In my experience C=1 is too little regularization. It seems like I always get the optimum C values at &lt;0.1. Same with SVMs, where C has the same meaning. For some reason, sklearn has default C=1 for both classifiers. I can think of several times I have gotten C &gt; 1 for SVMs after a grid search, but not once for LR.</p>",
      "rawMarkdown": "&gt; In my experience, if you standard scale your feature inputs default C=1 actually tends to work pretty well and isn't overkill for that scale.\n\nIn my experience C=1 is too little regularization. It seems like I always get the optimum C values at &lt;0.1. Same with SVMs, where C has the same meaning. For some reason, sklearn has default C=1 for both classifiers. I can think of several times I have gotten C &gt; 1 for SVMs after a grid search, but not once for LR.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 314170,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/14/2018 19:15:39",
      "content": "<blockquote>\n  <p>Even if the value C had some scale-independent meaning, it's not clear what would be a reasonable default.</p>\n</blockquote>\n\n<p>The answer to your question is the same as to many other questions on Kaggle: <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegressionCV.html\"><strong>cross-validation</strong></a>.</p>\n\n<p>By the way, I like that I am not the only one ranting about these issues O_O</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 314219,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "04/14/2018 22:19:13",
      "content": "<p>My intuition would be that it may not be as bad as you think. In my experience, if you standard scale your feature inputs default C=1 actually tends to work pretty well and isn't overkill for that scale. I can't find any documentation on it, but my guess for why sklearn chooses that default is empirical success with it.</p>\n\n<p>In addition, validating/CV-ing a choice of C shouldn't need to involve a wide parameter search space. Starting from C=1, you can check if your results improve or decline when you increase or decrease regularization (say C=0.1 vs. C=10), and then have a strong sense of the appropriate direction to go in for regularization strength - you just need to understand whether you are under or overfitting. From there, you can check different orders of magnitude in a binary search style then narrow down to a fine-tuning range if you want to get it really really precise. Realistically I think you can find very good parameters with only a small handful of runs this way, especially because there are only a few orders of magnitude around C=1 that are actually relevant. Even if you want to grid search across say 8 orders of magnitude (I pretty much would never recommend grid searching like this on non-trivially sized datasets), that's still not <em>that</em> many runs, especially if you're using a small set of model predictions as features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 314225,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/14/2018 22:43:55",
          "content": "<blockquote>\n  <p>Starting from C=1, you can check if your results improve or decline when you increase or decrease regularization (say C=0.1 vs. C=10), and then have a strong sense of the appropriate direction to go in for regularization strength - you just need to understand whether you are under or overfitting.</p>\n</blockquote>\n\n<p>That's one way of doing it. Alternatively, you can use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegressionCV.html\"><strong>LogisticRegressionCV</strong></a>, select how many different Cs you want tested, and sklearn will do everything else for you. Including cross-validation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 314227,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/14/2018 22:48:07",
          "content": "<blockquote>\n  <p>In my experience, if you standard scale your feature inputs default C=1 actually tends to work pretty well and isn't overkill for that scale.</p>\n</blockquote>\n\n<p>In my experience C=1 is too little regularization. It seems like I always get the optimum C values at &lt;0.1. Same with SVMs, where C has the same meaning. For some reason, sklearn has default C=1 for both classifiers. I can think of several times I have gotten C &gt; 1 for SVMs after a grid search, but not once for LR.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "314088": "So I'm trying to do blending the respectable way, by fitting a meta-learner on out-of-sample predictions. But I've just realized (well re-realized: it's not the first time I've had this realization, but I guess it's not sinking in) that there's a huge elephant sitting in the middle of the room where I'm trying to blend.  The elephant's name is C.  (Actually his name is One Over Lambda, but Scikit-Learn is a close personal friend of his and always refers to him by his nickname, C.)  \n\nBy default, he equals 1.  But that default is entirely arbitrary, one might even say meaningless, unless the regressors (base model predictions, in this case) are scaled.  If you express the inputs as raw ranks (in a range going from zero to several million), the effect of C=1 will provide only very modest regularization, whereas if you divide by N to get a range from 0 to 1, the effect of C=1 will give so much regularization that the results are often complete nonsense.  If you grasp the elephant's leg, you can say with great confidence that he's a tree, but if you grasp his trunk, you will be certain that he's a snake.\n\nEven if the value C had some scale-independent meaning, it's not clear what would be a reasonable default.  (Conversely, it's not clear what would be the appropriate way to scale your regressors so that the Scikit-Learn default of 1 becomes a reasonable one.)  As best I can tell, we're back to making intuitive judgments (looking at the weights implied by the regression coefficients and adjusting C until we get a set of weights that we like) or using trial and error (adjusting C until we get a LB score the we like).  \n\nMaybe the best approach (feasible in this case because there data are so large) is to divide the validation data in half, choose values of C for the first half, and optimize for the AUC value that results when the coefficients are applied to the second half.  In a simple world without a problematic time component we could just use k-fold CV, and it would be a straightforward case of optimizing stacking parameters with reusable training data.",
    "314170": "&gt; Even if the value C had some scale-independent meaning, it's not clear what would be a reasonable default.\n\nThe answer to your question is the same as to many other questions on Kaggle: [__cross-validation__](http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegressionCV.html).\n\nBy the way, I like that I am not the only one ranting about these issues O_O",
    "314219": "My intuition would be that it may not be as bad as you think. In my experience, if you standard scale your feature inputs default C=1 actually tends to work pretty well and isn't overkill for that scale. I can't find any documentation on it, but my guess for why sklearn chooses that default is empirical success with it.\n\nIn addition, validating/CV-ing a choice of C shouldn't need to involve a wide parameter search space. Starting from C=1, you can check if your results improve or decline when you increase or decrease regularization (say C=0.1 vs. C=10), and then have a strong sense of the appropriate direction to go in for regularization strength - you just need to understand whether you are under or overfitting. From there, you can check different orders of magnitude in a binary search style then narrow down to a fine-tuning range if you want to get it really really precise. Realistically I think you can find very good parameters with only a small handful of runs this way, especially because there are only a few orders of magnitude around C=1 that are actually relevant. Even if you want to grid search across say 8 orders of magnitude (I pretty much would never recommend grid searching like this on non-trivially sized datasets), that's still not *that* many runs, especially if you're using a small set of model predictions as features.",
    "314225": "&gt; Starting from C=1, you can check if your results improve or decline when you increase or decrease regularization (say C=0.1 vs. C=10), and then have a strong sense of the appropriate direction to go in for regularization strength - you just need to understand whether you are under or overfitting.\n\nThat's one way of doing it. Alternatively, you can use [__LogisticRegressionCV__](http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegressionCV.html), select how many different Cs you want tested, and sklearn will do everything else for you. Including cross-validation.",
    "314227": "&gt; In my experience, if you standard scale your feature inputs default C=1 actually tends to work pretty well and isn't overkill for that scale.\n\nIn my experience C=1 is too little regularization. It seems like I always get the optimum C values at &lt;0.1. Same with SVMs, where C has the same meaning. For some reason, sklearn has default C=1 for both classifiers. I can think of several times I have gotten C &gt; 1 for SVMs after a grid search, but not once for LR."
  },
  "source": "meta"
}