{
  "id": 55111,
  "title": "LGBM Bayesian Optimization",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55111",
  "author_name": "",
  "post_date": "2018-04-22T08:05:42.316528900Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi,\nWhich hyperparameters do you tune with Bayesian Optimization?</p>\n\n<p>1- It takes a veeeeery long time to run.</p>\n\n<p>2- I find that the AUC of most iterations of the optimization is somewhere around +- 0.0005 AUC\nIt gives no radical improvement or difference when using different parameters.</p>\n\n<p>3- Sometimes COMPLETELY different parameters give me the same \"best\" result, sometimes even \nnum_leaves = 8, max_depth = 3 \nappears up top  together with another iteration that has\nnum_leaves = 200, max_depth = 60</p>\n\n<p>So generally i'm not sure i'm using it properly.</p>\n\n<h2>These are the parameters I tune:</h2>\n\n<ul>\n<li>num.leaves = c(4L,40L)</li>\n<li>max.depth = c(2L,63L)</li>\n<li>min.child.samples = c(1L,500L)</li>\n<li>sub.sample = c(0.8,1)</li>\n<li>min.child.weight = c(0.01,57000)</li>\n<li>col.sample  = c(0.5,1)</li>\n<li>min.split.gain =  c(0,0.5)</li>\n<li>scale_pos_weight = c(1,400)</li>\n</ul>\n\n<h2>With this configuration (In R):</h2>\n\n<ul>\n<li>init_grid_dt = NULL</li>\n<li>init_points = 20</li>\n<li>n_iter = 150</li>\n<li>acq = \"ucb\"</li>\n<li>kappa = 2.576</li>\n<li>eps = 0.0,</li>\n</ul>\n\n<p>Do I need to run it with different configurations?</p>\n\n<p>Thanks!!</p>",
  "messages": [
    {
      "id": "317678",
      "postDate": "04/22/2018 08:05:42",
      "content": "<p>Hi,\nWhich hyperparameters do you tune with Bayesian Optimization?</p>\n\n<p>1- It takes a veeeeery long time to run.</p>\n\n<p>2- I find that the AUC of most iterations of the optimization is somewhere around +- 0.0005 AUC\nIt gives no radical improvement or difference when using different parameters.</p>\n\n<p>3- Sometimes COMPLETELY different parameters give me the same \"best\" result, sometimes even \nnum_leaves = 8, max_depth = 3 \nappears up top  together with another iteration that has\nnum_leaves = 200, max_depth = 60</p>\n\n<p>So generally i'm not sure i'm using it properly.</p>\n\n<h2>These are the parameters I tune:</h2>\n\n<ul>\n<li>num.leaves = c(4L,40L)</li>\n<li>max.depth = c(2L,63L)</li>\n<li>min.child.samples = c(1L,500L)</li>\n<li>sub.sample = c(0.8,1)</li>\n<li>min.child.weight = c(0.01,57000)</li>\n<li>col.sample  = c(0.5,1)</li>\n<li>min.split.gain =  c(0,0.5)</li>\n<li>scale_pos_weight = c(1,400)</li>\n</ul>\n\n<h2>With this configuration (In R):</h2>\n\n<ul>\n<li>init_grid_dt = NULL</li>\n<li>init_points = 20</li>\n<li>n_iter = 150</li>\n<li>acq = \"ucb\"</li>\n<li>kappa = 2.576</li>\n<li>eps = 0.0,</li>\n</ul>\n\n<p>Do I need to run it with different configurations?</p>\n\n<p>Thanks!!</p>",
      "rawMarkdown": "Hi,\nWhich hyperparameters do you tune with Bayesian Optimization?\n\n1- It takes a veeeeery long time to run.\n\n2- I find that the AUC of most iterations of the optimization is somewhere around +- 0.0005 AUC\nIt gives no radical improvement or difference when using different parameters.\n\n3- Sometimes COMPLETELY different parameters give me the same \"best\" result, sometimes even \nnum_leaves = 8, max_depth = 3 \nappears up top  together with another iteration that has\nnum_leaves = 200, max_depth = 60\n\nSo generally i'm not sure i'm using it properly.\n\nThese are the parameters I tune:\n--------------------------------\n\n - num.leaves = c(4L,40L)\n - max.depth = c(2L,63L)\n - min.child.samples = c(1L,500L)\n - sub.sample = c(0.8,1)\n - min.child.weight = c(0.01,57000)\n - col.sample  = c(0.5,1)\n - min.split.gain =  c(0,0.5)\n - scale_pos_weight = c(1,400)\n\nWith this configuration (In R):\n-------------------------------\n\n- init_grid_dt = NULL\n- init_points = 20\n- n_iter = 150\n- acq = \"ucb\"\n- kappa = 2.576\n- eps = 0.0,\n\nDo I need to run it with different configurations?\n\nThanks!!",
      "votes": null
    },
    {
      "id": "317721",
      "postDate": "04/22/2018 10:57:58",
      "content": "<p>How are you handling cross validation? This is often where things go wrong in time series.</p>\n\n<p>Check this out:</p>\n\n<p><a href=\"https://robjhyndman.com/hyndsight/tscv/\">https://robjhyndman.com/hyndsight/tscv/</a></p>",
      "rawMarkdown": "How are you handling cross validation? This is often where things go wrong in time series.\n\nCheck this out:\n\nhttps://robjhyndman.com/hyndsight/tscv/",
      "votes": null
    },
    {
      "id": "317729",
      "postDate": "04/22/2018 11:45:11",
      "content": "<p>Hi @Bryan Arnold,\nWhen I used Bayesian Optimization  I used day 9 hour 4 for CV.\nsince then I moved to using entire day 9 for CV.\nI just want to make sure if I did anything wrong before running Bayesian Optimization again on my new features and methods.</p>\n\n<p>I wouldn't treat this problem as a time series one as we have only 4 different days in training + test.\nI do have the \"hour\" as a feature of course.</p>\n\n<p>Can you elaborate what you meant?</p>",
      "rawMarkdown": "Hi @Bryan Arnold,\nWhen I used Bayesian Optimization  I used day 9 hour 4 for CV.\nsince then I moved to using entire day 9 for CV.\nI just want to make sure if I did anything wrong before running Bayesian Optimization again on my new features and methods.\n\nI wouldn't treat this problem as a time series one as we have only 4 different days in training + test.\nI do have the \"hour\" as a feature of course.\n\nCan you elaborate what you meant?",
      "votes": null
    },
    {
      "id": "317739",
      "postDate": "04/22/2018 12:34:15",
      "content": "<p>Are you using a fixed learning rate and number of trees? Also, are you using early stopping?</p>\n\n<p>To your 3rd point, in your specification with num leaves &lt;= 40 increasing the depth doesn't change anything, as the model complexity is limited by the leave number. Using 31 Leaves in my runs the depth usually doesnt go higher than 9 so i would assume that with 40 it will be around depth 10-11. So a max tree depth higher than 11 basically means that the algorithm will grow the maximum number of leaves without caring about depth.</p>",
      "rawMarkdown": "Are you using a fixed learning rate and number of trees? Also, are you using early stopping?\n\nTo your 3rd point, in your specification with num leaves &lt;= 40 increasing the depth doesn't change anything, as the model complexity is limited by the leave number. Using 31 Leaves in my runs the depth usually doesnt go higher than 9 so i would assume that with 40 it will be around depth 10-11. So a max tree depth higher than 11 basically means that the algorithm will grow the maximum number of leaves without caring about depth.",
      "votes": null
    },
    {
      "id": "317745",
      "postDate": "04/22/2018 12:54:44",
      "content": "<p>Hi @Malte Nalenz,\nThanks for your response.</p>\n\n<p>To your question - Yes, I do use a fixed learning rate of 0.3 while optimizing (0.2 for final training), and also early stopping of 10 rounds.</p>\n\n<p>I don't really understand your statement about Leaves and Depth,\nFrom what I understand, Depth sets how many levels are permitted in the tree, so if I set 10 leaves in 2 levels, the tree will only expand horizontally to allow up to 10 leaves across up to  2 levels.\nbut if i set 10 leaves with 10 Depth, it allows the algorithms to grow a tree 10 levels deep and grow vertically.\nAm I getting this wrong?</p>\n\n<p>Plus - my real life example of 8 leaves and 3 depths VS 200 leaves and 60 depth having the same AUC (along with about 10 more examples somewhere in the middle), seems a bit odd =/</p>\n\n<p>Could you maybe share the parameters you use in Bayesian Optimization? (if you use it)\nI found absolutely no documentation or articles about the Kappa, Epsilon and Acq parameters of the algorithm.</p>\n\n<p>Thanks again for your response =)</p>",
      "rawMarkdown": "Hi @Malte Nalenz,\nThanks for your response.\n\nTo your question - Yes, I do use a fixed learning rate of 0.3 while optimizing (0.2 for final training), and also early stopping of 10 rounds.\n\nI don't really understand your statement about Leaves and Depth,\nFrom what I understand, Depth sets how many levels are permitted in the tree, so if I set 10 leaves in 2 levels, the tree will only expand horizontally to allow up to 10 leaves across up to  2 levels.\nbut if i set 10 leaves with 10 Depth, it allows the algorithms to grow a tree 10 levels deep and grow vertically.\nAm I getting this wrong?\n\nPlus - my real life example of 8 leaves and 3 depths VS 200 leaves and 60 depth having the same AUC (along with about 10 more examples somewhere in the middle), seems a bit odd =/\n\nCould you maybe share the parameters you use in Bayesian Optimization? (if you use it)\nI found absolutely no documentation or articles about the Kappa, Epsilon and Acq parameters of the algorithm.\n\nThanks again for your response =)",
      "votes": null
    },
    {
      "id": "318318",
      "postDate": "04/23/2018 15:52:49",
      "content": "<blockquote>\n  <p>Can you elaborate what you meant?</p>\n</blockquote>\n\n<p>Sure thing. Let's forget about time series for the moment and think about a 5-fold cross-validation schema. It goes something like this:</p>\n\n<p>First, break up the training data into 5 bins:</p>\n\n<p>[--20%--][--20%--][--20%--][--20%--][--20%--]</p>\n\n<p>Then train your model on bins 1-4 and validate on *bin_5*.</p>\n\n<p>And repeat: train the model on bins excluding *bin_i* and validate on *bin_i*.</p>\n\n<p>Finally, average your validation scores.</p>\n\n<p>Ok. That's how cross-validation works normally. Two problems with that when you're dealing with time series:</p>\n\n<ol>\n<li>You may not want to use information from the future (could cause leakage depending on the scenario).</li>\n<li>Prediction accuracy falls off exponentially with respect to time for time series forecasts (i.e. you may predict 4 steps ahead well, but not 24 steps ahead).</li>\n</ol>\n\n<p>In our current problem, we don't have to worry about such eventualities though. Or do we? It depends on what your model is doing. If it is using features that encapsulates click rate or time duration between clicks, it might be a factor. If you're using daily aggregates, but only considering an hour, that could cause problems. Etc., etc.</p>\n\n<p>TL-DR; There is more to consider when dealing with time related data.</p>",
      "rawMarkdown": "&gt;Can you elaborate what you meant?\n\nSure thing. Let's forget about time series for the moment and think about a 5-fold cross-validation schema. It goes something like this:\n\nFirst, break up the training data into 5 bins:\n\n[--20%--][--20%--][--20%--][--20%--][--20%--]\n\nThen train your model on bins 1-4 and validate on *bin_5*.\n\nAnd repeat: train the model on bins excluding *bin_i* and validate on *bin_i*.\n\nFinally, average your validation scores.\n\nOk. That's how cross-validation works normally. Two problems with that when you're dealing with time series:\n\n 1. You may not want to use information from the future (could cause leakage depending on the scenario).\n 2. Prediction accuracy falls off exponentially with respect to time for time series forecasts (i.e. you may predict 4 steps ahead well, but not 24 steps ahead).\n\nIn our current problem, we don't have to worry about such eventualities though. Or do we? It depends on what your model is doing. If it is using features that encapsulates click rate or time duration between clicks, it might be a factor. If you're using daily aggregates, but only considering an hour, that could cause problems. Etc., etc.\n\nTL-DR; There is more to consider when dealing with time related data.",
      "votes": null
    },
    {
      "id": "318319",
      "postDate": "04/23/2018 15:55:36",
      "content": "<p>Your welcome! \nWhat you say about the depth is generally correct, but in the extrem cases with very few leaves the trees can not reach the maximum depth. On each level it will be at least 2 leaves (when expending vertically), which means that with 10 leaves the maximum depth that can be reached is 5. \nThat very shallow trees and very deep trees give the same result really seems odd. Might be connected to the features that you use.</p>\n\n<p>I didnt start with parameter optimization yet, still trying to find better features but will share it with you if i find good settings!</p>",
      "rawMarkdown": "Your welcome! \nWhat you say about the depth is generally correct, but in the extrem cases with very few leaves the trees can not reach the maximum depth. On each level it will be at least 2 leaves (when expending vertically), which means that with 10 leaves the maximum depth that can be reached is 5. \nThat very shallow trees and very deep trees give the same result really seems odd. Might be connected to the features that you use.\n\nI didnt start with parameter optimization yet, still trying to find better features but will share it with you if i find good settings!",
      "votes": null
    },
    {
      "id": "318357",
      "postDate": "04/23/2018 16:56:11",
      "content": "<p>Thanks a lot for your reply,\nI understand what you mean now about the leaves,\nLet me know if you find anything useful,\nIt's really odd that there are 0(!!) materials about RBayesianOptimization parameters out there, no examples, no instructions.\nEven the documentation itself is not very helpful.\nI'll probably revert to random search soon.</p>\n\n<p>Thanks again!</p>",
      "rawMarkdown": "Thanks a lot for your reply,\nI understand what you mean now about the leaves,\nLet me know if you find anything useful,\nIt's really odd that there are 0(!!) materials about RBayesianOptimization parameters out there, no examples, no instructions.\nEven the documentation itself is not very helpful.\nI'll probably revert to random search soon.\n\nThanks again!",
      "votes": null
    },
    {
      "id": "318360",
      "postDate": "04/23/2018 16:58:03",
      "content": "<p>Thanks for your reply,\nI agree with you but not sure it's related to this specific case, as we have only 3-4 days of data and\nI'm using entire day 9 for validation, not mixing anything between the hours\\days.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Thanks for your reply,\nI agree with you but not sure it's related to this specific case, as we have only 3-4 days of data and\nI'm using entire day 9 for validation, not mixing anything between the hours\\days.\n\nThanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 317721,
      "author_name": "puremath86",
      "author_url": "",
      "post_date": "04/22/2018 10:57:58",
      "content": "<p>How are you handling cross validation? This is often where things go wrong in time series.</p>\n\n<p>Check this out:</p>\n\n<p><a href=\"https://robjhyndman.com/hyndsight/tscv/\">https://robjhyndman.com/hyndsight/tscv/</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 317729,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "04/22/2018 11:45:11",
          "content": "<p>Hi @Bryan Arnold,\nWhen I used Bayesian Optimization  I used day 9 hour 4 for CV.\nsince then I moved to using entire day 9 for CV.\nI just want to make sure if I did anything wrong before running Bayesian Optimization again on my new features and methods.</p>\n\n<p>I wouldn't treat this problem as a time series one as we have only 4 different days in training + test.\nI do have the \"hour\" as a feature of course.</p>\n\n<p>Can you elaborate what you meant?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318318,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "04/23/2018 15:52:49",
          "content": "<blockquote>\n  <p>Can you elaborate what you meant?</p>\n</blockquote>\n\n<p>Sure thing. Let's forget about time series for the moment and think about a 5-fold cross-validation schema. It goes something like this:</p>\n\n<p>First, break up the training data into 5 bins:</p>\n\n<p>[--20%--][--20%--][--20%--][--20%--][--20%--]</p>\n\n<p>Then train your model on bins 1-4 and validate on *bin_5*.</p>\n\n<p>And repeat: train the model on bins excluding *bin_i* and validate on *bin_i*.</p>\n\n<p>Finally, average your validation scores.</p>\n\n<p>Ok. That's how cross-validation works normally. Two problems with that when you're dealing with time series:</p>\n\n<ol>\n<li>You may not want to use information from the future (could cause leakage depending on the scenario).</li>\n<li>Prediction accuracy falls off exponentially with respect to time for time series forecasts (i.e. you may predict 4 steps ahead well, but not 24 steps ahead).</li>\n</ol>\n\n<p>In our current problem, we don't have to worry about such eventualities though. Or do we? It depends on what your model is doing. If it is using features that encapsulates click rate or time duration between clicks, it might be a factor. If you're using daily aggregates, but only considering an hour, that could cause problems. Etc., etc.</p>\n\n<p>TL-DR; There is more to consider when dealing with time related data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318360,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "04/23/2018 16:58:03",
          "content": "<p>Thanks for your reply,\nI agree with you but not sure it's related to this specific case, as we have only 3-4 days of data and\nI'm using entire day 9 for validation, not mixing anything between the hours\\days.</p>\n\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 317739,
      "author_name": "malten",
      "author_url": "",
      "post_date": "04/22/2018 12:34:15",
      "content": "<p>Are you using a fixed learning rate and number of trees? Also, are you using early stopping?</p>\n\n<p>To your 3rd point, in your specification with num leaves &lt;= 40 increasing the depth doesn't change anything, as the model complexity is limited by the leave number. Using 31 Leaves in my runs the depth usually doesnt go higher than 9 so i would assume that with 40 it will be around depth 10-11. So a max tree depth higher than 11 basically means that the algorithm will grow the maximum number of leaves without caring about depth.</p>",
      "votes": null,
      "replies": [
        {
          "id": 317745,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "04/22/2018 12:54:44",
          "content": "<p>Hi @Malte Nalenz,\nThanks for your response.</p>\n\n<p>To your question - Yes, I do use a fixed learning rate of 0.3 while optimizing (0.2 for final training), and also early stopping of 10 rounds.</p>\n\n<p>I don't really understand your statement about Leaves and Depth,\nFrom what I understand, Depth sets how many levels are permitted in the tree, so if I set 10 leaves in 2 levels, the tree will only expand horizontally to allow up to 10 leaves across up to  2 levels.\nbut if i set 10 leaves with 10 Depth, it allows the algorithms to grow a tree 10 levels deep and grow vertically.\nAm I getting this wrong?</p>\n\n<p>Plus - my real life example of 8 leaves and 3 depths VS 200 leaves and 60 depth having the same AUC (along with about 10 more examples somewhere in the middle), seems a bit odd =/</p>\n\n<p>Could you maybe share the parameters you use in Bayesian Optimization? (if you use it)\nI found absolutely no documentation or articles about the Kappa, Epsilon and Acq parameters of the algorithm.</p>\n\n<p>Thanks again for your response =)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318319,
          "author_name": "malten",
          "author_url": "",
          "post_date": "04/23/2018 15:55:36",
          "content": "<p>Your welcome! \nWhat you say about the depth is generally correct, but in the extrem cases with very few leaves the trees can not reach the maximum depth. On each level it will be at least 2 leaves (when expending vertically), which means that with 10 leaves the maximum depth that can be reached is 5. \nThat very shallow trees and very deep trees give the same result really seems odd. Might be connected to the features that you use.</p>\n\n<p>I didnt start with parameter optimization yet, still trying to find better features but will share it with you if i find good settings!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318357,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "04/23/2018 16:56:11",
          "content": "<p>Thanks a lot for your reply,\nI understand what you mean now about the leaves,\nLet me know if you find anything useful,\nIt's really odd that there are 0(!!) materials about RBayesianOptimization parameters out there, no examples, no instructions.\nEven the documentation itself is not very helpful.\nI'll probably revert to random search soon.</p>\n\n<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "317678": "Hi,\nWhich hyperparameters do you tune with Bayesian Optimization?\n\n1- It takes a veeeeery long time to run.\n\n2- I find that the AUC of most iterations of the optimization is somewhere around +- 0.0005 AUC\nIt gives no radical improvement or difference when using different parameters.\n\n3- Sometimes COMPLETELY different parameters give me the same \"best\" result, sometimes even \nnum_leaves = 8, max_depth = 3 \nappears up top  together with another iteration that has\nnum_leaves = 200, max_depth = 60\n\nSo generally i'm not sure i'm using it properly.\n\nThese are the parameters I tune:\n--------------------------------\n\n - num.leaves = c(4L,40L)\n - max.depth = c(2L,63L)\n - min.child.samples = c(1L,500L)\n - sub.sample = c(0.8,1)\n - min.child.weight = c(0.01,57000)\n - col.sample  = c(0.5,1)\n - min.split.gain =  c(0,0.5)\n - scale_pos_weight = c(1,400)\n\nWith this configuration (In R):\n-------------------------------\n\n- init_grid_dt = NULL\n- init_points = 20\n- n_iter = 150\n- acq = \"ucb\"\n- kappa = 2.576\n- eps = 0.0,\n\nDo I need to run it with different configurations?\n\nThanks!!",
    "317721": "How are you handling cross validation? This is often where things go wrong in time series.\n\nCheck this out:\n\nhttps://robjhyndman.com/hyndsight/tscv/",
    "317729": "Hi @Bryan Arnold,\nWhen I used Bayesian Optimization  I used day 9 hour 4 for CV.\nsince then I moved to using entire day 9 for CV.\nI just want to make sure if I did anything wrong before running Bayesian Optimization again on my new features and methods.\n\nI wouldn't treat this problem as a time series one as we have only 4 different days in training + test.\nI do have the \"hour\" as a feature of course.\n\nCan you elaborate what you meant?",
    "317739": "Are you using a fixed learning rate and number of trees? Also, are you using early stopping?\n\nTo your 3rd point, in your specification with num leaves &lt;= 40 increasing the depth doesn't change anything, as the model complexity is limited by the leave number. Using 31 Leaves in my runs the depth usually doesnt go higher than 9 so i would assume that with 40 it will be around depth 10-11. So a max tree depth higher than 11 basically means that the algorithm will grow the maximum number of leaves without caring about depth.",
    "317745": "Hi @Malte Nalenz,\nThanks for your response.\n\nTo your question - Yes, I do use a fixed learning rate of 0.3 while optimizing (0.2 for final training), and also early stopping of 10 rounds.\n\nI don't really understand your statement about Leaves and Depth,\nFrom what I understand, Depth sets how many levels are permitted in the tree, so if I set 10 leaves in 2 levels, the tree will only expand horizontally to allow up to 10 leaves across up to  2 levels.\nbut if i set 10 leaves with 10 Depth, it allows the algorithms to grow a tree 10 levels deep and grow vertically.\nAm I getting this wrong?\n\nPlus - my real life example of 8 leaves and 3 depths VS 200 leaves and 60 depth having the same AUC (along with about 10 more examples somewhere in the middle), seems a bit odd =/\n\nCould you maybe share the parameters you use in Bayesian Optimization? (if you use it)\nI found absolutely no documentation or articles about the Kappa, Epsilon and Acq parameters of the algorithm.\n\nThanks again for your response =)",
    "318318": "&gt;Can you elaborate what you meant?\n\nSure thing. Let's forget about time series for the moment and think about a 5-fold cross-validation schema. It goes something like this:\n\nFirst, break up the training data into 5 bins:\n\n[--20%--][--20%--][--20%--][--20%--][--20%--]\n\nThen train your model on bins 1-4 and validate on *bin_5*.\n\nAnd repeat: train the model on bins excluding *bin_i* and validate on *bin_i*.\n\nFinally, average your validation scores.\n\nOk. That's how cross-validation works normally. Two problems with that when you're dealing with time series:\n\n 1. You may not want to use information from the future (could cause leakage depending on the scenario).\n 2. Prediction accuracy falls off exponentially with respect to time for time series forecasts (i.e. you may predict 4 steps ahead well, but not 24 steps ahead).\n\nIn our current problem, we don't have to worry about such eventualities though. Or do we? It depends on what your model is doing. If it is using features that encapsulates click rate or time duration between clicks, it might be a factor. If you're using daily aggregates, but only considering an hour, that could cause problems. Etc., etc.\n\nTL-DR; There is more to consider when dealing with time related data.",
    "318319": "Your welcome! \nWhat you say about the depth is generally correct, but in the extrem cases with very few leaves the trees can not reach the maximum depth. On each level it will be at least 2 leaves (when expending vertically), which means that with 10 leaves the maximum depth that can be reached is 5. \nThat very shallow trees and very deep trees give the same result really seems odd. Might be connected to the features that you use.\n\nI didnt start with parameter optimization yet, still trying to find better features but will share it with you if i find good settings!",
    "318357": "Thanks a lot for your reply,\nI understand what you mean now about the leaves,\nLet me know if you find anything useful,\nIt's really odd that there are 0(!!) materials about RBayesianOptimization parameters out there, no examples, no instructions.\nEven the documentation itself is not very helpful.\nI'll probably revert to random search soon.\n\nThanks again!",
    "318360": "Thanks for your reply,\nI agree with you but not sure it's related to this specific case, as we have only 3-4 days of data and\nI'm using entire day 9 for validation, not mixing anything between the hours\\days.\n\nThanks"
  },
  "source": "meta"
}