{
  "id": 94466,
  "title": "Top 10 - Solution - Giba and Amjad",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94466",
  "author_name": "Giba",
  "post_date": "2019-06-04T17:35:10.707000",
  "votes": 53,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I joined later this competition, but I found its interesting in the end. Fortunately my team mate had build hundreds of good features based on wavelet decomposition that worked very well, but I will let Amjad describe it in a separate post <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94467#latest-543680\">here</a> and feature engineering <a href=\"https://www.kaggle.com/amjad85/10th-place-feature-engineering?scriptVersionId=15179664\">here</a> . Thanks <a href=\"/amjad85\">@amjad85</a> !</p>\n\n<p>Our Top10 solution is basically a blend of 2 LGB models trained from 2 very similar datasets (build from different wavelet transformations). The trick we used is to adapt the mean of the predictions to match the mean of the testset as described in the p4677 paper. Using the paper we measured and calculated the mean to be around 6.51 and used it in our final submission. After the end we realized that the mean is a bit less than that.</p>\n\n<p>The equation used to post- process the top-10 solution was:\nypred = ypred * 6.51 / mean(ypred)</p>\n\n<p>Also we normalized all time_to_failure target of all quakes to 1.  It improved our CV score by 0.01+ but it showed the opposite behavior in Private LB. Using the original target scored 0.01 better in Private than normalized target.</p>\n\n<p>Also I found that LGB parameter must be very conservative and you must add a lot of regularization to work well. I finished using: </p>\n\n<p><code>objective=\"gamma\",\nmax_depth = -1,\nfeature_fraction = 0.025,\nbagging_fraction = 0.250,\nbagging_freq     = 1,\nnum_leaves = 7,\nmin_data_in_bin = 2,\nmax_bin = 25,\nmin_data_in_leaf = 4,\nlambda_l1 = 1.1,\nlambda_l2 = 0.1,\nlearning_rate=0.01</code></p>\n\n<p>To find best features and LGB parameters we used a nested based CV schema. \nWe split trainset in two random parts: a train with 10 quakes and a validation with 7 quakes.\nWe perform that split 30 times choosing 30 different pairs of train and valid sets.\nSo for each run we perform Kfold CV in the trainset and use early stop to find the best iteration only using train folds like lgb.cv function. Then using the models trained on train we predict on the valid set and apply the post processing knowing a-priori the validset mean (similar of what we have for testset thought p4677 paper).\nThen we average the 30 runs results. \nIts relatively stable and provided us a good way to evaluate the parameter and features. \nBased on that and other experiments we decided to drop from the datasets all the statistical features calculated before.</p>\n\n<p>Some experiment that didn't worked:\n- DeepLearning: our best DL models scored around 1.498\n- Reverse engineering testset segments order using DL.\n- Use lgb weights to try to match testset mean.</p>\n\n<p><a href=\"https://www.kaggle.com/titericz/top-1-lb-2-252-private\">This</a> is a kernel showing a single model performing top 1 with the post-processing trick</p>\n\n<p>Post mortem 1:  The validation technique we performed is very good to validate our parameters, features and post-processing approach. But once we finish doing that, its usual to refit the model using all data available in trainset. I just forgot to do that step and submitted our final solution using models trained on 60% of the trainset. Performing the full trainset fit we would have placed #2 :_(</p>\n\n<p>Post mortem 2:  Using the mean test of 6.3 would have placed us in the money :_(</p>\n\n<p>Post mortem 3:  Using the mean test of 6.3 and training on 100% of the trainset will bring a single model winner <a href=\"https://www.kaggle.com/titericz/top-1-lb-2-252-private\">here</a></p>",
  "messages": [
    {
      "id": 543679,
      "postDate": "2019-06-04T17:35:10.707Z",
      "content": "<p>I joined later this competition, but I found its interesting in the end. Fortunately my team mate had build hundreds of good features based on wavelet decomposition that worked very well, but I will let Amjad describe it in a separate post <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94467#latest-543680\">here</a> and feature engineering <a href=\"https://www.kaggle.com/amjad85/10th-place-feature-engineering?scriptVersionId=15179664\">here</a> . Thanks <a href=\"/amjad85\">@amjad85</a> !</p>\n\n<p>Our Top10 solution is basically a blend of 2 LGB models trained from 2 very similar datasets (build from different wavelet transformations). The trick we used is to adapt the mean of the predictions to match the mean of the testset as described in the p4677 paper. Using the paper we measured and calculated the mean to be around 6.51 and used it in our final submission. After the end we realized that the mean is a bit less than that.</p>\n\n<p>The equation used to post- process the top-10 solution was:\nypred = ypred * 6.51 / mean(ypred)</p>\n\n<p>Also we normalized all time_to_failure target of all quakes to 1.  It improved our CV score by 0.01+ but it showed the opposite behavior in Private LB. Using the original target scored 0.01 better in Private than normalized target.</p>\n\n<p>Also I found that LGB parameter must be very conservative and you must add a lot of regularization to work well. I finished using: </p>\n\n<p><code>objective=\"gamma\",\nmax_depth = -1,\nfeature_fraction = 0.025,\nbagging_fraction = 0.250,\nbagging_freq     = 1,\nnum_leaves = 7,\nmin_data_in_bin = 2,\nmax_bin = 25,\nmin_data_in_leaf = 4,\nlambda_l1 = 1.1,\nlambda_l2 = 0.1,\nlearning_rate=0.01</code></p>\n\n<p>To find best features and LGB parameters we used a nested based CV schema. \nWe split trainset in two random parts: a train with 10 quakes and a validation with 7 quakes.\nWe perform that split 30 times choosing 30 different pairs of train and valid sets.\nSo for each run we perform Kfold CV in the trainset and use early stop to find the best iteration only using train folds like lgb.cv function. Then using the models trained on train we predict on the valid set and apply the post processing knowing a-priori the validset mean (similar of what we have for testset thought p4677 paper).\nThen we average the 30 runs results. \nIts relatively stable and provided us a good way to evaluate the parameter and features. \nBased on that and other experiments we decided to drop from the datasets all the statistical features calculated before.</p>\n\n<p>Some experiment that didn't worked:\n- DeepLearning: our best DL models scored around 1.498\n- Reverse engineering testset segments order using DL.\n- Use lgb weights to try to match testset mean.</p>\n\n<p><a href=\"https://www.kaggle.com/titericz/top-1-lb-2-252-private\">This</a> is a kernel showing a single model performing top 1 with the post-processing trick</p>\n\n<p>Post mortem 1:  The validation technique we performed is very good to validate our parameters, features and post-processing approach. But once we finish doing that, its usual to refit the model using all data available in trainset. I just forgot to do that step and submitted our final solution using models trained on 60% of the trainset. Performing the full trainset fit we would have placed #2 :_(</p>\n\n<p>Post mortem 2:  Using the mean test of 6.3 would have placed us in the money :_(</p>\n\n<p>Post mortem 3:  Using the mean test of 6.3 and training on 100% of the trainset will bring a single model winner <a href=\"https://www.kaggle.com/titericz/top-1-lb-2-252-private\">here</a></p>",
      "rawMarkdown": "I joined later this competition, but I found its interesting in the end. Fortunately my team mate had build hundreds of good features based on wavelet decomposition that worked very well, but I will let Amjad describe it in a separate post [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94467#latest-543680) and feature engineering [here](https://www.kaggle.com/amjad85/10th-place-feature-engineering?scriptVersionId=15179664) . Thanks @amjad85 !\n\nOur Top10 solution is basically a blend of 2 LGB models trained from 2 very similar datasets (build from different wavelet transformations). The trick we used is to adapt the mean of the predictions to match the mean of the testset as described in the p4677 paper. Using the paper we measured and calculated the mean to be around 6.51 and used it in our final submission. After the end we realized that the mean is a bit less than that.\n\nThe equation used to post- process the top-10 solution was:\nypred = ypred * 6.51 / mean(ypred)\n\nAlso we normalized all time_to_failure target of all quakes to 1.  It improved our CV score by 0.01+ but it showed the opposite behavior in Private LB. Using the original target scored 0.01 better in Private than normalized target.\n\nAlso I found that LGB parameter must be very conservative and you must add a lot of regularization to work well. I finished using: \n\n`objective=\"gamma\",\nmax_depth = -1,\nfeature_fraction = 0.025,\nbagging_fraction = 0.250,\nbagging_freq     = 1,\nnum_leaves = 7,\nmin_data_in_bin = 2,\nmax_bin = 25,\nmin_data_in_leaf = 4,\nlambda_l1 = 1.1,\nlambda_l2 = 0.1,\nlearning_rate=0.01`\n\nTo find best features and LGB parameters we used a nested based CV schema. \nWe split trainset in two random parts: a train with 10 quakes and a validation with 7 quakes.\nWe perform that split 30 times choosing 30 different pairs of train and valid sets.\nSo for each run we perform Kfold CV in the trainset and use early stop to find the best iteration only using train folds like lgb.cv function. Then using the models trained on train we predict on the valid set and apply the post processing knowing a-priori the validset mean (similar of what we have for testset thought p4677 paper).\nThen we average the 30 runs results. \nIts relatively stable and provided us a good way to evaluate the parameter and features. \nBased on that and other experiments we decided to drop from the datasets all the statistical features calculated before.\n\nSome experiment that didn't worked:\n- DeepLearning: our best DL models scored around 1.498\n- Reverse engineering testset segments order using DL.\n- Use lgb weights to try to match testset mean.\n\n[This](https://www.kaggle.com/titericz/top-1-lb-2-252-private) is a kernel showing a single model performing top 1 with the post-processing trick\n\nPost mortem 1:  The validation technique we performed is very good to validate our parameters, features and post-processing approach. But once we finish doing that, its usual to refit the model using all data available in trainset. I just forgot to do that step and submitted our final solution using models trained on 60% of the trainset. Performing the full trainset fit we would have placed #2 :_(\n\nPost mortem 2:  Using the mean test of 6.3 would have placed us in the money :_(\n\nPost mortem 3:  Using the mean test of 6.3 and training on 100% of the trainset will bring a single model winner [here](https://www.kaggle.com/titericz/top-1-lb-2-252-private)",
      "votes": 52
    },
    {
      "id": 546568,
      "postDate": "2019-06-06T17:33:23.370Z",
      "content": "<p>Congratulations with the result and many thanks for the wonderful write-up, this is pure gold! It would be great if you could explain some details of your approach.</p>\n\n<p>Is there a particular reason why you used early stopping in the inner loop of your CV strategy as opposed to sweeping over a fixed number of boosting rounds? I can see how it could be slightly better at generalizing at the cost of approximately K times more computation. Also, I can see how there could be interactions with other hyperparameters but it would be great if you could share your view. What K did you use in the inner loop?</p>\n\n<p>How did you search for optimal hyperparameters? Given the costly inner loop (30*K model fits) I imagine you had to be smart about what settings to try.</p>",
      "rawMarkdown": "Congratulations with the result and many thanks for the wonderful write-up, this is pure gold! It would be great if you could explain some details of your approach.\n\nIs there a particular reason why you used early stopping in the inner loop of your CV strategy as opposed to sweeping over a fixed number of boosting rounds? I can see how it could be slightly better at generalizing at the cost of approximately K times more computation. Also, I can see how there could be interactions with other hyperparameters but it would be great if you could share your view. What K did you use in the inner loop?\n\nHow did you search for optimal hyperparameters? Given the costly inner loop (30*K model fits) I imagine you had to be smart about what settings to try.",
      "votes": 3
    },
    {
      "id": 543697,
      "postDate": "2019-06-04T17:52:14.420Z",
      "content": "<p>Thanks for sharing, and congrats to both of you.  Solid approach able to resist shakeup.</p>",
      "rawMarkdown": "Thanks for sharing, and congrats to both of you.  Solid approach able to resist shakeup.",
      "votes": 3,
      "replies": [
        {
          "id": 543699,
          "postDate": "2019-06-04T17:54:09.700Z",
          "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> . I believe most of the teams that realized that train and Private testset are different and worked to fix it resisted to the shakeup.</p>",
          "rawMarkdown": "Thanks @cpmpml . I believe most of the teams that realized that train and Private testset are different and worked to fix it resisted to the shakeup.",
          "votes": 3
        }
      ]
    },
    {
      "id": 544776,
      "postDate": "2019-06-05T22:48:43.073Z",
      "content": "<p>Awesome work! Congratulations on double GM as well. </p>",
      "rawMarkdown": "Awesome work! Congratulations on double GM as well. ",
      "votes": 1
    },
    {
      "id": 544034,
      "postDate": "2019-06-05T03:59:19.783Z",
      "content": "<p>Congratulations Giba! Thanks for sharing :) </p>",
      "rawMarkdown": "Congratulations Giba! Thanks for sharing :) ",
      "votes": 1
    },
    {
      "id": 543898,
      "postDate": "2019-06-04T23:25:35.270Z",
      "content": "<p>Thanks ! I like your post processing idea ! I wish I thought about that a few days ago :)</p>",
      "rawMarkdown": "Thanks ! I like your post processing idea ! I wish I thought about that a few days ago :)",
      "votes": 1
    },
    {
      "id": 543739,
      "postDate": "2019-06-04T18:40:13.367Z",
      "content": "<p>Congratulations to both for such a great result! Giba's magic + high quality feat. engineering by Amjad = solid bet. Thanks for sharing! :-)</p>",
      "rawMarkdown": "Congratulations to both for such a great result! Giba's magic + high quality feat. engineering by Amjad = solid bet. Thanks for sharing! :-)",
      "votes": 2,
      "replies": [
        {
          "id": 543786,
          "postDate": "2019-06-04T20:22:55.150Z",
          "content": "<p>Thanks <a href=\"/miguelpm\">@miguelpm</a> . The features built by <a href=\"/amjad85\">@amjad85</a> are great. </p>",
          "rawMarkdown": "Thanks @miguelpm . The features built by @amjad85 are great. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 543723,
      "postDate": "2019-06-04T18:11:37.767Z",
      "content": "<p>Congrats to both of you! \nI was not surprised to see you shook <em>up</em> and not down ;)</p>",
      "rawMarkdown": "Congrats to both of you! \nI was not surprised to see you shook *up* and not down ;)",
      "votes": 2,
      "replies": [
        {
          "id": 543729,
          "postDate": "2019-06-04T18:19:10.707Z",
          "content": "<p>Thank you <a href=\"/stecasasso\">@stecasasso</a> and congrats also for survive the shakeup.</p>",
          "rawMarkdown": "Thank you @stecasasso and congrats also for survive the shakeup."
        }
      ]
    },
    {
      "id": 544113,
      "postDate": "2019-06-05T06:17:34.157Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    },
    {
      "id": 552715,
      "postDate": "2019-06-14T11:50:59.920Z",
      "content": "<p>Thanks for sharing! </p>",
      "rawMarkdown": "Thanks for sharing! ",
      "votes": 1
    },
    {
      "id": 550169,
      "postDate": "2019-06-11T11:27:25.833Z",
      "content": "<p>Thank you for sharing <a href=\"/titericz\">@titericz</a> </p>",
      "rawMarkdown": "Thank you for sharing @titericz ",
      "votes": 1
    },
    {
      "id": 546962,
      "postDate": "2019-06-07T04:18:22.893Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 544290,
      "postDate": "2019-06-05T11:06:20.900Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 543716,
      "postDate": "2019-06-04T18:06:00.613Z",
      "content": "<p>Thank you for the great collaboration. </p>",
      "rawMarkdown": "Thank you for the great collaboration. ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 546568,
      "author_name": "Tom Van de Wiele",
      "author_url": "",
      "post_date": "2019-06-06T17:33:23.370000",
      "content": "<p>Congratulations with the result and many thanks for the wonderful write-up, this is pure gold! It would be great if you could explain some details of your approach.</p>\n\n<p>Is there a particular reason why you used early stopping in the inner loop of your CV strategy as opposed to sweeping over a fixed number of boosting rounds? I can see how it could be slightly better at generalizing at the cost of approximately K times more computation. Also, I can see how there could be interactions with other hyperparameters but it would be great if you could share your view. What K did you use in the inner loop?</p>\n\n<p>How did you search for optimal hyperparameters? Given the costly inner loop (30*K model fits) I imagine you had to be smart about what settings to try.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 543697,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-06-04T17:52:14.420000",
      "content": "<p>Thanks for sharing, and congrats to both of you.  Solid approach able to resist shakeup.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 543699,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-06-04T17:54:09.700000",
          "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> . I believe most of the teams that realized that train and Private testset are different and worked to fix it resisted to the shakeup.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 544776,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-06-05T22:48:43.073000",
      "content": "<p>Awesome work! Congratulations on double GM as well. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 544034,
      "author_name": "Prashanth Thangavel",
      "author_url": "",
      "post_date": "2019-06-05T03:59:19.783000",
      "content": "<p>Congratulations Giba! Thanks for sharing :) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543898,
      "author_name": "Antoine",
      "author_url": "",
      "post_date": "2019-06-04T23:25:35.270000",
      "content": "<p>Thanks ! I like your post processing idea ! I wish I thought about that a few days ago :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543739,
      "author_name": "miguel perez",
      "author_url": "",
      "post_date": "2019-06-04T18:40:13.367000",
      "content": "<p>Congratulations to both for such a great result! Giba's magic + high quality feat. engineering by Amjad = solid bet. Thanks for sharing! :-)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 543786,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-06-04T20:22:55.150000",
          "content": "<p>Thanks <a href=\"/miguelpm\">@miguelpm</a> . The features built by <a href=\"/amjad85\">@amjad85</a> are great. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543723,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2019-06-04T18:11:37.767000",
      "content": "<p>Congrats to both of you! \nI was not surprised to see you shook <em>up</em> and not down ;)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 543729,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-06-04T18:19:10.707000",
          "content": "<p>Thank you <a href=\"/stecasasso\">@stecasasso</a> and congrats also for survive the shakeup.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 544113,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-05T06:17:34.157000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 552715,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-14T11:50:59.920000",
      "content": "<p>Thanks for sharing! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 550169,
      "author_name": "Joel Hanson",
      "author_url": "",
      "post_date": "2019-06-11T11:27:25.833000",
      "content": "<p>Thank you for sharing <a href=\"/titericz\">@titericz</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 546962,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-06-07T04:18:22.893000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 544290,
      "author_name": "Timmmmmms",
      "author_url": "",
      "post_date": "2019-06-05T11:06:20.900000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543716,
      "author_name": "Amjad",
      "author_url": "",
      "post_date": "2019-06-04T18:06:00.613000",
      "content": "<p>Thank you for the great collaboration. </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543679": "I joined later this competition, but I found its interesting in the end. Fortunately my team mate had build hundreds of good features based on wavelet decomposition that worked very well, but I will let Amjad describe it in a separate post [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94467#latest-543680) and feature engineering [here](https://www.kaggle.com/amjad85/10th-place-feature-engineering?scriptVersionId=15179664) . Thanks @amjad85 !\n\nOur Top10 solution is basically a blend of 2 LGB models trained from 2 very similar datasets (build from different wavelet transformations). The trick we used is to adapt the mean of the predictions to match the mean of the testset as described in the p4677 paper. Using the paper we measured and calculated the mean to be around 6.51 and used it in our final submission. After the end we realized that the mean is a bit less than that.\n\nThe equation used to post- process the top-10 solution was:\nypred = ypred * 6.51 / mean(ypred)\n\nAlso we normalized all time_to_failure target of all quakes to 1.  It improved our CV score by 0.01+ but it showed the opposite behavior in Private LB. Using the original target scored 0.01 better in Private than normalized target.\n\nAlso I found that LGB parameter must be very conservative and you must add a lot of regularization to work well. I finished using: \n\n`objective=\"gamma\",\nmax_depth = -1,\nfeature_fraction = 0.025,\nbagging_fraction = 0.250,\nbagging_freq     = 1,\nnum_leaves = 7,\nmin_data_in_bin = 2,\nmax_bin = 25,\nmin_data_in_leaf = 4,\nlambda_l1 = 1.1,\nlambda_l2 = 0.1,\nlearning_rate=0.01`\n\nTo find best features and LGB parameters we used a nested based CV schema. \nWe split trainset in two random parts: a train with 10 quakes and a validation with 7 quakes.\nWe perform that split 30 times choosing 30 different pairs of train and valid sets.\nSo for each run we perform Kfold CV in the trainset and use early stop to find the best iteration only using train folds like lgb.cv function. Then using the models trained on train we predict on the valid set and apply the post processing knowing a-priori the validset mean (similar of what we have for testset thought p4677 paper).\nThen we average the 30 runs results. \nIts relatively stable and provided us a good way to evaluate the parameter and features. \nBased on that and other experiments we decided to drop from the datasets all the statistical features calculated before.\n\nSome experiment that didn't worked:\n- DeepLearning: our best DL models scored around 1.498\n- Reverse engineering testset segments order using DL.\n- Use lgb weights to try to match testset mean.\n\n[This](https://www.kaggle.com/titericz/top-1-lb-2-252-private) is a kernel showing a single model performing top 1 with the post-processing trick\n\nPost mortem 1:  The validation technique we performed is very good to validate our parameters, features and post-processing approach. But once we finish doing that, its usual to refit the model using all data available in trainset. I just forgot to do that step and submitted our final solution using models trained on 60% of the trainset. Performing the full trainset fit we would have placed #2 :_(\n\nPost mortem 2:  Using the mean test of 6.3 would have placed us in the money :_(\n\nPost mortem 3:  Using the mean test of 6.3 and training on 100% of the trainset will bring a single model winner [here](https://www.kaggle.com/titericz/top-1-lb-2-252-private)",
    "546568": "Congratulations with the result and many thanks for the wonderful write-up, this is pure gold! It would be great if you could explain some details of your approach.\n\nIs there a particular reason why you used early stopping in the inner loop of your CV strategy as opposed to sweeping over a fixed number of boosting rounds? I can see how it could be slightly better at generalizing at the cost of approximately K times more computation. Also, I can see how there could be interactions with other hyperparameters but it would be great if you could share your view. What K did you use in the inner loop?\n\nHow did you search for optimal hyperparameters? Given the costly inner loop (30*K model fits) I imagine you had to be smart about what settings to try.",
    "543697": "Thanks for sharing, and congrats to both of you.  Solid approach able to resist shakeup.",
    "544776": "Awesome work! Congratulations on double GM as well. ",
    "544034": "Congratulations Giba! Thanks for sharing :) ",
    "543898": "Thanks ! I like your post processing idea ! I wish I thought about that a few days ago :)",
    "543739": "Congratulations to both for such a great result! Giba's magic + high quality feat. engineering by Amjad = solid bet. Thanks for sharing! :-)",
    "543723": "Congrats to both of you! \nI was not surprised to see you shook *up* and not down ;)",
    "544113": "",
    "552715": "Thanks for sharing! ",
    "550169": "Thank you for sharing @titericz ",
    "546962": "Thanks for sharing.",
    "544290": "Thanks for sharing",
    "543716": "Thank you for the great collaboration. "
  }
}