{
  "id": 94324,
  "title": "Congrats, and my 5 min lottery tickets (81th place/silver medal)",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/shize-su-congrats-and-my-5-min-lottery-tickets-81t",
  "author_name": "",
  "post_date": "2019-06-04T03:59:47.523Z",
  "votes": 59,
  "comment_count": 22,
  "views": 0,
  "content": "<p>First, congrats to all the winners, especially Zoo team and my previous great teammate Danijel! All of you have done an amazing job!</p>\n\n<p>This competition is such an unstable competition that I previously didn't plan to spend any time on it. About 2 hour before the deadline, I got a bit of free time and decided to buy 2 lottery tickets for this competition and spent 5 mins to make 2 submissions. </p>\n\n<p>Basically, what I did is just: 1) take  the  old public LB1.456 benchmark solution (shared 2 months ago in the kernel's section <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\">https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction</a>, credits to Gabriel); 2) multiply the prediction by 1.08 and 0.96 respectively, and made these 2 submissions. That's it. </p>\n\n<p>As a result, the 1.08 adjusted submission gave private LB score 2.45 and 81th place.</p>\n\n<p>It seems that the scale/magnitude of prediction in the private LB part (compared with train data and public LB data) might probably be the largest uncertainty in this competition (in terms of public/private LB ranking), and applying a scale-up and scale-down factors adjustment might increase the chance to finish in the very top of the LB (though also increase the risk of finishing in the very bottom of the LB when factor adjustment is too aggressive)</p>\n\n<p>Btw, I also found that if we simply multiply 1.1 with the PublicLB 1.50 catboost benchmark solution (shared one month ago in kernel <a href=\"https://www.kaggle.com/hsinwenchang/mfcc-randomforestregressor-catboostregressor/output\">https://www.kaggle.com/hsinwenchang/mfcc-randomforestregressor-catboostregressor/output</a>), it will score 2.336 on the private LB (11th place) and gave a gold medal. Without any adjustment, the benchmark solution gave 2.45 private LB score.</p>\n\n<p>I guess that most of top 15 competitors might possibly win (or close to) top 1 place prize with some simple factor adjustment.</p>\n\n<p>Finally, I am in particular very interested in the top 1 solution, who seems to have a good winning margin on the private LB compared with all other teams. Good job!</p>",
  "messages": [
    {
      "id": "542567",
      "postDate": "06/04/2019 01:37:51",
      "content": "<p>First, congrats to all the winners, especially Zoo team and my previous great teammate Danijel! All of you have done an amazing job!</p>\n\n<p>This competition is such an unstable competition that I previously didn't plan to spend any time on it. About 2 hour before the deadline, I got a bit of free time and decided to buy 2 lottery tickets for this competition and spent 5 mins to make 2 submissions. </p>\n\n<p>Basically, what I did is just: 1) take  the  old public LB1.456 benchmark solution (shared 2 months ago in the kernel's section <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\">https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction</a>, credits to Gabriel); 2) multiply the prediction by 1.08 and 0.96 respectively, and made these 2 submissions. That's it. </p>\n\n<p>As a result, the 1.08 adjusted submission gave private LB score 2.45 and 81th place.</p>\n\n<p>It seems that the scale/magnitude of prediction in the private LB part (compared with train data and public LB data) might probably be the largest uncertainty in this competition (in terms of public/private LB ranking), and applying a scale-up and scale-down factors adjustment might increase the chance to finish in the very top of the LB (though also increase the risk of finishing in the very bottom of the LB when factor adjustment is too aggressive)</p>\n\n<p>Btw, I also found that if we simply multiply 1.1 with the PublicLB 1.50 catboost benchmark solution (shared one month ago in kernel <a href=\"https://www.kaggle.com/hsinwenchang/mfcc-randomforestregressor-catboostregressor/output\">https://www.kaggle.com/hsinwenchang/mfcc-randomforestregressor-catboostregressor/output</a>), it will score 2.336 on the private LB (11th place) and gave a gold medal. Without any adjustment, the benchmark solution gave 2.45 private LB score.</p>\n\n<p>I guess that most of top 15 competitors might possibly win (or close to) top 1 place prize with some simple factor adjustment.</p>\n\n<p>Finally, I am in particular very interested in the top 1 solution, who seems to have a good winning margin on the private LB compared with all other teams. Good job!</p>",
      "rawMarkdown": "First, congrats to all the winners, especially Zoo team and my previous great teammate Danijel! All of you have done an amazing job!\n\nThis competition is such an unstable competition that I previously didn't plan to spend any time on it. About 2 hour before the deadline, I got a bit of free time and decided to buy 2 lottery tickets for this competition and spent 5 mins to make 2 submissions. \n\nBasically, what I did is just: 1) take  the  old public LB1.456 benchmark solution (shared 2 months ago in the kernel's section https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction, credits to Gabriel); 2) multiply the prediction by 1.08 and 0.96 respectively, and made these 2 submissions. That's it. \n\nAs a result, the 1.08 adjusted submission gave private LB score 2.45 and 81th place.\n\nIt seems that the scale/magnitude of prediction in the private LB part (compared with train data and public LB data) might probably be the largest uncertainty in this competition (in terms of public/private LB ranking), and applying a scale-up and scale-down factors adjustment might increase the chance to finish in the very top of the LB (though also increase the risk of finishing in the very bottom of the LB when factor adjustment is too aggressive)\n\nBtw, I also found that if we simply multiply 1.1 with the PublicLB 1.50 catboost benchmark solution (shared one month ago in kernel https://www.kaggle.com/hsinwenchang/mfcc-randomforestregressor-catboostregressor/output), it will score 2.336 on the private LB (11th place) and gave a gold medal. Without any adjustment, the benchmark solution gave 2.45 private LB score.\n\nI guess that most of top 15 competitors might possibly win (or close to) top 1 place prize with some simple factor adjustment.\n\nFinally, I am in particular very interested in the top 1 solution, who seems to have a good winning margin on the private LB compared with all other teams. Good job!",
      "votes": null
    },
    {
      "id": "542590",
      "postDate": "06/04/2019 01:52:00",
      "content": "<p>Thanks for sharing. What would be the intuition of scaling up or down? Higher TTF for scale up and lower TTF for scale down?</p>",
      "rawMarkdown": "Thanks for sharing. What would be the intuition of scaling up or down? Higher TTF for scale up and lower TTF for scale down?",
      "votes": null
    },
    {
      "id": "542593",
      "postDate": "06/04/2019 01:53:32",
      "content": "<p>It happens because Public LB mean is around 4 and Private is around 6.7. So multiplying by 1.x or adding a constant to your predictions improves your Private LB. </p>\n\n<p>By the way, congrats <a href=\"/sushize\">@sushize</a> . This is the fastest silver medal I've ever seen ;)</p>",
      "rawMarkdown": "It happens because Public LB mean is around 4 and Private is around 6.7. So multiplying by 1.x or adding a constant to your predictions improves your Private LB. \n\nBy the way, congrats @sushize . This is the fastest silver medal I've ever seen ;)",
      "votes": null
    },
    {
      "id": "542597",
      "postDate": "06/04/2019 01:55:37",
      "content": "<p>Thanks,really amazing!</p>",
      "rawMarkdown": "Thanks,really amazing!",
      "votes": null
    },
    {
      "id": "542602",
      "postDate": "06/04/2019 02:00:19",
      "content": "<p>That's amazing. But how did you decide on 1.08 and 0.96 ? \nI just tried increasing the factor 1.085 and it's still 2.45430</p>",
      "rawMarkdown": "That's amazing. But how did you decide on 1.08 and 0.96 ? \nI just tried increasing the factor 1.085 and it's still 2.45430",
      "votes": null
    },
    {
      "id": "542616",
      "postDate": "06/04/2019 02:11:21",
      "content": "<p>Based on previous forum sharing, I just thought that it is more likely that the private LB might have higher(rather than lower) mean than public LB, and I decided to apply a bit more aggressive adjustment when scaling up and less aggressive adjustment when scaling down. The specific values 1.08 and 0.96 are just two random numbers from my intuition (though did some simple math calculation to make sure the adjustment would not be out of the range of \"making sense\", and also double check it by looking at the public LB score). For example, an adjustment factor of 8.0 would definitely blow up the score and won't make any sense (which could also be verified by looking at the public LB).</p>\n\n<p>Also, I made the 1.08 adjustment submission first, which gave me 1.68 public LB score, which I think it is in the range of \"makes sense\". If I found that submission to have too bad public LB score (e.g., 2.0+ public LB score), then for my 2nd submission I probably would apply a less aggressive 1.x+ adjustment (e.g., 1.02) rather than 0.96 adjustment, since I do think that there is a higher chance that private LB would have a higher mean (based on forum discussions). But given that my first submission looks reasonable, I made the 2nd scale-down submission with factor 0.96 as planned.</p>",
      "rawMarkdown": "Based on previous forum sharing, I just thought that it is more likely that the private LB might have higher(rather than lower) mean than public LB, and I decided to apply a bit more aggressive adjustment when scaling up and less aggressive adjustment when scaling down. The specific values 1.08 and 0.96 are just two random numbers from my intuition (though did some simple math calculation to make sure the adjustment would not be out of the range of \"making sense\", and also double check it by looking at the public LB score). For example, an adjustment factor of 8.0 would definitely blow up the score and won't make any sense (which could also be verified by looking at the public LB).\n\nAlso, I made the 1.08 adjustment submission first, which gave me 1.68 public LB score, which I think it is in the range of \"makes sense\". If I found that submission to have too bad public LB score (e.g., 2.0+ public LB score), then for my 2nd submission I probably would apply a less aggressive 1.x+ adjustment (e.g., 1.02) rather than 0.96 adjustment, since I do think that there is a higher chance that private LB would have a higher mean (based on forum discussions). But given that my first submission looks reasonable, I made the 2nd scale-down submission with factor 0.96 as planned.",
      "votes": null
    },
    {
      "id": "542705",
      "postDate": "06/04/2019 03:46:03",
      "content": "<p>Nice and clever way to try your luck</p>",
      "rawMarkdown": "Nice and clever way to try your luck",
      "votes": null
    },
    {
      "id": "542714",
      "postDate": "06/04/2019 03:55:45",
      "content": "<p>just check by submitting my best result ( private score 2.59332 ) time 1.1 and it gave me 2.35529, which is really amazing. </p>",
      "rawMarkdown": "just check by submitting my best result ( private score 2.59332 ) time 1.1 and it gave me 2.35529, which is really amazing.",
      "votes": null
    },
    {
      "id": "542758",
      "postDate": "06/04/2019 04:36:50",
      "content": "<p>I find this is interesting.  When I used mae for a separate data set for error analysis (even plotted it).  Wonder if mean error (without absolute) would turn out to be negative.  I think simply apply average target/predicted ratio would increase ranking. </p>",
      "rawMarkdown": "I find this is interesting.  When I used mae for a separate data set for error analysis (even plotted it).  Wonder if mean error (without absolute) would turn out to be negative.  I think simply apply average target/predicted ratio would increase ranking.",
      "votes": null
    },
    {
      "id": "542813",
      "postDate": "06/04/2019 05:38:40",
      "content": "<p>I was joking around a few days ago that doing exactly that might have good chances to score very high on public LB. Thanks for testing and thanks for the congratz!</p>",
      "rawMarkdown": "I was joking around a few days ago that doing exactly that might have good chances to score very high on public LB. Thanks for testing and thanks for the congratz!",
      "votes": null
    },
    {
      "id": "542822",
      "postDate": "06/04/2019 05:57:13",
      "content": "<p>ya, i tried this on our models just now and it improved our scores dramatically. i feel like almost any model would have been improved via scaling. i remember we had this discussion in our group whether the earthquakes would be longer in the test set than the training set. we weren't sure at the time. it seemed to have a big influence on the score -possibly because predicting earthquakes must be that difficult. i remember we were analyzing the earthquake errors and noticed that the model didn't seem to see the difference between longer and shorter earthquakes and seemed to make the mean prediction. We thought about scaling, but we ended up not reading the previous discussion post in detail, given time, and got bogged down doing other things. Should have looked into it. Rather than a conservative/risky approach for our submissions, we should have considered scaling our predictions, in consideration of the earthquake length.</p>",
      "rawMarkdown": "ya, i tried this on our models just now and it improved our scores dramatically. i feel like almost any model would have been improved via scaling. i remember we had this discussion in our group whether the earthquakes would be longer in the test set than the training set. we weren't sure at the time. it seemed to have a big influence on the score -possibly because predicting earthquakes must be that difficult. i remember we were analyzing the earthquake errors and noticed that the model didn't seem to see the difference between longer and shorter earthquakes and seemed to make the mean prediction. We thought about scaling, but we ended up not reading the previous discussion post in detail, given time, and got bogged down doing other things. Should have looked into it. Rather than a conservative/risky approach for our submissions, we should have considered scaling our predictions, in consideration of the earthquake length.",
      "votes": null
    },
    {
      "id": "542840",
      "postDate": "06/04/2019 06:21:41",
      "content": "<p>Thanks for the reply and seriously hats off for the sharp intuition with numbers. \nIf you don't mind another curiosity I have, why take Gabriel's kernel ? There were others with slightly better and slightly worse score at the time, and you manage to choose the one that works for the 2 submission. </p>",
      "rawMarkdown": "Thanks for the reply and seriously hats off for the sharp intuition with numbers. \nIf you don't mind another curiosity I have, why take Gabriel's kernel ? There were others with slightly better and slightly worse score at the time, and you manage to choose the one that works for the 2 submission.",
      "votes": null
    },
    {
      "id": "542877",
      "postDate": "06/04/2019 06:49:54",
      "content": "<p>what we did was eye-ball the training set and see if the earthquake lengths were increasing in time. based on our general observation, it didn't seem so (this is why we originally discarded this hypothesis). Maybe we should have checked this rigorously using a moving average. I didn't actually do this, but if this trend doesn't exist, then i guess it's a bit unusual that the earthquakes increased in length during the test set. If the earthquakes didn't increase in length in the the current test set (or for increasing time), the test set would have been an unusual sample, which is probably unlikely given the number of samples. </p>",
      "rawMarkdown": "what we did was eye-ball the training set and see if the earthquake lengths were increasing in time. based on our general observation, it didn't seem so (this is why we originally discarded this hypothesis). Maybe we should have checked this rigorously using a moving average. I didn't actually do this, but if this trend doesn't exist, then i guess it's a bit unusual that the earthquakes increased in length during the test set. If the earthquakes didn't increase in length in the the current test set (or for increasing time), the test set would have been an unusual sample, which is probably unlikely given the number of samples.",
      "votes": null
    },
    {
      "id": "542940",
      "postDate": "06/04/2019 08:07:12",
      "content": "<p>Congrats <a href=\"/sushize\">@sushize</a> for your fast and good solution. I was wondering, how did you come up with 1.08 and 0.96 in 5 minutes? :) My teammates deserve all the glory for this win. Unfortunately, in the last month of the competition I did not have the time to contribute anything at all.</p>",
      "rawMarkdown": "Congrats @sushize for your fast and good solution. I was wondering, how did you come up with 1.08 and 0.96 in 5 minutes? :) My teammates deserve all the glory for this win. Unfortunately, in the last month of the competition I did not have the time to contribute anything at all.",
      "votes": null
    },
    {
      "id": "542941",
      "postDate": "06/04/2019 08:07:26",
      "content": "<p>I tried your clever technique in one of my unselected best models  (unfortunately, it turned to be a silver medalist) and I got 2.35016 (gold zone medalist). Sometimes, cleverness (maybe I would say with a combination of luck also)beats hard work! Congrats! :-)</p>",
      "rawMarkdown": "I tried your clever technique in one of my unselected best models  (unfortunately, it turned to be a silver medalist) and I got 2.35016 (gold zone medalist). Sometimes, cleverness (maybe I would say with a combination of luck also)beats hard work! Congrats! :-)",
      "votes": null
    },
    {
      "id": "542970",
      "postDate": "06/04/2019 08:42:26",
      "content": "<p>The fastest hand on Kaggle :)</p>",
      "rawMarkdown": "The fastest hand on Kaggle :)",
      "votes": null
    },
    {
      "id": "543039",
      "postDate": "06/04/2019 09:43:48",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "rawMarkdown": "Congrats! Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "543262",
      "postDate": "06/04/2019 12:23:24",
      "content": "<p>Congrats! You are understanding the root of this competition that test time ttf difference is the most important point.\nYour idea wins the medal.</p>",
      "rawMarkdown": "Congrats! You are understanding the root of this competition that test time ttf difference is the most important point.\nYour idea wins the medal.",
      "votes": null
    },
    {
      "id": "543268",
      "postDate": "06/04/2019 12:28:20",
      "content": "<p>You beat everyone on time to value for sure!</p>",
      "rawMarkdown": "You beat everyone on time to value for sure!",
      "votes": null
    },
    {
      "id": "543338",
      "postDate": "06/04/2019 13:34:10",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "543552",
      "postDate": "06/04/2019 15:36:31",
      "content": "<p>I tried this game : multiplying my final submission by 1.19 gives me 2.296 on private LB : 2nd position!\nMean value is then 6.22. \nCan we play lottery again? :-)</p>",
      "rawMarkdown": "I tried this game : multiplying my final submission by 1.19 gives me 2.296 on private LB : 2nd position!\nMean value is then 6.22. \nCan we play lottery again? :-)",
      "votes": null
    },
    {
      "id": "543708",
      "postDate": "06/04/2019 18:01:14",
      "content": "<p>It is based on several simple thoughts:</p>\n\n<p>[1] The public LB score 1.45 is \"reasonably ok\".\n[2] The kernel solution was generated 2 months ago, at that time it focused more on the EDA part and didn't try every efforts to climb the public LB, and thus less prone to overfit to the public LB (compared with most recent kernels with slightly better public LB score)\n[3] I only have a couple of mins, and thus I just quickly take one benchmark solution that I \"like\" most at the first glance, rather than did a comprehensive search and comparison among all available \"good benchmark\". </p>\n\n<p>Btw, the same factor adjustment trick probably would work for most, if not all,  \"good public benchmarks\". And the one I selected performs reasonably well but definitely not the best/optimal one. As I also mentioned in the main post, if I selected the public LB 1.50 kernel solution shared 1 month ago and applying the same adjustment, it will gave a gold medal.</p>\n\n<p>In short, I just did sth simple that I think might have some nontrivial chance to finish very well on the private LB (by scaling up and scaling down), and the rest is luck. That's why I called my two submissions as \"lottery tickets\".</p>",
      "rawMarkdown": "It is based on several simple thoughts:\n\n[1] The public LB score 1.45 is \"reasonably ok\".\n[2] The kernel solution was generated 2 months ago, at that time it focused more on the EDA part and didn't try every efforts to climb the public LB, and thus less prone to overfit to the public LB (compared with most recent kernels with slightly better public LB score)\n[3] I only have a couple of mins, and thus I just quickly take one benchmark solution that I \"like\" most at the first glance, rather than did a comprehensive search and comparison among all available \"good benchmark\". \n\nBtw, the same factor adjustment trick probably would work for most, if not all,  \"good public benchmarks\". And the one I selected performs reasonably well but definitely not the best/optimal one. As I also mentioned in the main post, if I selected the public LB 1.50 kernel solution shared 1 month ago and applying the same adjustment, it will gave a gold medal.\n\nIn short, I just did sth simple that I think might have some nontrivial chance to finish very well on the private LB (by scaling up and scaling down), and the rest is luck. That's why I called my two submissions as \"lottery tickets\".",
      "votes": null
    },
    {
      "id": "552728",
      "postDate": "06/14/2019 12:22:11",
      "content": "<p>That's so cool lol</p>",
      "rawMarkdown": "That's so cool lol",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 542590,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "06/04/2019 01:52:00",
      "content": "<p>Thanks for sharing. What would be the intuition of scaling up or down? Higher TTF for scale up and lower TTF for scale down?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542593,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 01:53:32",
      "content": "<p>It happens because Public LB mean is around 4 and Private is around 6.7. So multiplying by 1.x or adding a constant to your predictions improves your Private LB. </p>\n\n<p>By the way, congrats <a href=\"/sushize\">@sushize</a> . This is the fastest silver medal I've ever seen ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542597,
      "author_name": "xiaohanzhi",
      "author_url": "",
      "post_date": "06/04/2019 01:55:37",
      "content": "<p>Thanks,really amazing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542602,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "06/04/2019 02:00:19",
      "content": "<p>That's amazing. But how did you decide on 1.08 and 0.96 ? \nI just tried increasing the factor 1.085 and it's still 2.45430</p>",
      "votes": null,
      "replies": [
        {
          "id": 542616,
          "author_name": "sushize",
          "author_url": "",
          "post_date": "06/04/2019 02:11:21",
          "content": "<p>Based on previous forum sharing, I just thought that it is more likely that the private LB might have higher(rather than lower) mean than public LB, and I decided to apply a bit more aggressive adjustment when scaling up and less aggressive adjustment when scaling down. The specific values 1.08 and 0.96 are just two random numbers from my intuition (though did some simple math calculation to make sure the adjustment would not be out of the range of \"making sense\", and also double check it by looking at the public LB score). For example, an adjustment factor of 8.0 would definitely blow up the score and won't make any sense (which could also be verified by looking at the public LB).</p>\n\n<p>Also, I made the 1.08 adjustment submission first, which gave me 1.68 public LB score, which I think it is in the range of \"makes sense\". If I found that submission to have too bad public LB score (e.g., 2.0+ public LB score), then for my 2nd submission I probably would apply a less aggressive 1.x+ adjustment (e.g., 1.02) rather than 0.96 adjustment, since I do think that there is a higher chance that private LB would have a higher mean (based on forum discussions). But given that my first submission looks reasonable, I made the 2nd scale-down submission with factor 0.96 as planned.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 542840,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "06/04/2019 06:21:41",
          "content": "<p>Thanks for the reply and seriously hats off for the sharp intuition with numbers. \nIf you don't mind another curiosity I have, why take Gabriel's kernel ? There were others with slightly better and slightly worse score at the time, and you manage to choose the one that works for the 2 submission. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543708,
          "author_name": "sushize",
          "author_url": "",
          "post_date": "06/04/2019 18:01:14",
          "content": "<p>It is based on several simple thoughts:</p>\n\n<p>[1] The public LB score 1.45 is \"reasonably ok\".\n[2] The kernel solution was generated 2 months ago, at that time it focused more on the EDA part and didn't try every efforts to climb the public LB, and thus less prone to overfit to the public LB (compared with most recent kernels with slightly better public LB score)\n[3] I only have a couple of mins, and thus I just quickly take one benchmark solution that I \"like\" most at the first glance, rather than did a comprehensive search and comparison among all available \"good benchmark\". </p>\n\n<p>Btw, the same factor adjustment trick probably would work for most, if not all,  \"good public benchmarks\". And the one I selected performs reasonably well but definitely not the best/optimal one. As I also mentioned in the main post, if I selected the public LB 1.50 kernel solution shared 1 month ago and applying the same adjustment, it will gave a gold medal.</p>\n\n<p>In short, I just did sth simple that I think might have some nontrivial chance to finish very well on the private LB (by scaling up and scaling down), and the rest is luck. That's why I called my two submissions as \"lottery tickets\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 542705,
      "author_name": "johnkyvetos",
      "author_url": "",
      "post_date": "06/04/2019 03:46:03",
      "content": "<p>Nice and clever way to try your luck</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542714,
      "author_name": "zhixuanliu",
      "author_url": "",
      "post_date": "06/04/2019 03:55:45",
      "content": "<p>just check by submitting my best result ( private score 2.59332 ) time 1.1 and it gave me 2.35529, which is really amazing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542758,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "06/04/2019 04:36:50",
      "content": "<p>I find this is interesting.  When I used mae for a separate data set for error analysis (even plotted it).  Wonder if mean error (without absolute) would turn out to be negative.  I think simply apply average target/predicted ratio would increase ranking. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542813,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "06/04/2019 05:38:40",
      "content": "<p>I was joking around a few days ago that doing exactly that might have good chances to score very high on public LB. Thanks for testing and thanks for the congratz!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542822,
      "author_name": "arvindmvepa",
      "author_url": "",
      "post_date": "06/04/2019 05:57:13",
      "content": "<p>ya, i tried this on our models just now and it improved our scores dramatically. i feel like almost any model would have been improved via scaling. i remember we had this discussion in our group whether the earthquakes would be longer in the test set than the training set. we weren't sure at the time. it seemed to have a big influence on the score -possibly because predicting earthquakes must be that difficult. i remember we were analyzing the earthquake errors and noticed that the model didn't seem to see the difference between longer and shorter earthquakes and seemed to make the mean prediction. We thought about scaling, but we ended up not reading the previous discussion post in detail, given time, and got bogged down doing other things. Should have looked into it. Rather than a conservative/risky approach for our submissions, we should have considered scaling our predictions, in consideration of the earthquake length.</p>",
      "votes": null,
      "replies": [
        {
          "id": 542877,
          "author_name": "arvindmvepa",
          "author_url": "",
          "post_date": "06/04/2019 06:49:54",
          "content": "<p>what we did was eye-ball the training set and see if the earthquake lengths were increasing in time. based on our general observation, it didn't seem so (this is why we originally discarded this hypothesis). Maybe we should have checked this rigorously using a moving average. I didn't actually do this, but if this trend doesn't exist, then i guess it's a bit unusual that the earthquakes increased in length during the test set. If the earthquakes didn't increase in length in the the current test set (or for increasing time), the test set would have been an unusual sample, which is probably unlikely given the number of samples. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 542940,
      "author_name": "danijelk",
      "author_url": "",
      "post_date": "06/04/2019 08:07:12",
      "content": "<p>Congrats <a href=\"/sushize\">@sushize</a> for your fast and good solution. I was wondering, how did you come up with 1.08 and 0.96 in 5 minutes? :) My teammates deserve all the glory for this win. Unfortunately, in the last month of the competition I did not have the time to contribute anything at all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542941,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "06/04/2019 08:07:26",
      "content": "<p>I tried your clever technique in one of my unselected best models  (unfortunately, it turned to be a silver medalist) and I got 2.35016 (gold zone medalist). Sometimes, cleverness (maybe I would say with a combination of luck also)beats hard work! Congrats! :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542970,
      "author_name": "stanislavblinov",
      "author_url": "",
      "post_date": "06/04/2019 08:42:26",
      "content": "<p>The fastest hand on Kaggle :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543039,
      "author_name": "prashanththangavel",
      "author_url": "",
      "post_date": "06/04/2019 09:43:48",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543262,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "06/04/2019 12:23:24",
      "content": "<p>Congrats! You are understanding the root of this competition that test time ttf difference is the most important point.\nYour idea wins the medal.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543268,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 12:28:20",
      "content": "<p>You beat everyone on time to value for sure!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543338,
      "author_name": "akhileshrai",
      "author_url": "",
      "post_date": "06/04/2019 13:34:10",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543552,
      "author_name": "zidmie",
      "author_url": "",
      "post_date": "06/04/2019 15:36:31",
      "content": "<p>I tried this game : multiplying my final submission by 1.19 gives me 2.296 on private LB : 2nd position!\nMean value is then 6.22. \nCan we play lottery again? :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 552728,
      "author_name": "",
      "author_url": "",
      "post_date": "06/14/2019 12:22:11",
      "content": "<p>That's so cool lol</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542567": "First, congrats to all the winners, especially Zoo team and my previous great teammate Danijel! All of you have done an amazing job!\n\nThis competition is such an unstable competition that I previously didn't plan to spend any time on it. About 2 hour before the deadline, I got a bit of free time and decided to buy 2 lottery tickets for this competition and spent 5 mins to make 2 submissions. \n\nBasically, what I did is just: 1) take  the  old public LB1.456 benchmark solution (shared 2 months ago in the kernel's section https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction, credits to Gabriel); 2) multiply the prediction by 1.08 and 0.96 respectively, and made these 2 submissions. That's it. \n\nAs a result, the 1.08 adjusted submission gave private LB score 2.45 and 81th place.\n\nIt seems that the scale/magnitude of prediction in the private LB part (compared with train data and public LB data) might probably be the largest uncertainty in this competition (in terms of public/private LB ranking), and applying a scale-up and scale-down factors adjustment might increase the chance to finish in the very top of the LB (though also increase the risk of finishing in the very bottom of the LB when factor adjustment is too aggressive)\n\nBtw, I also found that if we simply multiply 1.1 with the PublicLB 1.50 catboost benchmark solution (shared one month ago in kernel https://www.kaggle.com/hsinwenchang/mfcc-randomforestregressor-catboostregressor/output), it will score 2.336 on the private LB (11th place) and gave a gold medal. Without any adjustment, the benchmark solution gave 2.45 private LB score.\n\nI guess that most of top 15 competitors might possibly win (or close to) top 1 place prize with some simple factor adjustment.\n\nFinally, I am in particular very interested in the top 1 solution, who seems to have a good winning margin on the private LB compared with all other teams. Good job!",
    "542590": "Thanks for sharing. What would be the intuition of scaling up or down? Higher TTF for scale up and lower TTF for scale down?",
    "542593": "It happens because Public LB mean is around 4 and Private is around 6.7. So multiplying by 1.x or adding a constant to your predictions improves your Private LB. \n\nBy the way, congrats @sushize . This is the fastest silver medal I've ever seen ;)",
    "542597": "Thanks,really amazing!",
    "542602": "That's amazing. But how did you decide on 1.08 and 0.96 ? \nI just tried increasing the factor 1.085 and it's still 2.45430",
    "542616": "Based on previous forum sharing, I just thought that it is more likely that the private LB might have higher(rather than lower) mean than public LB, and I decided to apply a bit more aggressive adjustment when scaling up and less aggressive adjustment when scaling down. The specific values 1.08 and 0.96 are just two random numbers from my intuition (though did some simple math calculation to make sure the adjustment would not be out of the range of \"making sense\", and also double check it by looking at the public LB score). For example, an adjustment factor of 8.0 would definitely blow up the score and won't make any sense (which could also be verified by looking at the public LB).\n\nAlso, I made the 1.08 adjustment submission first, which gave me 1.68 public LB score, which I think it is in the range of \"makes sense\". If I found that submission to have too bad public LB score (e.g., 2.0+ public LB score), then for my 2nd submission I probably would apply a less aggressive 1.x+ adjustment (e.g., 1.02) rather than 0.96 adjustment, since I do think that there is a higher chance that private LB would have a higher mean (based on forum discussions). But given that my first submission looks reasonable, I made the 2nd scale-down submission with factor 0.96 as planned.",
    "542705": "Nice and clever way to try your luck",
    "542714": "just check by submitting my best result ( private score 2.59332 ) time 1.1 and it gave me 2.35529, which is really amazing.",
    "542758": "I find this is interesting.  When I used mae for a separate data set for error analysis (even plotted it).  Wonder if mean error (without absolute) would turn out to be negative.  I think simply apply average target/predicted ratio would increase ranking.",
    "542813": "I was joking around a few days ago that doing exactly that might have good chances to score very high on public LB. Thanks for testing and thanks for the congratz!",
    "542822": "ya, i tried this on our models just now and it improved our scores dramatically. i feel like almost any model would have been improved via scaling. i remember we had this discussion in our group whether the earthquakes would be longer in the test set than the training set. we weren't sure at the time. it seemed to have a big influence on the score -possibly because predicting earthquakes must be that difficult. i remember we were analyzing the earthquake errors and noticed that the model didn't seem to see the difference between longer and shorter earthquakes and seemed to make the mean prediction. We thought about scaling, but we ended up not reading the previous discussion post in detail, given time, and got bogged down doing other things. Should have looked into it. Rather than a conservative/risky approach for our submissions, we should have considered scaling our predictions, in consideration of the earthquake length.",
    "542840": "Thanks for the reply and seriously hats off for the sharp intuition with numbers. \nIf you don't mind another curiosity I have, why take Gabriel's kernel ? There were others with slightly better and slightly worse score at the time, and you manage to choose the one that works for the 2 submission.",
    "542877": "what we did was eye-ball the training set and see if the earthquake lengths were increasing in time. based on our general observation, it didn't seem so (this is why we originally discarded this hypothesis). Maybe we should have checked this rigorously using a moving average. I didn't actually do this, but if this trend doesn't exist, then i guess it's a bit unusual that the earthquakes increased in length during the test set. If the earthquakes didn't increase in length in the the current test set (or for increasing time), the test set would have been an unusual sample, which is probably unlikely given the number of samples.",
    "542940": "Congrats @sushize for your fast and good solution. I was wondering, how did you come up with 1.08 and 0.96 in 5 minutes? :) My teammates deserve all the glory for this win. Unfortunately, in the last month of the competition I did not have the time to contribute anything at all.",
    "542941": "I tried your clever technique in one of my unselected best models  (unfortunately, it turned to be a silver medalist) and I got 2.35016 (gold zone medalist). Sometimes, cleverness (maybe I would say with a combination of luck also)beats hard work! Congrats! :-)",
    "542970": "The fastest hand on Kaggle :)",
    "543039": "Congrats! Thanks for sharing :)",
    "543262": "Congrats! You are understanding the root of this competition that test time ttf difference is the most important point.\nYour idea wins the medal.",
    "543268": "You beat everyone on time to value for sure!",
    "543338": "Congratulations!",
    "543552": "I tried this game : multiplying my final submission by 1.19 gives me 2.296 on private LB : 2nd position!\nMean value is then 6.22. \nCan we play lottery again? :-)",
    "543708": "It is based on several simple thoughts:\n\n[1] The public LB score 1.45 is \"reasonably ok\".\n[2] The kernel solution was generated 2 months ago, at that time it focused more on the EDA part and didn't try every efforts to climb the public LB, and thus less prone to overfit to the public LB (compared with most recent kernels with slightly better public LB score)\n[3] I only have a couple of mins, and thus I just quickly take one benchmark solution that I \"like\" most at the first glance, rather than did a comprehensive search and comparison among all available \"good benchmark\". \n\nBtw, the same factor adjustment trick probably would work for most, if not all,  \"good public benchmarks\". And the one I selected performs reasonably well but definitely not the best/optimal one. As I also mentioned in the main post, if I selected the public LB 1.50 kernel solution shared 1 month ago and applying the same adjustment, it will gave a gold medal.\n\nIn short, I just did sth simple that I think might have some nontrivial chance to finish very well on the private LB (by scaling up and scaling down), and the rest is luck. That's why I called my two submissions as \"lottery tickets\".",
    "552728": "That's so cool lol"
  },
  "source": "meta"
}