{
  "id": 94446,
  "title": "My no-test-leak 23rd solution",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/kha-vo-my-no-test-leak-23rd-solution",
  "author_name": "",
  "post_date": "2019-06-04T14:44:51.829791800Z",
  "votes": 35,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hello guys,</p>\n\n<p>After a day both agonising my lost solo gold medal, as well as enjoy my luck to survive the shake, I am now be able to do this mini write-up. </p>\n\n<p>Although I admitted I exploited the advantage of test leak to survive, my 23rd position DID NOT use any kind of test leak information. What I mean is that as the leak test estimates are known, it just made me less suspect my final submission with a very bad public LB (1.350) score as well as bad local MAE score, and select it as the \"risky\" submission. That submission does not use any kind of leak information.</p>\n\n<p><strong>Feature Engineering</strong>\n+ The top feature is the 1st coefficient of numpy.polyfit  of order 3 on the columns of STFT spectrogram. \n+ The 2nd top feature is just the magnitude of a specific band of FFT frequency, after some smoothing techniques on the FFT spectrum.\n+ Other weaker features are some welch coefficents, roll std, STFT contrast level, C3 unlinearity....</p>\n\n<p><strong>Models</strong>\nI used LGB, XGB, CatBoost, and RF. All 4 of them score almost identical in my CV as well as LB (around 1.31 to 1.33). A simple blend of them gave me 1.290 on public LB. Hyperparameters were tuned by random search and later by Gaussian optimization. LGB with 40 leaves, 20 min_data_in_leave, depth 7, subsample 0.5. XGB with depth 7, subsample 0.5. </p>\n\n<p><strong>CV</strong>\nThis is the part that made me disappointed at the end. I was disappointed as I spent a lot of time building this trustworthy CV framework, but in the end to get a good score, one would not need it at all. However, this CV framework can serve me as a legacy for future competitions.\n- First, I bet so many people may wonder why using sqrt(y) gives better result than raw y. That's because the varied variance of predictions. Higher ttf samples are more unpredictable. So we need to focus more on low ttf samples, because they are achievable. If we use raw y, CV will overfit on high ttf samples, and underfit (not converged) in low ttf samples. </p>\n\n<ul>\n<li><p>After some careful CV iterations (which will be described next), I chose y^(0.55) to be the best transformation.</p></li>\n<li><p>Leave-one-quake-out (LOQO) or shuffled CV will give the same result, if the distributions of true train data and early stopping data is the same, with a SEPARATE validation data. As a result, I used 10 stratified folds based on 16 quake_id. For each fold train, I further use another CV of another 10 stratified folds of the training fold (which has 9/10 data of the whole original training set).  So, in the end each innermost loop of training session consists of (9/10)*(9/10) samples from the whole original train set. The inner loop CV serves to get the best averaged iteration for each outer fold train.</p></li>\n<li><p>I train on 10 different seeds of the outer CV.</p></li>\n<li><p>It turns out that using shuffled CV in general gives better result than LOQO, because each sample is validated by 10 different subsets of seeds. In LOQO, each data sample is validated by just 1 set of seed, that is the combination of all the quakes other than the one containing that sample.</p></li>\n</ul>\n\n<p><strong>Stack</strong></p>\n\n<ul>\n<li><p>Optimizing MSE: Using a simple ridge regression sample on all OOF predictions of 4 models (each with 10 seeds, as said), I got my best submission which score 1.350 public LB and 2.409 private (23rd position). Please note that, this is my best CV only in terms of MSE. The stacked MAE is much worse than any of the 1st level model. This is a gamble that I made, given that the private set has significantly higher ttf samples. Without that knowledge, I would have not selected this submission. This is my \"risky\" submission, and turned out to be the saviour.</p></li>\n<li><p>Optimizing MAE: Ridge regression does not allow us to directly optimize MAE. So as usual, transform y to y^(0.55), and do MSE. By this, my stacked MAE is much better, which gives me 1.29x on public LB. This is my \"safe\" submission, and turned out to be the disaster.</p></li>\n</ul>\n\n<p><strong>What I regret</strong>\nI did the exact estimate of private LB based on the image of the p4677 paper. Then, I use the training weights obtained by getting the binned normalised histogram of private set divide the binned normalised histogram of the training set. Then I use just 1 seed split of 1 simple XGB model. That model gives 1.380 on public LB, and 2.375 on private (which could earn me a gold). If I blend more models with these weights, and finally stacking, I will end of in the top 3. I decided not to use this leak at the end, and I regret it. However, it's still lucky to me that my non-leak submission secured my 23rd position.</p>\n\n<p>Thanks for reading.</p>",
  "messages": [
    {
      "id": "543469",
      "postDate": "06/04/2019 14:44:51",
      "content": "<p>Hello guys,</p>\n\n<p>After a day both agonising my lost solo gold medal, as well as enjoy my luck to survive the shake, I am now be able to do this mini write-up. </p>\n\n<p>Although I admitted I exploited the advantage of test leak to survive, my 23rd position DID NOT use any kind of test leak information. What I mean is that as the leak test estimates are known, it just made me less suspect my final submission with a very bad public LB (1.350) score as well as bad local MAE score, and select it as the \"risky\" submission. That submission does not use any kind of leak information.</p>\n\n<p><strong>Feature Engineering</strong>\n+ The top feature is the 1st coefficient of numpy.polyfit  of order 3 on the columns of STFT spectrogram. \n+ The 2nd top feature is just the magnitude of a specific band of FFT frequency, after some smoothing techniques on the FFT spectrum.\n+ Other weaker features are some welch coefficents, roll std, STFT contrast level, C3 unlinearity....</p>\n\n<p><strong>Models</strong>\nI used LGB, XGB, CatBoost, and RF. All 4 of them score almost identical in my CV as well as LB (around 1.31 to 1.33). A simple blend of them gave me 1.290 on public LB. Hyperparameters were tuned by random search and later by Gaussian optimization. LGB with 40 leaves, 20 min_data_in_leave, depth 7, subsample 0.5. XGB with depth 7, subsample 0.5. </p>\n\n<p><strong>CV</strong>\nThis is the part that made me disappointed at the end. I was disappointed as I spent a lot of time building this trustworthy CV framework, but in the end to get a good score, one would not need it at all. However, this CV framework can serve me as a legacy for future competitions.\n- First, I bet so many people may wonder why using sqrt(y) gives better result than raw y. That's because the varied variance of predictions. Higher ttf samples are more unpredictable. So we need to focus more on low ttf samples, because they are achievable. If we use raw y, CV will overfit on high ttf samples, and underfit (not converged) in low ttf samples. </p>\n\n<ul>\n<li><p>After some careful CV iterations (which will be described next), I chose y^(0.55) to be the best transformation.</p></li>\n<li><p>Leave-one-quake-out (LOQO) or shuffled CV will give the same result, if the distributions of true train data and early stopping data is the same, with a SEPARATE validation data. As a result, I used 10 stratified folds based on 16 quake_id. For each fold train, I further use another CV of another 10 stratified folds of the training fold (which has 9/10 data of the whole original training set).  So, in the end each innermost loop of training session consists of (9/10)*(9/10) samples from the whole original train set. The inner loop CV serves to get the best averaged iteration for each outer fold train.</p></li>\n<li><p>I train on 10 different seeds of the outer CV.</p></li>\n<li><p>It turns out that using shuffled CV in general gives better result than LOQO, because each sample is validated by 10 different subsets of seeds. In LOQO, each data sample is validated by just 1 set of seed, that is the combination of all the quakes other than the one containing that sample.</p></li>\n</ul>\n\n<p><strong>Stack</strong></p>\n\n<ul>\n<li><p>Optimizing MSE: Using a simple ridge regression sample on all OOF predictions of 4 models (each with 10 seeds, as said), I got my best submission which score 1.350 public LB and 2.409 private (23rd position). Please note that, this is my best CV only in terms of MSE. The stacked MAE is much worse than any of the 1st level model. This is a gamble that I made, given that the private set has significantly higher ttf samples. Without that knowledge, I would have not selected this submission. This is my \"risky\" submission, and turned out to be the saviour.</p></li>\n<li><p>Optimizing MAE: Ridge regression does not allow us to directly optimize MAE. So as usual, transform y to y^(0.55), and do MSE. By this, my stacked MAE is much better, which gives me 1.29x on public LB. This is my \"safe\" submission, and turned out to be the disaster.</p></li>\n</ul>\n\n<p><strong>What I regret</strong>\nI did the exact estimate of private LB based on the image of the p4677 paper. Then, I use the training weights obtained by getting the binned normalised histogram of private set divide the binned normalised histogram of the training set. Then I use just 1 seed split of 1 simple XGB model. That model gives 1.380 on public LB, and 2.375 on private (which could earn me a gold). If I blend more models with these weights, and finally stacking, I will end of in the top 3. I decided not to use this leak at the end, and I regret it. However, it's still lucky to me that my non-leak submission secured my 23rd position.</p>\n\n<p>Thanks for reading.</p>",
      "rawMarkdown": "Hello guys,\n\nAfter a day both agonising my lost solo gold medal, as well as enjoy my luck to survive the shake, I am now be able to do this mini write-up. \n\nAlthough I admitted I exploited the advantage of test leak to survive, my 23rd position DID NOT use any kind of test leak information. What I mean is that as the leak test estimates are known, it just made me less suspect my final submission with a very bad public LB (1.350) score as well as bad local MAE score, and select it as the \"risky\" submission. That submission does not use any kind of leak information.\n\n**Feature Engineering**\n+ The top feature is the 1st coefficient of numpy.polyfit  of order 3 on the columns of STFT spectrogram. \n+ The 2nd top feature is just the magnitude of a specific band of FFT frequency, after some smoothing techniques on the FFT spectrum.\n+ Other weaker features are some welch coefficents, roll std, STFT contrast level, C3 unlinearity....\n\n**Models**\nI used LGB, XGB, CatBoost, and RF. All 4 of them score almost identical in my CV as well as LB (around 1.31 to 1.33). A simple blend of them gave me 1.290 on public LB. Hyperparameters were tuned by random search and later by Gaussian optimization. LGB with 40 leaves, 20 min_data_in_leave, depth 7, subsample 0.5. XGB with depth 7, subsample 0.5. \n\n**CV**\nThis is the part that made me disappointed at the end. I was disappointed as I spent a lot of time building this trustworthy CV framework, but in the end to get a good score, one would not need it at all. However, this CV framework can serve me as a legacy for future competitions.\n- First, I bet so many people may wonder why using sqrt(y) gives better result than raw y. That's because the varied variance of predictions. Higher ttf samples are more unpredictable. So we need to focus more on low ttf samples, because they are achievable. If we use raw y, CV will overfit on high ttf samples, and underfit (not converged) in low ttf samples. \n\n- After some careful CV iterations (which will be described next), I chose y^(0.55) to be the best transformation.\n\n- Leave-one-quake-out (LOQO) or shuffled CV will give the same result, if the distributions of true train data and early stopping data is the same, with a SEPARATE validation data. As a result, I used 10 stratified folds based on 16 quake_id. For each fold train, I further use another CV of another 10 stratified folds of the training fold (which has 9/10 data of the whole original training set).  So, in the end each innermost loop of training session consists of (9/10)*(9/10) samples from the whole original train set. The inner loop CV serves to get the best averaged iteration for each outer fold train.\n\n- I train on 10 different seeds of the outer CV.\n\n- It turns out that using shuffled CV in general gives better result than LOQO, because each sample is validated by 10 different subsets of seeds. In LOQO, each data sample is validated by just 1 set of seed, that is the combination of all the quakes other than the one containing that sample.\n\n**Stack**\n\n- Optimizing MSE: Using a simple ridge regression sample on all OOF predictions of 4 models (each with 10 seeds, as said), I got my best submission which score 1.350 public LB and 2.409 private (23rd position). Please note that, this is my best CV only in terms of MSE. The stacked MAE is much worse than any of the 1st level model. This is a gamble that I made, given that the private set has significantly higher ttf samples. Without that knowledge, I would have not selected this submission. This is my \"risky\" submission, and turned out to be the saviour.\n\n- Optimizing MAE: Ridge regression does not allow us to directly optimize MAE. So as usual, transform y to y^(0.55), and do MSE. By this, my stacked MAE is much better, which gives me 1.29x on public LB. This is my \"safe\" submission, and turned out to be the disaster.\n\n**What I regret**\nI did the exact estimate of private LB based on the image of the p4677 paper. Then, I use the training weights obtained by getting the binned normalised histogram of private set divide the binned normalised histogram of the training set. Then I use just 1 seed split of 1 simple XGB model. That model gives 1.380 on public LB, and 2.375 on private (which could earn me a gold). If I blend more models with these weights, and finally stacking, I will end of in the top 3. I decided not to use this leak at the end, and I regret it. However, it's still lucky to me that my non-leak submission secured my 23rd position.\n\nThanks for reading.",
      "votes": null
    },
    {
      "id": "543487",
      "postDate": "06/04/2019 14:57:22",
      "content": "<p>Thanks for explaining sqrt(y)! </p>",
      "rawMarkdown": "Thanks for explaining sqrt(y)!",
      "votes": null
    },
    {
      "id": "543541",
      "postDate": "06/04/2019 15:27:30",
      "content": "<p>Congrats and thanks for sharing. Now it makes sense that a power transformation on target performs better due unpredictable high ttf values.</p>",
      "rawMarkdown": "Congrats and thanks for sharing. Now it makes sense that a power transformation on target performs better due unpredictable high ttf values.",
      "votes": null
    },
    {
      "id": "543572",
      "postDate": "06/04/2019 15:46:35",
      "content": "<p>Thanks for sharingn, and congrats on the great result.  I tried sqrt(y) as well, but found that using huber objective was better at the time.  Maybe I should have revisited it!</p>",
      "rawMarkdown": "Thanks for sharingn, and congrats on the great result.  I tried sqrt(y) as well, but found that using huber objective was better at the time.  Maybe I should have revisited it!",
      "votes": null
    },
    {
      "id": "543725",
      "postDate": "06/04/2019 18:14:05",
      "content": "<p>Congrats and thanks for sharing. It was fun watching you climbing the LB again and again since the beginning of the competition. You deserved a gold :( Better luck in upcoming competitions.</p>",
      "rawMarkdown": "Congrats and thanks for sharing. It was fun watching you climbing the LB again and again since the beginning of the competition. You deserved a gold :( Better luck in upcoming competitions.",
      "votes": null
    },
    {
      "id": "543868",
      "postDate": "06/04/2019 22:36:38",
      "content": "<p>We also tried sqrt(y), but we were not able to think about this as deeply as you did. Great job, thank you for sharing.</p>",
      "rawMarkdown": "We also tried sqrt(y), but we were not able to think about this as deeply as you did. Great job, thank you for sharing.",
      "votes": null
    },
    {
      "id": "543900",
      "postDate": "06/04/2019 23:31:29",
      "content": "<p>Thanks Amjad <a href=\"/amjad85\">@amjad85</a> and congrats your first gold medal. </p>",
      "rawMarkdown": "Thanks Amjad @amjad85 and congrats your first gold medal.",
      "votes": null
    },
    {
      "id": "543902",
      "postDate": "06/04/2019 23:32:25",
      "content": "<p>Thanks, and let me share the feelings of being dropped due to leak. Indeed, other transformations also work well, such as log(y+5). Gamma or huber regression is other types of regression that assumes the variance increases w.r.t target. They all work better than the nightmare MAE loss or MSE.</p>",
      "rawMarkdown": "Thanks, and let me share the feelings of being dropped due to leak. Indeed, other transformations also work well, such as log(y+5). Gamma or huber regression is other types of regression that assumes the variance increases w.r.t target. They all work better than the nightmare MAE loss or MSE.",
      "votes": null
    },
    {
      "id": "543913",
      "postDate": "06/04/2019 23:45:28",
      "content": "<p>Wow, you really worked hard regarding transformations! I am so happy to know such kinds of solutions. Thank you.</p>",
      "rawMarkdown": "Wow, you really worked hard regarding transformations! I am so happy to know such kinds of solutions. Thank you.",
      "votes": null
    },
    {
      "id": "543926",
      "postDate": "06/05/2019 00:15:05",
      "content": "<p>Given your late join, it is of course impossible for you to do everything, but I see you have much joy in this competition. Congrats on your impressive survival and leading the public LB.</p>",
      "rawMarkdown": "Given your late join, it is of course impossible for you to do everything, but I see you have much joy in this competition. Congrats on your impressive survival and leading the public LB.",
      "votes": null
    },
    {
      "id": "544013",
      "postDate": "06/05/2019 03:15:59",
      "content": "<p>Our team also survived without test leak information. we have analyzed many test prediction distributions from many models (lgb,xgb with different params, different dataset) and selected the submissions that have shape of distribution look most similar to train  distribution. If only we knew the leak, especially mean of test set, we could end up with better result  :)</p>",
      "rawMarkdown": "Our team also survived without test leak information. we have analyzed many test prediction distributions from many models (lgb,xgb with different params, different dataset) and selected the submissions that have shape of distribution look most similar to train  distribution. If only we knew the leak, especially mean of test set, we could end up with better result  :)",
      "votes": null
    },
    {
      "id": "544019",
      "postDate": "06/05/2019 03:26:03",
      "content": "<p>nice!</p>",
      "rawMarkdown": "nice!",
      "votes": null
    },
    {
      "id": "544033",
      "postDate": "06/05/2019 03:58:22",
      "content": "<p>Congratulations! Thanks for sharing :) </p>",
      "rawMarkdown": "Congratulations! Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "544462",
      "postDate": "06/05/2019 14:32:57",
      "content": "<p>Great explanation, now I see I should have spent more time exploring target transformation. Thanks for sharing and congratulations for great position!</p>",
      "rawMarkdown": "Great explanation, now I see I should have spent more time exploring target transformation. Thanks for sharing and congratulations for great position!",
      "votes": null
    },
    {
      "id": "544507",
      "postDate": "06/05/2019 15:26:46",
      "content": "<p>Not to make you regret it more, but why did you do this : </p>\n\n<blockquote>\n  <p>I decided not to use this leak at the end</p>\n</blockquote>",
      "rawMarkdown": "Not to make you regret it more, but why did you do this : \n\n&gt;  I decided not to use this leak at the end",
      "votes": null
    },
    {
      "id": "544510",
      "postDate": "06/05/2019 15:30:57",
      "content": "<p>My non leaked model has the same mean with the leaked one. I just did not want to gamble too much. </p>",
      "rawMarkdown": "My non leaked model has the same mean with the leaked one. I just did not want to gamble too much.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543487,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "06/04/2019 14:57:22",
      "content": "<p>Thanks for explaining sqrt(y)! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543541,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 15:27:30",
      "content": "<p>Congrats and thanks for sharing. Now it makes sense that a power transformation on target performs better due unpredictable high ttf values.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543572,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 15:46:35",
      "content": "<p>Thanks for sharingn, and congrats on the great result.  I tried sqrt(y) as well, but found that using huber objective was better at the time.  Maybe I should have revisited it!</p>",
      "votes": null,
      "replies": [
        {
          "id": 543926,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "06/05/2019 00:15:05",
          "content": "<p>Given your late join, it is of course impossible for you to do everything, but I see you have much joy in this competition. Congrats on your impressive survival and leading the public LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543725,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "06/04/2019 18:14:05",
      "content": "<p>Congrats and thanks for sharing. It was fun watching you climbing the LB again and again since the beginning of the competition. You deserved a gold :( Better luck in upcoming competitions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 543900,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "06/04/2019 23:31:29",
          "content": "<p>Thanks Amjad <a href=\"/amjad85\">@amjad85</a> and congrats your first gold medal. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543868,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "06/04/2019 22:36:38",
      "content": "<p>We also tried sqrt(y), but we were not able to think about this as deeply as you did. Great job, thank you for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 543902,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "06/04/2019 23:32:25",
          "content": "<p>Thanks, and let me share the feelings of being dropped due to leak. Indeed, other transformations also work well, such as log(y+5). Gamma or huber regression is other types of regression that assumes the variance increases w.r.t target. They all work better than the nightmare MAE loss or MSE.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543913,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "06/04/2019 23:45:28",
          "content": "<p>Wow, you really worked hard regarding transformations! I am so happy to know such kinds of solutions. Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544013,
      "author_name": "hung102",
      "author_url": "",
      "post_date": "06/05/2019 03:15:59",
      "content": "<p>Our team also survived without test leak information. we have analyzed many test prediction distributions from many models (lgb,xgb with different params, different dataset) and selected the submissions that have shape of distribution look most similar to train  distribution. If only we knew the leak, especially mean of test set, we could end up with better result  :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 544019,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "06/05/2019 03:26:03",
          "content": "<p>nice!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544033,
      "author_name": "prashanththangavel",
      "author_url": "",
      "post_date": "06/05/2019 03:58:22",
      "content": "<p>Congratulations! Thanks for sharing :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544462,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "06/05/2019 14:32:57",
      "content": "<p>Great explanation, now I see I should have spent more time exploring target transformation. Thanks for sharing and congratulations for great position!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544507,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/05/2019 15:26:46",
      "content": "<p>Not to make you regret it more, but why did you do this : </p>\n\n<blockquote>\n  <p>I decided not to use this leak at the end</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 544510,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "06/05/2019 15:30:57",
          "content": "<p>My non leaked model has the same mean with the leaked one. I just did not want to gamble too much. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "543469": "Hello guys,\n\nAfter a day both agonising my lost solo gold medal, as well as enjoy my luck to survive the shake, I am now be able to do this mini write-up. \n\nAlthough I admitted I exploited the advantage of test leak to survive, my 23rd position DID NOT use any kind of test leak information. What I mean is that as the leak test estimates are known, it just made me less suspect my final submission with a very bad public LB (1.350) score as well as bad local MAE score, and select it as the \"risky\" submission. That submission does not use any kind of leak information.\n\n**Feature Engineering**\n+ The top feature is the 1st coefficient of numpy.polyfit  of order 3 on the columns of STFT spectrogram. \n+ The 2nd top feature is just the magnitude of a specific band of FFT frequency, after some smoothing techniques on the FFT spectrum.\n+ Other weaker features are some welch coefficents, roll std, STFT contrast level, C3 unlinearity....\n\n**Models**\nI used LGB, XGB, CatBoost, and RF. All 4 of them score almost identical in my CV as well as LB (around 1.31 to 1.33). A simple blend of them gave me 1.290 on public LB. Hyperparameters were tuned by random search and later by Gaussian optimization. LGB with 40 leaves, 20 min_data_in_leave, depth 7, subsample 0.5. XGB with depth 7, subsample 0.5. \n\n**CV**\nThis is the part that made me disappointed at the end. I was disappointed as I spent a lot of time building this trustworthy CV framework, but in the end to get a good score, one would not need it at all. However, this CV framework can serve me as a legacy for future competitions.\n- First, I bet so many people may wonder why using sqrt(y) gives better result than raw y. That's because the varied variance of predictions. Higher ttf samples are more unpredictable. So we need to focus more on low ttf samples, because they are achievable. If we use raw y, CV will overfit on high ttf samples, and underfit (not converged) in low ttf samples. \n\n- After some careful CV iterations (which will be described next), I chose y^(0.55) to be the best transformation.\n\n- Leave-one-quake-out (LOQO) or shuffled CV will give the same result, if the distributions of true train data and early stopping data is the same, with a SEPARATE validation data. As a result, I used 10 stratified folds based on 16 quake_id. For each fold train, I further use another CV of another 10 stratified folds of the training fold (which has 9/10 data of the whole original training set).  So, in the end each innermost loop of training session consists of (9/10)*(9/10) samples from the whole original train set. The inner loop CV serves to get the best averaged iteration for each outer fold train.\n\n- I train on 10 different seeds of the outer CV.\n\n- It turns out that using shuffled CV in general gives better result than LOQO, because each sample is validated by 10 different subsets of seeds. In LOQO, each data sample is validated by just 1 set of seed, that is the combination of all the quakes other than the one containing that sample.\n\n**Stack**\n\n- Optimizing MSE: Using a simple ridge regression sample on all OOF predictions of 4 models (each with 10 seeds, as said), I got my best submission which score 1.350 public LB and 2.409 private (23rd position). Please note that, this is my best CV only in terms of MSE. The stacked MAE is much worse than any of the 1st level model. This is a gamble that I made, given that the private set has significantly higher ttf samples. Without that knowledge, I would have not selected this submission. This is my \"risky\" submission, and turned out to be the saviour.\n\n- Optimizing MAE: Ridge regression does not allow us to directly optimize MAE. So as usual, transform y to y^(0.55), and do MSE. By this, my stacked MAE is much better, which gives me 1.29x on public LB. This is my \"safe\" submission, and turned out to be the disaster.\n\n**What I regret**\nI did the exact estimate of private LB based on the image of the p4677 paper. Then, I use the training weights obtained by getting the binned normalised histogram of private set divide the binned normalised histogram of the training set. Then I use just 1 seed split of 1 simple XGB model. That model gives 1.380 on public LB, and 2.375 on private (which could earn me a gold). If I blend more models with these weights, and finally stacking, I will end of in the top 3. I decided not to use this leak at the end, and I regret it. However, it's still lucky to me that my non-leak submission secured my 23rd position.\n\nThanks for reading.",
    "543487": "Thanks for explaining sqrt(y)!",
    "543541": "Congrats and thanks for sharing. Now it makes sense that a power transformation on target performs better due unpredictable high ttf values.",
    "543572": "Thanks for sharingn, and congrats on the great result.  I tried sqrt(y) as well, but found that using huber objective was better at the time.  Maybe I should have revisited it!",
    "543725": "Congrats and thanks for sharing. It was fun watching you climbing the LB again and again since the beginning of the competition. You deserved a gold :( Better luck in upcoming competitions.",
    "543868": "We also tried sqrt(y), but we were not able to think about this as deeply as you did. Great job, thank you for sharing.",
    "543900": "Thanks Amjad @amjad85 and congrats your first gold medal.",
    "543902": "Thanks, and let me share the feelings of being dropped due to leak. Indeed, other transformations also work well, such as log(y+5). Gamma or huber regression is other types of regression that assumes the variance increases w.r.t target. They all work better than the nightmare MAE loss or MSE.",
    "543913": "Wow, you really worked hard regarding transformations! I am so happy to know such kinds of solutions. Thank you.",
    "543926": "Given your late join, it is of course impossible for you to do everything, but I see you have much joy in this competition. Congrats on your impressive survival and leading the public LB.",
    "544013": "Our team also survived without test leak information. we have analyzed many test prediction distributions from many models (lgb,xgb with different params, different dataset) and selected the submissions that have shape of distribution look most similar to train  distribution. If only we knew the leak, especially mean of test set, we could end up with better result  :)",
    "544019": "nice!",
    "544033": "Congratulations! Thanks for sharing :)",
    "544462": "Great explanation, now I see I should have spent more time exploring target transformation. Thanks for sharing and congratulations for great position!",
    "544507": "Not to make you regret it more, but why did you do this : \n\n&gt;  I decided not to use this leak at the end",
    "544510": "My non leaked model has the same mean with the leaked one. I just did not want to gamble too much."
  },
  "source": "meta"
}