{
  "id": 94433,
  "title": "#6 Solution ",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/carlospk-6-solution",
  "author_name": "",
  "post_date": "2019-06-05T02:23:05.890Z",
  "votes": 44,
  "comment_count": 16,
  "views": 0,
  "content": "<p>The main thing I took from my previous 2 competitions was that spending time in really understanding the problem and analyzing data was key to succeed. So in this one I started by carefully reading all papers, and I quickly realized that the data was from P4677. Actually, I realized it a couple of days before it was disclosed. The size of the train and test sets were 60-40% in both, the competition and the paper, so there was a first reason to think that the test set of the competition was the same. However, there was still a chance that the test set of the competition was different.</p>\n\n<p>So my first step was to visually select the experiments (cycles) from the training set, trying to match the length of those in the paper (with a ruler :smile:). At first I picked 9 from the 15 full cycles of the training data. Then I compared some features and realized that they became much more similar between train and test for the selected subset of experiments, which supported that the test set was the one of the paper:\n<img src=\"https://i.postimg.cc/DzbwGrjg/Feature-Train-Vs-Test.png\" alt=\"Comparison AbsMean Train vs Test\"></p>\n\n<p>Using those selected cycles, I realized an inverse correlation between my CV and the public LB, so I totally forgot about the LB. Actually, my 2 selected submissions scored 1.96799 and 1.82806.</p>\n\n<p>Nevertheless, I wanted to use more data than just the 9 experiments (the data was already small!), while trying to match train and test distributions. To solved this, I picked 4 different subsets from the training data, I built a model for each of them and I finally blended their predictions. I used 11 cycles from the training sets in total but in different combinations. For each of those 4 training data subsets I followed the following steps:\n- Feature selection (using shap most important features and also forward and backward feature elimination). I elimianted all features with different distributions in train-test.\n- Transform features to standard normal distribution, since I saw a shift between train and test:\n<img src=\"https://i.postimg.cc/6pPQTj0q/NS-Transformation.png\" alt=\"Transformation to Standard Normal Distribution\">\n- Build 3 models: LGBM, XGB and a shallow NNet (1 or 2 layers with less than 7 nodes). This shallow NNets worked well (on transformed data). LGBM and XGB tuned using bayesian optimization.\n- Blend the 3 models selecting the weights by hand (trying to avoid overfitting).</p>\n\n<p>This is the summary of the model:\n<img src=\"https://i.postimg.cc/Vkv6fQX2/6-Solution-Summary.png\" alt=\"Summary\"></p>\n\n<p>The oof predictions of the model with the largest weight in the final blending:\n<img src=\"https://i.postimg.cc/JnBjHB1V/oof.png\" alt=\"oof\"></p>\n\n<p>Finally, congrats to all the winners and to those who survived the leaderboard earthquake! I guess my experience with earthquakes as a Chilean helped me a bit :smiley:</p>",
  "messages": [
    {
      "id": "543363",
      "postDate": "06/04/2019 13:51:54",
      "content": "<p>The main thing I took from my previous 2 competitions was that spending time in really understanding the problem and analyzing data was key to succeed. So in this one I started by carefully reading all papers, and I quickly realized that the data was from P4677. Actually, I realized it a couple of days before it was disclosed. The size of the train and test sets were 60-40% in both, the competition and the paper, so there was a first reason to think that the test set of the competition was the same. However, there was still a chance that the test set of the competition was different.</p>\n\n<p>So my first step was to visually select the experiments (cycles) from the training set, trying to match the length of those in the paper (with a ruler :smile:). At first I picked 9 from the 15 full cycles of the training data. Then I compared some features and realized that they became much more similar between train and test for the selected subset of experiments, which supported that the test set was the one of the paper:\n<img src=\"https://i.postimg.cc/DzbwGrjg/Feature-Train-Vs-Test.png\" alt=\"Comparison AbsMean Train vs Test\"></p>\n\n<p>Using those selected cycles, I realized an inverse correlation between my CV and the public LB, so I totally forgot about the LB. Actually, my 2 selected submissions scored 1.96799 and 1.82806.</p>\n\n<p>Nevertheless, I wanted to use more data than just the 9 experiments (the data was already small!), while trying to match train and test distributions. To solved this, I picked 4 different subsets from the training data, I built a model for each of them and I finally blended their predictions. I used 11 cycles from the training sets in total but in different combinations. For each of those 4 training data subsets I followed the following steps:\n- Feature selection (using shap most important features and also forward and backward feature elimination). I elimianted all features with different distributions in train-test.\n- Transform features to standard normal distribution, since I saw a shift between train and test:\n<img src=\"https://i.postimg.cc/6pPQTj0q/NS-Transformation.png\" alt=\"Transformation to Standard Normal Distribution\">\n- Build 3 models: LGBM, XGB and a shallow NNet (1 or 2 layers with less than 7 nodes). This shallow NNets worked well (on transformed data). LGBM and XGB tuned using bayesian optimization.\n- Blend the 3 models selecting the weights by hand (trying to avoid overfitting).</p>\n\n<p>This is the summary of the model:\n<img src=\"https://i.postimg.cc/Vkv6fQX2/6-Solution-Summary.png\" alt=\"Summary\"></p>\n\n<p>The oof predictions of the model with the largest weight in the final blending:\n<img src=\"https://i.postimg.cc/JnBjHB1V/oof.png\" alt=\"oof\"></p>\n\n<p>Finally, congrats to all the winners and to those who survived the leaderboard earthquake! I guess my experience with earthquakes as a Chilean helped me a bit :smiley:</p>",
      "rawMarkdown": "The main thing I took from my previous 2 competitions was that spending time in really understanding the problem and analyzing data was key to succeed. So in this one I started by carefully reading all papers, and I quickly realized that the data was from P4677. Actually, I realized it a couple of days before it was disclosed. The size of the train and test sets were 60-40% in both, the competition and the paper, so there was a first reason to think that the test set of the competition was the same. However, there was still a chance that the test set of the competition was different.\n\nSo my first step was to visually select the experiments (cycles) from the training set, trying to match the length of those in the paper (with a ruler :smile:). At first I picked 9 from the 15 full cycles of the training data. Then I compared some features and realized that they became much more similar between train and test for the selected subset of experiments, which supported that the test set was the one of the paper:\n![Comparison AbsMean Train vs Test](https://i.postimg.cc/DzbwGrjg/Feature-Train-Vs-Test.png)\n\nUsing those selected cycles, I realized an inverse correlation between my CV and the public LB, so I totally forgot about the LB. Actually, my 2 selected submissions scored 1.96799 and 1.82806.\n\nNevertheless, I wanted to use more data than just the 9 experiments (the data was already small!), while trying to match train and test distributions. To solved this, I picked 4 different subsets from the training data, I built a model for each of them and I finally blended their predictions. I used 11 cycles from the training sets in total but in different combinations. For each of those 4 training data subsets I followed the following steps:\n- Feature selection (using shap most important features and also forward and backward feature elimination). I elimianted all features with different distributions in train-test.\n- Transform features to standard normal distribution, since I saw a shift between train and test:\n![Transformation to Standard Normal Distribution](https://i.postimg.cc/6pPQTj0q/NS-Transformation.png)\n- Build 3 models: LGBM, XGB and a shallow NNet (1 or 2 layers with less than 7 nodes). This shallow NNets worked well (on transformed data). LGBM and XGB tuned using bayesian optimization.\n- Blend the 3 models selecting the weights by hand (trying to avoid overfitting).\n\nThis is the summary of the model:\n![Summary](https://i.postimg.cc/Vkv6fQX2/6-Solution-Summary.png)\n\nThe oof predictions of the model with the largest weight in the final blending:\n![oof](https://i.postimg.cc/JnBjHB1V/oof.png)\n\nFinally, congrats to all the winners and to those who survived the leaderboard earthquake! I guess my experience with earthquakes as a Chilean helped me a bit :smiley:",
      "votes": null
    },
    {
      "id": "543367",
      "postDate": "06/04/2019 13:54:21",
      "content": "<p>Thanks for sharing, and congrats on the solo gold!</p>\n\n<blockquote>\n  <p>spending time in really understanding the problem and analyzing data was key to succeed</p>\n</blockquote>\n\n<p>Words of wisdom.</p>",
      "rawMarkdown": "Thanks for sharing, and congrats on the solo gold!\n\n&gt; spending time in really understanding the problem and analyzing data was key to succeed\n\nWords of wisdom.",
      "votes": null
    },
    {
      "id": "543501",
      "postDate": "06/04/2019 15:09:20",
      "content": "<p>Congrats! Thanks for sharing your solution!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution!",
      "votes": null
    },
    {
      "id": "543528",
      "postDate": "06/04/2019 15:22:04",
      "content": "<p>Double congrats for the golden solo Carlos! Thanks for sharing the solution.</p>",
      "rawMarkdown": "Double congrats for the golden solo Carlos! Thanks for sharing the solution.",
      "votes": null
    },
    {
      "id": "543565",
      "postDate": "06/04/2019 15:43:01",
      "content": "<p>Great job!</p>",
      "rawMarkdown": "Great job!",
      "votes": null
    },
    {
      "id": "543732",
      "postDate": "06/04/2019 18:22:29",
      "content": "<p>Ah.. some important features for me were related to picks. I calculated them on data in original units and also on acoustic_signal transformed to gaussian distribution:\n<code>\nx_roll_std = x.rolling(1000).std().dropna().values\nfor i in [0.7,0.75,0.8,0.85,0.9,1.0,1.5,1.75,2.0,2.25,2.5]:\n        peaks, h = find_peaks(x_roll_std, height=[i], distance=2000)\n        if(h['peak_heights'].shape[0]&amp;gt;1):\n            X_tr.loc[segment, 'peaks_count_' + str(i)] = h['peak_heights'].shape[0]\n            X_tr.loc[segment, 'peaks_mean_' + str(i)] = h['peak_heights'].mean()\n            X_tr.loc[segment, 'peaks_std_' + str(i)] = h['peak_heights'].std()\n            X_tr.loc[segment, 'peaks_max_' + str(i)] = h['peak_heights'].max()\n            X_tr.loc[segment, 'peaks_min_' + str(i)] = h['peak_heights'].min()\n            #Peak prominences\n            prominences = peak_prominences(x_roll_std, peaks)[0]\n            contour_heights = x_roll_std[peaks] - prominences\n            X_tr.loc[segment, 'peaks_prom_mean_' + str(i)] = contour_heights.mean()\n            X_tr.loc[segment, 'peaks_prom_std_' + str(i)] = contour_heights.std()\n            X_tr.loc[segment, 'peaks_prom_max_' + str(i)] = contour_heights.max()\n            X_tr.loc[segment, 'peaks_prom_min_' + str(i)] = contour_heights.min()\n            #distance between peaks\n            X_tr.loc[segment, 'peaks_dist_mean_' + str(i)] = np.diff(peaks).mean()\n            X_tr.loc[segment, 'peaks_dist_std_' + str(i)] = np.diff(peaks).std()\n            X_tr.loc[segment, 'peaks_dist_max_' + str(i)] = np.diff(peaks).max()\n            X_tr.loc[segment, 'peaks_dist_min_' + str(i)] = np.diff(peaks).min()\n</code></p>",
      "rawMarkdown": "Ah.. some important features for me were related to picks. I calculated them on data in original units and also on acoustic_signal transformed to gaussian distribution:\n```\nx_roll_std = x.rolling(1000).std().dropna().values\nfor i in [0.7,0.75,0.8,0.85,0.9,1.0,1.5,1.75,2.0,2.25,2.5]:\n        peaks, h = find_peaks(x_roll_std, height=[i], distance=2000)\n        if(h['peak_heights'].shape[0]&gt;1):\n            X_tr.loc[segment, 'peaks_count_' + str(i)] = h['peak_heights'].shape[0]\n            X_tr.loc[segment, 'peaks_mean_' + str(i)] = h['peak_heights'].mean()\n            X_tr.loc[segment, 'peaks_std_' + str(i)] = h['peak_heights'].std()\n            X_tr.loc[segment, 'peaks_max_' + str(i)] = h['peak_heights'].max()\n            X_tr.loc[segment, 'peaks_min_' + str(i)] = h['peak_heights'].min()\n            #Peak prominences\n            prominences = peak_prominences(x_roll_std, peaks)[0]\n            contour_heights = x_roll_std[peaks] - prominences\n            X_tr.loc[segment, 'peaks_prom_mean_' + str(i)] = contour_heights.mean()\n            X_tr.loc[segment, 'peaks_prom_std_' + str(i)] = contour_heights.std()\n            X_tr.loc[segment, 'peaks_prom_max_' + str(i)] = contour_heights.max()\n            X_tr.loc[segment, 'peaks_prom_min_' + str(i)] = contour_heights.min()\n            #distance between peaks\n            X_tr.loc[segment, 'peaks_dist_mean_' + str(i)] = np.diff(peaks).mean()\n            X_tr.loc[segment, 'peaks_dist_std_' + str(i)] = np.diff(peaks).std()\n            X_tr.loc[segment, 'peaks_dist_max_' + str(i)] = np.diff(peaks).max()\n            X_tr.loc[segment, 'peaks_dist_min_' + str(i)] = np.diff(peaks).min()\n```",
      "votes": null
    },
    {
      "id": "543894",
      "postDate": "06/04/2019 23:19:22",
      "content": "<p>Congrats! I admire your survey to find P4677 and hand tuning ability of weights. Is there any tips for avoiding overfitting in selecting weights?</p>",
      "rawMarkdown": "Congrats! I admire your survey to find P4677 and hand tuning ability of weights. Is there any tips for avoiding overfitting in selecting weights?",
      "votes": null
    },
    {
      "id": "543915",
      "postDate": "06/04/2019 23:48:25",
      "content": "<p>Thanks <a href=\"/sishihara\">@sishihara</a>. I just started with equal weights between lgbm and xgb (both 0.5), and started adding/substracting 0.1 until the oof score was the minimum (finally adding/substracting 0.05). And then I started adding 0.1 NNet and substracting to lgbm or xgb until getting the minimum oof score. I had some fun doing that.. haha. As you can see XGB worked better for me. I added NNets at the end and they helped to improve the score. \nI just thought that using an optimization algorithm to find the optimal weights could lead to overfitting given the characteristics of the data. </p>",
      "rawMarkdown": "Thanks @sishihara. I just started with equal weights between lgbm and xgb (both 0.5), and started adding/substracting 0.1 until the oof score was the minimum (finally adding/substracting 0.05). And then I started adding 0.1 NNet and substracting to lgbm or xgb until getting the minimum oof score. I had some fun doing that.. haha. As you can see XGB worked better for me. I added NNets at the end and they helped to improve the score. \nI just thought that using an optimization algorithm to find the optimal weights could lead to overfitting given the characteristics of the data.",
      "votes": null
    },
    {
      "id": "543967",
      "postDate": "06/05/2019 01:46:05",
      "content": "<p>Cool! We tried an optimization algorithm and found it overfitting...</p>",
      "rawMarkdown": "Cool! We tried an optimization algorithm and found it overfitting...",
      "votes": null
    },
    {
      "id": "543974",
      "postDate": "06/05/2019 02:10:52",
      "content": "<p>Congratulations <a href=\"/carlospk\">@carlospk</a> </p>",
      "rawMarkdown": "Congratulations @carlospk",
      "votes": null
    },
    {
      "id": "543977",
      "postDate": "06/05/2019 02:12:33",
      "content": "<p>That's a tip I learned from Grandmaster <a href=\"/titericz\">@titericz</a> some years ago to try to avoid overfitting when working in problems with few or noisy data ;)</p>",
      "rawMarkdown": "That's a tip I learned from Grandmaster @titericz some years ago to try to avoid overfitting when working in problems with few or noisy data ;)",
      "votes": null
    },
    {
      "id": "543978",
      "postDate": "06/05/2019 02:13:32",
      "content": "<p>Good work, thanks for sharing</p>",
      "rawMarkdown": "Good work, thanks for sharing",
      "votes": null
    },
    {
      "id": "550037",
      "postDate": "06/11/2019 08:37:29",
      "content": "<p>Congratulations! Thanks for sharing, I learned a lot from this.</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing, I learned a lot from this.",
      "votes": null
    },
    {
      "id": "552696",
      "postDate": "06/14/2019 11:03:57",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "561164",
      "postDate": "06/26/2019 07:40:59",
      "content": "<p>Thank you for sharing your work. I learned a lot from you. </p>\n\n<p>I have one question. What were reasons for you to choose those specific experiments for each model from 11 cycles? Did you kind of make them different to maximize the benefit of blending?</p>",
      "rawMarkdown": "Thank you for sharing your work. I learned a lot from you. \n\nI have one question. What were reasons for you to choose those specific experiments for each model from 11 cycles? Did you kind of make them different to maximize the benefit of blending?",
      "votes": null
    },
    {
      "id": "561488",
      "postDate": "06/26/2019 13:38:58",
      "content": "<p>Nice that it helped you to learn something <a href=\"/hatomugi\">@hatomugi</a> :)</p>\n\n<p>I selected different subsets of experiments in order to try to match the test distribution while trying to maximize the total number of experiments used to train (matching distributions while avoiding to use too few training data). \nFor example, you can notice that the subset of experiments with the highest weight in the blending (Val Set3), contains 7 experiments used as training data (1, 2, 4, 7, 10, 11 and 14). But, if you consider the experiments I used in total, the different models are trained on 11 experiments (1,2,3,4,7,8,9,10,11,12 and 14).</p>\n\n<p>Besides, I didn't know which of those subsets was the best one to use, so I thought it was better to use all of them ;)</p>",
      "rawMarkdown": "Nice that it helped you to learn something @hatomugi :)\n\nI selected different subsets of experiments in order to try to match the test distribution while trying to maximize the total number of experiments used to train (matching distributions while avoiding to use too few training data). \nFor example, you can notice that the subset of experiments with the highest weight in the blending (Val Set3), contains 7 experiments used as training data (1, 2, 4, 7, 10, 11 and 14). But, if you consider the experiments I used in total, the different models are trained on 11 experiments (1,2,3,4,7,8,9,10,11,12 and 14).\n\nBesides, I didn't know which of those subsets was the best one to use, so I thought it was better to use all of them ;)",
      "votes": null
    },
    {
      "id": "561517",
      "postDate": "06/26/2019 14:01:18",
      "content": "<p>Oh, I got it. That makes super sense. You tried to adjust the distribution between train and test set and also maximized the range of training data at the same time. Cool idea :) Thanks for the kind explanation.</p>",
      "rawMarkdown": "Oh, I got it. That makes super sense. You tried to adjust the distribution between train and test set and also maximized the range of training data at the same time. Cool idea :) Thanks for the kind explanation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543367,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 13:54:21",
      "content": "<p>Thanks for sharing, and congrats on the solo gold!</p>\n\n<blockquote>\n  <p>spending time in really understanding the problem and analyzing data was key to succeed</p>\n</blockquote>\n\n<p>Words of wisdom.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543501,
      "author_name": "dhaqui",
      "author_url": "",
      "post_date": "06/04/2019 15:09:20",
      "content": "<p>Congrats! Thanks for sharing your solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543528,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 15:22:04",
      "content": "<p>Double congrats for the golden solo Carlos! Thanks for sharing the solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543565,
      "author_name": "a45632",
      "author_url": "",
      "post_date": "06/04/2019 15:43:01",
      "content": "<p>Great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543732,
      "author_name": "carlospk",
      "author_url": "",
      "post_date": "06/04/2019 18:22:29",
      "content": "<p>Ah.. some important features for me were related to picks. I calculated them on data in original units and also on acoustic_signal transformed to gaussian distribution:\n<code>\nx_roll_std = x.rolling(1000).std().dropna().values\nfor i in [0.7,0.75,0.8,0.85,0.9,1.0,1.5,1.75,2.0,2.25,2.5]:\n        peaks, h = find_peaks(x_roll_std, height=[i], distance=2000)\n        if(h['peak_heights'].shape[0]&amp;gt;1):\n            X_tr.loc[segment, 'peaks_count_' + str(i)] = h['peak_heights'].shape[0]\n            X_tr.loc[segment, 'peaks_mean_' + str(i)] = h['peak_heights'].mean()\n            X_tr.loc[segment, 'peaks_std_' + str(i)] = h['peak_heights'].std()\n            X_tr.loc[segment, 'peaks_max_' + str(i)] = h['peak_heights'].max()\n            X_tr.loc[segment, 'peaks_min_' + str(i)] = h['peak_heights'].min()\n            #Peak prominences\n            prominences = peak_prominences(x_roll_std, peaks)[0]\n            contour_heights = x_roll_std[peaks] - prominences\n            X_tr.loc[segment, 'peaks_prom_mean_' + str(i)] = contour_heights.mean()\n            X_tr.loc[segment, 'peaks_prom_std_' + str(i)] = contour_heights.std()\n            X_tr.loc[segment, 'peaks_prom_max_' + str(i)] = contour_heights.max()\n            X_tr.loc[segment, 'peaks_prom_min_' + str(i)] = contour_heights.min()\n            #distance between peaks\n            X_tr.loc[segment, 'peaks_dist_mean_' + str(i)] = np.diff(peaks).mean()\n            X_tr.loc[segment, 'peaks_dist_std_' + str(i)] = np.diff(peaks).std()\n            X_tr.loc[segment, 'peaks_dist_max_' + str(i)] = np.diff(peaks).max()\n            X_tr.loc[segment, 'peaks_dist_min_' + str(i)] = np.diff(peaks).min()\n</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543894,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "06/04/2019 23:19:22",
      "content": "<p>Congrats! I admire your survey to find P4677 and hand tuning ability of weights. Is there any tips for avoiding overfitting in selecting weights?</p>",
      "votes": null,
      "replies": [
        {
          "id": 543915,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "06/04/2019 23:48:25",
          "content": "<p>Thanks <a href=\"/sishihara\">@sishihara</a>. I just started with equal weights between lgbm and xgb (both 0.5), and started adding/substracting 0.1 until the oof score was the minimum (finally adding/substracting 0.05). And then I started adding 0.1 NNet and substracting to lgbm or xgb until getting the minimum oof score. I had some fun doing that.. haha. As you can see XGB worked better for me. I added NNets at the end and they helped to improve the score. \nI just thought that using an optimization algorithm to find the optimal weights could lead to overfitting given the characteristics of the data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543967,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "06/05/2019 01:46:05",
          "content": "<p>Cool! We tried an optimization algorithm and found it overfitting...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543977,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "06/05/2019 02:12:33",
          "content": "<p>That's a tip I learned from Grandmaster <a href=\"/titericz\">@titericz</a> some years ago to try to avoid overfitting when working in problems with few or noisy data ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 552696,
          "author_name": "",
          "author_url": "",
          "post_date": "06/14/2019 11:03:57",
          "content": "<p>Thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543974,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "06/05/2019 02:10:52",
      "content": "<p>Congratulations <a href=\"/carlospk\">@carlospk</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543978,
      "author_name": "shemskurtoglu",
      "author_url": "",
      "post_date": "06/05/2019 02:13:32",
      "content": "<p>Good work, thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 550037,
      "author_name": "zhj4655",
      "author_url": "",
      "post_date": "06/11/2019 08:37:29",
      "content": "<p>Congratulations! Thanks for sharing, I learned a lot from this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 561164,
      "author_name": "hatomugi",
      "author_url": "",
      "post_date": "06/26/2019 07:40:59",
      "content": "<p>Thank you for sharing your work. I learned a lot from you. </p>\n\n<p>I have one question. What were reasons for you to choose those specific experiments for each model from 11 cycles? Did you kind of make them different to maximize the benefit of blending?</p>",
      "votes": null,
      "replies": [
        {
          "id": 561488,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "06/26/2019 13:38:58",
          "content": "<p>Nice that it helped you to learn something <a href=\"/hatomugi\">@hatomugi</a> :)</p>\n\n<p>I selected different subsets of experiments in order to try to match the test distribution while trying to maximize the total number of experiments used to train (matching distributions while avoiding to use too few training data). \nFor example, you can notice that the subset of experiments with the highest weight in the blending (Val Set3), contains 7 experiments used as training data (1, 2, 4, 7, 10, 11 and 14). But, if you consider the experiments I used in total, the different models are trained on 11 experiments (1,2,3,4,7,8,9,10,11,12 and 14).</p>\n\n<p>Besides, I didn't know which of those subsets was the best one to use, so I thought it was better to use all of them ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 561517,
          "author_name": "hatomugi",
          "author_url": "",
          "post_date": "06/26/2019 14:01:18",
          "content": "<p>Oh, I got it. That makes super sense. You tried to adjust the distribution between train and test set and also maximized the range of training data at the same time. Cool idea :) Thanks for the kind explanation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "543363": "The main thing I took from my previous 2 competitions was that spending time in really understanding the problem and analyzing data was key to succeed. So in this one I started by carefully reading all papers, and I quickly realized that the data was from P4677. Actually, I realized it a couple of days before it was disclosed. The size of the train and test sets were 60-40% in both, the competition and the paper, so there was a first reason to think that the test set of the competition was the same. However, there was still a chance that the test set of the competition was different.\n\nSo my first step was to visually select the experiments (cycles) from the training set, trying to match the length of those in the paper (with a ruler :smile:). At first I picked 9 from the 15 full cycles of the training data. Then I compared some features and realized that they became much more similar between train and test for the selected subset of experiments, which supported that the test set was the one of the paper:\n![Comparison AbsMean Train vs Test](https://i.postimg.cc/DzbwGrjg/Feature-Train-Vs-Test.png)\n\nUsing those selected cycles, I realized an inverse correlation between my CV and the public LB, so I totally forgot about the LB. Actually, my 2 selected submissions scored 1.96799 and 1.82806.\n\nNevertheless, I wanted to use more data than just the 9 experiments (the data was already small!), while trying to match train and test distributions. To solved this, I picked 4 different subsets from the training data, I built a model for each of them and I finally blended their predictions. I used 11 cycles from the training sets in total but in different combinations. For each of those 4 training data subsets I followed the following steps:\n- Feature selection (using shap most important features and also forward and backward feature elimination). I elimianted all features with different distributions in train-test.\n- Transform features to standard normal distribution, since I saw a shift between train and test:\n![Transformation to Standard Normal Distribution](https://i.postimg.cc/6pPQTj0q/NS-Transformation.png)\n- Build 3 models: LGBM, XGB and a shallow NNet (1 or 2 layers with less than 7 nodes). This shallow NNets worked well (on transformed data). LGBM and XGB tuned using bayesian optimization.\n- Blend the 3 models selecting the weights by hand (trying to avoid overfitting).\n\nThis is the summary of the model:\n![Summary](https://i.postimg.cc/Vkv6fQX2/6-Solution-Summary.png)\n\nThe oof predictions of the model with the largest weight in the final blending:\n![oof](https://i.postimg.cc/JnBjHB1V/oof.png)\n\nFinally, congrats to all the winners and to those who survived the leaderboard earthquake! I guess my experience with earthquakes as a Chilean helped me a bit :smiley:",
    "543367": "Thanks for sharing, and congrats on the solo gold!\n\n&gt; spending time in really understanding the problem and analyzing data was key to succeed\n\nWords of wisdom.",
    "543501": "Congrats! Thanks for sharing your solution!",
    "543528": "Double congrats for the golden solo Carlos! Thanks for sharing the solution.",
    "543565": "Great job!",
    "543732": "Ah.. some important features for me were related to picks. I calculated them on data in original units and also on acoustic_signal transformed to gaussian distribution:\n```\nx_roll_std = x.rolling(1000).std().dropna().values\nfor i in [0.7,0.75,0.8,0.85,0.9,1.0,1.5,1.75,2.0,2.25,2.5]:\n        peaks, h = find_peaks(x_roll_std, height=[i], distance=2000)\n        if(h['peak_heights'].shape[0]&gt;1):\n            X_tr.loc[segment, 'peaks_count_' + str(i)] = h['peak_heights'].shape[0]\n            X_tr.loc[segment, 'peaks_mean_' + str(i)] = h['peak_heights'].mean()\n            X_tr.loc[segment, 'peaks_std_' + str(i)] = h['peak_heights'].std()\n            X_tr.loc[segment, 'peaks_max_' + str(i)] = h['peak_heights'].max()\n            X_tr.loc[segment, 'peaks_min_' + str(i)] = h['peak_heights'].min()\n            #Peak prominences\n            prominences = peak_prominences(x_roll_std, peaks)[0]\n            contour_heights = x_roll_std[peaks] - prominences\n            X_tr.loc[segment, 'peaks_prom_mean_' + str(i)] = contour_heights.mean()\n            X_tr.loc[segment, 'peaks_prom_std_' + str(i)] = contour_heights.std()\n            X_tr.loc[segment, 'peaks_prom_max_' + str(i)] = contour_heights.max()\n            X_tr.loc[segment, 'peaks_prom_min_' + str(i)] = contour_heights.min()\n            #distance between peaks\n            X_tr.loc[segment, 'peaks_dist_mean_' + str(i)] = np.diff(peaks).mean()\n            X_tr.loc[segment, 'peaks_dist_std_' + str(i)] = np.diff(peaks).std()\n            X_tr.loc[segment, 'peaks_dist_max_' + str(i)] = np.diff(peaks).max()\n            X_tr.loc[segment, 'peaks_dist_min_' + str(i)] = np.diff(peaks).min()\n```",
    "543894": "Congrats! I admire your survey to find P4677 and hand tuning ability of weights. Is there any tips for avoiding overfitting in selecting weights?",
    "543915": "Thanks @sishihara. I just started with equal weights between lgbm and xgb (both 0.5), and started adding/substracting 0.1 until the oof score was the minimum (finally adding/substracting 0.05). And then I started adding 0.1 NNet and substracting to lgbm or xgb until getting the minimum oof score. I had some fun doing that.. haha. As you can see XGB worked better for me. I added NNets at the end and they helped to improve the score. \nI just thought that using an optimization algorithm to find the optimal weights could lead to overfitting given the characteristics of the data.",
    "543967": "Cool! We tried an optimization algorithm and found it overfitting...",
    "543974": "Congratulations @carlospk",
    "543977": "That's a tip I learned from Grandmaster @titericz some years ago to try to avoid overfitting when working in problems with few or noisy data ;)",
    "543978": "Good work, thanks for sharing",
    "550037": "Congratulations! Thanks for sharing, I learned a lot from this.",
    "552696": "Thanks for sharing!",
    "561164": "Thank you for sharing your work. I learned a lot from you. \n\nI have one question. What were reasons for you to choose those specific experiments for each model from 11 cycles? Did you kind of make them different to maximize the benefit of blending?",
    "561488": "Nice that it helped you to learn something @hatomugi :)\n\nI selected different subsets of experiments in order to try to match the test distribution while trying to maximize the total number of experiments used to train (matching distributions while avoiding to use too few training data). \nFor example, you can notice that the subset of experiments with the highest weight in the blending (Val Set3), contains 7 experiments used as training data (1, 2, 4, 7, 10, 11 and 14). But, if you consider the experiments I used in total, the different models are trained on 11 experiments (1,2,3,4,7,8,9,10,11,12 and 14).\n\nBesides, I didn't know which of those subsets was the best one to use, so I thought it was better to use all of them ;)",
    "561517": "Oh, I got it. That makes super sense. You tried to adjust the distribution between train and test set and also maximized the range of training data at the same time. Cool idea :) Thanks for the kind explanation."
  },
  "source": "meta"
}