{
  "id": 94665,
  "title": "179th place / top 5% (not submitted) solution - first competition",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/aline-almeida-179th-place-top-5-not-submitted-solu",
  "author_name": "",
  "post_date": "2019-06-06T03:23:26.170330100Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am not new to Data Science, but that was my first ML competition (apart from 'Homesite Quote Conversion' for testing purposes, three years ago).</p>\n\n<p>I will share the process regarding my best (not submitted) solution. I did not trust my CV and got in the [Earth]shakeup!! </p>\n\n<p><strong>Scores</strong></p>\n\n<p>Public score: 1,86776\nPrivate score: 2,48755 (it would be #179, top 5%; it is worth noting, though, that if everyone else was evaluated by their best solution in private board, the rank would be different). \nAverage CV score: 2,04663 </p>\n\n<p><strong>Features</strong></p>\n\n<ul>\n<li>Got about ~150 basic features from public kernel. </li>\n<li>Generated ~800 new features using tsfresh package. </li>\n<li>Dropped correlated columns with threshold 0.95. </li>\n<li>Decided to work with only 50 features. </li>\n<li>Combined different techniques to rank features, including RFECV.</li>\n<li>Blended my personal random taste with this rank and obtained the following 50 features:\n21 from public kernel: \n<code>\nabs_max_roll_mean_1000, avg_first_10000, avg_last_10000, avg_last_50000, iqr, min_first_10000, min_roll_mean_1000, min_roll_std_100, min_roll_std_1000, q01_roll_mean_100, q01_roll_mean_1000, q01_roll_std_10, q01_roll_std_1000, q05_roll_mean_100, q05_roll_std_100, q05_roll_std_1000, q95_roll_mean_100, q95_roll_mean_1000, q99_roll_mean_100, q99_roll_mean_1000, std_roll_mean_1000\n</code>\n29 from those generated with tsfresh:\n<code>\nabs_energy, agg_linear_trend__f_agg_mean__chunk_len_500__attr_stderr, ar_coefficient__k_10__coeff_0, augmented_dickey_fuller__attr_teststat, autocorrelation__lag_4, c3__lag_2, change_quantiles__f_agg_var__isabs_False__qh_0.4__ql_0.2, change_quantiles__f_agg_var__isabs_False__qh_0.8__ql_0.2, count_below_mean, cwt_coefficients__widths_(2, 5, 10, 20)__coeff_0__w_2, energy_ratio_by_chunks__num_segments_10__segment_focus_4, fft_coefficient__coeff_37__attr_imag, fft_coefficient__coeff_66__attr_abs, fft_coefficient__coeff_95__attr_angle, fft_coefficient__coeff_99__attr_real, number_crossing_m__m_0, number_crossing_m__m_-1, number_cwt_peaks__n_5, number_peaks__n_1, number_peaks__n_500, partial_autocorrelation__lag_8, quantile__q_0.1, quantile__q_0.2, range_count__max_1__min_-1, ratio_beyond_r_sigma__r_5, ratio_value_number_to_time_series_length, spkt_welch_density__coeff_2, sum_of_reoccurring_values, value_count__value_1\n</code></li>\n</ul>\n\n<p><strong>Model</strong>\nObtained the following pipeline, optimized with TPOT:\n<code>\nmake_pipeline(\n    StackingEstimator(estimator=ElasticNetCV(l1_ratio=0.1, tol=0.001)),\n    Normalizer(norm=\"l2\"),\n    Normalizer(norm=\"l1\"),\n    StackingEstimator(estimator=LassoLarsCV(normalize=True)),\n    LinearSVR(C=15.0, dual=True, epsilon=1.0, loss=\"epsilon_insensitive\", tol=1e-05)\n)\n</code></p>\n\n<p><strong>Things I tried that did not work</strong>\nAugmentation, as shared in this kernel: \n<a href=\"https://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting\">https://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting</a></p>\n\n<p>(but I should have tried it again after the new features I generated...)</p>",
  "messages": [
    {
      "id": "545912",
      "postDate": "06/06/2019 03:23:26",
      "content": "<p>I am not new to Data Science, but that was my first ML competition (apart from 'Homesite Quote Conversion' for testing purposes, three years ago).</p>\n\n<p>I will share the process regarding my best (not submitted) solution. I did not trust my CV and got in the [Earth]shakeup!! </p>\n\n<p><strong>Scores</strong></p>\n\n<p>Public score: 1,86776\nPrivate score: 2,48755 (it would be #179, top 5%; it is worth noting, though, that if everyone else was evaluated by their best solution in private board, the rank would be different). \nAverage CV score: 2,04663 </p>\n\n<p><strong>Features</strong></p>\n\n<ul>\n<li>Got about ~150 basic features from public kernel. </li>\n<li>Generated ~800 new features using tsfresh package. </li>\n<li>Dropped correlated columns with threshold 0.95. </li>\n<li>Decided to work with only 50 features. </li>\n<li>Combined different techniques to rank features, including RFECV.</li>\n<li>Blended my personal random taste with this rank and obtained the following 50 features:\n21 from public kernel: \n<code>\nabs_max_roll_mean_1000, avg_first_10000, avg_last_10000, avg_last_50000, iqr, min_first_10000, min_roll_mean_1000, min_roll_std_100, min_roll_std_1000, q01_roll_mean_100, q01_roll_mean_1000, q01_roll_std_10, q01_roll_std_1000, q05_roll_mean_100, q05_roll_std_100, q05_roll_std_1000, q95_roll_mean_100, q95_roll_mean_1000, q99_roll_mean_100, q99_roll_mean_1000, std_roll_mean_1000\n</code>\n29 from those generated with tsfresh:\n<code>\nabs_energy, agg_linear_trend__f_agg_mean__chunk_len_500__attr_stderr, ar_coefficient__k_10__coeff_0, augmented_dickey_fuller__attr_teststat, autocorrelation__lag_4, c3__lag_2, change_quantiles__f_agg_var__isabs_False__qh_0.4__ql_0.2, change_quantiles__f_agg_var__isabs_False__qh_0.8__ql_0.2, count_below_mean, cwt_coefficients__widths_(2, 5, 10, 20)__coeff_0__w_2, energy_ratio_by_chunks__num_segments_10__segment_focus_4, fft_coefficient__coeff_37__attr_imag, fft_coefficient__coeff_66__attr_abs, fft_coefficient__coeff_95__attr_angle, fft_coefficient__coeff_99__attr_real, number_crossing_m__m_0, number_crossing_m__m_-1, number_cwt_peaks__n_5, number_peaks__n_1, number_peaks__n_500, partial_autocorrelation__lag_8, quantile__q_0.1, quantile__q_0.2, range_count__max_1__min_-1, ratio_beyond_r_sigma__r_5, ratio_value_number_to_time_series_length, spkt_welch_density__coeff_2, sum_of_reoccurring_values, value_count__value_1\n</code></li>\n</ul>\n\n<p><strong>Model</strong>\nObtained the following pipeline, optimized with TPOT:\n<code>\nmake_pipeline(\n    StackingEstimator(estimator=ElasticNetCV(l1_ratio=0.1, tol=0.001)),\n    Normalizer(norm=\"l2\"),\n    Normalizer(norm=\"l1\"),\n    StackingEstimator(estimator=LassoLarsCV(normalize=True)),\n    LinearSVR(C=15.0, dual=True, epsilon=1.0, loss=\"epsilon_insensitive\", tol=1e-05)\n)\n</code></p>\n\n<p><strong>Things I tried that did not work</strong>\nAugmentation, as shared in this kernel: \n<a href=\"https://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting\">https://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting</a></p>\n\n<p>(but I should have tried it again after the new features I generated...)</p>",
      "rawMarkdown": "I am not new to Data Science, but that was my first ML competition (apart from 'Homesite Quote Conversion' for testing purposes, three years ago).\n\nI will share the process regarding my best (not submitted) solution. I did not trust my CV and got in the [Earth]shakeup!! \n\n**Scores**\n\nPublic score: 1,86776\nPrivate score: 2,48755 (it would be #179, top 5%; it is worth noting, though, that if everyone else was evaluated by their best solution in private board, the rank would be different). \nAverage CV score: 2,04663 \n\n\n**Features**\n\n- Got about ~150 basic features from public kernel. \n- Generated ~800 new features using tsfresh package. \n- Dropped correlated columns with threshold 0.95. \n- Decided to work with only 50 features. \n- Combined different techniques to rank features, including RFECV.\n- Blended my personal random taste with this rank and obtained the following 50 features:\n21 from public kernel: \n```\nabs_max_roll_mean_1000, avg_first_10000, avg_last_10000, avg_last_50000, iqr, min_first_10000, min_roll_mean_1000, min_roll_std_100, min_roll_std_1000, q01_roll_mean_100, q01_roll_mean_1000, q01_roll_std_10, q01_roll_std_1000, q05_roll_mean_100, q05_roll_std_100, q05_roll_std_1000, q95_roll_mean_100, q95_roll_mean_1000, q99_roll_mean_100, q99_roll_mean_1000, std_roll_mean_1000\n```\n29 from those generated with tsfresh:\n```\nabs_energy, agg_linear_trend__f_agg_mean__chunk_len_500__attr_stderr, ar_coefficient__k_10__coeff_0, augmented_dickey_fuller__attr_teststat, autocorrelation__lag_4, c3__lag_2, change_quantiles__f_agg_var__isabs_False__qh_0.4__ql_0.2, change_quantiles__f_agg_var__isabs_False__qh_0.8__ql_0.2, count_below_mean, cwt_coefficients__widths_(2, 5, 10, 20)__coeff_0__w_2, energy_ratio_by_chunks__num_segments_10__segment_focus_4, fft_coefficient__coeff_37__attr_imag, fft_coefficient__coeff_66__attr_abs, fft_coefficient__coeff_95__attr_angle, fft_coefficient__coeff_99__attr_real, number_crossing_m__m_0, number_crossing_m__m_-1, number_cwt_peaks__n_5, number_peaks__n_1, number_peaks__n_500, partial_autocorrelation__lag_8, quantile__q_0.1, quantile__q_0.2, range_count__max_1__min_-1, ratio_beyond_r_sigma__r_5, ratio_value_number_to_time_series_length, spkt_welch_density__coeff_2, sum_of_reoccurring_values, value_count__value_1\n```\n\n**Model**\nObtained the following pipeline, optimized with TPOT:\n```\nmake_pipeline(\n    StackingEstimator(estimator=ElasticNetCV(l1_ratio=0.1, tol=0.001)),\n    Normalizer(norm=\"l2\"),\n    Normalizer(norm=\"l1\"),\n    StackingEstimator(estimator=LassoLarsCV(normalize=True)),\n    LinearSVR(C=15.0, dual=True, epsilon=1.0, loss=\"epsilon_insensitive\", tol=1e-05)\n)\n```\n\n\n\n**Things I tried that did not work**\nAugmentation, as shared in this kernel: \nhttps://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting\n\n(but I should have tried it again after the new features I generated...)",
      "votes": null
    },
    {
      "id": "546044",
      "postDate": "06/06/2019 07:42:06",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "546142",
      "postDate": "06/06/2019 09:44:31",
      "content": "<p>Well done. Good luck for your next competitions. </p>",
      "rawMarkdown": "Well done. Good luck for your next competitions.",
      "votes": null
    },
    {
      "id": "546352",
      "postDate": "06/06/2019 14:02:15",
      "content": "<p>Thanks <a href=\"/abhinand05\">@abhinand05</a> ! See you in the future</p>",
      "rawMarkdown": "Thanks @abhinand05 ! See you in the future",
      "votes": null
    },
    {
      "id": "546353",
      "postDate": "06/06/2019 14:02:44",
      "content": "<p>Thanks  <a href=\"/dhaqui\">@dhaqui</a> !  See you in the future</p>",
      "rawMarkdown": "Thanks  @dhaqui !  See you in the future",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 546044,
      "author_name": "dhaqui",
      "author_url": "",
      "post_date": "06/06/2019 07:42:06",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 546353,
          "author_name": "alinealmeida",
          "author_url": "",
          "post_date": "06/06/2019 14:02:44",
          "content": "<p>Thanks  <a href=\"/dhaqui\">@dhaqui</a> !  See you in the future</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546142,
      "author_name": "abhinand05",
      "author_url": "",
      "post_date": "06/06/2019 09:44:31",
      "content": "<p>Well done. Good luck for your next competitions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 546352,
          "author_name": "alinealmeida",
          "author_url": "",
          "post_date": "06/06/2019 14:02:15",
          "content": "<p>Thanks <a href=\"/abhinand05\">@abhinand05</a> ! See you in the future</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "545912": "I am not new to Data Science, but that was my first ML competition (apart from 'Homesite Quote Conversion' for testing purposes, three years ago).\n\nI will share the process regarding my best (not submitted) solution. I did not trust my CV and got in the [Earth]shakeup!! \n\n**Scores**\n\nPublic score: 1,86776\nPrivate score: 2,48755 (it would be #179, top 5%; it is worth noting, though, that if everyone else was evaluated by their best solution in private board, the rank would be different). \nAverage CV score: 2,04663 \n\n\n**Features**\n\n- Got about ~150 basic features from public kernel. \n- Generated ~800 new features using tsfresh package. \n- Dropped correlated columns with threshold 0.95. \n- Decided to work with only 50 features. \n- Combined different techniques to rank features, including RFECV.\n- Blended my personal random taste with this rank and obtained the following 50 features:\n21 from public kernel: \n```\nabs_max_roll_mean_1000, avg_first_10000, avg_last_10000, avg_last_50000, iqr, min_first_10000, min_roll_mean_1000, min_roll_std_100, min_roll_std_1000, q01_roll_mean_100, q01_roll_mean_1000, q01_roll_std_10, q01_roll_std_1000, q05_roll_mean_100, q05_roll_std_100, q05_roll_std_1000, q95_roll_mean_100, q95_roll_mean_1000, q99_roll_mean_100, q99_roll_mean_1000, std_roll_mean_1000\n```\n29 from those generated with tsfresh:\n```\nabs_energy, agg_linear_trend__f_agg_mean__chunk_len_500__attr_stderr, ar_coefficient__k_10__coeff_0, augmented_dickey_fuller__attr_teststat, autocorrelation__lag_4, c3__lag_2, change_quantiles__f_agg_var__isabs_False__qh_0.4__ql_0.2, change_quantiles__f_agg_var__isabs_False__qh_0.8__ql_0.2, count_below_mean, cwt_coefficients__widths_(2, 5, 10, 20)__coeff_0__w_2, energy_ratio_by_chunks__num_segments_10__segment_focus_4, fft_coefficient__coeff_37__attr_imag, fft_coefficient__coeff_66__attr_abs, fft_coefficient__coeff_95__attr_angle, fft_coefficient__coeff_99__attr_real, number_crossing_m__m_0, number_crossing_m__m_-1, number_cwt_peaks__n_5, number_peaks__n_1, number_peaks__n_500, partial_autocorrelation__lag_8, quantile__q_0.1, quantile__q_0.2, range_count__max_1__min_-1, ratio_beyond_r_sigma__r_5, ratio_value_number_to_time_series_length, spkt_welch_density__coeff_2, sum_of_reoccurring_values, value_count__value_1\n```\n\n**Model**\nObtained the following pipeline, optimized with TPOT:\n```\nmake_pipeline(\n    StackingEstimator(estimator=ElasticNetCV(l1_ratio=0.1, tol=0.001)),\n    Normalizer(norm=\"l2\"),\n    Normalizer(norm=\"l1\"),\n    StackingEstimator(estimator=LassoLarsCV(normalize=True)),\n    LinearSVR(C=15.0, dual=True, epsilon=1.0, loss=\"epsilon_insensitive\", tol=1e-05)\n)\n```\n\n\n\n**Things I tried that did not work**\nAugmentation, as shared in this kernel: \nhttps://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting\n\n(but I should have tried it again after the new features I generated...)",
    "546044": "Congrats!",
    "546142": "Well done. Good luck for your next competitions.",
    "546352": "Thanks @abhinand05 ! See you in the future",
    "546353": "Thanks  @dhaqui !  See you in the future"
  },
  "source": "meta"
}