{
  "id": 94609,
  "title": "9th place solution",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/mykper-9th-place-solution",
  "author_name": "",
  "post_date": "2019-06-05T18:36:00.503Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>After discovery of possible (at that point) p4677 related leak, I decided to drop from competition. Heading into last week of competition I had clear idea how I would exploit it. This post <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94086\">Will this decide who wins?</a> and the fact I had nothing to lose (dropped to ~1K public place) if assumption about test set distribution was wrong I decided to implement the idea.</p>\n\n<p>I wanted to use Leave K-EQ out type of CV scheme, but thanks to <a href=\"/cpmpml\">@cpmpml</a> we know that difference between mean of validation and train could lead to under/overfit, so if mean of all EQ is the same, I think, it might work.</p>\n\n<p>I transformed target to normalized time to failure multiplied by twice the estimated mean of private set. I had 2 values (2 submissions): 6.3 and 5.9. In hindsight 5.9 was way off, but my estimation process was quiet crude. Late submission with higher value have slightly better score.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544535/13414/target.png\" alt=\"target\"></p>\n\n<p>I chose LGBM, as I had been using it prior to leak discovery. \nFeatures were peaked from public kernel <a href=\"https://www.kaggle.com/artgor/even-more-features\">\"Even more features\"</a> by <a href=\"/artgor\">@artgor</a>, according to past models feature importance and <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93148\">\"My Top 30 Features\"</a> by <a href=\"/scirpus\">@scirpus</a>\n<code>\nnum_peaks_10,\nnum_crossing_0,\npercentile_roll_std_5_window_100,\nabs_percentile_80, \nfftr_percentile_roll_std_80_window_10000\n</code>\nFolds were generated using this code:\n<code>for val_t in itertools.combinations(range(15), 2):</code>\nwhere <code>val_t</code> is two EQ indexes for validation. Only full EQ were used.\nAll these were enough for 9th place.</p>\n\n<p>More interesting for me were ideas that didn't work.\nOne of them was effort to predict not ttf, but a pair of period of EQ and normalized time to failure. I did try siamese network based on CNN1D with similarity metric:\n<code>np.exp(-np.abs(right_period - left_period)/c1) * np.exp(-np.abs(right_norm_ttf - left_norm_ttf)/c2)</code>\nconstants <code>c1</code> and <code>c2</code> were to limit points that are close enough (hyperparameters). I did not fully explore this solution, but it was fun to work on.\nI did try other approaches but I strongly believe that period of EQ is not predictable.</p>",
  "messages": [
    {
      "id": "544535",
      "postDate": "06/05/2019 16:14:48",
      "content": "<p>After discovery of possible (at that point) p4677 related leak, I decided to drop from competition. Heading into last week of competition I had clear idea how I would exploit it. This post <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94086\">Will this decide who wins?</a> and the fact I had nothing to lose (dropped to ~1K public place) if assumption about test set distribution was wrong I decided to implement the idea.</p>\n\n<p>I wanted to use Leave K-EQ out type of CV scheme, but thanks to <a href=\"/cpmpml\">@cpmpml</a> we know that difference between mean of validation and train could lead to under/overfit, so if mean of all EQ is the same, I think, it might work.</p>\n\n<p>I transformed target to normalized time to failure multiplied by twice the estimated mean of private set. I had 2 values (2 submissions): 6.3 and 5.9. In hindsight 5.9 was way off, but my estimation process was quiet crude. Late submission with higher value have slightly better score.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544535/13414/target.png\" alt=\"target\"></p>\n\n<p>I chose LGBM, as I had been using it prior to leak discovery. \nFeatures were peaked from public kernel <a href=\"https://www.kaggle.com/artgor/even-more-features\">\"Even more features\"</a> by <a href=\"/artgor\">@artgor</a>, according to past models feature importance and <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93148\">\"My Top 30 Features\"</a> by <a href=\"/scirpus\">@scirpus</a>\n<code>\nnum_peaks_10,\nnum_crossing_0,\npercentile_roll_std_5_window_100,\nabs_percentile_80, \nfftr_percentile_roll_std_80_window_10000\n</code>\nFolds were generated using this code:\n<code>for val_t in itertools.combinations(range(15), 2):</code>\nwhere <code>val_t</code> is two EQ indexes for validation. Only full EQ were used.\nAll these were enough for 9th place.</p>\n\n<p>More interesting for me were ideas that didn't work.\nOne of them was effort to predict not ttf, but a pair of period of EQ and normalized time to failure. I did try siamese network based on CNN1D with similarity metric:\n<code>np.exp(-np.abs(right_period - left_period)/c1) * np.exp(-np.abs(right_norm_ttf - left_norm_ttf)/c2)</code>\nconstants <code>c1</code> and <code>c2</code> were to limit points that are close enough (hyperparameters). I did not fully explore this solution, but it was fun to work on.\nI did try other approaches but I strongly believe that period of EQ is not predictable.</p>",
      "rawMarkdown": "After discovery of possible (at that point) p4677 related leak, I decided to drop from competition. Heading into last week of competition I had clear idea how I would exploit it. This post [Will this decide who wins?](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94086) and the fact I had nothing to lose (dropped to ~1K public place) if assumption about test set distribution was wrong I decided to implement the idea.\n\nI wanted to use Leave K-EQ out type of CV scheme, but thanks to @cpmpml we know that difference between mean of validation and train could lead to under/overfit, so if mean of all EQ is the same, I think, it might work.\n\nI transformed target to normalized time to failure multiplied by twice the estimated mean of private set. I had 2 values (2 submissions): 6.3 and 5.9. In hindsight 5.9 was way off, but my estimation process was quiet crude. Late submission with higher value have slightly better score.\n\n![target](https://storage.googleapis.com/kaggle-forum-message-attachments/544535/13414/target.png)\n\nI chose LGBM, as I had been using it prior to leak discovery. \nFeatures were peaked from public kernel [\"Even more features\"](https://www.kaggle.com/artgor/even-more-features) by @artgor, according to past models feature importance and [\"My Top 30 Features\"](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93148) by @scirpus\n```\nnum_peaks_10,\nnum_crossing_0,\npercentile_roll_std_5_window_100,\nabs_percentile_80, \nfftr_percentile_roll_std_80_window_10000\n```\nFolds were generated using this code:\n`for val_t in itertools.combinations(range(15), 2):`\nwhere `val_t` is two EQ indexes for validation. Only full EQ were used.\nAll these were enough for 9th place.\n\nMore interesting for me were ideas that didn't work.\nOne of them was effort to predict not ttf, but a pair of period of EQ and normalized time to failure. I did try siamese network based on CNN1D with similarity metric:\n`np.exp(-np.abs(right_period - left_period)/c1) * np.exp(-np.abs(right_norm_ttf - left_norm_ttf)/c2)`\nconstants `c1` and `c2` were to limit points that are close enough (hyperparameters). I did not fully explore this solution, but it was fun to work on.\nI did try other approaches but I strongly believe that period of EQ is not predictable.",
      "votes": null
    },
    {
      "id": "545993",
      "postDate": "06/06/2019 05:46:38",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 545993,
      "author_name": "timmmmmms",
      "author_url": "",
      "post_date": "06/06/2019 05:46:38",
      "content": "<p>Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "544535": "After discovery of possible (at that point) p4677 related leak, I decided to drop from competition. Heading into last week of competition I had clear idea how I would exploit it. This post [Will this decide who wins?](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94086) and the fact I had nothing to lose (dropped to ~1K public place) if assumption about test set distribution was wrong I decided to implement the idea.\n\nI wanted to use Leave K-EQ out type of CV scheme, but thanks to @cpmpml we know that difference between mean of validation and train could lead to under/overfit, so if mean of all EQ is the same, I think, it might work.\n\nI transformed target to normalized time to failure multiplied by twice the estimated mean of private set. I had 2 values (2 submissions): 6.3 and 5.9. In hindsight 5.9 was way off, but my estimation process was quiet crude. Late submission with higher value have slightly better score.\n\n![target](https://storage.googleapis.com/kaggle-forum-message-attachments/544535/13414/target.png)\n\nI chose LGBM, as I had been using it prior to leak discovery. \nFeatures were peaked from public kernel [\"Even more features\"](https://www.kaggle.com/artgor/even-more-features) by @artgor, according to past models feature importance and [\"My Top 30 Features\"](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93148) by @scirpus\n```\nnum_peaks_10,\nnum_crossing_0,\npercentile_roll_std_5_window_100,\nabs_percentile_80, \nfftr_percentile_roll_std_80_window_10000\n```\nFolds were generated using this code:\n`for val_t in itertools.combinations(range(15), 2):`\nwhere `val_t` is two EQ indexes for validation. Only full EQ were used.\nAll these were enough for 9th place.\n\nMore interesting for me were ideas that didn't work.\nOne of them was effort to predict not ttf, but a pair of period of EQ and normalized time to failure. I did try siamese network based on CNN1D with similarity metric:\n`np.exp(-np.abs(right_period - left_period)/c1) * np.exp(-np.abs(right_norm_ttf - left_norm_ttf)/c2)`\nconstants `c1` and `c2` were to limit points that are close enough (hyperparameters). I did not fully explore this solution, but it was fun to work on.\nI did try other approaches but I strongly believe that period of EQ is not predictable.",
    "545993": "Thanks!"
  },
  "source": "meta"
}