{
  "id": 94357,
  "title": "Data augmentation tricks",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94357",
  "author_name": "",
  "post_date": "2019-06-04T04:06:33.269495400Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I used a few a quick and dirty post processing trick to get a few extra points (not used in final submission). </p>\n\n<ol>\n<li>make sure to np.clip(<code>test_predictions</code>, a_min=0, a_max=None) ← I used this a lot</li>\n<li>set test segments containing max acoustic_data values &gt; 3000 to 0.29 ← I used this once</li>\n</ol>\n\n<blockquote>\n  <p><code>max_values = np.apply_along_axis(np.max, 1, test_data)</code>\n   <code>mask = np.argwhere(max_values &amp;gt; 3000)</code>\n   <code>for idx in mask:</code>\n   &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<code>submission['time_to_failure'][idx] = 0.29</code></p>\n</blockquote>\n\n<p>Just thought I'd share a few of my tricks. Why 3000? I noticed that anything greater than 3000 was likely low TTF. I then zoomed into train segments and each spike &gt; 3000 acoustic was between 0.292 - 0.286. I thought about adding those 9 segments into training but I figured hard post processing would work just as well. Why clip the lowest test TTF to 0? Well I saw some of my models predicting &lt; 0 and I thought it would make sense to keep predicted TTF &gt;=0. </p>\n\n<p>Did anyone else use fun or useful tricks? Please share!</p>",
  "messages": [
    {
      "id": "542724",
      "postDate": "06/04/2019 04:06:33",
      "content": "<p>I used a few a quick and dirty post processing trick to get a few extra points (not used in final submission). </p>\n\n<ol>\n<li>make sure to np.clip(<code>test_predictions</code>, a_min=0, a_max=None) ← I used this a lot</li>\n<li>set test segments containing max acoustic_data values &gt; 3000 to 0.29 ← I used this once</li>\n</ol>\n\n<blockquote>\n  <p><code>max_values = np.apply_along_axis(np.max, 1, test_data)</code>\n   <code>mask = np.argwhere(max_values &amp;gt; 3000)</code>\n   <code>for idx in mask:</code>\n   &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<code>submission['time_to_failure'][idx] = 0.29</code></p>\n</blockquote>\n\n<p>Just thought I'd share a few of my tricks. Why 3000? I noticed that anything greater than 3000 was likely low TTF. I then zoomed into train segments and each spike &gt; 3000 acoustic was between 0.292 - 0.286. I thought about adding those 9 segments into training but I figured hard post processing would work just as well. Why clip the lowest test TTF to 0? Well I saw some of my models predicting &lt; 0 and I thought it would make sense to keep predicted TTF &gt;=0. </p>\n\n<p>Did anyone else use fun or useful tricks? Please share!</p>",
      "rawMarkdown": "I used a few a quick and dirty post processing trick to get a few extra points (not used in final submission). \n\n1. make sure to np.clip(`test_predictions`, a_min=0, a_max=None) ← I used this a lot\n2. set test segments containing max acoustic_data values &gt; 3000 to 0.29 ← I used this once\n\n&gt;  `max_values = np.apply_along_axis(np.max, 1, test_data)`\n&gt;  `mask = np.argwhere(max_values &gt; 3000)`\n&gt;  `for idx in mask:`\n&gt;  &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;`submission['time_to_failure'][idx] = 0.29`\n\nJust thought I'd share a few of my tricks. Why 3000? I noticed that anything greater than 3000 was likely low TTF. I then zoomed into train segments and each spike &gt; 3000 acoustic was between 0.292 - 0.286. I thought about adding those 9 segments into training but I figured hard post processing would work just as well. Why clip the lowest test TTF to 0? Well I saw some of my models predicting &lt; 0 and I thought it would make sense to keep predicted TTF &gt;=0. \n\nDid anyone else use fun or useful tricks? Please share!",
      "votes": null
    },
    {
      "id": "542839",
      "postDate": "06/04/2019 06:18:56",
      "content": "<p>Nice tips.\nOr you could use ELU (exponential linear unit) to cap your output for better training, or manually do this (somehow I am not so comfortably doing so based on mathematical reasons) by <code>pred[pred&lt;0] = np.exp(pred[pred&lt;0]) - 1</code>.</p>",
      "rawMarkdown": "Nice tips.\nOr you could use ELU (exponential linear unit) to cap your output for better training, or manually do this (somehow I am not so comfortably doing so based on mathematical reasons) by `pred[pred&lt;0] = np.exp(pred[pred&lt;0]) - 1`.",
      "votes": null
    },
    {
      "id": "544091",
      "postDate": "06/05/2019 05:43:18",
      "content": "<p>Thanks for sharing. I also double-checked that all the predictions were positive. However, I didn't dare to set a value when a huge peak was present (<code>max_values &gt; 3000</code>).</p>",
      "rawMarkdown": "Thanks for sharing. I also double-checked that all the predictions were positive. However, I didn't dare to set a value when a huge peak was present (`max_values &gt; 3000`).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 542839,
      "author_name": "scaomath",
      "author_url": "",
      "post_date": "06/04/2019 06:18:56",
      "content": "<p>Nice tips.\nOr you could use ELU (exponential linear unit) to cap your output for better training, or manually do this (somehow I am not so comfortably doing so based on mathematical reasons) by <code>pred[pred&lt;0] = np.exp(pred[pred&lt;0]) - 1</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544091,
      "author_name": "ricarddelgado",
      "author_url": "",
      "post_date": "06/05/2019 05:43:18",
      "content": "<p>Thanks for sharing. I also double-checked that all the predictions were positive. However, I didn't dare to set a value when a huge peak was present (<code>max_values &gt; 3000</code>).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542724": "I used a few a quick and dirty post processing trick to get a few extra points (not used in final submission). \n\n1. make sure to np.clip(`test_predictions`, a_min=0, a_max=None) ← I used this a lot\n2. set test segments containing max acoustic_data values &gt; 3000 to 0.29 ← I used this once\n\n&gt;  `max_values = np.apply_along_axis(np.max, 1, test_data)`\n&gt;  `mask = np.argwhere(max_values &gt; 3000)`\n&gt;  `for idx in mask:`\n&gt;  &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;`submission['time_to_failure'][idx] = 0.29`\n\nJust thought I'd share a few of my tricks. Why 3000? I noticed that anything greater than 3000 was likely low TTF. I then zoomed into train segments and each spike &gt; 3000 acoustic was between 0.292 - 0.286. I thought about adding those 9 segments into training but I figured hard post processing would work just as well. Why clip the lowest test TTF to 0? Well I saw some of my models predicting &lt; 0 and I thought it would make sense to keep predicted TTF &gt;=0. \n\nDid anyone else use fun or useful tricks? Please share!",
    "542839": "Nice tips.\nOr you could use ELU (exponential linear unit) to cap your output for better training, or manually do this (somehow I am not so comfortably doing so based on mathematical reasons) by `pred[pred&lt;0] = np.exp(pred[pred&lt;0]) - 1`.",
    "544091": "Thanks for sharing. I also double-checked that all the predictions were positive. However, I didn't dare to set a value when a huge peak was present (`max_values &gt; 3000`)."
  },
  "source": "meta"
}