{
  "id": 83746,
  "title": "Validation - splitting on earthquakes",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/83746",
  "author_name": "",
  "post_date": "2019-03-12T15:18:51.733720600Z",
  "votes": 32,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I might be little late to the party, but I've been looking into the validation setup. Several people floated the idea of splitting by earthquake, so the following snippet might come in handy (haven't seen it anywhere in the kernels):</p>\n\n<p>```\nxtrain = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})</p>\n\n<h1>detect points corresponding to quake occurrence</h1>\n\n<p>quake_list = np.where(np.diff(xtrain['time_to_failure']) &gt; 5)[0]\nxtrain['ind'] = 0; xtrain['ind'].iloc[quake_list] =1</p>\n\n<h1>quake number in sequence</h1>\n\n<p>xtrain['quake_ind'] = xtrain['ind'].cumsum()\nxtrain.drop('ind', axis = 1, inplace = True)</p>\n\n<h1>fold for validation</h1>\n\n<p>xtrain['fold_id'] = (xtrain['quake_ind'] % 5 ) \n```</p>",
  "messages": [
    {
      "id": "488511",
      "postDate": "03/12/2019 15:18:51",
      "content": "<p>I might be little late to the party, but I've been looking into the validation setup. Several people floated the idea of splitting by earthquake, so the following snippet might come in handy (haven't seen it anywhere in the kernels):</p>\n\n<p>```\nxtrain = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})</p>\n\n<h1>detect points corresponding to quake occurrence</h1>\n\n<p>quake_list = np.where(np.diff(xtrain['time_to_failure']) &gt; 5)[0]\nxtrain['ind'] = 0; xtrain['ind'].iloc[quake_list] =1</p>\n\n<h1>quake number in sequence</h1>\n\n<p>xtrain['quake_ind'] = xtrain['ind'].cumsum()\nxtrain.drop('ind', axis = 1, inplace = True)</p>\n\n<h1>fold for validation</h1>\n\n<p>xtrain['fold_id'] = (xtrain['quake_ind'] % 5 ) \n```</p>",
      "rawMarkdown": "I might be little late to the party, but I've been looking into the validation setup. Several people floated the idea of splitting by earthquake, so the following snippet might come in handy (haven't seen it anywhere in the kernels):\n\n```\nxtrain = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\n# detect points corresponding to quake occurrence\nquake_list = np.where(np.diff(xtrain['time_to_failure']) &gt; 5)[0]\nxtrain['ind'] = 0; xtrain['ind'].iloc[quake_list] =1\n# quake number in sequence\nxtrain['quake_ind'] = xtrain['ind'].cumsum()\nxtrain.drop('ind', axis = 1, inplace = True)\n# fold for validation\nxtrain['fold_id'] = (xtrain['quake_ind'] % 5 ) \n```",
      "votes": null
    },
    {
      "id": "488651",
      "postDate": "03/12/2019 20:08:04",
      "content": "<p>Thanks for the snippet Konrad!</p>\n\n<p>I first expected to have an important change on the low part of the distribution of <code>time_to_failure</code>. But actually, after cutting the signal in contiguous segments within each earthquake, the final distribution was pretty similar and the low values are still well represented.</p>",
      "rawMarkdown": "Thanks for the snippet Konrad!\n\nI first expected to have an important change on the low part of the distribution of `time_to_failure`. But actually, after cutting the signal in contiguous segments within each earthquake, the final distribution was pretty similar and the low values are still well represented.",
      "votes": null
    },
    {
      "id": "489152",
      "postDate": "03/13/2019 15:05:11",
      "content": "<p>Can you share this with us?</p>",
      "rawMarkdown": "Can you share this with us?",
      "votes": null
    },
    {
      "id": "489211",
      "postDate": "03/13/2019 16:23:37",
      "content": "<p>Sure! I just put it here: <a href=\"https://www.kaggle.com/ricarddelgado/lanl-sampling-schemes\">https://www.kaggle.com/ricarddelgado/lanl-sampling-schemes</a></p>",
      "rawMarkdown": "Sure! I just put it here: https://www.kaggle.com/ricarddelgado/lanl-sampling-schemes",
      "votes": null
    },
    {
      "id": "516495",
      "postDate": "04/14/2019 09:45:22",
      "content": "<p>Thanks for sharing the snippet <a href=\"/konradb\">@konradb</a> !</p>",
      "rawMarkdown": "Thanks for sharing the snippet @konradb !",
      "votes": null
    },
    {
      "id": "517942",
      "postDate": "04/16/2019 17:26:47",
      "content": "<p>This is exactly what I've been searching for.  Thank you so much for sharing, Konrad!</p>",
      "rawMarkdown": "This is exactly what I've been searching for.  Thank you so much for sharing, Konrad!",
      "votes": null
    },
    {
      "id": "520298",
      "postDate": "04/20/2019 16:41:44",
      "content": "<p>Thanks for sharing this snippet. 👍 </p>",
      "rawMarkdown": "Thanks for sharing this snippet. 👍",
      "votes": null
    },
    {
      "id": "526620",
      "postDate": "05/03/2019 12:24:20",
      "content": "<p>Hi! np.float32 data type does not provide a sufficient precision in order to discriminate between many subsequent samples. Did someone notice the same issue?</p>",
      "rawMarkdown": "Hi! np.float32 data type does not provide a sufficient precision in order to discriminate between many subsequent samples. Did someone notice the same issue?",
      "votes": null
    },
    {
      "id": "526630",
      "postDate": "05/03/2019 12:35:12",
      "content": "<p>Why is it an issue?  And if it is an issue then use float64 ;)</p>",
      "rawMarkdown": "Why is it an issue?  And if it is an issue then use float64 ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 488651,
      "author_name": "ricarddelgado",
      "author_url": "",
      "post_date": "03/12/2019 20:08:04",
      "content": "<p>Thanks for the snippet Konrad!</p>\n\n<p>I first expected to have an important change on the low part of the distribution of <code>time_to_failure</code>. But actually, after cutting the signal in contiguous segments within each earthquake, the final distribution was pretty similar and the low values are still well represented.</p>",
      "votes": null,
      "replies": [
        {
          "id": 489152,
          "author_name": "hmcranbercourt",
          "author_url": "",
          "post_date": "03/13/2019 15:05:11",
          "content": "<p>Can you share this with us?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 489211,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "03/13/2019 16:23:37",
          "content": "<p>Sure! I just put it here: <a href=\"https://www.kaggle.com/ricarddelgado/lanl-sampling-schemes\">https://www.kaggle.com/ricarddelgado/lanl-sampling-schemes</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 516495,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "04/14/2019 09:45:22",
      "content": "<p>Thanks for sharing the snippet <a href=\"/konradb\">@konradb</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517942,
      "author_name": "adeprince3",
      "author_url": "",
      "post_date": "04/16/2019 17:26:47",
      "content": "<p>This is exactly what I've been searching for.  Thank you so much for sharing, Konrad!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 520298,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "04/20/2019 16:41:44",
      "content": "<p>Thanks for sharing this snippet. 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526620,
      "author_name": "matteosangiorgio",
      "author_url": "",
      "post_date": "05/03/2019 12:24:20",
      "content": "<p>Hi! np.float32 data type does not provide a sufficient precision in order to discriminate between many subsequent samples. Did someone notice the same issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 526630,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 12:35:12",
          "content": "<p>Why is it an issue?  And if it is an issue then use float64 ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "488511": "I might be little late to the party, but I've been looking into the validation setup. Several people floated the idea of splitting by earthquake, so the following snippet might come in handy (haven't seen it anywhere in the kernels):\n\n```\nxtrain = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\n# detect points corresponding to quake occurrence\nquake_list = np.where(np.diff(xtrain['time_to_failure']) &gt; 5)[0]\nxtrain['ind'] = 0; xtrain['ind'].iloc[quake_list] =1\n# quake number in sequence\nxtrain['quake_ind'] = xtrain['ind'].cumsum()\nxtrain.drop('ind', axis = 1, inplace = True)\n# fold for validation\nxtrain['fold_id'] = (xtrain['quake_ind'] % 5 ) \n```",
    "488651": "Thanks for the snippet Konrad!\n\nI first expected to have an important change on the low part of the distribution of `time_to_failure`. But actually, after cutting the signal in contiguous segments within each earthquake, the final distribution was pretty similar and the low values are still well represented.",
    "489152": "Can you share this with us?",
    "489211": "Sure! I just put it here: https://www.kaggle.com/ricarddelgado/lanl-sampling-schemes",
    "516495": "Thanks for sharing the snippet @konradb !",
    "517942": "This is exactly what I've been searching for.  Thank you so much for sharing, Konrad!",
    "520298": "Thanks for sharing this snippet. 👍",
    "526620": "Hi! np.float32 data type does not provide a sufficient precision in order to discriminate between many subsequent samples. Did someone notice the same issue?",
    "526630": "Why is it an issue?  And if it is an issue then use float64 ;)"
  },
  "source": "meta"
}