{
  "id": 90330,
  "title": "LANL training data in .ft feather format",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90330",
  "author_name": "",
  "post_date": "2019-04-22T23:53:06.374811900Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>LANL Training Dataset in feather format - loads within 10s (compared to 2m40s using .csv)</p>\n\n<blockquote>\n  <p>import feather\n  train = feather.read_dataframe('../input/lanl-ft/train.ft')</p>\n</blockquote>\n\n<p><a href=\"https://www.kaggle.com/teeyee314/lanl-ft\">https://www.kaggle.com/teeyee314/lanl-ft</a></p>\n\n<ul>\n<li><code>acoustic_data</code>: int16</li>\n<li><code>time_to_failure</code>: float32</li>\n</ul>",
  "messages": [
    {
      "id": "521447",
      "postDate": "04/22/2019 23:53:06",
      "content": "<p>LANL Training Dataset in feather format - loads within 10s (compared to 2m40s using .csv)</p>\n\n<blockquote>\n  <p>import feather\n  train = feather.read_dataframe('../input/lanl-ft/train.ft')</p>\n</blockquote>\n\n<p><a href=\"https://www.kaggle.com/teeyee314/lanl-ft\">https://www.kaggle.com/teeyee314/lanl-ft</a></p>\n\n<ul>\n<li><code>acoustic_data</code>: int16</li>\n<li><code>time_to_failure</code>: float32</li>\n</ul>",
      "rawMarkdown": "LANL Training Dataset in feather format - loads within 10s (compared to 2m40s using .csv)\n\n&gt; import feather\ntrain = feather.read_dataframe('../input/lanl-ft/train.ft')\n\n[https://www.kaggle.com/teeyee314/lanl-ft](https://www.kaggle.com/teeyee314/lanl-ft)\n\n- `acoustic_data`: int16\n- `time_to_failure`: float32",
      "votes": null
    },
    {
      "id": "521573",
      "postDate": "04/23/2019 03:59:08",
      "content": "<p>I think you need float64 for the time to failure or you might loose precision.</p>",
      "rawMarkdown": "I think you need float64 for the time to failure or you might loose precision.",
      "votes": null
    },
    {
      "id": "521630",
      "postDate": "04/23/2019 06:57:32",
      "content": "<p>The predictions will always have errors much larger than this precision. How would it be relevant? You could even ditch the column entirely and just store the indices of <code>time_to_failure = 0</code> (or actually lowest values before they rise again, which is not zero) and reconstruct <code>time_to_failure</code>.</p>",
      "rawMarkdown": "The predictions will always have errors much larger than this precision. How would it be relevant? You could even ditch the column entirely and just store the indices of `time_to_failure = 0` (or actually lowest values before they rise again, which is not zero) and reconstruct `time_to_failure`.",
      "votes": null
    },
    {
      "id": "521640",
      "postDate": "04/23/2019 07:22:22",
      "content": "<p><code>np.sqrt(mean_squared_error(train64.values, train32.values))</code></p>\n\n<p><code>mean_absolute_error(train64.values, train32.values)</code></p>\n\n<p>I'm not a math expert, but the RMSE and MAE calculated above:  1.6711039630961481e-07 and 1.206283880006011e-07 respectively. </p>",
      "rawMarkdown": "`np.sqrt(mean_squared_error(train64.values, train32.values))`\n\n`mean_absolute_error(train64.values, train32.values)`\n\nI'm not a math expert, but the RMSE and MAE calculated above:  1.6711039630961481e-07 and 1.206283880006011e-07 respectively.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 521573,
      "author_name": "jsaguiar",
      "author_url": "",
      "post_date": "04/23/2019 03:59:08",
      "content": "<p>I think you need float64 for the time to failure or you might loose precision.</p>",
      "votes": null,
      "replies": [
        {
          "id": 521630,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "04/23/2019 06:57:32",
          "content": "<p>The predictions will always have errors much larger than this precision. How would it be relevant? You could even ditch the column entirely and just store the indices of <code>time_to_failure = 0</code> (or actually lowest values before they rise again, which is not zero) and reconstruct <code>time_to_failure</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521640,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "04/23/2019 07:22:22",
          "content": "<p><code>np.sqrt(mean_squared_error(train64.values, train32.values))</code></p>\n\n<p><code>mean_absolute_error(train64.values, train32.values)</code></p>\n\n<p>I'm not a math expert, but the RMSE and MAE calculated above:  1.6711039630961481e-07 and 1.206283880006011e-07 respectively. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "521447": "LANL Training Dataset in feather format - loads within 10s (compared to 2m40s using .csv)\n\n&gt; import feather\ntrain = feather.read_dataframe('../input/lanl-ft/train.ft')\n\n[https://www.kaggle.com/teeyee314/lanl-ft](https://www.kaggle.com/teeyee314/lanl-ft)\n\n- `acoustic_data`: int16\n- `time_to_failure`: float32",
    "521573": "I think you need float64 for the time to failure or you might loose precision.",
    "521630": "The predictions will always have errors much larger than this precision. How would it be relevant? You could even ditch the column entirely and just store the indices of `time_to_failure = 0` (or actually lowest values before they rise again, which is not zero) and reconstruct `time_to_failure`.",
    "521640": "`np.sqrt(mean_squared_error(train64.values, train32.values))`\n\n`mean_absolute_error(train64.values, train32.values)`\n\nI'm not a math expert, but the RMSE and MAE calculated above:  1.6711039630961481e-07 and 1.206283880006011e-07 respectively."
  },
  "source": "meta"
}