{
  "id": 89819,
  "title": "Total time of training set",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89819",
  "author_name": "",
  "post_date": "2019-04-17T22:33:09.339216700Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Has anyone calculated the total time duration of the training set? Just curious.</p>",
  "messages": [
    {
      "id": "518832",
      "postDate": "04/17/2019 22:33:09",
      "content": "<p>Has anyone calculated the total time duration of the training set? Just curious.</p>",
      "rawMarkdown": "Has anyone calculated the total time duration of the training set? Just curious.",
      "votes": null
    },
    {
      "id": "518956",
      "postDate": "04/18/2019 06:19:59",
      "content": "<p>length of test * duration of 1 frame:\n0.0392*2562=100.4 seconds</p>",
      "rawMarkdown": "length of test * duration of 1 frame:\n0.0392*2562=100.4 seconds",
      "votes": null
    },
    {
      "id": "518989",
      "postDate": "04/18/2019 07:46:37",
      "content": "<p>You forgot to add the delay between each chunk of 4096 observations.</p>",
      "rawMarkdown": "You forgot to add the delay between each chunk of 4096 observations.",
      "votes": null
    },
    {
      "id": "519005",
      "postDate": "04/18/2019 08:10:40",
      "content": "<p>Nope! 0.0392 is even a bit too long: I sample this time directly from the bigger chunks.\n<code>\ntrain = pd.read_hdf('../input/train.hdf', 'table',dtype={'acoustic_data': np.float32, 'time_to_failure': np.float64})\nprint((train[\"time_to_failure\"][14*150_000]-train[\"time_to_failure\"][0*150_000])/14)\n-0.03890736662714285\n</code></p>",
      "rawMarkdown": "Nope! 0.0392 is even a bit too long: I sample this time directly from the bigger chunks.\n```\ntrain = pd.read_hdf('../input/train.hdf', 'table',dtype={'acoustic_data': np.float32, 'time_to_failure': np.float64})\nprint((train[\"time_to_failure\"][14*150_000]-train[\"time_to_failure\"][0*150_000])/14)\n-0.03890736662714285\n```",
      "votes": null
    },
    {
      "id": "519030",
      "postDate": "04/18/2019 08:55:15",
      "content": "<p>Here is the correct way to compute it:</p>\n\n<pre><code>df = train[train.time_to_failure &amp;gt; train.time_to_failure.shift().fillna(0)]\ndf.time_to_failure.sum() - train.time_to_failure.values[-1]\n</code></pre>\n\n<p>result is 163.42061</p>",
      "rawMarkdown": "Here is the correct way to compute it:\n\n    df = train[train.time_to_failure &gt; train.time_to_failure.shift().fillna(0)]\n    df.time_to_failure.sum() - train.time_to_failure.values[-1]\n\nresult is 163.42061",
      "votes": null
    },
    {
      "id": "519032",
      "postDate": "04/18/2019 09:00:35",
      "content": "<p>oops I confused test with train</p>",
      "rawMarkdown": "oops I confused test with train",
      "votes": null
    },
    {
      "id": "519045",
      "postDate": "04/18/2019 09:19:07",
      "content": "<p>Even with that your computation is off by 1 second or so.</p>",
      "rawMarkdown": "Even with that your computation is off by 1 second or so.",
      "votes": null
    },
    {
      "id": "519247",
      "postDate": "04/18/2019 15:55:08",
      "content": "<p>Thanks, I wonder if training should be optimized on few second intervals rather than the exact same size as the test set. What would be the pros and cons in such case..  </p>",
      "rawMarkdown": "Thanks, I wonder if training should be optimized on few second intervals rather than the exact same size as the test set. What would be the pros and cons in such case..",
      "votes": null
    },
    {
      "id": "519287",
      "postDate": "04/18/2019 17:07:24",
      "content": "<p>There is only one way to know: try and see what works best.</p>",
      "rawMarkdown": "There is only one way to know: try and see what works best.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 518956,
      "author_name": "ralphy",
      "author_url": "",
      "post_date": "04/18/2019 06:19:59",
      "content": "<p>length of test * duration of 1 frame:\n0.0392*2562=100.4 seconds</p>",
      "votes": null,
      "replies": [
        {
          "id": 518989,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/18/2019 07:46:37",
          "content": "<p>You forgot to add the delay between each chunk of 4096 observations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519005,
          "author_name": "ralphy",
          "author_url": "",
          "post_date": "04/18/2019 08:10:40",
          "content": "<p>Nope! 0.0392 is even a bit too long: I sample this time directly from the bigger chunks.\n<code>\ntrain = pd.read_hdf('../input/train.hdf', 'table',dtype={'acoustic_data': np.float32, 'time_to_failure': np.float64})\nprint((train[\"time_to_failure\"][14*150_000]-train[\"time_to_failure\"][0*150_000])/14)\n-0.03890736662714285\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519030,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/18/2019 08:55:15",
          "content": "<p>Here is the correct way to compute it:</p>\n\n<pre><code>df = train[train.time_to_failure &amp;gt; train.time_to_failure.shift().fillna(0)]\ndf.time_to_failure.sum() - train.time_to_failure.values[-1]\n</code></pre>\n\n<p>result is 163.42061</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 519032,
      "author_name": "ralphy",
      "author_url": "",
      "post_date": "04/18/2019 09:00:35",
      "content": "<p>oops I confused test with train</p>",
      "votes": null,
      "replies": [
        {
          "id": 519045,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/18/2019 09:19:07",
          "content": "<p>Even with that your computation is off by 1 second or so.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 519247,
      "author_name": "rsaund",
      "author_url": "",
      "post_date": "04/18/2019 15:55:08",
      "content": "<p>Thanks, I wonder if training should be optimized on few second intervals rather than the exact same size as the test set. What would be the pros and cons in such case..  </p>",
      "votes": null,
      "replies": [
        {
          "id": 519287,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/18/2019 17:07:24",
          "content": "<p>There is only one way to know: try and see what works best.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "518832": "Has anyone calculated the total time duration of the training set? Just curious.",
    "518956": "length of test * duration of 1 frame:\n0.0392*2562=100.4 seconds",
    "518989": "You forgot to add the delay between each chunk of 4096 observations.",
    "519005": "Nope! 0.0392 is even a bit too long: I sample this time directly from the bigger chunks.\n```\ntrain = pd.read_hdf('../input/train.hdf', 'table',dtype={'acoustic_data': np.float32, 'time_to_failure': np.float64})\nprint((train[\"time_to_failure\"][14*150_000]-train[\"time_to_failure\"][0*150_000])/14)\n-0.03890736662714285\n```",
    "519030": "Here is the correct way to compute it:\n\n    df = train[train.time_to_failure &gt; train.time_to_failure.shift().fillna(0)]\n    df.time_to_failure.sum() - train.time_to_failure.values[-1]\n\nresult is 163.42061",
    "519032": "oops I confused test with train",
    "519045": "Even with that your computation is off by 1 second or so.",
    "519247": "Thanks, I wonder if training should be optimized on few second intervals rather than the exact same size as the test set. What would be the pros and cons in such case..",
    "519287": "There is only one way to know: try and see what works best."
  },
  "source": "meta"
}