{
  "id": 89151,
  "title": "Inconsistent Sampling Times?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89151",
  "author_name": "",
  "post_date": "2019-04-11T20:53:51.125081800Z",
  "votes": 10,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Greetings everyone.\nI just joined and started to do some EDA in this competition and tried to read up on everything. </p>\n\n<p>In the description the hosts state, that the sample rate is 4 MHz (a new measurement value each 250 ns). To me, that means each row stands for a 250 ns slot and the TTF should represent that. </p>\n\n<p>To my surprise, TTF has very inconsistent deltas throughout the train set (which the hosts stated is continous). <strong>UPDATE</strong>: <a href=\"/zaharch\">@zaharch</a> pointed out this is only a float32 precision error.</p>\n\n<p>This is the very beginning with a float32 precision:\n<code>\n     acoustic_data  time_to_failure\n0               12      1.469099998\n10               5      1.469099998\n20               4      1.469099998\n30               4      1.469099998\n40               0      1.469099998\n50              11      1.469099879\n60               4      1.469099879\n70               6      1.469099879\n80               6      1.469099879\n90               1      1.469099879\n100             -1      1.469099879\n110              6      1.469099879\n120              4      1.469099879\n130              3      1.469099879\n140              5      1.469099879\n150              3      1.469099760\n160             11      1.469099760\n170              3      1.469099760\n180              4      1.469099760\n190              1      1.469099760\n200              7      1.469099760\n210             10      1.469099760\n220              1      1.469099760\n230              6      1.469099760\n240             -1      1.469099760\n250              7      1.469099760\n260              7      1.469099641\n270              6      1.469099641\n280              5      1.469099641\n290              9      1.469099641\n</code>\nTTF is not updated each row but every about 100 rows. </p>\n\n<p>In contrast, this is the very beginning with a float64 precision:\n<code>\n     acoustic_data  time_to_failure\n0               12     1.4690999832\n10               5     1.4690999722\n20               4     1.4690999612\n30               4     1.4690999502\n40               0     1.4690999392\n50              11     1.4690999282\n60               4     1.4690999172\n70               6     1.4690999062\n80               6     1.4690998952\n90               1     1.4690998842\n100             -1     1.4690998732\n110              6     1.4690998622\n120              4     1.4690998512\n130              3     1.4690998402\n140              5     1.4690998292\n150              3     1.4690998182\n160             11     1.4690998072\n170              3     1.4690997962\n180              4     1.4690997852\n190              1     1.4690997742\n200              7     1.4690997632\n210             10     1.4690997522\n220              1     1.4690997412\n230              6     1.4690997302\n240             -1     1.4690997192\n250              7     1.4690997082\n260              7     1.4690996972\n270              6     1.4690996862\n280              5     1.4690996752\n290              9     1.4690996642\n</code></p>\n\n<p>We are seeing a delta of about 110 ns in 100 rows. That is roughly 1.1 ns for one row. How can that match to the 4 MHz and 250 ns for each row?</p>\n\n<p>A bit further in the set, coming near the earthquake event:</p>\n\n<p><code>\n         acoustic_data  time_to_failure\n5644270              4     0.0039954963\n5644271              2     0.0039954952\n5644272             10     0.0039954941\n5644273              3     0.0039954930\n5644274              2     0.0039954919\n5644275              2     0.0039954908\n5644276              4     0.0039954897\n5644277              8     0.0039954886\n5644278              3     0.0039954875\n5644279              0     0.0039954864\n5644280              3     0.0039954853\n5644281              7     0.0039954842\n5644282              6     0.0039954831\n5644283              6     0.0039954820\n5644284              5     0.0039954809\n5644285              3     0.0039954798\n5644286              1     0.0029999843\n5644287              5     0.0029999832\n5644288              6     0.0029999821\n5644289              2     0.0029999810\n</code> \nHere, the TTF values are updated every single row (for both, float32 and float64) and again we see a delta of 1.1 ns (row 5644270 to 5644285). How can that match to the 4 MHz and 250 ns for each row? It is almost 1 GHz!</p>\n\n<p>More surprisingly, we find a large gap at 5644285 to 5644286 of about 1 ms. What is that?\nThe hosts were talking about binning the measurements and a gap of 12 µs.</p>",
  "messages": [
    {
      "id": "514678",
      "postDate": "04/11/2019 20:53:51",
      "content": "<p>Greetings everyone.\nI just joined and started to do some EDA in this competition and tried to read up on everything. </p>\n\n<p>In the description the hosts state, that the sample rate is 4 MHz (a new measurement value each 250 ns). To me, that means each row stands for a 250 ns slot and the TTF should represent that. </p>\n\n<p>To my surprise, TTF has very inconsistent deltas throughout the train set (which the hosts stated is continous). <strong>UPDATE</strong>: <a href=\"/zaharch\">@zaharch</a> pointed out this is only a float32 precision error.</p>\n\n<p>This is the very beginning with a float32 precision:\n<code>\n     acoustic_data  time_to_failure\n0               12      1.469099998\n10               5      1.469099998\n20               4      1.469099998\n30               4      1.469099998\n40               0      1.469099998\n50              11      1.469099879\n60               4      1.469099879\n70               6      1.469099879\n80               6      1.469099879\n90               1      1.469099879\n100             -1      1.469099879\n110              6      1.469099879\n120              4      1.469099879\n130              3      1.469099879\n140              5      1.469099879\n150              3      1.469099760\n160             11      1.469099760\n170              3      1.469099760\n180              4      1.469099760\n190              1      1.469099760\n200              7      1.469099760\n210             10      1.469099760\n220              1      1.469099760\n230              6      1.469099760\n240             -1      1.469099760\n250              7      1.469099760\n260              7      1.469099641\n270              6      1.469099641\n280              5      1.469099641\n290              9      1.469099641\n</code>\nTTF is not updated each row but every about 100 rows. </p>\n\n<p>In contrast, this is the very beginning with a float64 precision:\n<code>\n     acoustic_data  time_to_failure\n0               12     1.4690999832\n10               5     1.4690999722\n20               4     1.4690999612\n30               4     1.4690999502\n40               0     1.4690999392\n50              11     1.4690999282\n60               4     1.4690999172\n70               6     1.4690999062\n80               6     1.4690998952\n90               1     1.4690998842\n100             -1     1.4690998732\n110              6     1.4690998622\n120              4     1.4690998512\n130              3     1.4690998402\n140              5     1.4690998292\n150              3     1.4690998182\n160             11     1.4690998072\n170              3     1.4690997962\n180              4     1.4690997852\n190              1     1.4690997742\n200              7     1.4690997632\n210             10     1.4690997522\n220              1     1.4690997412\n230              6     1.4690997302\n240             -1     1.4690997192\n250              7     1.4690997082\n260              7     1.4690996972\n270              6     1.4690996862\n280              5     1.4690996752\n290              9     1.4690996642\n</code></p>\n\n<p>We are seeing a delta of about 110 ns in 100 rows. That is roughly 1.1 ns for one row. How can that match to the 4 MHz and 250 ns for each row?</p>\n\n<p>A bit further in the set, coming near the earthquake event:</p>\n\n<p><code>\n         acoustic_data  time_to_failure\n5644270              4     0.0039954963\n5644271              2     0.0039954952\n5644272             10     0.0039954941\n5644273              3     0.0039954930\n5644274              2     0.0039954919\n5644275              2     0.0039954908\n5644276              4     0.0039954897\n5644277              8     0.0039954886\n5644278              3     0.0039954875\n5644279              0     0.0039954864\n5644280              3     0.0039954853\n5644281              7     0.0039954842\n5644282              6     0.0039954831\n5644283              6     0.0039954820\n5644284              5     0.0039954809\n5644285              3     0.0039954798\n5644286              1     0.0029999843\n5644287              5     0.0029999832\n5644288              6     0.0029999821\n5644289              2     0.0029999810\n</code> \nHere, the TTF values are updated every single row (for both, float32 and float64) and again we see a delta of 1.1 ns (row 5644270 to 5644285). How can that match to the 4 MHz and 250 ns for each row? It is almost 1 GHz!</p>\n\n<p>More surprisingly, we find a large gap at 5644285 to 5644286 of about 1 ms. What is that?\nThe hosts were talking about binning the measurements and a gap of 12 µs.</p>",
      "rawMarkdown": "Greetings everyone.\nI just joined and started to do some EDA in this competition and tried to read up on everything. \n\nIn the description the hosts state, that the sample rate is 4 MHz (a new measurement value each 250 ns). To me, that means each row stands for a 250 ns slot and the TTF should represent that. \n\nTo my surprise, TTF has very inconsistent deltas throughout the train set (which the hosts stated is continous). **UPDATE**: @zaharch pointed out this is only a float32 precision error.\n\nThis is the very beginning with a float32 precision:\n```\n     acoustic_data  time_to_failure\n0               12      1.469099998\n10               5      1.469099998\n20               4      1.469099998\n30               4      1.469099998\n40               0      1.469099998\n50              11      1.469099879\n60               4      1.469099879\n70               6      1.469099879\n80               6      1.469099879\n90               1      1.469099879\n100             -1      1.469099879\n110              6      1.469099879\n120              4      1.469099879\n130              3      1.469099879\n140              5      1.469099879\n150              3      1.469099760\n160             11      1.469099760\n170              3      1.469099760\n180              4      1.469099760\n190              1      1.469099760\n200              7      1.469099760\n210             10      1.469099760\n220              1      1.469099760\n230              6      1.469099760\n240             -1      1.469099760\n250              7      1.469099760\n260              7      1.469099641\n270              6      1.469099641\n280              5      1.469099641\n290              9      1.469099641\n```\nTTF is not updated each row but every about 100 rows. \n\nIn contrast, this is the very beginning with a float64 precision:\n```\n     acoustic_data  time_to_failure\n0               12     1.4690999832\n10               5     1.4690999722\n20               4     1.4690999612\n30               4     1.4690999502\n40               0     1.4690999392\n50              11     1.4690999282\n60               4     1.4690999172\n70               6     1.4690999062\n80               6     1.4690998952\n90               1     1.4690998842\n100             -1     1.4690998732\n110              6     1.4690998622\n120              4     1.4690998512\n130              3     1.4690998402\n140              5     1.4690998292\n150              3     1.4690998182\n160             11     1.4690998072\n170              3     1.4690997962\n180              4     1.4690997852\n190              1     1.4690997742\n200              7     1.4690997632\n210             10     1.4690997522\n220              1     1.4690997412\n230              6     1.4690997302\n240             -1     1.4690997192\n250              7     1.4690997082\n260              7     1.4690996972\n270              6     1.4690996862\n280              5     1.4690996752\n290              9     1.4690996642\n```\n\nWe are seeing a delta of about 110 ns in 100 rows. That is roughly 1.1 ns for one row. How can that match to the 4 MHz and 250 ns for each row?\n\nA bit further in the set, coming near the earthquake event:\n\n```\n         acoustic_data  time_to_failure\n5644270              4     0.0039954963\n5644271              2     0.0039954952\n5644272             10     0.0039954941\n5644273              3     0.0039954930\n5644274              2     0.0039954919\n5644275              2     0.0039954908\n5644276              4     0.0039954897\n5644277              8     0.0039954886\n5644278              3     0.0039954875\n5644279              0     0.0039954864\n5644280              3     0.0039954853\n5644281              7     0.0039954842\n5644282              6     0.0039954831\n5644283              6     0.0039954820\n5644284              5     0.0039954809\n5644285              3     0.0039954798\n5644286              1     0.0029999843\n5644287              5     0.0029999832\n5644288              6     0.0029999821\n5644289              2     0.0029999810\n``` \nHere, the TTF values are updated every single row (for both, float32 and float64) and again we see a delta of 1.1 ns (row 5644270 to 5644285). How can that match to the 4 MHz and 250 ns for each row? It is almost 1 GHz!\n\nMore surprisingly, we find a large gap at 5644285 to 5644286 of about 1 ms. What is that?\nThe hosts were talking about binning the measurements and a gap of 12 µs.",
      "votes": null
    },
    {
      "id": "515313",
      "postDate": "04/12/2019 12:25:10",
      "content": "<p>How about resampling as in  <a href=\"https://nl.mathworks.com/help/signal/ref/resample.html\">https://nl.mathworks.com/help/signal/ref/resample.html</a>\nMight work !!</p>",
      "rawMarkdown": "How about resampling as in  https://nl.mathworks.com/help/signal/ref/resample.html\nMight work !!",
      "votes": null
    },
    {
      "id": "515434",
      "postDate": "04/12/2019 15:37:36",
      "content": "<p>How do you read the data? For me the values are different</p>\n\n<p><code>\ndd = pd.read_csv(PATH/'train.csv', nrows=20)\ndd['time_to_failure'][0]\n1.4690999832\ndd['time_to_failure'][10]\n1.4690999722\n</code></p>\n\n<p>I guess you read it with reduced precision, something along the lines\n<code>\ndd = pd.read_csv(PATH/'train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\n</code></p>\n\n<p>Anyway, this is how the data is structured. Batches of 4096 samples with 1.1 ns average difference and then approximately 1ms jumps between the batches. The samples are sampled at 4MHz (slightly less), but logged in batches, and the times are the logging time, not the actual sampling time.</p>",
      "rawMarkdown": "How do you read the data? For me the values are different\n\n```\ndd = pd.read_csv(PATH/'train.csv', nrows=20)\ndd['time_to_failure'][0]\n1.4690999832\ndd['time_to_failure'][10]\n1.4690999722\n```\n\nI guess you read it with reduced precision, something along the lines\n```\ndd = pd.read_csv(PATH/'train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\n```\n\nAnyway, this is how the data is structured. Batches of 4096 samples with 1.1 ns average difference and then approximately 1ms jumps between the batches. The samples are sampled at 4MHz (slightly less), but logged in batches, and the times are the logging time, not the actual sampling time.",
      "votes": null
    },
    {
      "id": "515485",
      "postDate": "04/12/2019 17:28:12",
      "content": "<p>Thanks a lot for your comment. \nYou are correct, the precision created the inconsistency in the TTF values. I did use <code>np.float32</code>. <code>np.float64</code> should be used here. With that, I am getting a pretty much constant TTF delta of about 1.1 ns. However, the gap (5644285 to 5644286 of about 1 ms) is still visible.\nI have updated the original post accordingly. </p>\n\n<p>Nevertheless, I have a problem with your second explanation. \n<code>\nAnyway, this is how the data is structured. Batches of 4096 samples with 1.1 ns average difference and then approximately 1ms jumps between the batches. The samples are sampled at 4MHz (slightly less), but logged in batches, and the times are the logging time, not the actual sampling time.\n</code>\nSo, you are saying we have about 4096*1.1 ns = ~4,5 µs of data logging and then a pause of 1 ms?\nWhere do you have the 4 MHz then? Sampling would still be at almost 1 GHz. With rather long pauses 3 orders of magnitude longer than data aquisition. </p>",
      "rawMarkdown": "Thanks a lot for your comment. \nYou are correct, the precision created the inconsistency in the TTF values. I did use `np.float32`. `np.float64` should be used here. With that, I am getting a pretty much constant TTF delta of about 1.1 ns. However, the gap (5644285 to 5644286 of about 1 ms) is still visible.\nI have updated the original post accordingly. \n\nNevertheless, I have a problem with your second explanation. \n```\nAnyway, this is how the data is structured. Batches of 4096 samples with 1.1 ns average difference and then approximately 1ms jumps between the batches. The samples are sampled at 4MHz (slightly less), but logged in batches, and the times are the logging time, not the actual sampling time.\n```\nSo, you are saying we have about 4096*1.1 ns = ~4,5 µs of data logging and then a pause of 1 ms?\nWhere do you have the 4 MHz then? Sampling would still be at almost 1 GHz. With rather long pauses 3 orders of magnitude longer than data aquisition.",
      "votes": null
    },
    {
      "id": "515514",
      "postDate": "04/12/2019 18:02:19",
      "content": "<p>I imagine it the following way. The device makes measurements every 250ns, 4MHz but it stores it into internal buffer. Then every 4096 samples it allocates 12us time for logging/sending the data out. It does it quite fast, 1.1 ns per sample, so it actually takes only about 4us in total for that. The timestamp reported with a sample point is that of the logging, not the actual time of the measurement. This is an artifact of the device.</p>",
      "rawMarkdown": "I imagine it the following way. The device makes measurements every 250ns, 4MHz but it stores it into internal buffer. Then every 4096 samples it allocates 12us time for logging/sending the data out. It does it quite fast, 1.1 ns per sample, so it actually takes only about 4us in total for that. The timestamp reported with a sample point is that of the logging, not the actual time of the measurement. This is an artifact of the device.",
      "votes": null
    },
    {
      "id": "515530",
      "postDate": "04/12/2019 18:18:48",
      "content": "<p>Bertrand (Host):\n<code>\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\n</code></p>\n\n<p>I still have a very hard time to see this in the actual data. </p>\n\n<p>We have:\n- 1.1 ns delta in each row.\n- Assuming TTF values are correct, that correlates to a sampling rate of 909 MHz (much higher than the stated 4 MHz).\n- 1 ms gap between two bins (much higher than the stated 12 µs)</p>\n\n<p>What am I missing?</p>",
      "rawMarkdown": "Bertrand (Host):\n```\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\n```\n\nI still have a very hard time to see this in the actual data. \n\nWe have:\n- 1.1 ns delta in each row.\n- Assuming TTF values are correct, that correlates to a sampling rate of 909 MHz (much higher than the stated 4 MHz).\n- 1 ms gap between two bins (much higher than the stated 12 µs)\n\nWhat am I missing?",
      "votes": null
    },
    {
      "id": "515539",
      "postDate": "04/12/2019 18:41:02",
      "content": "<p>I started to write a new explanation, but then understood that I had already explained the same thing twice :). Please read through my previous messages carefully, I think it should be clear and answers you question. </p>",
      "rawMarkdown": "I started to write a new explanation, but then understood that I had already explained the same thing twice :). Please read through my previous messages carefully, I think it should be clear and answers you question.",
      "votes": null
    },
    {
      "id": "515551",
      "postDate": "04/12/2019 18:58:03",
      "content": "<p>Your comments imply that TTF values are not really TTF values but actually timestamps of the data transfer. \nIf so, why are the hosts not saying that?</p>\n\n<p>In the sense of analysis it would be extremly important to know if the pauses are actually 1 ms or rather short 4 µs or 12 µs. </p>\n\n<p>I carefully read the additional information discussion again: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526</a>\nThat critical piece of information is still missing. </p>",
      "rawMarkdown": "Your comments imply that TTF values are not really TTF values but actually timestamps of the data transfer. \nIf so, why are the hosts not saying that?\n\nIn the sense of analysis it would be extremly important to know if the pauses are actually 1 ms or rather short 4 µs or 12 µs. \n\nI carefully read the additional information discussion again: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526\nThat critical piece of information is still missing.",
      "votes": null
    },
    {
      "id": "515558",
      "postDate": "04/12/2019 19:10:39",
      "content": "<p>I agree, the organizers don't make it clear. Many people asked Bertrand that same question but he decided not to comment on that for some reason. Probably he wants to make as few comments as possible, and other information he provided is already (almost) enough to make the correct deduction. </p>",
      "rawMarkdown": "I agree, the organizers don't make it clear. Many people asked Bertrand that same question but he decided not to comment on that for some reason. Probably he wants to make as few comments as possible, and other information he provided is already (almost) enough to make the correct deduction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 515313,
      "author_name": "hmcranbercourt",
      "author_url": "",
      "post_date": "04/12/2019 12:25:10",
      "content": "<p>How about resampling as in  <a href=\"https://nl.mathworks.com/help/signal/ref/resample.html\">https://nl.mathworks.com/help/signal/ref/resample.html</a>\nMight work !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 515434,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "04/12/2019 15:37:36",
      "content": "<p>How do you read the data? For me the values are different</p>\n\n<p><code>\ndd = pd.read_csv(PATH/'train.csv', nrows=20)\ndd['time_to_failure'][0]\n1.4690999832\ndd['time_to_failure'][10]\n1.4690999722\n</code></p>\n\n<p>I guess you read it with reduced precision, something along the lines\n<code>\ndd = pd.read_csv(PATH/'train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\n</code></p>\n\n<p>Anyway, this is how the data is structured. Batches of 4096 samples with 1.1 ns average difference and then approximately 1ms jumps between the batches. The samples are sampled at 4MHz (slightly less), but logged in batches, and the times are the logging time, not the actual sampling time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 515485,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/12/2019 17:28:12",
          "content": "<p>Thanks a lot for your comment. \nYou are correct, the precision created the inconsistency in the TTF values. I did use <code>np.float32</code>. <code>np.float64</code> should be used here. With that, I am getting a pretty much constant TTF delta of about 1.1 ns. However, the gap (5644285 to 5644286 of about 1 ms) is still visible.\nI have updated the original post accordingly. </p>\n\n<p>Nevertheless, I have a problem with your second explanation. \n<code>\nAnyway, this is how the data is structured. Batches of 4096 samples with 1.1 ns average difference and then approximately 1ms jumps between the batches. The samples are sampled at 4MHz (slightly less), but logged in batches, and the times are the logging time, not the actual sampling time.\n</code>\nSo, you are saying we have about 4096*1.1 ns = ~4,5 µs of data logging and then a pause of 1 ms?\nWhere do you have the 4 MHz then? Sampling would still be at almost 1 GHz. With rather long pauses 3 orders of magnitude longer than data aquisition. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515514,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "04/12/2019 18:02:19",
          "content": "<p>I imagine it the following way. The device makes measurements every 250ns, 4MHz but it stores it into internal buffer. Then every 4096 samples it allocates 12us time for logging/sending the data out. It does it quite fast, 1.1 ns per sample, so it actually takes only about 4us in total for that. The timestamp reported with a sample point is that of the logging, not the actual time of the measurement. This is an artifact of the device.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515530,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/12/2019 18:18:48",
          "content": "<p>Bertrand (Host):\n<code>\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\n</code></p>\n\n<p>I still have a very hard time to see this in the actual data. </p>\n\n<p>We have:\n- 1.1 ns delta in each row.\n- Assuming TTF values are correct, that correlates to a sampling rate of 909 MHz (much higher than the stated 4 MHz).\n- 1 ms gap between two bins (much higher than the stated 12 µs)</p>\n\n<p>What am I missing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515539,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "04/12/2019 18:41:02",
          "content": "<p>I started to write a new explanation, but then understood that I had already explained the same thing twice :). Please read through my previous messages carefully, I think it should be clear and answers you question. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515551,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/12/2019 18:58:03",
          "content": "<p>Your comments imply that TTF values are not really TTF values but actually timestamps of the data transfer. \nIf so, why are the hosts not saying that?</p>\n\n<p>In the sense of analysis it would be extremly important to know if the pauses are actually 1 ms or rather short 4 µs or 12 µs. </p>\n\n<p>I carefully read the additional information discussion again: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526</a>\nThat critical piece of information is still missing. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515558,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "04/12/2019 19:10:39",
          "content": "<p>I agree, the organizers don't make it clear. Many people asked Bertrand that same question but he decided not to comment on that for some reason. Probably he wants to make as few comments as possible, and other information he provided is already (almost) enough to make the correct deduction. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "514678": "Greetings everyone.\nI just joined and started to do some EDA in this competition and tried to read up on everything. \n\nIn the description the hosts state, that the sample rate is 4 MHz (a new measurement value each 250 ns). To me, that means each row stands for a 250 ns slot and the TTF should represent that. \n\nTo my surprise, TTF has very inconsistent deltas throughout the train set (which the hosts stated is continous). **UPDATE**: @zaharch pointed out this is only a float32 precision error.\n\nThis is the very beginning with a float32 precision:\n```\n     acoustic_data  time_to_failure\n0               12      1.469099998\n10               5      1.469099998\n20               4      1.469099998\n30               4      1.469099998\n40               0      1.469099998\n50              11      1.469099879\n60               4      1.469099879\n70               6      1.469099879\n80               6      1.469099879\n90               1      1.469099879\n100             -1      1.469099879\n110              6      1.469099879\n120              4      1.469099879\n130              3      1.469099879\n140              5      1.469099879\n150              3      1.469099760\n160             11      1.469099760\n170              3      1.469099760\n180              4      1.469099760\n190              1      1.469099760\n200              7      1.469099760\n210             10      1.469099760\n220              1      1.469099760\n230              6      1.469099760\n240             -1      1.469099760\n250              7      1.469099760\n260              7      1.469099641\n270              6      1.469099641\n280              5      1.469099641\n290              9      1.469099641\n```\nTTF is not updated each row but every about 100 rows. \n\nIn contrast, this is the very beginning with a float64 precision:\n```\n     acoustic_data  time_to_failure\n0               12     1.4690999832\n10               5     1.4690999722\n20               4     1.4690999612\n30               4     1.4690999502\n40               0     1.4690999392\n50              11     1.4690999282\n60               4     1.4690999172\n70               6     1.4690999062\n80               6     1.4690998952\n90               1     1.4690998842\n100             -1     1.4690998732\n110              6     1.4690998622\n120              4     1.4690998512\n130              3     1.4690998402\n140              5     1.4690998292\n150              3     1.4690998182\n160             11     1.4690998072\n170              3     1.4690997962\n180              4     1.4690997852\n190              1     1.4690997742\n200              7     1.4690997632\n210             10     1.4690997522\n220              1     1.4690997412\n230              6     1.4690997302\n240             -1     1.4690997192\n250              7     1.4690997082\n260              7     1.4690996972\n270              6     1.4690996862\n280              5     1.4690996752\n290              9     1.4690996642\n```\n\nWe are seeing a delta of about 110 ns in 100 rows. That is roughly 1.1 ns for one row. How can that match to the 4 MHz and 250 ns for each row?\n\nA bit further in the set, coming near the earthquake event:\n\n```\n         acoustic_data  time_to_failure\n5644270              4     0.0039954963\n5644271              2     0.0039954952\n5644272             10     0.0039954941\n5644273              3     0.0039954930\n5644274              2     0.0039954919\n5644275              2     0.0039954908\n5644276              4     0.0039954897\n5644277              8     0.0039954886\n5644278              3     0.0039954875\n5644279              0     0.0039954864\n5644280              3     0.0039954853\n5644281              7     0.0039954842\n5644282              6     0.0039954831\n5644283              6     0.0039954820\n5644284              5     0.0039954809\n5644285              3     0.0039954798\n5644286              1     0.0029999843\n5644287              5     0.0029999832\n5644288              6     0.0029999821\n5644289              2     0.0029999810\n``` \nHere, the TTF values are updated every single row (for both, float32 and float64) and again we see a delta of 1.1 ns (row 5644270 to 5644285). How can that match to the 4 MHz and 250 ns for each row? It is almost 1 GHz!\n\nMore surprisingly, we find a large gap at 5644285 to 5644286 of about 1 ms. What is that?\nThe hosts were talking about binning the measurements and a gap of 12 µs.",
    "515313": "How about resampling as in  https://nl.mathworks.com/help/signal/ref/resample.html\nMight work !!",
    "515434": "How do you read the data? For me the values are different\n\n```\ndd = pd.read_csv(PATH/'train.csv', nrows=20)\ndd['time_to_failure'][0]\n1.4690999832\ndd['time_to_failure'][10]\n1.4690999722\n```\n\nI guess you read it with reduced precision, something along the lines\n```\ndd = pd.read_csv(PATH/'train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\n```\n\nAnyway, this is how the data is structured. Batches of 4096 samples with 1.1 ns average difference and then approximately 1ms jumps between the batches. The samples are sampled at 4MHz (slightly less), but logged in batches, and the times are the logging time, not the actual sampling time.",
    "515485": "Thanks a lot for your comment. \nYou are correct, the precision created the inconsistency in the TTF values. I did use `np.float32`. `np.float64` should be used here. With that, I am getting a pretty much constant TTF delta of about 1.1 ns. However, the gap (5644285 to 5644286 of about 1 ms) is still visible.\nI have updated the original post accordingly. \n\nNevertheless, I have a problem with your second explanation. \n```\nAnyway, this is how the data is structured. Batches of 4096 samples with 1.1 ns average difference and then approximately 1ms jumps between the batches. The samples are sampled at 4MHz (slightly less), but logged in batches, and the times are the logging time, not the actual sampling time.\n```\nSo, you are saying we have about 4096*1.1 ns = ~4,5 µs of data logging and then a pause of 1 ms?\nWhere do you have the 4 MHz then? Sampling would still be at almost 1 GHz. With rather long pauses 3 orders of magnitude longer than data aquisition.",
    "515514": "I imagine it the following way. The device makes measurements every 250ns, 4MHz but it stores it into internal buffer. Then every 4096 samples it allocates 12us time for logging/sending the data out. It does it quite fast, 1.1 ns per sample, so it actually takes only about 4us in total for that. The timestamp reported with a sample point is that of the logging, not the actual time of the measurement. This is an artifact of the device.",
    "515530": "Bertrand (Host):\n```\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\n```\n\nI still have a very hard time to see this in the actual data. \n\nWe have:\n- 1.1 ns delta in each row.\n- Assuming TTF values are correct, that correlates to a sampling rate of 909 MHz (much higher than the stated 4 MHz).\n- 1 ms gap between two bins (much higher than the stated 12 µs)\n\nWhat am I missing?",
    "515539": "I started to write a new explanation, but then understood that I had already explained the same thing twice :). Please read through my previous messages carefully, I think it should be clear and answers you question.",
    "515551": "Your comments imply that TTF values are not really TTF values but actually timestamps of the data transfer. \nIf so, why are the hosts not saying that?\n\nIn the sense of analysis it would be extremly important to know if the pauses are actually 1 ms or rather short 4 µs or 12 µs. \n\nI carefully read the additional information discussion again: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526\nThat critical piece of information is still missing.",
    "515558": "I agree, the organizers don't make it clear. Many people asked Bertrand that same question but he decided not to comment on that for some reason. Probably he wants to make as few comments as possible, and other information he provided is already (almost) enough to make the correct deduction."
  },
  "source": "meta"
}