{
  "id": 77526,
  "title": "Additional info",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77526",
  "author_name": "Bertrand RL",
  "post_date": "2019-01-13T21:10:32.658000",
  "votes": 161,
  "comment_count": 112,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n\n<p>I will try to answer questions in the forum as things go. To answer a few recurring questions:</p>\n\n<ul>\n<li><p>The goal of the challenge is to capture the physical state of the laboratory fault and how close it is from failure from a snapshot of the seismic data it is emitting. You will have to build a model that predicts the time remaining before failure from a chunk of seismic data, like we have done in our first paper above on easier data.</p></li>\n<li><p>The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds. </p></li>\n<li><p>The seismic data is recorded using a piezoceramic sensor, which outputs a voltage upon deformation by incoming seismic waves. The seismic data of the input is this recorded voltage, in integers.</p></li>\n<li><p>Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.</p></li>\n<li><p>Time to failure is based on a measure of fault strength (shear stress, not part of the data for the competition). When a labquake occurs this stress drops unambiguously.</p></li>\n<li><p>The data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device. </p></li>\n</ul>\n\n<p>Bertrand</p>\n\n<p><em>Edited to answer additional questions</em></p>",
  "messages": [
    {
      "id": 455431,
      "postDate": "2019-01-13T21:10:32.657Z",
      "content": "<p>Hi everyone,</p>\n\n<p>I will try to answer questions in the forum as things go. To answer a few recurring questions:</p>\n\n<ul>\n<li><p>The goal of the challenge is to capture the physical state of the laboratory fault and how close it is from failure from a snapshot of the seismic data it is emitting. You will have to build a model that predicts the time remaining before failure from a chunk of seismic data, like we have done in our first paper above on easier data.</p></li>\n<li><p>The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds. </p></li>\n<li><p>The seismic data is recorded using a piezoceramic sensor, which outputs a voltage upon deformation by incoming seismic waves. The seismic data of the input is this recorded voltage, in integers.</p></li>\n<li><p>Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.</p></li>\n<li><p>Time to failure is based on a measure of fault strength (shear stress, not part of the data for the competition). When a labquake occurs this stress drops unambiguously.</p></li>\n<li><p>The data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device. </p></li>\n</ul>\n\n<p>Bertrand</p>\n\n<p><em>Edited to answer additional questions</em></p>",
      "rawMarkdown": "Hi everyone,\n\nI will try to answer questions in the forum as things go. To answer a few recurring questions:\n\n- The goal of the challenge is to capture the physical state of the laboratory fault and how close it is from failure from a snapshot of the seismic data it is emitting. You will have to build a model that predicts the time remaining before failure from a chunk of seismic data, like we have done in our first paper above on easier data.\n\n- The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds. \n\n- The seismic data is recorded using a piezoceramic sensor, which outputs a voltage upon deformation by incoming seismic waves. The seismic data of the input is this recorded voltage, in integers.\n\n- Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\n\n- Time to failure is based on a measure of fault strength (shear stress, not part of the data for the competition). When a labquake occurs this stress drops unambiguously.\n\n- The data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device. \n\n\nBertrand\n\n*Edited to answer additional questions*",
      "votes": 160
    },
    {
      "id": 468488,
      "postDate": "2019-02-09T01:02:33.783Z",
      "content": "<p>I think the time-to-failure data is not particularly clear, for the reasons at the bottom.</p>\n\n<p>If the data is even spaced, as denoted, a simpler approach to setting up the challenge would be to just provide a list of the acoustic data for training (no time-to-failure), and provide indices into that list as to where the labquakes appeared.  Then the challenge is to predict how long (how many time points) after a test set that a lab quake would occur.</p>\n\n<p>Please advise if there is useful structure in the ttf data, or if it can simply be assumed to be regularly spaced?</p>\n\n<p>The ttf data is confusing because:\n - each sample drops ~1ns over a ~4096 sample period, even though the samples should be 250ns apart (4MHz)\n - the time series takes large jumps every ~4096 period to \"catch up\" to the real sampling rate\n - the period between large jumps isn't actually always 4096 -- it's looks to be less sometimes\n - even interpolating over the whole dataset, it looks like the sampling is closer to 3.85MHz, not the 4MHz denoted</p>",
      "rawMarkdown": "I think the time-to-failure data is not particularly clear, for the reasons at the bottom.\n\nIf the data is even spaced, as denoted, a simpler approach to setting up the challenge would be to just provide a list of the acoustic data for training (no time-to-failure), and provide indices into that list as to where the labquakes appeared.  Then the challenge is to predict how long (how many time points) after a test set that a lab quake would occur.\n\nPlease advise if there is useful structure in the ttf data, or if it can simply be assumed to be regularly spaced?\n\nThe ttf data is confusing because:\n - each sample drops ~1ns over a ~4096 sample period, even though the samples should be 250ns apart (4MHz)\n - the time series takes large jumps every ~4096 period to \"catch up\" to the real sampling rate\n - the period between large jumps isn't actually always 4096 -- it's looks to be less sometimes\n - even interpolating over the whole dataset, it looks like the sampling is closer to 3.85MHz, not the 4MHz denoted",
      "votes": 21,
      "replies": [
        {
          "id": 468877,
          "postDate": "2019-02-09T23:04:32.013Z",
          "content": "<p>Thanks. The best clear explanation. </p>",
          "rawMarkdown": "Thanks. The best clear explanation. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 456854,
      "postDate": "2019-01-16T16:59:29.997Z",
      "content": "<p>Can you verify the following assumption by <a href=\"/chiqun\">@chiqun</a> for BOTH test and train data: acoustic sampling rate is 4MHz with a timeoffailure sample rate ~1000Hz? as we notice that in each chunk (4096 rows) of frames, the last two frames have a time difference ~ 0.001s, all the others have time difference ~1e-9 s. Can you also explain why it is like that physically.  Thanks. </p>",
      "rawMarkdown": "Can you verify the following assumption by @chiqun for BOTH test and train data: acoustic sampling rate is 4MHz with a timeoffailure sample rate ~1000Hz? as we notice that in each chunk (4096 rows) of frames, the last two frames have a time difference ~ 0.001s, all the others have time difference ~1e-9 s. Can you also explain why it is like that physically.  Thanks. ",
      "votes": 11,
      "replies": [
        {
          "id": 457472,
          "postDate": "2019-01-17T13:57:21.997Z",
          "content": "<p>Hi Bertrand!\nMy additional request for clarification: At 4 MHz sampling rate, the time difference between two successive data points (i.e. lines in the CSV file) is 1/(4 MHz) = 250 ns, right? Therefore, time_to_failure should correctly decrease by 250 ns each step. What we actually see in the time_to_failure column in train.csv is an artifact of one of the physical devices used in the experiment.\nIn that case, basically all that is relevant is a 1D time series of acoustic data and a list of the indices of those 16 data points where an earthquake occurred. Are my assumptions correct?\nThanks!</p>",
          "rawMarkdown": "Hi Bertrand!\nMy additional request for clarification: At 4 MHz sampling rate, the time difference between two successive data points (i.e. lines in the CSV file) is 1/(4 MHz) = 250 ns, right? Therefore, time_to_failure should correctly decrease by 250 ns each step. What we actually see in the time_to_failure column in train.csv is an artifact of one of the physical devices used in the experiment.\nIn that case, basically all that is relevant is a 1D time series of acoustic data and a list of the indices of those 16 data points where an earthquake occurred. Are my assumptions correct?\nThanks!",
          "votes": 5
        },
        {
          "id": 466150,
          "postDate": "2019-02-04T19:24:27.097Z",
          "content": "<p>Hi! The periodic 0.001s of missing data is an artifact of the recording device, I will add a comment about it in the main post, thanks for bringing it up!</p>",
          "rawMarkdown": "Hi! The periodic 0.001s of missing data is an artifact of the recording device, I will add a comment about it in the main post, thanks for bringing it up!",
          "votes": 2
        },
        {
          "id": 508175,
          "postDate": "2019-04-05T18:22:56.090Z",
          "content": "<p>Hi, I'm new here and i want to insert the time in the samples, it's more easy for me but i want to be sure. In the Andrews Script example his submission is like this:\n1   seg_id  time_to_failure</p>\n\n<p>2   seg_00030f  3.16345030153463</p>\n\n<p>3   seg_0012b5  5.1229034920016705</p>\n\n<p>4   seg_00184e  4.878316564579939</p>\n\n<p>5   seg_003339  7.867255038555255\nWhat are the 3.16345030153463, 5.1229034920016705 intervals? Nanoseconds or Microseconds? Or the number of lines until his expected event? I hope you understand my point and what i mean. Thank you.\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/88119#latest-508265\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/88119#latest-508265</a></p>",
          "rawMarkdown": "Hi, I'm new here and i want to insert the time in the samples, it's more easy for me but i want to be sure. In the Andrews Script example his submission is like this:\n1\tseg_id\ttime_to_failure\n\n2\tseg_00030f\t3.16345030153463\n\n3\tseg_0012b5\t5.1229034920016705\n\n4\tseg_00184e\t4.878316564579939\n\n5\tseg_003339\t7.867255038555255\nWhat are the 3.16345030153463, 5.1229034920016705 intervals? Nanoseconds or Microseconds? Or the number of lines until his expected event? I hope you understand my point and what i mean. Thank you.\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/88119#latest-508265"
        },
        {
          "id": 508306,
          "postDate": "2019-04-06T00:29:50.637Z",
          "content": "<p>It is seconds, time_to_failure is measured in seconds in IO files.</p>",
          "rawMarkdown": "It is seconds, time_to_failure is measured in seconds in IO files."
        }
      ]
    },
    {
      "id": 467334,
      "postDate": "2019-02-06T22:21:29.133Z",
      "content": "<p>Just to repeat this question and bring to the top, where are these recording gaps located in the test data?  It looks like each test segment is 150001 samples, which is not evenly divisible by 4096.</p>",
      "rawMarkdown": "Just to repeat this question and bring to the top, where are these recording gaps located in the test data?  It looks like each test segment is 150001 samples, which is not evenly divisible by 4096.",
      "votes": 9,
      "replies": [
        {
          "id": 477138,
          "postDate": "2019-02-24T01:04:41.030Z",
          "content": "<p>it is 150,000 exactly, the header of the csv might be the extra row... But it is also not divisible. Also there are sometimes only 4095 ticks between gaps...</p>",
          "rawMarkdown": "it is 150,000 exactly, the header of the csv might be the extra row... But it is also not divisible. Also there are sometimes only 4095 ticks between gaps...",
          "votes": 2
        }
      ]
    },
    {
      "id": 461314,
      "postDate": "2019-01-25T18:44:02.080Z",
      "content": "<p>It was mentioned that \"there are several earthquake cycles in the test set as well\". Do any of these earthquakes present within the duration of a test set segment?</p>\n\n<p>For example, said if a earthquake presents in the middle of a 150,000 steps segment, then only the signals from later half of the segment (75000 to 149000) will be useful to predict the next earthquake. </p>\n\n<p>It is important to know, because if this exists in the test set, then the model will need to be trained to handle this, i.e. identify if and when an earthquake happen within the segment, and then either ignore the signals before earthquake, or extract information from signals before earthquake differently than from signals after earthquake in order to predict next earthquake - a more much difficult task.</p>\n\n<p>Just want to ask to make sure.</p>",
      "rawMarkdown": "It was mentioned that \"there are several earthquake cycles in the test set as well\". Do any of these earthquakes present within the duration of a test set segment?\n\nFor example, said if a earthquake presents in the middle of a 150,000 steps segment, then only the signals from later half of the segment (75000 to 149000) will be useful to predict the next earthquake. \n\nIt is important to know, because if this exists in the test set, then the model will need to be trained to handle this, i.e. identify if and when an earthquake happen within the segment, and then either ignore the signals before earthquake, or extract information from signals before earthquake differently than from signals after earthquake in order to predict next earthquake - a more much difficult task.\n\nJust want to ask to make sure.",
      "votes": 9,
      "replies": [
        {
          "id": 461364,
          "postDate": "2019-01-25T21:04:55.263Z",
          "content": "<p>There’re only a few samples like that. A few! So basically it does not affect anything. </p>",
          "rawMarkdown": "There’re only a few samples like that. A few! So basically it does not affect anything. ",
          "votes": 5
        },
        {
          "id": 461872,
          "postDate": "2019-01-27T08:23:15.090Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 524621,
      "postDate": "2019-04-29T07:50:26.160Z",
      "content": "<p>We have a month until the end.</p>\n\n<p>If I had one wish it would be more data - a lot of my effort is in trying not to overfit - could you consider having a supplementary data set added in the Data section so we can increase our training data?</p>\n\n<p>Obviously I am not sure if the goal of this competition is to obtain ideas on features or actual decent models.  The variance w.r.t. MAE is huge and this can be reduced with more data.</p>",
      "rawMarkdown": "We have a month until the end.\n\nIf I had one wish it would be more data - a lot of my effort is in trying not to overfit - could you consider having a supplementary data set added in the Data section so we can increase our training data?\n\nObviously I am not sure if the goal of this competition is to obtain ideas on features or actual decent models.  The variance w.r.t. MAE is huge and this can be reduced with more data.",
      "votes": 7
    },
    {
      "id": 480454,
      "postDate": "2019-02-28T08:00:44.053Z",
      "content": "<p>Hi Bertrand!\nThank you for organizing this contest. Could you kindly answer a few questions about the data?</p>\n\n<ol>\n<li>Why was the sampling frequency chosen to be 4 MHz? Is there anywhere in your articles an estimate of the possible aliasing of data based on the speed and wavelength of the sound propagating in the sample's material? If there's aliasing, we may be missing some important signal characteristics.</li>\n<li>Has the analog signal been filtered at the output of the piezo sensor? The data looks noisy, like there is a low-pass filter missing.</li>\n<li>Am I right that since the 150000 samples are not multiples of 4096, then the 12 microsecond space can be placed in an arbitrary place of the chunk? And two parts of one bin can be in different chunks? This means that the data inside the chunk is not continuous, and this significantly limits the use of standard signal processing methods. Such pauses, where we lose 48 samples, are high-frequency noise in the signal. It would be more convenient and close to the real field scenario to break the data by bins so that there would be no pauses inside each record.</li>\n</ol>",
      "rawMarkdown": "Hi Bertrand!\nThank you for organizing this contest. Could you kindly answer a few questions about the data?\n\n1. Why was the sampling frequency chosen to be 4 MHz? Is there anywhere in your articles an estimate of the possible aliasing of data based on the speed and wavelength of the sound propagating in the sample's material? If there's aliasing, we may be missing some important signal characteristics.\n2. Has the analog signal been filtered at the output of the piezo sensor? The data looks noisy, like there is a low-pass filter missing.\n3. Am I right that since the 150000 samples are not multiples of 4096, then the 12 microsecond space can be placed in an arbitrary place of the chunk? And two parts of one bin can be in different chunks? This means that the data inside the chunk is not continuous, and this significantly limits the use of standard signal processing methods. Such pauses, where we lose 48 samples, are high-frequency noise in the signal. It would be more convenient and close to the real field scenario to break the data by bins so that there would be no pauses inside each record.",
      "votes": 7
    },
    {
      "id": 465224,
      "postDate": "2019-02-02T16:44:51.193Z",
      "content": "<p>Hi Bertrand</p>\n\n<p>Are the test segments contiguous or picked at random?</p>",
      "rawMarkdown": "Hi Bertrand\n\nAre the test segments contiguous or picked at random?",
      "votes": 7
    },
    {
      "id": 469301,
      "postDate": "2019-02-10T23:01:39.497Z",
      "content": "<p>Perchance I am wrong, but the additional info does not seem accurate. After carefully analysing the time gaps I believe the following is happening:</p>\n\n<p>There are two phases that is correct: recording, and let's say IO sequence, when there is no recording.</p>\n\n<p>When recording happens it happens at ~0.9GSamples/s for 4096 samples (sometimes a sample is missing, so 4095 samples are recorded in one go etc.)</p>\n\n<p>After the recording phase the IO phase takes not 12ms but 1.06ms on average.</p>\n\n<p>It is true that in some sense if you combine these you get near 3.85MSamples/s. </p>\n\n<p>To sum up recording takes approximately 4.5us after this there is 1060us silence. This repeats. This is important because frequency analysis is greatly affected by this, especially because we have no idea, where the recording gaps are in the test set.</p>\n\n<p>Please someone confirm...</p>\n\n<p>Note: I am wondering also if it is possible that the recording was in fact contignuous, but the times are somehow crowded up into these chunks? This can be possible if in fact the times were assigned not on measurement but at some processing level.</p>",
      "rawMarkdown": "Perchance I am wrong, but the additional info does not seem accurate. After carefully analysing the time gaps I believe the following is happening:\n\nThere are two phases that is correct: recording, and let's say IO sequence, when there is no recording.\n\nWhen recording happens it happens at ~0.9GSamples/s for 4096 samples (sometimes a sample is missing, so 4095 samples are recorded in one go etc.)\n\nAfter the recording phase the IO phase takes not 12ms but 1.06ms on average.\n\nIt is true that in some sense if you combine these you get near 3.85MSamples/s. \n\nTo sum up recording takes approximately 4.5us after this there is 1060us silence. This repeats. This is important because frequency analysis is greatly affected by this, especially because we have no idea, where the recording gaps are in the test set.\n\nPlease someone confirm...\n\nNote: I am wondering also if it is possible that the recording was in fact contignuous, but the times are somehow crowded up into these chunks? This can be possible if in fact the times were assigned not on measurement but at some processing level.\n",
      "votes": 6,
      "replies": [
        {
          "id": 469371,
          "postDate": "2019-02-11T04:17:37.547Z",
          "content": "<p>I was wandering the same thing. As far as I get from the papers, they calculate ttf, so I was hoping for really contiguous acoustic data, just twisted ttf. \nBut the team confirmed about the gap and now that scenario looks more right, just the 12ms looks like 1 ms. \nHope they give us more details soon!</p>",
          "rawMarkdown": "I was wandering the same thing. As far as I get from the papers, they calculate ttf, so I was hoping for really contiguous acoustic data, just twisted ttf. \nBut the team confirmed about the gap and now that scenario looks more right, just the 12ms looks like 1 ms. \nHope they give us more details soon!",
          "votes": 1
        },
        {
          "id": 470419,
          "postDate": "2019-02-12T23:02:43.340Z",
          "content": "<p>Since I analysed the gaps and can confirm there is no continuity between the bins, in other words the gap is real. Mind the gap!</p>",
          "rawMarkdown": "Since I analysed the gaps and can confirm there is no continuity between the bins, in other words the gap is real. Mind the gap!",
          "votes": 3
        },
        {
          "id": 470471,
          "postDate": "2019-02-13T03:09:58.020Z",
          "content": "<p>Thank you! I am still waiting some official clarification. \nMeanwhile I am thinking about the extreme ratio between the length of the bin and the length of the gap.  </p>",
          "rawMarkdown": "Thank you! I am still waiting some official clarification. \nMeanwhile I am thinking about the extreme ratio between the length of the bin and the length of the gap.  "
        },
        {
          "id": 471019,
          "postDate": "2019-02-13T23:11:26.460Z",
          "content": "<p>Well it is aproximately .004 (bin/gap)... If I understand you correctly. The gap is pretty constant... ...also the bin</p>",
          "rawMarkdown": "Well it is aproximately .004 (bin/gap)... If I understand you correctly. The gap is pretty constant... ...also the bin"
        },
        {
          "id": 471220,
          "postDate": "2019-02-14T06:44:40.013Z",
          "content": "<p>Just very long gaps compared to the bins, so much we don't see. </p>",
          "rawMarkdown": "Just very long gaps compared to the bins, so much we don't see. "
        },
        {
          "id": 483418,
          "postDate": "2019-03-04T15:33:52.897Z",
          "content": "<p>Your summary of the gaps looks right to me, but I have to assume that because the time to fault is <strong>not</strong> evenly spaced (both within a 4096 chunk and between chunks), that the data reflects actual measurements rather than being assigned arbitrarily after the fact. </p>",
          "rawMarkdown": "Your summary of the gaps looks right to me, but I have to assume that because the time to fault is **not** evenly spaced (both within a 4096 chunk and between chunks), that the data reflects actual measurements rather than being assigned arbitrarily after the fact. "
        }
      ]
    },
    {
      "id": 455965,
      "postDate": "2019-01-14T22:01:12.263Z",
      "content": "<p>Thanks.\nStill, is <code>time_to_failure</code> measured in seconds in training dataset?\nThis kind of might explain 0.001 or 0.0011 time difference between every 4095/4096 points. But time difference between 2 measurements is 1.1e-9. How can that be?\nMight it happen, that intermediate <code>time_to_failure</code> values in traning data are miscalculated?</p>",
      "rawMarkdown": "Thanks.\nStill, is `time_to_failure` measured in seconds in training dataset?\nThis kind of might explain 0.001 or 0.0011 time difference between every 4095/4096 points. But time difference between 2 measurements is 1.1e-9. How can that be?\nMight it happen, that intermediate `time_to_failure` values in traning data are miscalculated?",
      "votes": 6,
      "replies": [
        {
          "id": 456364,
          "postDate": "2019-01-15T17:13:18.390Z",
          "content": "<p>May be it is a 4096 ADC that starting consequentally with 1 ns delay.</p>",
          "rawMarkdown": "May be it is a 4096 ADC that starting consequentally with 1 ns delay.",
          "votes": 2
        },
        {
          "id": 456496,
          "postDate": "2019-01-15T23:29:36.307Z",
          "content": "<p>Would this be a way of trying to reconstruct the whole signal from the seperate ones?</p>",
          "rawMarkdown": "Would this be a way of trying to reconstruct the whole signal from the seperate ones?",
          "votes": 1
        },
        {
          "id": 458580,
          "postDate": "2019-01-20T00:40:24.460Z",
          "content": "<p>Ok, now I'm pretty sure that device used in the experiment sends data every T = 0.001064s in chunks of N = 4096 measurements.\n&gt; T = (total time of training data) / ((number of data packets) - 1) = 163.4294s / (153600 - 1)</p>\n\n<p>Single time stamp returned with each chunk is rounded to 4 digits during time_to_failure calculation, producing 0.001 and 0.0011 time differences between chunks, in exactly same sequence which we can see in training data.</p>\n\n<p>120 packets have only 4095 measurements and short packets spreaded evenly accross training data (first packet in every block of 1280 packets is short). Probably, it has something to do with data transmission protocol.</p>\n\n<p>Anyway, frame rate that can be obtained from time_to_failure values is approximately:\n&gt; F = N/T = 3.849622 MHz</p>\n\n<p>4MHz would correspond to 0.001024s intervals between data packets.\nThere can be number of tasks, where 0.04ms might be spent inside measuring device.</p>\n\n<p>So, I understand why @BertrandRL claims measurement frequency to be 4MHz.</p>\n\n<p>Training data looks pretty much continuous without gaps and claimed to be so. And time_to_failure looks like really well aligned with acoustic_data. At least, acoustic peaks stay around time_to_failure=0.315s for all earthquakes.</p>",
          "rawMarkdown": "Ok, now I'm pretty sure that device used in the experiment sends data every T = 0.001064s in chunks of N = 4096 measurements.\n&gt; T = (total time of training data) / ((number of data packets) - 1) = 163.4294s / (153600 - 1)\n\nSingle time stamp returned with each chunk is rounded to 4 digits during time_to_failure calculation, producing 0.001 and 0.0011 time differences between chunks, in exactly same sequence which we can see in training data.\n\n120 packets have only 4095 measurements and short packets spreaded evenly accross training data (first packet in every block of 1280 packets is short). Probably, it has something to do with data transmission protocol.\n\nAnyway, frame rate that can be obtained from time_to_failure values is approximately:\n&gt; F = N/T = 3.849622 MHz\n\n4MHz would correspond to 0.001024s intervals between data packets.\nThere can be number of tasks, where 0.04ms might be spent inside measuring device.\n\nSo, I understand why @BertrandRL claims measurement frequency to be 4MHz.\n\nTraining data looks pretty much continuous without gaps and claimed to be so. And time_to_failure looks like really well aligned with acoustic_data. At least, acoustic peaks stay around time_to_failure=0.315s for all earthquakes.",
          "votes": 12
        },
        {
          "id": 466152,
          "postDate": "2019-02-04T19:30:36.150Z",
          "content": "<p>Thanks for bringing that up, there is indeed a gap every 4096 samples, an artifact of the recording device. I added clarification to my initial post, thanks!</p>",
          "rawMarkdown": "Thanks for bringing that up, there is indeed a gap every 4096 samples, an artifact of the recording device. I added clarification to my initial post, thanks!",
          "votes": 3
        },
        {
          "id": 469308,
          "postDate": "2019-02-10T23:18:48.807Z",
          "content": "<p>@BertrandRL, is it possible that the recording is contignuos, but at recording somehow the time data is \"compressed\" into this 4.5us bins?\nOr the other possibility is that there is recording for 4.5us and then there is no recording for 1060us.</p>\n\n<p>Which one do you think is happening in reality while recording? Thank you.</p>",
          "rawMarkdown": "@BertrandRL, is it possible that the recording is contignuos, but at recording somehow the time data is \"compressed\" into this 4.5us bins?\nOr the other possibility is that there is recording for 4.5us and then there is no recording for 1060us.\n\nWhich one do you think is happening in reality while recording? Thank you."
        },
        {
          "id": 471053,
          "postDate": "2019-02-14T00:29:47.513Z",
          "content": "<p>I'd love to believe in 12ms gaps between bins as stated in clarification, but it doesn't fit in my brain in any possible way with 1.064ms TTF difference between bins in training data.</p>\n\n<p>I used to work a bit with wireless low energy measuring devices. And that's why I believe, that 4096 sample bin represents data collected over previous 1.024ms then followed by 0.04ms gap, when measuring device is busy with some other job, probably with network activity.</p>",
          "rawMarkdown": "I'd love to believe in 12ms gaps between bins as stated in clarification, but it doesn't fit in my brain in any possible way with 1.064ms TTF difference between bins in training data.\n \nI used to work a bit with wireless low energy measuring devices. And that's why I believe, that 4096 sample bin represents data collected over previous 1.024ms then followed by 0.04ms gap, when measuring device is busy with some other job, probably with network activity.",
          "votes": 1
        }
      ]
    },
    {
      "id": 455498,
      "postDate": "2019-01-14T03:59:43.367Z",
      "content": "<p>Thanks the clarification. It is amazing that people are already getting models accurate to 1.5s using only 3/80 of a second!  It seems like trying to play \"Name That Tune\" with 1 note.</p>",
      "rawMarkdown": "Thanks the clarification. It is amazing that people are already getting models accurate to 1.5s using only 3/80 of a second!  It seems like trying to play \"Name That Tune\" with 1 note.",
      "votes": 6,
      "replies": [
        {
          "id": 457517,
          "postDate": "2019-01-17T15:44:51.630Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 519879,
      "postDate": "2019-04-19T19:25:02.350Z",
      "content": "<p>Bertrand RL:\n&gt; \nThe data is recorded in bins of 4096 samples. Within those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.</p>\n\n<p>From the time_to_failure column in the data, I measure a delta t of 1.1 ns (like many others) which is inconsistent with a sample rate of 4 MHz (250 ns?) by a huge factor. For the \"gap\", I get two possibilities: 1 millisecond and 1.1 milliseconds (1000 and 1100 microseconds) , again inconsistent with the 12 *<em>micro</em>*second quote.</p>\n\n<p>I assume that time_to_failure is measured in seconds, consistent with:\n&gt; The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds.</p>\n\n<p>But the discrepancies exists no matter the time units.</p>\n\n<p>I haven't looked at this for a while so I might be making a dumb error. If not, why the huge discrepancies?</p>\n\n<p>edit: reading other comments, this seems to be a regular question that hasn't been properly addressed. I think a paragraph or two explaining why all the Kaggler's measurements are incorrect might be in order.</p>",
      "rawMarkdown": "Bertrand RL:\n&gt; \nThe data is recorded in bins of 4096 samples. Within those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\n\nFrom the time_to_failure column in the data, I measure a delta t of 1.1 ns (like many others) which is inconsistent with a sample rate of 4 MHz (250 ns?) by a huge factor. For the \"gap\", I get two possibilities: 1 millisecond and 1.1 milliseconds (1000 and 1100 microseconds) , again inconsistent with the 12 **micro**second quote.\n\nI assume that time_to_failure is measured in seconds, consistent with:\n&gt; The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds.\n\nBut the discrepancies exists no matter the time units.\n\nI haven't looked at this for a while so I might be making a dumb error. If not, why the huge discrepancies?\n\nedit: reading other comments, this seems to be a regular question that hasn't been properly addressed. I think a paragraph or two explaining why all the Kaggler's measurements are incorrect might be in order.",
      "votes": 3,
      "replies": [
        {
          "id": 520105,
          "postDate": "2019-04-20T07:16:33.863Z",
          "content": "<p>Thanks for raising this again.  I am new here, having entered 2 days ago, but I am a bit puzzled to not see any response since 3 months form organizers. </p>\n\n<p>Here is some additional analysis.  I apologize it someone else have already published similar views.  </p>\n\n<p>Using ttf, one can estimate the total training time to be 163.42061, using:</p>\n\n<pre><code>df = train[train.time_to_failure &gt; train.time_to_failure.shift().fillna(0)]\ndf.time_to_failure.sum() - train.time_to_failure.values[-1]\n</code></pre>\n\n<p>If the data was indeed sampled at 4 MHz without any interruption,  then, as you pointed out, the time delta between each measure should be 0.25 microsecond, a.  Total time would be 157.28637, as computed by:</p>\n\n<pre><code>train.shape[0] / 4e6\n</code></pre>\n\n<p>The delta in total time is due to the interruptions.  If there is one every 4096 then the interruption is 40 microseconds:</p>\n\n<pre><code>(163.42061 - 157.28637) / (train.shape[0] / 4096)\n</code></pre>\n\n<p>If, simple hypothesis, someone mistakenly thought there was an interruption every 150000 observations, then the interruption is 1.4 millisecond, close to what was announced in the first place...:</p>\n\n<pre><code>(163.42061 - train.shape[0] / 4e6) / (train.shape[0] / 150000)\n</code></pre>",
          "rawMarkdown": "Thanks for raising this again.  I am new here, having entered 2 days ago, but I am a bit puzzled to not see any response since 3 months form organizers. \n\nHere is some additional analysis.  I apologize it someone else have already published similar views.  \n\nUsing ttf, one can estimate the total training time to be 163.42061, using:\n\n    df = train[train.time_to_failure &gt; train.time_to_failure.shift().fillna(0)]\n    df.time_to_failure.sum() - train.time_to_failure.values[-1]\n\nIf the data was indeed sampled at 4 MHz without any interruption,  then, as you pointed out, the time delta between each measure should be 0.25 microsecond, a.  Total time would be 157.28637, as computed by:\n\n    train.shape[0] / 4e6\n\nThe delta in total time is due to the interruptions.  If there is one every 4096 then the interruption is 40 microseconds:\n\n    (163.42061 - 157.28637) / (train.shape[0] / 4096)\n\nIf, simple hypothesis, someone mistakenly thought there was an interruption every 150000 observations, then the interruption is 1.4 millisecond, close to what was announced in the first place...:\n\n    (163.42061 - train.shape[0] / 4e6) / (train.shape[0] / 150000)\n\n",
          "votes": 6
        },
        {
          "id": 520200,
          "postDate": "2019-04-20T12:08:15.110Z",
          "content": "<p>According to data from p4581, sampling frequency is 3.968 MHz.\nIf we take into account 12e-6 s gap (might be not exactly 12e-6):\n<code>\n4096 / 3.968e6 + 12e-6\n0.001044258\n</code>\nTime vector has delta of <code>0.001044</code> (presumably seconds)\nTime vector is separated from acoustic data, so might as well be that discrepancy in our case is the result of data combining.</p>\n\n<p>Bertrand RL (from \"Introduction\"):\n&gt; p4581 is an experiment that has the same setup in terms of machine, material and recording apparatus.</p>",
          "rawMarkdown": "According to data from p4581, sampling frequency is 3.968 MHz.\nIf we take into account 12e-6 s gap (might be not exactly 12e-6):\n```\n4096 / 3.968e6 + 12e-6\n0.001044258\n```\nTime vector has delta of `0.001044` (presumably seconds)\nTime vector is separated from acoustic data, so might as well be that discrepancy in our case is the result of data combining.\n\nBertrand RL (from \"Introduction\"):\n&gt; p4581 is an experiment that has the same setup in terms of machine, material and recording apparatus.",
          "votes": 2
        },
        {
          "id": 520344,
          "postDate": "2019-04-20T18:47:09.633Z",
          "content": "<p>Issue is that you get a different ttf  that way.</p>",
          "rawMarkdown": "Issue is that you get a different ttf  that way."
        },
        {
          "id": 520492,
          "postDate": "2019-04-21T05:30:37.193Z",
          "content": "<p>You raides a good point. As I mentioned <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#458580\">in my another comment</a> there are 120 larger blocks of 1280 bins each in training dataset. First bin of each block has only 4095 samples. I would imagine there is a gap between these bigger blocks, rather then between bins. It gives:\n<code>\n(163.42061 - 157.28637) / 120 = 0.051118(6) ~ 0.0511\n</code>\nwhich is (512 - 1), which in turn points me to alternative way of thinking.</p>\n\n<p>Earlier <a href=\"/merepoule\">@merepoule</a> was saying about 12 milliseconds gap but later edited it based on comments. Lets imagine that initial statement was in fact accurate. Then we can calculate number of gaps:\n<code>\n(163.42061 - 157.28637) / 0.012 = 511.18(6) ~ 511\n</code>\nAnd number of blocks: <code>511 + 1 = 512</code>\nLooks pretty likely to me. Then blocks are of <code>153600 / 512 = 300</code> bins and <code>300 * 4096 = 1228800</code> samples each.\nThe problem with this theory though is that time_to_failure values assigned to training data are in fact contiguous (don't show any noticeable gaps).</p>\n\n<p>And if we are talking about microseconds, I don't think solutions accuracy will be affected in any way.</p>",
          "rawMarkdown": "You raides a good point. As I mentioned [in my another comment](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#458580) there are 120 larger blocks of 1280 bins each in training dataset. First bin of each block has only 4095 samples. I would imagine there is a gap between these bigger blocks, rather then between bins. It gives:\n```\n(163.42061 - 157.28637) / 120 = 0.051118(6) ~ 0.0511\n```\nwhich is (512 - 1), which in turn points me to alternative way of thinking.\n\nEarlier @merepoule was saying about 12 milliseconds gap but later edited it based on comments. Lets imagine that initial statement was in fact accurate. Then we can calculate number of gaps:\n```\n(163.42061 - 157.28637) / 0.012 = 511.18(6) ~ 511\n```\nAnd number of blocks: `511 + 1 = 512`\nLooks pretty likely to me. Then blocks are of `153600 / 512 = 300` bins and `300 * 4096 = 1228800` samples each.\nThe problem with this theory though is that time_to_failure values assigned to training data are in fact contiguous (don't show any noticeable gaps).\n\nAnd if we are talking about microseconds, I don't think solutions accuracy will be affected in any way.",
          "votes": 3
        },
        {
          "id": 521354,
          "postDate": "2019-04-22T20:22:52.993Z",
          "content": "<p>I guess that 4 MHz is simply a proxy. The data is actually recorded at 3.851 MHz. One bin contains 4096 samples measured in 4096*1.1 nanoseconds plus the artifact of 1 or 1.1 milliseconds. On average the artifact is 1.059 milliseconds. This means:\n<code>\n4096 / (4096 * 1.1*10**-9 + 1.059*10**-3) ~ 3851413 ~ 3.851 MHz\n</code></p>",
          "rawMarkdown": "I guess that 4 MHz is simply a proxy. The data is actually recorded at 3.851 MHz. One bin contains 4096 samples measured in 4096*1.1 nanoseconds plus the artifact of 1 or 1.1 milliseconds. On average the artifact is 1.059 milliseconds. This means:\n```\n4096 / (4096 * 1.1*10**-9 + 1.059*10**-3) ~ 3851413 ~ 3.851 MHz\n```",
          "votes": 8
        },
        {
          "id": 521685,
          "postDate": "2019-04-23T09:07:11.993Z",
          "content": "<p><a href=\"/elvenmonk\">@elvenmonk</a> You nailed it IMHO.</p>",
          "rawMarkdown": "@elvenmonk You nailed it IMHO."
        },
        {
          "id": 522025,
          "postDate": "2019-04-23T19:15:58.593Z",
          "content": "<p><a href=\"/danijelk\">@danijelk</a> I don't believe that data can be and actually is measured every 1.1 nanoseconds. There are no any visible shifts in acoustic data on the edge of adjacent 4096 bins. You can find accurate average frequency measurement in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#458580\">my another comment</a>\nMeantime, I shared <a href=\"https://www.kaggle.com/elvenmonk/lanl-ttf-error-validation\">Kernel</a> that approximates <code>time_to_failure</code> values in training data set by linear time distribution with 0.6 milliseconds precision.</p>",
          "rawMarkdown": "@danijelk I don't believe that data can be and actually is measured every 1.1 nanoseconds. There are no any visible shifts in acoustic data on the edge of adjacent 4096 bins. You can find accurate average frequency measurement in [my another comment](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#458580)\nMeantime, I shared [Kernel](https://www.kaggle.com/elvenmonk/lanl-ttf-error-validation) that approximates `time_to_failure` values in training data set by linear time distribution with 0.6 milliseconds precision.",
          "votes": 2
        },
        {
          "id": 522063,
          "postDate": "2019-04-23T20:16:43.210Z",
          "content": "<p><a href=\"/elvenmonk\">@elvenmonk</a> Data in p4581 are chunks of size 20971520 bytes.\n<code>1280 * 4096 * 2 (channels) * 2 (bytes per sample) == 20971520</code></p>",
          "rawMarkdown": "@elvenmonk Data in p4581 are chunks of size 20971520 bytes.\n`1280 * 4096 * 2 (channels) * 2 (bytes per sample) == 20971520`"
        },
        {
          "id": 522095,
          "postDate": "2019-04-23T21:21:04.243Z",
          "content": "<p><a href=\"/elvenmonk\">@elvenmonk</a> I guess you are right. Your numbers are more precise than mine</p>",
          "rawMarkdown": "@elvenmonk I guess you are right. Your numbers are more precise than mine"
        },
        {
          "id": 522163,
          "postDate": "2019-04-23T23:59:10.030Z",
          "content": "<p><a href=\"/mykper\">@mykper</a> Awesome! Thanks for referencing P4581 experiment again. I'll take a deeper look into related materials, specifically to <a href=\"https://permalink.lanl.gov/object/tr?what=info%3Alanl-repo%2Flareport%2FLA-UR-17-29312\">this earlier LANL research</a></p>",
          "rawMarkdown": "@mykper Awesome! Thanks for referencing P4581 experiment again. I'll take a deeper look into related materials, specifically to [this earlier LANL research](https://permalink.lanl.gov/object/tr?what=info%3Alanl-repo%2Flareport%2FLA-UR-17-29312)"
        }
      ]
    },
    {
      "id": 504511,
      "postDate": "2019-03-31T18:53:19.310Z",
      "content": "<p>Hello Bertrand,</p>\n\n<p>I have a question about the behaviour of the earthquake machine during an experiment.</p>\n\n<p>I suppose the physical characteristics of the three blocks remain unchanged during the test, but what about the two granular layers?</p>\n\n<p>Do they not suffer damage that could alter their physical characteristics during the experiment?</p>\n\n<p>Regards,</p>\n\n<p>Daniel</p>",
      "rawMarkdown": "Hello Bertrand,\n\nI have a question about the behaviour of the earthquake machine during an experiment.\n\nI suppose the physical characteristics of the three blocks remain unchanged during the test, but what about the two granular layers?\n\nDo they not suffer damage that could alter their physical characteristics during the experiment?\n\nRegards,\n\nDaniel",
      "votes": 3
    },
    {
      "id": 490591,
      "postDate": "2019-03-14T17:27:43.083Z",
      "content": "<p>Hi Bertrand, Thank you for this context.\nI have two questions about data.\n1) My first question is about test datasets. Each sequence of 150000 data precedes an earthquake: the 2624 test sets are precursors of 2624 different earthquakes?\n2) The variable time_to_failure in train dataset has an average frequency of 3.85 Mhz, given 4095 or 4096 intervals of 10^(-9) seconds followed by an interval of about 0.001 seconds. The sequence is decreasing, except for 16 jumps that correspond to 16 earthquakes.\nMy second question is: test data times are claimed to be at 4 Mhz. Are they evenly spaced, or do they follow the irregular pattern of time_to_failure in train dataset?</p>\n\n<p>Thank you for any possible help.</p>\n\n<p>Marcello</p>",
      "rawMarkdown": "Hi Bertrand, Thank you for this context.\nI have two questions about data.\n1) My first question is about test datasets. Each sequence of 150000 data precedes an earthquake: the 2624 test sets are precursors of 2624 different earthquakes?\n2) The variable time_to_failure in train dataset has an average frequency of 3.85 Mhz, given 4095 or 4096 intervals of 10^(-9) seconds followed by an interval of about 0.001 seconds. The sequence is decreasing, except for 16 jumps that correspond to 16 earthquakes.\nMy second question is: test data times are claimed to be at 4 Mhz. Are they evenly spaced, or do they follow the irregular pattern of time_to_failure in train dataset?\n\nThank you for any possible help.\n\nMarcello",
      "votes": 3
    },
    {
      "id": 476110,
      "postDate": "2019-02-21T16:10:15.857Z",
      "content": "<p>I case someone missed that too, 12 ms was changed to 12 MICROseconds</p>",
      "rawMarkdown": "I case someone missed that too, 12 ms was changed to 12 MICROseconds",
      "votes": 3,
      "replies": [
        {
          "id": 476311,
          "postDate": "2019-02-22T00:29:17.177Z",
          "content": "<p>I noticed, but that confuses me even more.</p>",
          "rawMarkdown": "I noticed, but that confuses me even more.",
          "votes": 1
        },
        {
          "id": 477136,
          "postDate": "2019-02-24T01:02:36.663Z",
          "content": "<p>Which is funny since it is about 1 millisecond in reality...</p>",
          "rawMarkdown": "Which is funny since it is about 1 millisecond in reality...",
          "votes": 3
        }
      ]
    },
    {
      "id": 475373,
      "postDate": "2019-02-20T17:22:18.527Z",
      "content": "<p>Is it possible that failure dectection <em>fail</em> to detect some events?</p>\n\n<p>Or that audio signal spikes are related to a some kind of relaxation of the internal state, also if not leading to a failure?</p>\n\n<p>Please find in this <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/81302#latest-475311\">discussion</a> some consideration about that.</p>\n\n<p><img src=\"https://drive.google.com/uc?id=1UbCr2m1SBxNlnLm9Frd1zi2CRiVcaU-e\" alt=\"Graph of prediction over the entire training set\"></p>",
      "rawMarkdown": "Is it possible that failure dectection *fail* to detect some events?\n\nOr that audio signal spikes are related to a some kind of relaxation of the internal state, also if not leading to a failure?\n\nPlease find in this [discussion][1] some consideration about that.\n\n![Graph of prediction over the entire training set][2]\n\n[1]:https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/81302#latest-475311\n[2]: https://drive.google.com/uc?id=1UbCr2m1SBxNlnLm9Frd1zi2CRiVcaU-e",
      "votes": 3,
      "replies": [
        {
          "id": 475576,
          "postDate": "2019-02-20T23:22:40.363Z",
          "content": "<p>You've got it) The biggest challenge of this challenge is to figure out whether next spike will be powerful enough to be considered failure and if not how to know time between next spike and real failure (given that relaxation after \"fake\" spike is not yet known). I guess if test data has 2 consequent non-failure spikes, predictions will fail badly on those segments.</p>",
          "rawMarkdown": "You've got it) The biggest challenge of this challenge is to figure out whether next spike will be powerful enough to be considered failure and if not how to know time between next spike and real failure (given that relaxation after \"fake\" spike is not yet known). I guess if test data has 2 consequent non-failure spikes, predictions will fail badly on those segments.",
          "votes": 3
        },
        {
          "id": 477780,
          "postDate": "2019-02-25T09:00:14.943Z",
          "content": "<p>I wouldn't say it is the biggest challenge though, because how would you use that info? In the test set you don't know how 150k bin's are connected, so you have to guess it on an individual basis.</p>\n\n<p>It is most likely that this info is not inherent in the 150k bins, well it was not assigned to them as per measurement, but through their timely order retrospectively.</p>",
          "rawMarkdown": "I wouldn't say it is the biggest challenge though, because how would you use that info? In the test set you don't know how 150k bin's are connected, so you have to guess it on an individual basis.\n\nIt is most likely that this info is not inherent in the 150k bins, well it was not assigned to them as per measurement, but through their timely order retrospectively.\n"
        }
      ]
    },
    {
      "id": 467226,
      "postDate": "2019-02-06T17:01:33.317Z",
      "content": "<p>The 12ms-gap (see \"Additional Info\") between chunks is in contradictory to the time_to_failure (see train.csv) !</p>",
      "rawMarkdown": "The 12ms-gap (see \"Additional Info\") between chunks is in contradictory to the time_to_failure (see train.csv) !",
      "votes": 3
    },
    {
      "id": 456814,
      "postDate": "2019-01-16T15:30:10.673Z",
      "content": "<p>Could you please verify to following guess/statement as mentioned bij <a href=\"/chiqun\">@chiqun</a> :My guess is the acoustic sampling rate is 4MHz, but the timeoffailure sample rate is 1000Hz. Need to be verified by the author.</p>",
      "rawMarkdown": "Could you please verify to following guess/statement as mentioned bij @chiqun :My guess is the acoustic sampling rate is 4MHz, but the timeoffailure sample rate is 1000Hz. Need to be verified by the author.",
      "votes": 3
    },
    {
      "id": 456001,
      "postDate": "2019-01-14T23:35:02.403Z",
      "content": "<p>I have some short questions for clarification.</p>\n\n<p>good to know that the input data is recorded in volts. Is that true for both training and test data, making them BOTH in precisely the same units, unnormalized? From the description, it seems that it is but I want to make sure.</p>\n\n<p>you say that training and testing are from the same experiment. I assume that there is NO overlap between the two data sets - is that true? Also, does the entire test dataset come from one earthquake cycle or many?</p>\n\n<p>finally, is the training set time series \"stationary\" meaning that each earthquake cycle is an equivalent time series modulo statistical fluctuations? Or does the physics change as the experiment progresses?</p>",
      "rawMarkdown": "I have some short questions for clarification.\n\n good to know that the input data is recorded in volts. Is that true for both training and test data, making them BOTH in precisely the same units, unnormalized? From the description, it seems that it is but I want to make sure.\n\nyou say that training and testing are from the same experiment. I assume that there is NO overlap between the two data sets - is that true? Also, does the entire test dataset come from one earthquake cycle or many?\n\nfinally, is the training set time series \"stationary\" meaning that each earthquake cycle is an equivalent time series modulo statistical fluctuations? Or does the physics change as the experiment progresses?\n",
      "votes": 3,
      "replies": [
        {
          "id": 456504,
          "postDate": "2019-01-15T23:54:47.027Z",
          "content": "<p>Hi Pete, the answer to all your questions is yes: the recording device is the same throughout the experiment, including for the training and the testing set. There is no overlap between the sets, that are contiguous. There are several earthquake cycles in the test set as well. The time series is stationary and the physics unchanged except for a small experimental artifact which is that some material from the fault is progressively lost as the experiment proceeds.</p>",
          "rawMarkdown": "Hi Pete, the answer to all your questions is yes: the recording device is the same throughout the experiment, including for the training and the testing set. There is no overlap between the sets, that are contiguous. There are several earthquake cycles in the test set as well. The time series is stationary and the physics unchanged except for a small experimental artifact which is that some material from the fault is progressively lost as the experiment proceeds.",
          "votes": 6
        },
        {
          "id": 456557,
          "postDate": "2019-01-16T03:43:14.707Z",
          "content": "<p>@BertrandRL: Thanks for your answer. May I ask you that do all the test chunks constitute a full and contiguous signal (by some order) ? Or they cannot be assembled into a contiguous signal?</p>",
          "rawMarkdown": "@BertrandRL: Thanks for your answer. May I ask you that do all the test chunks constitute a full and contiguous signal (by some order) ? Or they cannot be assembled into a contiguous signal?",
          "votes": 3
        },
        {
          "id": 462065,
          "postDate": "2019-01-27T15:33:41.317Z",
          "content": "<p>Bertrand,\nI understand that the test data occurs just after the train data.\nOne information is also important IMO: how is the test data split into public and private test data? \n1. The 13% public data (used for the public Leaderboard), corresponds to the first events of the test data? \n2. Or is it randomly split over the 100% test data?\nThanks</p>",
          "rawMarkdown": "Bertrand,\nI understand that the test data occurs just after the train data.\nOne information is also important IMO: how is the test data split into public and private test data? \n1. The 13% public data (used for the public Leaderboard), corresponds to the first events of the test data? \n2. Or is it randomly split over the 100% test data?\nThanks",
          "votes": 2
        }
      ]
    },
    {
      "id": 466481,
      "postDate": "2019-02-05T13:19:10.200Z",
      "content": "<p>contiguous == sharing a common border; touching.\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12ms gap between each bin, an artifact of the recording device. </p>\n\n<p>LOL</p>",
      "rawMarkdown": "contiguous == sharing a common border; touching.\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12ms gap between each bin, an artifact of the recording device. \n\nLOL",
      "votes": 3,
      "replies": [
        {
          "id": 466780,
          "postDate": "2019-02-05T23:16:08.527Z",
          "content": "<p>How did you get it's an artifact?</p>",
          "rawMarkdown": "How did you get it's an artifact?"
        },
        {
          "id": 467263,
          "postDate": "2019-02-06T18:12:50.767Z",
          "content": "<p>Read point 6 of his comment</p>",
          "rawMarkdown": "Read point 6 of his comment"
        },
        {
          "id": 467339,
          "postDate": "2019-02-06T22:30:15.447Z",
          "content": "<p>Ups, I missed it somehow, thank you. Now I see a lot of posts spotting that </p>\n\n<p>lol indeed</p>",
          "rawMarkdown": "Ups, I missed it somehow, thank you. Now I see a lot of posts spotting that \n\nlol indeed",
          "votes": 1
        }
      ]
    },
    {
      "id": 505489,
      "postDate": "2019-04-02T04:20:15.533Z",
      "content": "<p>Hello Mr Bertrand RL,</p>\n\n<p><strong>You said the data is recorded in bins of 4096 \"samples\"</strong>, does sample mean data points or the collection of 150000 data points? </p>\n\n<p><strong>And after each bin there is a 12 microseconds gap.</strong> Is the data still being recorded during these 12MS gap? If so, is this noise recorded at 4MHz?</p>",
      "rawMarkdown": "Hello Mr Bertrand RL,\n\n**You said the data is recorded in bins of 4096 \"samples\"**, does sample mean data points or the collection of 150000 data points? \n\n**And after each bin there is a 12 microseconds gap.** Is the data still being recorded during these 12MS gap? If so, is this noise recorded at 4MHz?",
      "votes": 1,
      "replies": [
        {
          "id": 507299,
          "postDate": "2019-04-04T14:30:18.943Z",
          "content": "<p>Hi! I would assume that: \n1) If 1 sample was 150_000 data points, 4096 samples would contain 614.4 millions of data points. And the whole data set is 629.145 millions, so it would be less than 2 bins, and Bertrand says \"The data is recorded in bins\" - which is plural. If we somehow estimate the approximate time duration of the whole training dataset, we could clarify this question completely.\n2) In a comment below Bertrand said that in the gap the data is missing, so, most likely, it was not recorded.</p>",
          "rawMarkdown": "Hi! I would assume that: \n1) If 1 sample was 150_000 data points, 4096 samples would contain 614.4 millions of data points. And the whole data set is 629.145 millions, so it would be less than 2 bins, and Bertrand says \"The data is recorded in bins\" - which is plural. If we somehow estimate the approximate time duration of the whole training dataset, we could clarify this question completely.\n2) In a comment below Bertrand said that in the gap the data is missing, so, most likely, it was not recorded."
        }
      ]
    },
    {
      "id": 491698,
      "postDate": "2019-03-16T02:30:33.983Z",
      "content": "<p>Hi Bertrand,</p>\n\n<p>If test set are contiguous, participants will track segments from future to past after finding obvious earthquake. Is this solution acceptable for you?</p>\n\n<p>Thanks, </p>",
      "rawMarkdown": "Hi Bertrand,\n\nIf test set are contiguous, participants will track segments from future to past after finding obvious earthquake. Is this solution acceptable for you?\n\nThanks, ",
      "votes": 1
    },
    {
      "id": 464314,
      "postDate": "2019-01-31T15:30:46.407Z",
      "content": "<p>What about problem with size of train and test files? Do you have some problems with data processing on Kaggle instance?</p>",
      "rawMarkdown": "What about problem with size of train and test files? Do you have some problems with data processing on Kaggle instance?",
      "votes": 1
    },
    {
      "id": 463771,
      "postDate": "2019-01-30T15:37:28.707Z",
      "content": "<p>Hello,\nCould we get more detailed information on which piezoceramic sensor you are using to gather the seismic data? \nThanks in advance.</p>",
      "rawMarkdown": "Hello,\nCould we get more detailed information on which piezoceramic sensor you are using to gather the seismic data? \nThanks in advance.",
      "votes": 1
    },
    {
      "id": 522038,
      "postDate": "2019-04-23T19:38:19.393Z",
      "content": "<p>Here is a picture of what  get for a series of length twice the test data files. The plot is of delta_t between samples. Each red mark marks the end of 4096 samples. The black line at the top looks like zero but it is really -1.1 ns - way too small of course. The negative values, the \"gaps\", are -1ms and -1.1 ms. They come in threes with the first smaller in magnitude than the next two. Everything is separated by 4096 samples (you would have to be able to zoom to see this)\nIt's a pretty raw plot - little or no manipulation but comments on whether it is correct are welcome.</p>\n\n<p>I see the gaps as, roughly, making up for the fact that the majority of intervals have the (incorrect) 1.1 ns sampling. The error is mitigated at the gaps. If this were exact, then every gap would be ~1.018 ms (instead of the 1ms/1.1ms mixture), which is why I characterize it as \"rough\".</p>",
      "rawMarkdown": "Here is a picture of what  get for a series of length twice the test data files. The plot is of delta_t between samples. Each red mark marks the end of 4096 samples. The black line at the top looks like zero but it is really -1.1 ns - way too small of course. The negative values, the \"gaps\", are -1ms and -1.1 ms. They come in threes with the first smaller in magnitude than the next two. Everything is separated by 4096 samples (you would have to be able to zoom to see this)\nIt's a pretty raw plot - little or no manipulation but comments on whether it is correct are welcome.\n\nI see the gaps as, roughly, making up for the fact that the majority of intervals have the (incorrect) 1.1 ns sampling. The error is mitigated at the gaps. If this were exact, then every gap would be ~1.018 ms (instead of the 1ms/1.1ms mixture), which is why I characterize it as \"rough\".",
      "votes": 2
    },
    {
      "id": 476288,
      "postDate": "2019-02-21T22:58:37.267Z",
      "content": "<p>I would be very interested to know whether Public portion of test data is contiguous slice in time rather than a random sample.</p>\n\n<p>My guess it is, and so it might be absolutely worthless to snoop on the leaderboard.</p>\n\n<p>Also am I right to assume that the whole test data in not in timely order?</p>\n\n<p>Thanks for the input...</p>",
      "rawMarkdown": "I would be very interested to know whether Public portion of test data is contiguous slice in time rather than a random sample.\n\nMy guess it is, and so it might be absolutely worthless to snoop on the leaderboard.\n\nAlso am I right to assume that the whole test data in not in timely order?\n\nThanks for the input...",
      "votes": 2
    },
    {
      "id": 466169,
      "postDate": "2019-02-04T19:52:44.750Z",
      "content": "<p>Hi Bertrand, \nIs it possible to share more information on the splitting of the test set between public and private LB? Was the split at random with similar distribution, random without looking at the distribution, shared earthquakes, or different earthquakes? In other words, how really useful is the public LB?</p>",
      "rawMarkdown": "Hi Bertrand, \nIs it possible to share more information on the splitting of the test set between public and private LB? Was the split at random with similar distribution, random without looking at the distribution, shared earthquakes, or different earthquakes? In other words, how really useful is the public LB?",
      "votes": 2
    },
    {
      "id": 460559,
      "postDate": "2019-01-24T00:33:23.937Z",
      "content": "<p>A few have mentioned this point:</p>\n\n<p>In the training data, measurements seem to have been taken every 1 ns. However, every 4096 measurements, there is a gap of 1ms in data.</p>\n\n<p>I assume this is an experimental constraint, and as such the test data should contain the same gaps.</p>\n\n<p>However, can we be given a time indication of the measurements, so that we could know where these gaps appear in the test data?\nSomething like the time difference between the first and the n-th measurement maybe (for all test measurements)?</p>",
      "rawMarkdown": "A few have mentioned this point:\n\nIn the training data, measurements seem to have been taken every 1 ns. However, every 4096 measurements, there is a gap of 1ms in data.\n\nI assume this is an experimental constraint, and as such the test data should contain the same gaps.\n\nHowever, can we be given a time indication of the measurements, so that we could know where these gaps appear in the test data?\nSomething like the time difference between the first and the n-th measurement maybe (for all test measurements)?",
      "votes": 2,
      "replies": [
        {
          "id": 460678,
          "postDate": "2019-01-24T08:27:13.333Z",
          "content": "<p>Yeah, because it's quite possible that precisely that millisecond of data holds the key to predicting a quake that will occur about 4 seconds later.</p>",
          "rawMarkdown": "Yeah, because it's quite possible that precisely that millisecond of data holds the key to predicting a quake that will occur about 4 seconds later."
        },
        {
          "id": 467170,
          "postDate": "2019-02-06T15:26:12.387Z",
          "content": "<p>This would be quite useful if we are seriously trying to find a way to predict earthquakes, since any method based on frequencies will need to know where the gaps are in the data.  Without it we are mostly stuck using magnitude of the signal which is only one part of the information that we should be able to use.  Of course, the ideal solution is to get better measuring equipment that can record data at the same time as measuring it, but I don't expect that to happen during the challenge.</p>",
          "rawMarkdown": "This would be quite useful if we are seriously trying to find a way to predict earthquakes, since any method based on frequencies will need to know where the gaps are in the data.  Without it we are mostly stuck using magnitude of the signal which is only one part of the information that we should be able to use.  Of course, the ideal solution is to get better measuring equipment that can record data at the same time as measuring it, but I don't expect that to happen during the challenge.",
          "votes": 1
        },
        {
          "id": 467299,
          "postDate": "2019-02-06T20:20:40.867Z",
          "content": "<p>Lomscargle or Gaussian Process could be options?</p>",
          "rawMarkdown": "Lomscargle or Gaussian Process could be options?"
        },
        {
          "id": 469140,
          "postDate": "2019-02-10T15:20:17.987Z",
          "content": "<p>Daniel Gerigk: if you don't know where these gaps ar,e frequency analysis (atleast for some frequencies) becomes useless. So Yes, precisely that millisecond is very-very important...</p>",
          "rawMarkdown": "Daniel Gerigk: if you don't know where these gaps ar,e frequency analysis (atleast for some frequencies) becomes useless. So Yes, precisely that millisecond is very-very important..."
        },
        {
          "id": 469253,
          "postDate": "2019-02-10T19:36:17.887Z",
          "content": "<p>It is possible that you are right.</p>",
          "rawMarkdown": "It is possible that you are right."
        }
      ]
    },
    {
      "id": 457098,
      "postDate": "2019-01-16T23:59:04.553Z",
      "content": "<p>Hi Bertrand,</p>\n\n<p>the train dataset is contiguous in time, so are the test dataset segments. Concerning the test datasets segments, are they contiguous by the order they appear on the data section?\nFinally, is the whole test dataset contiguous in time with the train dataset?</p>\n\n<p>thanks in advance</p>",
      "rawMarkdown": "Hi Bertrand,\n\nthe train dataset is contiguous in time, so are the test dataset segments. Concerning the test datasets segments, are they contiguous by the order they appear on the data section?\nFinally, is the whole test dataset contiguous in time with the train dataset?\n\nthanks in advance",
      "votes": 2,
      "replies": [
        {
          "id": 476289,
          "postDate": "2019-02-21T23:00:26.320Z",
          "content": "<p>Most important question here, and nobody answering it...</p>\n\n<p>My guess it is not contignuos as they appear...</p>",
          "rawMarkdown": "Most important question here, and nobody answering it...\n\nMy guess it is not contignuos as they appear...",
          "votes": 2
        }
      ]
    },
    {
      "id": 456330,
      "postDate": "2019-01-15T15:29:15.260Z",
      "content": "<p>How do you decide when there has been an earthquake? Specifically, is the classification based on reaching a threshold in the acoustic signal?</p>",
      "rawMarkdown": "How do you decide when there has been an earthquake? Specifically, is the classification based on reaching a threshold in the acoustic signal?",
      "votes": 2,
      "replies": [
        {
          "id": 456376,
          "postDate": "2019-01-15T17:38:59.087Z",
          "content": "<p>There has been an earthquake when time_to_failure reaches 0, there are 16 earthquake in the training set.</p>",
          "rawMarkdown": "There has been an earthquake when time_to_failure reaches 0, there are 16 earthquake in the training set."
        },
        {
          "id": 456502,
          "postDate": "2019-01-15T23:50:41.857Z",
          "content": "<p>Hi Misp, I added the answer in my original post, thanks for asking.</p>",
          "rawMarkdown": "Hi Misp, I added the answer in my original post, thanks for asking.",
          "votes": 2
        },
        {
          "id": 456663,
          "postDate": "2019-01-16T09:02:11.520Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        },
        {
          "id": 457726,
          "postDate": "2019-01-18T01:34:51.343Z",
          "content": "<p>\"There has been an earthquake when timetofailure reaches 0\"\nThe first example (5656573) has nothing looking like a quake starting at that point.  Far as I can see the quake happens shortly beforehand.</p>",
          "rawMarkdown": "\"There has been an earthquake when timetofailure reaches 0\"\nThe first example (5656573) has nothing looking like a quake starting at that point.  Far as I can see the quake happens shortly beforehand.",
          "votes": 2
        },
        {
          "id": 463166,
          "postDate": "2019-01-29T15:03:50.963Z",
          "content": "<p>0 is not part of the data though, so you will never actually have an earth quake in the acoustic data as it would make the whole purpose of that challenge pointless. \nIf you can predict it when it happens, it's too late already.\nThe other events that are happening before time 0 are precursors.</p>",
          "rawMarkdown": "0 is not part of the data though, so you will never actually have an earth quake in the acoustic data as it would make the whole purpose of that challenge pointless. \nIf you can predict it when it happens, it's too late already.\nThe other events that are happening before time 0 are precursors.",
          "votes": 2
        }
      ]
    },
    {
      "id": 455994,
      "postDate": "2019-01-14T22:58:04.460Z",
      "content": "<p>If the test folder contains the files that are chunks of the single big train file, than why each seg file does not include the time_to_failure data as a second column? Just to keep us busy  finding each chunk in the train file instead of working on a model?  </p>",
      "rawMarkdown": "If the test folder contains the files that are chunks of the single big train file, than why each seg file does not include the time_to_failure data as a second column? Just to keep us busy  finding each chunk in the train file instead of working on a model?  ",
      "replies": [
        {
          "id": 456515,
          "postDate": "2019-01-16T00:07:24.510Z",
          "content": "<p>Hi i262666, the test file comes from a separate portion of the experiment, it is not from the training set.</p>",
          "rawMarkdown": "Hi i262666, the test file comes from a separate portion of the experiment, it is not from the training set.",
          "votes": 1
        }
      ]
    },
    {
      "id": 541180,
      "postDate": "2019-06-02T01:23:05.517Z",
      "content": "<p>I found that there is a correlation coefficient -0.6 between (time margin of the input peaks )/abs(input peak's value) and the answer's peak value.</p>",
      "rawMarkdown": "I found that there is a correlation coefficient -0.6 between (time margin of the input peaks )/abs(input peak's value) and the answer's peak value."
    },
    {
      "id": 535069,
      "postDate": "2019-05-22T08:36:30.600Z",
      "content": "<p>Can you tell me the relationship between voltage and seismic data? I also find that the data in train is not enough 4096 samples, it's true?</p>",
      "rawMarkdown": "Can you tell me the relationship between voltage and seismic data? I also find that the data in train is not enough 4096 samples, it's true?",
      "replies": [
        {
          "id": 536588,
          "postDate": "2019-05-24T18:26:18.420Z",
          "content": "<p>About your first question- a sensor records the <em>acoustic emission</em> during the experiment.</p>",
          "rawMarkdown": "About your first question- a sensor records the *acoustic emission* during the experiment."
        }
      ]
    },
    {
      "id": 527744,
      "postDate": "2019-05-06T07:59:51.863Z",
      "content": "<p>Hi Bertrand. Could you say what type of instrumentation do you use? in order to perform the instrumental correction. If you have the poles, zeros and the constant, would be great, since I read the article you posted and you use obspy, thanks.</p>",
      "rawMarkdown": "Hi Bertrand. Could you say what type of instrumentation do you use? in order to perform the instrumental correction. If you have the poles, zeros and the constant, would be great, since I read the article you posted and you use obspy, thanks."
    },
    {
      "id": 527492,
      "postDate": "2019-05-05T16:46:38.290Z",
      "content": "<p>In the kernel, whenever I try to load the training data in dataframe, the kernel crashes. What should i do ?</p>",
      "rawMarkdown": "In the kernel, whenever I try to load the training data in dataframe, the kernel crashes. What should i do ?",
      "replies": [
        {
          "id": 527506,
          "postDate": "2019-05-05T17:18:33.610Z",
          "content": "<p>You can try loading the training data from my feather format loaded on float32 instead of float64.  float32 should be sufficient with getting started. See below.</p>\n\n<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330#latest-521640\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330#latest-521640</a></p>",
          "rawMarkdown": "You can try loading the training data from my feather format loaded on float32 instead of float64.  float32 should be sufficient with getting started. See below.\n\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330#latest-521640",
          "votes": 1
        }
      ]
    },
    {
      "id": 525853,
      "postDate": "2019-05-01T20:29:49.877Z",
      "content": "<p>Here for the level upgrade.</p>",
      "rawMarkdown": "Here for the level upgrade."
    },
    {
      "id": 525327,
      "postDate": "2019-04-30T18:05:46.260Z",
      "content": "<p>This really helped in understanding the problem as the overview is not very lucid. Cheers !!</p>",
      "rawMarkdown": "This really helped in understanding the problem as the overview is not very lucid. Cheers !!"
    },
    {
      "id": 522727,
      "postDate": "2019-04-24T22:44:56.480Z",
      "content": "<p>Several have wondered why we are recording at ~4 MHz when google searches are not finding accelerometers that can measure such a high acoustic frequency. In fact, in the \"supporting information\" referenced, we see:\n&gt; The sampling rate of the acoustic data is 330kHz. We are band-limited by the accelerometer (the frequency response is poor above about 50kHz).</p>\n\n<p>This suggests that state-of-the-art fails above 50 kHz. The spectrum we see in the competition data  goes much higher (see the attached image for a spectral measurement).</p>\n\n<p>The question I have is: was there a significant upgrade in the acoustic sensor for the data we are working with? Should we trust all of that high frequency \"signal\". Are we looking at a set of spurious detector resonances?</p>\n\n<p>It is important information! </p>\n\n<p>I hope the organizers can help with this.</p>",
      "rawMarkdown": "Several have wondered why we are recording at ~4 MHz when google searches are not finding accelerometers that can measure such a high acoustic frequency. In fact, in the \"supporting information\" referenced, we see:\n&gt; The sampling rate of the acoustic data is 330kHz. We are band-limited by the accelerometer (the frequency response is poor above about 50kHz).\n\nThis suggests that state-of-the-art fails above 50 kHz. The spectrum we see in the competition data  goes much higher (see the attached image for a spectral measurement).\n\nThe question I have is: was there a significant upgrade in the acoustic sensor for the data we are working with? Should we trust all of that high frequency \"signal\". Are we looking at a set of spurious detector resonances?\n\n It is important information! \n\nI hope the organizers can help with this.",
      "replies": [
        {
          "id": 522757,
          "postDate": "2019-04-25T00:40:55.140Z",
          "content": "<p>OK I looked at the references again and found a discussion of the piezo transducer in:</p>\n\n<p><a href=\"https://www.nature.com/articles/s41561-018-0272-8.epdf?author_access_token=8kRQlLxIzyiAKxzY4NhmldRgN0jAjWel9jnR3ZoTv0OT5J4u5SavZNgoMEsNMRwe5SW-57Q_oC2JdZbNN9Lxfx_YfiU00o_eAGb1b0hj9Tol9jOT0XgifHp4eFMfNY6k4BbIBtEnrhdPAcp6m-CjoQ%3D%3D\">https://www.nature.com/articles/s41561-018-0272-8.epdf?author_access_token=8kRQlLxIzyiAKxzY4NhmldRgN0jAjWel9jnR3ZoTv0OT5J4u5SavZNgoMEsNMRwe5SW-57Q_oC2JdZbNN9Lxfx_YfiU00o_eAGb1b0hj9Tol9jOT0XgifHp4eFMfNY6k4BbIBtEnrhdPAcp6m-CjoQ%3D%3D</a></p>\n\n<p>It is the third paper referenced in \"Introduction\".\n Bottom of page 1 states that the transducer is accurate from 0.02-2 MHz. That explains their choice of 4 MHz. So I'll just take that at face value.</p>",
          "rawMarkdown": "OK I looked at the references again and found a discussion of the piezo transducer in:\n\nhttps://www.nature.com/articles/s41561-018-0272-8.epdf?author_access_token=8kRQlLxIzyiAKxzY4NhmldRgN0jAjWel9jnR3ZoTv0OT5J4u5SavZNgoMEsNMRwe5SW-57Q_oC2JdZbNN9Lxfx_YfiU00o_eAGb1b0hj9Tol9jOT0XgifHp4eFMfNY6k4BbIBtEnrhdPAcp6m-CjoQ%3D%3D\n\nIt is the third paper referenced in \"Introduction\".\n Bottom of page 1 states that the transducer is accurate from 0.02-2 MHz. That explains their choice of 4 MHz. So I'll just take that at face value.",
          "votes": 1
        }
      ]
    },
    {
      "id": 504433,
      "postDate": "2019-03-31T16:03:33.120Z",
      "content": "<p>Does it help to apply a low pass filter to the data? </p>",
      "rawMarkdown": "Does it help to apply a low pass filter to the data? "
    },
    {
      "id": 472130,
      "postDate": "2019-02-15T11:58:35.153Z",
      "content": "<p>Hi Bertrand,</p>\n\n<p>1) the acoustic sensor is attached to the central plunger, or to one of the 2 side plates that provide normal force?</p>\n\n<p>2) in the experiment corresponding to the kaggle data, is instantaneous pressure or instantaneous velocity constant? my gut feeling tells me that on longer timescales both are constant, but this can obviously not be the case instantaneously, is the vertical central plunger say actuated by a screw or hydraulically or ...?</p>\n\n<p>EDIT: regarding question 2, I am trying to estimate what the \"output resistance\" is of the driver, if we make an analogy of pressure = voltage, then when there is movement = current the pressure is lower, until friction has taken over and pressure is restored... in the electrical case this is equivalent to stating that in practice every power source has an intrinsic or output resistance...</p>",
      "rawMarkdown": "Hi Bertrand,\n\n1) the acoustic sensor is attached to the central plunger, or to one of the 2 side plates that provide normal force?\n\n2) in the experiment corresponding to the kaggle data, is instantaneous pressure or instantaneous velocity constant? my gut feeling tells me that on longer timescales both are constant, but this can obviously not be the case instantaneously, is the vertical central plunger say actuated by a screw or hydraulically or ...?\n\nEDIT: regarding question 2, I am trying to estimate what the \"output resistance\" is of the driver, if we make an analogy of pressure = voltage, then when there is movement = current the pressure is lower, until friction has taken over and pressure is restored... in the electrical case this is equivalent to stating that in practice every power source has an intrinsic or output resistance..."
    },
    {
      "id": 469866,
      "postDate": "2019-02-11T23:43:27.387Z",
      "content": "<p>12ms gap does not seem be big to ttf, but it actually matters a lot for frequency (e.g., Fourier) analysis. </p>",
      "rawMarkdown": "12ms gap does not seem be big to ttf, but it actually matters a lot for frequency (e.g., Fourier) analysis. "
    },
    {
      "id": 468712,
      "postDate": "2019-02-09T14:12:56.807Z",
      "content": "<p>As I understand, time_to_failure is calculated somehow and the true remaining time for each bin is the value corresponding to it's last point. Is that correct?</p>\n\n<p>Is the acoustic emission data contiguous, is the last point of each bin (with  about 4096 points/rows in the file) just before the first of the next bin, are all the points recorded with the same rate?\nOr, there is 0.012 s interval with no acoustic emission data after each bin and the recorded bins are just concatenated at the end?</p>",
      "rawMarkdown": "As I understand, time_to_failure is calculated somehow and the true remaining time for each bin is the value corresponding to it's last point. Is that correct?\n\nIs the acoustic emission data contiguous, is the last point of each bin (with  about 4096 points/rows in the file) just before the first of the next bin, are all the points recorded with the same rate?\nOr, there is 0.012 s interval with no acoustic emission data after each bin and the recorded bins are just concatenated at the end?",
      "replies": [
        {
          "id": 468850,
          "postDate": "2019-02-09T20:52:15.657Z",
          "content": "<p>I think the latter is true.</p>",
          "rawMarkdown": "I think the latter is true."
        },
        {
          "id": 468898,
          "postDate": "2019-02-10T01:32:00.720Z",
          "content": "<p>Maybe it is. But than, the 12 ms doesn't match.  It is much less than 12 ms. Everything is recorded every 1e-9 s and we can see that decreasing in ttf within the bins. The first point of the new bin has  ~ 0.001 s difference in ttf with the last of the previous - a gap of 1 ms. It really looks like we does not have all the data, just concatenated bins. </p>",
          "rawMarkdown": "Maybe it is. But than, the 12 ms doesn't match.  It is much less than 12 ms. Everything is recorded every 1e-9 s and we can see that decreasing in ttf within the bins. The first point of the new bin has  ~ 0.001 s difference in ttf with the last of the previous - a gap of 1 ms. It really looks like we does not have all the data, just concatenated bins. "
        },
        {
          "id": 468952,
          "postDate": "2019-02-10T05:16:32.097Z",
          "content": "<p>I agree.</p>",
          "rawMarkdown": "I agree."
        }
      ]
    },
    {
      "id": 464429,
      "postDate": "2019-01-31T21:21:40.157Z",
      "content": "<p>Hi, many thanks for these details.\nI have a question, maybe obvious, but a confirmation would be great for the dummy I am :-) :\nthe reference point to time to failure is the end of the chunk, right ?</p>",
      "rawMarkdown": "Hi, many thanks for these details.\nI have a question, maybe obvious, but a confirmation would be great for the dummy I am :-) :\nthe reference point to time to failure is the end of the chunk, right ?",
      "replies": [
        {
          "id": 464799,
          "postDate": "2019-02-01T14:23:13.137Z",
          "content": "<p>Yes.</p>",
          "rawMarkdown": "Yes.",
          "votes": 1
        },
        {
          "id": 464888,
          "postDate": "2019-02-01T18:04:03.073Z",
          "content": "<p>Thanks 😊</p>",
          "rawMarkdown": "Thanks 😊"
        }
      ]
    },
    {
      "id": 494667,
      "postDate": "2019-03-20T05:16:20.093Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 466000,
      "postDate": "2019-02-04T13:46:24.177Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 460274,
      "postDate": "2019-01-23T10:39:04.143Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 526451,
      "postDate": "2019-05-03T04:48:40.713Z",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.\n"
    },
    {
      "id": 526102,
      "postDate": "2019-05-02T11:01:05.883Z",
      "content": "<p>Thank you a lot!</p>",
      "rawMarkdown": "Thank you a lot!"
    },
    {
      "id": 496097,
      "postDate": "2019-03-21T22:18:45.837Z",
      "content": "<p>Thank you for this context!</p>",
      "rawMarkdown": "Thank you for this context!"
    },
    {
      "id": 469123,
      "postDate": "2019-02-10T14:44:13.223Z",
      "content": "<p>Thank you, this has been very helpful.</p>",
      "rawMarkdown": "Thank you, this has been very helpful."
    },
    {
      "id": 458336,
      "postDate": "2019-01-19T12:04:19.843Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    }
  ],
  "comments": [
    {
      "id": 468488,
      "author_name": "Chris Roat",
      "author_url": "",
      "post_date": "2019-02-09T01:02:33.783000",
      "content": "<p>I think the time-to-failure data is not particularly clear, for the reasons at the bottom.</p>\n\n<p>If the data is even spaced, as denoted, a simpler approach to setting up the challenge would be to just provide a list of the acoustic data for training (no time-to-failure), and provide indices into that list as to where the labquakes appeared.  Then the challenge is to predict how long (how many time points) after a test set that a lab quake would occur.</p>\n\n<p>Please advise if there is useful structure in the ttf data, or if it can simply be assumed to be regularly spaced?</p>\n\n<p>The ttf data is confusing because:\n - each sample drops ~1ns over a ~4096 sample period, even though the samples should be 250ns apart (4MHz)\n - the time series takes large jumps every ~4096 period to \"catch up\" to the real sampling rate\n - the period between large jumps isn't actually always 4096 -- it's looks to be less sometimes\n - even interpolating over the whole dataset, it looks like the sampling is closer to 3.85MHz, not the 4MHz denoted</p>",
      "votes": 21,
      "replies": [
        {
          "id": 468877,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-02-09T23:04:32.013000",
          "content": "<p>Thanks. The best clear explanation. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 456854,
      "author_name": "SmallColdTension",
      "author_url": "",
      "post_date": "2019-01-16T16:59:29.997000",
      "content": "<p>Can you verify the following assumption by <a href=\"/chiqun\">@chiqun</a> for BOTH test and train data: acoustic sampling rate is 4MHz with a timeoffailure sample rate ~1000Hz? as we notice that in each chunk (4096 rows) of frames, the last two frames have a time difference ~ 0.001s, all the others have time difference ~1e-9 s. Can you also explain why it is like that physically.  Thanks. </p>",
      "votes": 11,
      "replies": [
        {
          "id": 457472,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2019-01-17T13:57:21.997000",
          "content": "<p>Hi Bertrand!\nMy additional request for clarification: At 4 MHz sampling rate, the time difference between two successive data points (i.e. lines in the CSV file) is 1/(4 MHz) = 250 ns, right? Therefore, time_to_failure should correctly decrease by 250 ns each step. What we actually see in the time_to_failure column in train.csv is an artifact of one of the physical devices used in the experiment.\nIn that case, basically all that is relevant is a 1D time series of acoustic data and a list of the indices of those 16 data points where an earthquake occurred. Are my assumptions correct?\nThanks!</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 466150,
          "author_name": "Bertrand RL",
          "author_url": "",
          "post_date": "2019-02-04T19:24:27.097000",
          "content": "<p>Hi! The periodic 0.001s of missing data is an artifact of the recording device, I will add a comment about it in the main post, thanks for bringing it up!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 508175,
          "author_name": "GeoX",
          "author_url": "",
          "post_date": "2019-04-05T18:22:56.090000",
          "content": "<p>Hi, I'm new here and i want to insert the time in the samples, it's more easy for me but i want to be sure. In the Andrews Script example his submission is like this:\n1   seg_id  time_to_failure</p>\n\n<p>2   seg_00030f  3.16345030153463</p>\n\n<p>3   seg_0012b5  5.1229034920016705</p>\n\n<p>4   seg_00184e  4.878316564579939</p>\n\n<p>5   seg_003339  7.867255038555255\nWhat are the 3.16345030153463, 5.1229034920016705 intervals? Nanoseconds or Microseconds? Or the number of lines until his expected event? I hope you understand my point and what i mean. Thank you.\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/88119#latest-508265\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/88119#latest-508265</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 508306,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-04-06T00:29:50.637000",
          "content": "<p>It is seconds, time_to_failure is measured in seconds in IO files.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 467334,
      "author_name": "macbend",
      "author_url": "",
      "post_date": "2019-02-06T22:21:29.133000",
      "content": "<p>Just to repeat this question and bring to the top, where are these recording gaps located in the test data?  It looks like each test segment is 150001 samples, which is not evenly divisible by 4096.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 477138,
          "author_name": "PéterSoltész",
          "author_url": "",
          "post_date": "2019-02-24T01:04:41.030000",
          "content": "<p>it is 150,000 exactly, the header of the csv might be the extra row... But it is also not divisible. Also there are sometimes only 4095 ticks between gaps...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 461314,
      "author_name": "Pandagg",
      "author_url": "",
      "post_date": "2019-01-25T18:44:02.080000",
      "content": "<p>It was mentioned that \"there are several earthquake cycles in the test set as well\". Do any of these earthquakes present within the duration of a test set segment?</p>\n\n<p>For example, said if a earthquake presents in the middle of a 150,000 steps segment, then only the signals from later half of the segment (75000 to 149000) will be useful to predict the next earthquake. </p>\n\n<p>It is important to know, because if this exists in the test set, then the model will need to be trained to handle this, i.e. identify if and when an earthquake happen within the segment, and then either ignore the signals before earthquake, or extract information from signals before earthquake differently than from signals after earthquake in order to predict next earthquake - a more much difficult task.</p>\n\n<p>Just want to ask to make sure.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 461364,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-01-25T21:04:55.263000",
          "content": "<p>There’re only a few samples like that. A few! So basically it does not affect anything. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 461872,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-27T08:23:15.090000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 524621,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-04-29T07:50:26.160000",
      "content": "<p>We have a month until the end.</p>\n\n<p>If I had one wish it would be more data - a lot of my effort is in trying not to overfit - could you consider having a supplementary data set added in the Data section so we can increase our training data?</p>\n\n<p>Obviously I am not sure if the goal of this competition is to obtain ideas on features or actual decent models.  The variance w.r.t. MAE is huge and this can be reduced with more data.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 480454,
      "author_name": "Margarita Chizh",
      "author_url": "",
      "post_date": "2019-02-28T08:00:44.053000",
      "content": "<p>Hi Bertrand!\nThank you for organizing this contest. Could you kindly answer a few questions about the data?</p>\n\n<ol>\n<li>Why was the sampling frequency chosen to be 4 MHz? Is there anywhere in your articles an estimate of the possible aliasing of data based on the speed and wavelength of the sound propagating in the sample's material? If there's aliasing, we may be missing some important signal characteristics.</li>\n<li>Has the analog signal been filtered at the output of the piezo sensor? The data looks noisy, like there is a low-pass filter missing.</li>\n<li>Am I right that since the 150000 samples are not multiples of 4096, then the 12 microsecond space can be placed in an arbitrary place of the chunk? And two parts of one bin can be in different chunks? This means that the data inside the chunk is not continuous, and this significantly limits the use of standard signal processing methods. Such pauses, where we lose 48 samples, are high-frequency noise in the signal. It would be more convenient and close to the real field scenario to break the data by bins so that there would be no pauses inside each record.</li>\n</ol>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 465224,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-02-02T16:44:51.193000",
      "content": "<p>Hi Bertrand</p>\n\n<p>Are the test segments contiguous or picked at random?</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 469301,
      "author_name": "PéterSoltész",
      "author_url": "",
      "post_date": "2019-02-10T23:01:39.497000",
      "content": "<p>Perchance I am wrong, but the additional info does not seem accurate. After carefully analysing the time gaps I believe the following is happening:</p>\n\n<p>There are two phases that is correct: recording, and let's say IO sequence, when there is no recording.</p>\n\n<p>When recording happens it happens at ~0.9GSamples/s for 4096 samples (sometimes a sample is missing, so 4095 samples are recorded in one go etc.)</p>\n\n<p>After the recording phase the IO phase takes not 12ms but 1.06ms on average.</p>\n\n<p>It is true that in some sense if you combine these you get near 3.85MSamples/s. </p>\n\n<p>To sum up recording takes approximately 4.5us after this there is 1060us silence. This repeats. This is important because frequency analysis is greatly affected by this, especially because we have no idea, where the recording gaps are in the test set.</p>\n\n<p>Please someone confirm...</p>\n\n<p>Note: I am wondering also if it is possible that the recording was in fact contignuous, but the times are somehow crowded up into these chunks? This can be possible if in fact the times were assigned not on measurement but at some processing level.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 469371,
          "author_name": "petyaAngelova",
          "author_url": "",
          "post_date": "2019-02-11T04:17:37.547000",
          "content": "<p>I was wandering the same thing. As far as I get from the papers, they calculate ttf, so I was hoping for really contiguous acoustic data, just twisted ttf. \nBut the team confirmed about the gap and now that scenario looks more right, just the 12ms looks like 1 ms. \nHope they give us more details soon!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 470419,
          "author_name": "PéterSoltész",
          "author_url": "",
          "post_date": "2019-02-12T23:02:43.340000",
          "content": "<p>Since I analysed the gaps and can confirm there is no continuity between the bins, in other words the gap is real. Mind the gap!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 470471,
          "author_name": "petyaAngelova",
          "author_url": "",
          "post_date": "2019-02-13T03:09:58.020000",
          "content": "<p>Thank you! I am still waiting some official clarification. \nMeanwhile I am thinking about the extreme ratio between the length of the bin and the length of the gap.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 471019,
          "author_name": "PéterSoltész",
          "author_url": "",
          "post_date": "2019-02-13T23:11:26.460000",
          "content": "<p>Well it is aproximately .004 (bin/gap)... If I understand you correctly. The gap is pretty constant... ...also the bin</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 471220,
          "author_name": "petyaAngelova",
          "author_url": "",
          "post_date": "2019-02-14T06:44:40.013000",
          "content": "<p>Just very long gaps compared to the bins, so much we don't see. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 483418,
          "author_name": "Umang",
          "author_url": "",
          "post_date": "2019-03-04T15:33:52.897000",
          "content": "<p>Your summary of the gaps looks right to me, but I have to assume that because the time to fault is <strong>not</strong> evenly spaced (both within a 4096 chunk and between chunks), that the data reflects actual measurements rather than being assigned arbitrarily after the fact. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 455965,
      "author_name": "Serhii Hrynko",
      "author_url": "",
      "post_date": "2019-01-14T22:01:12.263000",
      "content": "<p>Thanks.\nStill, is <code>time_to_failure</code> measured in seconds in training dataset?\nThis kind of might explain 0.001 or 0.0011 time difference between every 4095/4096 points. But time difference between 2 measurements is 1.1e-9. How can that be?\nMight it happen, that intermediate <code>time_to_failure</code> values in traning data are miscalculated?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 456364,
          "author_name": "Sergey Lebedev",
          "author_url": "",
          "post_date": "2019-01-15T17:13:18.390000",
          "content": "<p>May be it is a 4096 ADC that starting consequentally with 1 ns delay.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 456496,
          "author_name": "HorizonPicking2k18",
          "author_url": "",
          "post_date": "2019-01-15T23:29:36.307000",
          "content": "<p>Would this be a way of trying to reconstruct the whole signal from the seperate ones?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 458580,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-01-20T00:40:24.460000",
          "content": "<p>Ok, now I'm pretty sure that device used in the experiment sends data every T = 0.001064s in chunks of N = 4096 measurements.\n&gt; T = (total time of training data) / ((number of data packets) - 1) = 163.4294s / (153600 - 1)</p>\n\n<p>Single time stamp returned with each chunk is rounded to 4 digits during time_to_failure calculation, producing 0.001 and 0.0011 time differences between chunks, in exactly same sequence which we can see in training data.</p>\n\n<p>120 packets have only 4095 measurements and short packets spreaded evenly accross training data (first packet in every block of 1280 packets is short). Probably, it has something to do with data transmission protocol.</p>\n\n<p>Anyway, frame rate that can be obtained from time_to_failure values is approximately:\n&gt; F = N/T = 3.849622 MHz</p>\n\n<p>4MHz would correspond to 0.001024s intervals between data packets.\nThere can be number of tasks, where 0.04ms might be spent inside measuring device.</p>\n\n<p>So, I understand why @BertrandRL claims measurement frequency to be 4MHz.</p>\n\n<p>Training data looks pretty much continuous without gaps and claimed to be so. And time_to_failure looks like really well aligned with acoustic_data. At least, acoustic peaks stay around time_to_failure=0.315s for all earthquakes.</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 466152,
          "author_name": "Bertrand RL",
          "author_url": "",
          "post_date": "2019-02-04T19:30:36.150000",
          "content": "<p>Thanks for bringing that up, there is indeed a gap every 4096 samples, an artifact of the recording device. I added clarification to my initial post, thanks!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 469308,
          "author_name": "PéterSoltész",
          "author_url": "",
          "post_date": "2019-02-10T23:18:48.807000",
          "content": "<p>@BertrandRL, is it possible that the recording is contignuos, but at recording somehow the time data is \"compressed\" into this 4.5us bins?\nOr the other possibility is that there is recording for 4.5us and then there is no recording for 1060us.</p>\n\n<p>Which one do you think is happening in reality while recording? Thank you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 471053,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-02-14T00:29:47.513000",
          "content": "<p>I'd love to believe in 12ms gaps between bins as stated in clarification, but it doesn't fit in my brain in any possible way with 1.064ms TTF difference between bins in training data.</p>\n\n<p>I used to work a bit with wireless low energy measuring devices. And that's why I believe, that 4096 sample bin represents data collected over previous 1.024ms then followed by 0.04ms gap, when measuring device is busy with some other job, probably with network activity.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 455498,
      "author_name": "Mike Holcomb",
      "author_url": "",
      "post_date": "2019-01-14T03:59:43.367000",
      "content": "<p>Thanks the clarification. It is amazing that people are already getting models accurate to 1.5s using only 3/80 of a second!  It seems like trying to play \"Name That Tune\" with 1 note.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 457517,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-17T15:44:51.630000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 519879,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2019-04-19T19:25:02.350000",
      "content": "<p>Bertrand RL:\n&gt; \nThe data is recorded in bins of 4096 samples. Within those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.</p>\n\n<p>From the time_to_failure column in the data, I measure a delta t of 1.1 ns (like many others) which is inconsistent with a sample rate of 4 MHz (250 ns?) by a huge factor. For the \"gap\", I get two possibilities: 1 millisecond and 1.1 milliseconds (1000 and 1100 microseconds) , again inconsistent with the 12 *<em>micro</em>*second quote.</p>\n\n<p>I assume that time_to_failure is measured in seconds, consistent with:\n&gt; The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds.</p>\n\n<p>But the discrepancies exists no matter the time units.</p>\n\n<p>I haven't looked at this for a while so I might be making a dumb error. If not, why the huge discrepancies?</p>\n\n<p>edit: reading other comments, this seems to be a regular question that hasn't been properly addressed. I think a paragraph or two explaining why all the Kaggler's measurements are incorrect might be in order.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 520105,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-20T07:16:33.863000",
          "content": "<p>Thanks for raising this again.  I am new here, having entered 2 days ago, but I am a bit puzzled to not see any response since 3 months form organizers. </p>\n\n<p>Here is some additional analysis.  I apologize it someone else have already published similar views.  </p>\n\n<p>Using ttf, one can estimate the total training time to be 163.42061, using:</p>\n\n<pre><code>df = train[train.time_to_failure &gt; train.time_to_failure.shift().fillna(0)]\ndf.time_to_failure.sum() - train.time_to_failure.values[-1]\n</code></pre>\n\n<p>If the data was indeed sampled at 4 MHz without any interruption,  then, as you pointed out, the time delta between each measure should be 0.25 microsecond, a.  Total time would be 157.28637, as computed by:</p>\n\n<pre><code>train.shape[0] / 4e6\n</code></pre>\n\n<p>The delta in total time is due to the interruptions.  If there is one every 4096 then the interruption is 40 microseconds:</p>\n\n<pre><code>(163.42061 - 157.28637) / (train.shape[0] / 4096)\n</code></pre>\n\n<p>If, simple hypothesis, someone mistakenly thought there was an interruption every 150000 observations, then the interruption is 1.4 millisecond, close to what was announced in the first place...:</p>\n\n<pre><code>(163.42061 - train.shape[0] / 4e6) / (train.shape[0] / 150000)\n</code></pre>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 520200,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "2019-04-20T12:08:15.110000",
          "content": "<p>According to data from p4581, sampling frequency is 3.968 MHz.\nIf we take into account 12e-6 s gap (might be not exactly 12e-6):\n<code>\n4096 / 3.968e6 + 12e-6\n0.001044258\n</code>\nTime vector has delta of <code>0.001044</code> (presumably seconds)\nTime vector is separated from acoustic data, so might as well be that discrepancy in our case is the result of data combining.</p>\n\n<p>Bertrand RL (from \"Introduction\"):\n&gt; p4581 is an experiment that has the same setup in terms of machine, material and recording apparatus.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 520344,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-20T18:47:09.633000",
          "content": "<p>Issue is that you get a different ttf  that way.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 520492,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-04-21T05:30:37.193000",
          "content": "<p>You raides a good point. As I mentioned <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#458580\">in my another comment</a> there are 120 larger blocks of 1280 bins each in training dataset. First bin of each block has only 4095 samples. I would imagine there is a gap between these bigger blocks, rather then between bins. It gives:\n<code>\n(163.42061 - 157.28637) / 120 = 0.051118(6) ~ 0.0511\n</code>\nwhich is (512 - 1), which in turn points me to alternative way of thinking.</p>\n\n<p>Earlier <a href=\"/merepoule\">@merepoule</a> was saying about 12 milliseconds gap but later edited it based on comments. Lets imagine that initial statement was in fact accurate. Then we can calculate number of gaps:\n<code>\n(163.42061 - 157.28637) / 0.012 = 511.18(6) ~ 511\n</code>\nAnd number of blocks: <code>511 + 1 = 512</code>\nLooks pretty likely to me. Then blocks are of <code>153600 / 512 = 300</code> bins and <code>300 * 4096 = 1228800</code> samples each.\nThe problem with this theory though is that time_to_failure values assigned to training data are in fact contiguous (don't show any noticeable gaps).</p>\n\n<p>And if we are talking about microseconds, I don't think solutions accuracy will be affected in any way.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 521354,
          "author_name": "Danijel Kivaranovic",
          "author_url": "",
          "post_date": "2019-04-22T20:22:52.993000",
          "content": "<p>I guess that 4 MHz is simply a proxy. The data is actually recorded at 3.851 MHz. One bin contains 4096 samples measured in 4096*1.1 nanoseconds plus the artifact of 1 or 1.1 milliseconds. On average the artifact is 1.059 milliseconds. This means:\n<code>\n4096 / (4096 * 1.1*10**-9 + 1.059*10**-3) ~ 3851413 ~ 3.851 MHz\n</code></p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 521685,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-23T09:07:11.993000",
          "content": "<p><a href=\"/elvenmonk\">@elvenmonk</a> You nailed it IMHO.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522025,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-04-23T19:15:58.593000",
          "content": "<p><a href=\"/danijelk\">@danijelk</a> I don't believe that data can be and actually is measured every 1.1 nanoseconds. There are no any visible shifts in acoustic data on the edge of adjacent 4096 bins. You can find accurate average frequency measurement in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#458580\">my another comment</a>\nMeantime, I shared <a href=\"https://www.kaggle.com/elvenmonk/lanl-ttf-error-validation\">Kernel</a> that approximates <code>time_to_failure</code> values in training data set by linear time distribution with 0.6 milliseconds precision.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 522063,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "2019-04-23T20:16:43.210000",
          "content": "<p><a href=\"/elvenmonk\">@elvenmonk</a> Data in p4581 are chunks of size 20971520 bytes.\n<code>1280 * 4096 * 2 (channels) * 2 (bytes per sample) == 20971520</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522095,
          "author_name": "Danijel Kivaranovic",
          "author_url": "",
          "post_date": "2019-04-23T21:21:04.243000",
          "content": "<p><a href=\"/elvenmonk\">@elvenmonk</a> I guess you are right. Your numbers are more precise than mine</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522163,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-04-23T23:59:10.030000",
          "content": "<p><a href=\"/mykper\">@mykper</a> Awesome! Thanks for referencing P4581 experiment again. I'll take a deeper look into related materials, specifically to <a href=\"https://permalink.lanl.gov/object/tr?what=info%3Alanl-repo%2Flareport%2FLA-UR-17-29312\">this earlier LANL research</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 504511,
      "author_name": "Dan8262",
      "author_url": "",
      "post_date": "2019-03-31T18:53:19.310000",
      "content": "<p>Hello Bertrand,</p>\n\n<p>I have a question about the behaviour of the earthquake machine during an experiment.</p>\n\n<p>I suppose the physical characteristics of the three blocks remain unchanged during the test, but what about the two granular layers?</p>\n\n<p>Do they not suffer damage that could alter their physical characteristics during the experiment?</p>\n\n<p>Regards,</p>\n\n<p>Daniel</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 490591,
      "author_name": "Marcello Chiodi",
      "author_url": "",
      "post_date": "2019-03-14T17:27:43.083000",
      "content": "<p>Hi Bertrand, Thank you for this context.\nI have two questions about data.\n1) My first question is about test datasets. Each sequence of 150000 data precedes an earthquake: the 2624 test sets are precursors of 2624 different earthquakes?\n2) The variable time_to_failure in train dataset has an average frequency of 3.85 Mhz, given 4095 or 4096 intervals of 10^(-9) seconds followed by an interval of about 0.001 seconds. The sequence is decreasing, except for 16 jumps that correspond to 16 earthquakes.\nMy second question is: test data times are claimed to be at 4 Mhz. Are they evenly spaced, or do they follow the irregular pattern of time_to_failure in train dataset?</p>\n\n<p>Thank you for any possible help.</p>\n\n<p>Marcello</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 476110,
      "author_name": "DavidS",
      "author_url": "",
      "post_date": "2019-02-21T16:10:15.857000",
      "content": "<p>I case someone missed that too, 12 ms was changed to 12 MICROseconds</p>",
      "votes": 3,
      "replies": [
        {
          "id": 476311,
          "author_name": "petyaAngelova",
          "author_url": "",
          "post_date": "2019-02-22T00:29:17.177000",
          "content": "<p>I noticed, but that confuses me even more.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 477136,
          "author_name": "PéterSoltész",
          "author_url": "",
          "post_date": "2019-02-24T01:02:36.663000",
          "content": "<p>Which is funny since it is about 1 millisecond in reality...</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 475373,
      "author_name": "Simone Genta",
      "author_url": "",
      "post_date": "2019-02-20T17:22:18.527000",
      "content": "<p>Is it possible that failure dectection <em>fail</em> to detect some events?</p>\n\n<p>Or that audio signal spikes are related to a some kind of relaxation of the internal state, also if not leading to a failure?</p>\n\n<p>Please find in this <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/81302#latest-475311\">discussion</a> some consideration about that.</p>\n\n<p><img src=\"https://drive.google.com/uc?id=1UbCr2m1SBxNlnLm9Frd1zi2CRiVcaU-e\" alt=\"Graph of prediction over the entire training set\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 475576,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-02-20T23:22:40.363000",
          "content": "<p>You've got it) The biggest challenge of this challenge is to figure out whether next spike will be powerful enough to be considered failure and if not how to know time between next spike and real failure (given that relaxation after \"fake\" spike is not yet known). I guess if test data has 2 consequent non-failure spikes, predictions will fail badly on those segments.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 477780,
          "author_name": "PéterSoltész",
          "author_url": "",
          "post_date": "2019-02-25T09:00:14.943000",
          "content": "<p>I wouldn't say it is the biggest challenge though, because how would you use that info? In the test set you don't know how 150k bin's are connected, so you have to guess it on an individual basis.</p>\n\n<p>It is most likely that this info is not inherent in the 150k bins, well it was not assigned to them as per measurement, but through their timely order retrospectively.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 467226,
      "author_name": "max(Kolbe)",
      "author_url": "",
      "post_date": "2019-02-06T17:01:33.317000",
      "content": "<p>The 12ms-gap (see \"Additional Info\") between chunks is in contradictory to the time_to_failure (see train.csv) !</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 456814,
      "author_name": "Van de Rakt",
      "author_url": "",
      "post_date": "2019-01-16T15:30:10.673000",
      "content": "<p>Could you please verify to following guess/statement as mentioned bij <a href=\"/chiqun\">@chiqun</a> :My guess is the acoustic sampling rate is 4MHz, but the timeoffailure sample rate is 1000Hz. Need to be verified by the author.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 456001,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2019-01-14T23:35:02.403000",
      "content": "<p>I have some short questions for clarification.</p>\n\n<p>good to know that the input data is recorded in volts. Is that true for both training and test data, making them BOTH in precisely the same units, unnormalized? From the description, it seems that it is but I want to make sure.</p>\n\n<p>you say that training and testing are from the same experiment. I assume that there is NO overlap between the two data sets - is that true? Also, does the entire test dataset come from one earthquake cycle or many?</p>\n\n<p>finally, is the training set time series \"stationary\" meaning that each earthquake cycle is an equivalent time series modulo statistical fluctuations? Or does the physics change as the experiment progresses?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 456504,
          "author_name": "Bertrand RL",
          "author_url": "",
          "post_date": "2019-01-15T23:54:47.027000",
          "content": "<p>Hi Pete, the answer to all your questions is yes: the recording device is the same throughout the experiment, including for the training and the testing set. There is no overlap between the sets, that are contiguous. There are several earthquake cycles in the test set as well. The time series is stationary and the physics unchanged except for a small experimental artifact which is that some material from the fault is progressively lost as the experiment proceeds.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 456557,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-01-16T03:43:14.707000",
          "content": "<p>@BertrandRL: Thanks for your answer. May I ask you that do all the test chunks constitute a full and contiguous signal (by some order) ? Or they cannot be assembled into a contiguous signal?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 462065,
          "author_name": "Zidmie",
          "author_url": "",
          "post_date": "2019-01-27T15:33:41.317000",
          "content": "<p>Bertrand,\nI understand that the test data occurs just after the train data.\nOne information is also important IMO: how is the test data split into public and private test data? \n1. The 13% public data (used for the public Leaderboard), corresponds to the first events of the test data? \n2. Or is it randomly split over the 100% test data?\nThanks</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 466481,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-02-05T13:19:10.200000",
      "content": "<p>contiguous == sharing a common border; touching.\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12ms gap between each bin, an artifact of the recording device. </p>\n\n<p>LOL</p>",
      "votes": 3,
      "replies": [
        {
          "id": 466780,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-02-05T23:16:08.527000",
          "content": "<p>How did you get it's an artifact?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 467263,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-02-06T18:12:50.767000",
          "content": "<p>Read point 6 of his comment</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 467339,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-02-06T22:30:15.447000",
          "content": "<p>Ups, I missed it somehow, thank you. Now I see a lot of posts spotting that </p>\n\n<p>lol indeed</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 505489,
      "author_name": "Manoj Raman",
      "author_url": "",
      "post_date": "2019-04-02T04:20:15.533000",
      "content": "<p>Hello Mr Bertrand RL,</p>\n\n<p><strong>You said the data is recorded in bins of 4096 \"samples\"</strong>, does sample mean data points or the collection of 150000 data points? </p>\n\n<p><strong>And after each bin there is a 12 microseconds gap.</strong> Is the data still being recorded during these 12MS gap? If so, is this noise recorded at 4MHz?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 507299,
          "author_name": "Margarita Chizh",
          "author_url": "",
          "post_date": "2019-04-04T14:30:18.943000",
          "content": "<p>Hi! I would assume that: \n1) If 1 sample was 150_000 data points, 4096 samples would contain 614.4 millions of data points. And the whole data set is 629.145 millions, so it would be less than 2 bins, and Bertrand says \"The data is recorded in bins\" - which is plural. If we somehow estimate the approximate time duration of the whole training dataset, we could clarify this question completely.\n2) In a comment below Bertrand said that in the gap the data is missing, so, most likely, it was not recorded.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 491698,
      "author_name": "AkiraSosa",
      "author_url": "",
      "post_date": "2019-03-16T02:30:33.983000",
      "content": "<p>Hi Bertrand,</p>\n\n<p>If test set are contiguous, participants will track segments from future to past after finding obvious earthquake. Is this solution acceptable for you?</p>\n\n<p>Thanks, </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 464314,
      "author_name": " Igor Krasovskiy",
      "author_url": "",
      "post_date": "2019-01-31T15:30:46.407000",
      "content": "<p>What about problem with size of train and test files? Do you have some problems with data processing on Kaggle instance?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 463771,
      "author_name": "Jesus Leon",
      "author_url": "",
      "post_date": "2019-01-30T15:37:28.707000",
      "content": "<p>Hello,\nCould we get more detailed information on which piezoceramic sensor you are using to gather the seismic data? \nThanks in advance.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 522038,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2019-04-23T19:38:19.393000",
      "content": "<p>Here is a picture of what  get for a series of length twice the test data files. The plot is of delta_t between samples. Each red mark marks the end of 4096 samples. The black line at the top looks like zero but it is really -1.1 ns - way too small of course. The negative values, the \"gaps\", are -1ms and -1.1 ms. They come in threes with the first smaller in magnitude than the next two. Everything is separated by 4096 samples (you would have to be able to zoom to see this)\nIt's a pretty raw plot - little or no manipulation but comments on whether it is correct are welcome.</p>\n\n<p>I see the gaps as, roughly, making up for the fact that the majority of intervals have the (incorrect) 1.1 ns sampling. The error is mitigated at the gaps. If this were exact, then every gap would be ~1.018 ms (instead of the 1ms/1.1ms mixture), which is why I characterize it as \"rough\".</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 476288,
      "author_name": "PéterSoltész",
      "author_url": "",
      "post_date": "2019-02-21T22:58:37.267000",
      "content": "<p>I would be very interested to know whether Public portion of test data is contiguous slice in time rather than a random sample.</p>\n\n<p>My guess it is, and so it might be absolutely worthless to snoop on the leaderboard.</p>\n\n<p>Also am I right to assume that the whole test data in not in timely order?</p>\n\n<p>Thanks for the input...</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 466169,
      "author_name": "Amjad",
      "author_url": "",
      "post_date": "2019-02-04T19:52:44.750000",
      "content": "<p>Hi Bertrand, \nIs it possible to share more information on the splitting of the test set between public and private LB? Was the split at random with similar distribution, random without looking at the distribution, shared earthquakes, or different earthquakes? In other words, how really useful is the public LB?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 460559,
      "author_name": "atlas",
      "author_url": "",
      "post_date": "2019-01-24T00:33:23.937000",
      "content": "<p>A few have mentioned this point:</p>\n\n<p>In the training data, measurements seem to have been taken every 1 ns. However, every 4096 measurements, there is a gap of 1ms in data.</p>\n\n<p>I assume this is an experimental constraint, and as such the test data should contain the same gaps.</p>\n\n<p>However, can we be given a time indication of the measurements, so that we could know where these gaps appear in the test data?\nSomething like the time difference between the first and the n-th measurement maybe (for all test measurements)?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 460678,
          "author_name": "Daniel Gerigk",
          "author_url": "",
          "post_date": "2019-01-24T08:27:13.333000",
          "content": "<p>Yeah, because it's quite possible that precisely that millisecond of data holds the key to predicting a quake that will occur about 4 seconds later.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 467170,
          "author_name": "Ben Nye",
          "author_url": "",
          "post_date": "2019-02-06T15:26:12.387000",
          "content": "<p>This would be quite useful if we are seriously trying to find a way to predict earthquakes, since any method based on frequencies will need to know where the gaps are in the data.  Without it we are mostly stuck using magnitude of the signal which is only one part of the information that we should be able to use.  Of course, the ideal solution is to get better measuring equipment that can record data at the same time as measuring it, but I don't expect that to happen during the challenge.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 467299,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-02-06T20:20:40.867000",
          "content": "<p>Lomscargle or Gaussian Process could be options?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 469140,
          "author_name": "PéterSoltész",
          "author_url": "",
          "post_date": "2019-02-10T15:20:17.987000",
          "content": "<p>Daniel Gerigk: if you don't know where these gaps ar,e frequency analysis (atleast for some frequencies) becomes useless. So Yes, precisely that millisecond is very-very important...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 469253,
          "author_name": "Daniel Gerigk",
          "author_url": "",
          "post_date": "2019-02-10T19:36:17.887000",
          "content": "<p>It is possible that you are right.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 457098,
      "author_name": "Dimosthenis Karaflos",
      "author_url": "",
      "post_date": "2019-01-16T23:59:04.553000",
      "content": "<p>Hi Bertrand,</p>\n\n<p>the train dataset is contiguous in time, so are the test dataset segments. Concerning the test datasets segments, are they contiguous by the order they appear on the data section?\nFinally, is the whole test dataset contiguous in time with the train dataset?</p>\n\n<p>thanks in advance</p>",
      "votes": 2,
      "replies": [
        {
          "id": 476289,
          "author_name": "PéterSoltész",
          "author_url": "",
          "post_date": "2019-02-21T23:00:26.320000",
          "content": "<p>Most important question here, and nobody answering it...</p>\n\n<p>My guess it is not contignuos as they appear...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 456330,
      "author_name": "misp",
      "author_url": "",
      "post_date": "2019-01-15T15:29:15.260000",
      "content": "<p>How do you decide when there has been an earthquake? Specifically, is the classification based on reaching a threshold in the acoustic signal?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 456376,
          "author_name": "Antoine Pare",
          "author_url": "",
          "post_date": "2019-01-15T17:38:59.087000",
          "content": "<p>There has been an earthquake when time_to_failure reaches 0, there are 16 earthquake in the training set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 456502,
          "author_name": "Bertrand RL",
          "author_url": "",
          "post_date": "2019-01-15T23:50:41.857000",
          "content": "<p>Hi Misp, I added the answer in my original post, thanks for asking.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 456663,
          "author_name": "misp",
          "author_url": "",
          "post_date": "2019-01-16T09:02:11.520000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 457726,
          "author_name": "gumbot",
          "author_url": "",
          "post_date": "2019-01-18T01:34:51.343000",
          "content": "<p>\"There has been an earthquake when timetofailure reaches 0\"\nThe first example (5656573) has nothing looking like a quake starting at that point.  Far as I can see the quake happens shortly beforehand.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 463166,
          "author_name": "Antoine Pare",
          "author_url": "",
          "post_date": "2019-01-29T15:03:50.963000",
          "content": "<p>0 is not part of the data though, so you will never actually have an earth quake in the acoustic data as it would make the whole purpose of that challenge pointless. \nIf you can predict it when it happens, it's too late already.\nThe other events that are happening before time 0 are precursors.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 455994,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-14T22:58:04.460000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 456515,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-16T00:07:24.510000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 541180,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-02T01:23:05.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 535069,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-22T08:36:30.600000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 536588,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-24T18:26:18.420000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 527744,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-06T07:59:51.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 527492,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-05T16:46:38.290000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 527506,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-05T17:18:33.610000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 525853,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-01T20:29:49.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 525327,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-30T18:05:46.260000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 522727,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-24T22:44:56.480000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 522757,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-25T00:40:55.140000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 504433,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-31T16:03:33.120000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 472130,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-15T11:58:35.153000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 469866,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-11T23:43:27.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 468712,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-09T14:12:56.807000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 468850,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-09T20:52:15.657000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 468898,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-10T01:32:00.720000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 468952,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-10T05:16:32.097000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 464429,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-31T21:21:40.157000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 464799,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-01T14:23:13.137000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 464888,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-01T18:04:03.073000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 494667,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-20T05:16:20.093000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 466000,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-04T13:46:24.177000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 460274,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-23T10:39:04.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 526451,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-03T04:48:40.713000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 526102,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-02T11:01:05.883000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 496097,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-21T22:18:45.837000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 469123,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-10T14:44:13.223000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 458336,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-19T12:04:19.843000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "455431": "Hi everyone,\n\nI will try to answer questions in the forum as things go. To answer a few recurring questions:\n\n- The goal of the challenge is to capture the physical state of the laboratory fault and how close it is from failure from a snapshot of the seismic data it is emitting. You will have to build a model that predicts the time remaining before failure from a chunk of seismic data, like we have done in our first paper above on easier data.\n\n- The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds. \n\n- The seismic data is recorded using a piezoceramic sensor, which outputs a voltage upon deformation by incoming seismic waves. The seismic data of the input is this recorded voltage, in integers.\n\n- Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\n\n- Time to failure is based on a measure of fault strength (shear stress, not part of the data for the competition). When a labquake occurs this stress drops unambiguously.\n\n- The data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device. \n\n\nBertrand\n\n*Edited to answer additional questions*",
    "468488": "I think the time-to-failure data is not particularly clear, for the reasons at the bottom.\n\nIf the data is even spaced, as denoted, a simpler approach to setting up the challenge would be to just provide a list of the acoustic data for training (no time-to-failure), and provide indices into that list as to where the labquakes appeared.  Then the challenge is to predict how long (how many time points) after a test set that a lab quake would occur.\n\nPlease advise if there is useful structure in the ttf data, or if it can simply be assumed to be regularly spaced?\n\nThe ttf data is confusing because:\n - each sample drops ~1ns over a ~4096 sample period, even though the samples should be 250ns apart (4MHz)\n - the time series takes large jumps every ~4096 period to \"catch up\" to the real sampling rate\n - the period between large jumps isn't actually always 4096 -- it's looks to be less sometimes\n - even interpolating over the whole dataset, it looks like the sampling is closer to 3.85MHz, not the 4MHz denoted",
    "456854": "Can you verify the following assumption by @chiqun for BOTH test and train data: acoustic sampling rate is 4MHz with a timeoffailure sample rate ~1000Hz? as we notice that in each chunk (4096 rows) of frames, the last two frames have a time difference ~ 0.001s, all the others have time difference ~1e-9 s. Can you also explain why it is like that physically.  Thanks. ",
    "467334": "Just to repeat this question and bring to the top, where are these recording gaps located in the test data?  It looks like each test segment is 150001 samples, which is not evenly divisible by 4096.",
    "461314": "It was mentioned that \"there are several earthquake cycles in the test set as well\". Do any of these earthquakes present within the duration of a test set segment?\n\nFor example, said if a earthquake presents in the middle of a 150,000 steps segment, then only the signals from later half of the segment (75000 to 149000) will be useful to predict the next earthquake. \n\nIt is important to know, because if this exists in the test set, then the model will need to be trained to handle this, i.e. identify if and when an earthquake happen within the segment, and then either ignore the signals before earthquake, or extract information from signals before earthquake differently than from signals after earthquake in order to predict next earthquake - a more much difficult task.\n\nJust want to ask to make sure.",
    "524621": "We have a month until the end.\n\nIf I had one wish it would be more data - a lot of my effort is in trying not to overfit - could you consider having a supplementary data set added in the Data section so we can increase our training data?\n\nObviously I am not sure if the goal of this competition is to obtain ideas on features or actual decent models.  The variance w.r.t. MAE is huge and this can be reduced with more data.",
    "480454": "Hi Bertrand!\nThank you for organizing this contest. Could you kindly answer a few questions about the data?\n\n1. Why was the sampling frequency chosen to be 4 MHz? Is there anywhere in your articles an estimate of the possible aliasing of data based on the speed and wavelength of the sound propagating in the sample's material? If there's aliasing, we may be missing some important signal characteristics.\n2. Has the analog signal been filtered at the output of the piezo sensor? The data looks noisy, like there is a low-pass filter missing.\n3. Am I right that since the 150000 samples are not multiples of 4096, then the 12 microsecond space can be placed in an arbitrary place of the chunk? And two parts of one bin can be in different chunks? This means that the data inside the chunk is not continuous, and this significantly limits the use of standard signal processing methods. Such pauses, where we lose 48 samples, are high-frequency noise in the signal. It would be more convenient and close to the real field scenario to break the data by bins so that there would be no pauses inside each record.",
    "465224": "Hi Bertrand\n\nAre the test segments contiguous or picked at random?",
    "469301": "Perchance I am wrong, but the additional info does not seem accurate. After carefully analysing the time gaps I believe the following is happening:\n\nThere are two phases that is correct: recording, and let's say IO sequence, when there is no recording.\n\nWhen recording happens it happens at ~0.9GSamples/s for 4096 samples (sometimes a sample is missing, so 4095 samples are recorded in one go etc.)\n\nAfter the recording phase the IO phase takes not 12ms but 1.06ms on average.\n\nIt is true that in some sense if you combine these you get near 3.85MSamples/s. \n\nTo sum up recording takes approximately 4.5us after this there is 1060us silence. This repeats. This is important because frequency analysis is greatly affected by this, especially because we have no idea, where the recording gaps are in the test set.\n\nPlease someone confirm...\n\nNote: I am wondering also if it is possible that the recording was in fact contignuous, but the times are somehow crowded up into these chunks? This can be possible if in fact the times were assigned not on measurement but at some processing level.\n",
    "455965": "Thanks.\nStill, is `time_to_failure` measured in seconds in training dataset?\nThis kind of might explain 0.001 or 0.0011 time difference between every 4095/4096 points. But time difference between 2 measurements is 1.1e-9. How can that be?\nMight it happen, that intermediate `time_to_failure` values in traning data are miscalculated?",
    "455498": "Thanks the clarification. It is amazing that people are already getting models accurate to 1.5s using only 3/80 of a second!  It seems like trying to play \"Name That Tune\" with 1 note.",
    "519879": "Bertrand RL:\n&gt; \nThe data is recorded in bins of 4096 samples. Within those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\n\nFrom the time_to_failure column in the data, I measure a delta t of 1.1 ns (like many others) which is inconsistent with a sample rate of 4 MHz (250 ns?) by a huge factor. For the \"gap\", I get two possibilities: 1 millisecond and 1.1 milliseconds (1000 and 1100 microseconds) , again inconsistent with the 12 **micro**second quote.\n\nI assume that time_to_failure is measured in seconds, consistent with:\n&gt; The input is a chunk of 0.0375 seconds of seismic data (ordered in time), which is recorded at 4MHz, hence 150'000 data points, and the output is time remaining until the following lab earthquake, in seconds.\n\nBut the discrepancies exists no matter the time units.\n\nI haven't looked at this for a while so I might be making a dumb error. If not, why the huge discrepancies?\n\nedit: reading other comments, this seems to be a regular question that hasn't been properly addressed. I think a paragraph or two explaining why all the Kaggler's measurements are incorrect might be in order.",
    "504511": "Hello Bertrand,\n\nI have a question about the behaviour of the earthquake machine during an experiment.\n\nI suppose the physical characteristics of the three blocks remain unchanged during the test, but what about the two granular layers?\n\nDo they not suffer damage that could alter their physical characteristics during the experiment?\n\nRegards,\n\nDaniel",
    "490591": "Hi Bertrand, Thank you for this context.\nI have two questions about data.\n1) My first question is about test datasets. Each sequence of 150000 data precedes an earthquake: the 2624 test sets are precursors of 2624 different earthquakes?\n2) The variable time_to_failure in train dataset has an average frequency of 3.85 Mhz, given 4095 or 4096 intervals of 10^(-9) seconds followed by an interval of about 0.001 seconds. The sequence is decreasing, except for 16 jumps that correspond to 16 earthquakes.\nMy second question is: test data times are claimed to be at 4 Mhz. Are they evenly spaced, or do they follow the irregular pattern of time_to_failure in train dataset?\n\nThank you for any possible help.\n\nMarcello",
    "476110": "I case someone missed that too, 12 ms was changed to 12 MICROseconds",
    "475373": "Is it possible that failure dectection *fail* to detect some events?\n\nOr that audio signal spikes are related to a some kind of relaxation of the internal state, also if not leading to a failure?\n\nPlease find in this [discussion][1] some consideration about that.\n\n![Graph of prediction over the entire training set][2]\n\n[1]:https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/81302#latest-475311\n[2]: https://drive.google.com/uc?id=1UbCr2m1SBxNlnLm9Frd1zi2CRiVcaU-e",
    "467226": "The 12ms-gap (see \"Additional Info\") between chunks is in contradictory to the time_to_failure (see train.csv) !",
    "456814": "Could you please verify to following guess/statement as mentioned bij @chiqun :My guess is the acoustic sampling rate is 4MHz, but the timeoffailure sample rate is 1000Hz. Need to be verified by the author.",
    "456001": "I have some short questions for clarification.\n\n good to know that the input data is recorded in volts. Is that true for both training and test data, making them BOTH in precisely the same units, unnormalized? From the description, it seems that it is but I want to make sure.\n\nyou say that training and testing are from the same experiment. I assume that there is NO overlap between the two data sets - is that true? Also, does the entire test dataset come from one earthquake cycle or many?\n\nfinally, is the training set time series \"stationary\" meaning that each earthquake cycle is an equivalent time series modulo statistical fluctuations? Or does the physics change as the experiment progresses?\n",
    "466481": "contiguous == sharing a common border; touching.\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12ms gap between each bin, an artifact of the recording device. \n\nLOL",
    "505489": "Hello Mr Bertrand RL,\n\n**You said the data is recorded in bins of 4096 \"samples\"**, does sample mean data points or the collection of 150000 data points? \n\n**And after each bin there is a 12 microseconds gap.** Is the data still being recorded during these 12MS gap? If so, is this noise recorded at 4MHz?",
    "491698": "Hi Bertrand,\n\nIf test set are contiguous, participants will track segments from future to past after finding obvious earthquake. Is this solution acceptable for you?\n\nThanks, ",
    "464314": "What about problem with size of train and test files? Do you have some problems with data processing on Kaggle instance?",
    "463771": "Hello,\nCould we get more detailed information on which piezoceramic sensor you are using to gather the seismic data? \nThanks in advance.",
    "522038": "Here is a picture of what  get for a series of length twice the test data files. The plot is of delta_t between samples. Each red mark marks the end of 4096 samples. The black line at the top looks like zero but it is really -1.1 ns - way too small of course. The negative values, the \"gaps\", are -1ms and -1.1 ms. They come in threes with the first smaller in magnitude than the next two. Everything is separated by 4096 samples (you would have to be able to zoom to see this)\nIt's a pretty raw plot - little or no manipulation but comments on whether it is correct are welcome.\n\nI see the gaps as, roughly, making up for the fact that the majority of intervals have the (incorrect) 1.1 ns sampling. The error is mitigated at the gaps. If this were exact, then every gap would be ~1.018 ms (instead of the 1ms/1.1ms mixture), which is why I characterize it as \"rough\".",
    "476288": "I would be very interested to know whether Public portion of test data is contiguous slice in time rather than a random sample.\n\nMy guess it is, and so it might be absolutely worthless to snoop on the leaderboard.\n\nAlso am I right to assume that the whole test data in not in timely order?\n\nThanks for the input...",
    "466169": "Hi Bertrand, \nIs it possible to share more information on the splitting of the test set between public and private LB? Was the split at random with similar distribution, random without looking at the distribution, shared earthquakes, or different earthquakes? In other words, how really useful is the public LB?",
    "460559": "A few have mentioned this point:\n\nIn the training data, measurements seem to have been taken every 1 ns. However, every 4096 measurements, there is a gap of 1ms in data.\n\nI assume this is an experimental constraint, and as such the test data should contain the same gaps.\n\nHowever, can we be given a time indication of the measurements, so that we could know where these gaps appear in the test data?\nSomething like the time difference between the first and the n-th measurement maybe (for all test measurements)?",
    "457098": "Hi Bertrand,\n\nthe train dataset is contiguous in time, so are the test dataset segments. Concerning the test datasets segments, are they contiguous by the order they appear on the data section?\nFinally, is the whole test dataset contiguous in time with the train dataset?\n\nthanks in advance",
    "456330": "How do you decide when there has been an earthquake? Specifically, is the classification based on reaching a threshold in the acoustic signal?",
    "455994": "If the test folder contains the files that are chunks of the single big train file, than why each seg file does not include the time_to_failure data as a second column? Just to keep us busy  finding each chunk in the train file instead of working on a model?  ",
    "541180": "I found that there is a correlation coefficient -0.6 between (time margin of the input peaks )/abs(input peak's value) and the answer's peak value.",
    "535069": "Can you tell me the relationship between voltage and seismic data? I also find that the data in train is not enough 4096 samples, it's true?",
    "527744": "Hi Bertrand. Could you say what type of instrumentation do you use? in order to perform the instrumental correction. If you have the poles, zeros and the constant, would be great, since I read the article you posted and you use obspy, thanks.",
    "527492": "In the kernel, whenever I try to load the training data in dataframe, the kernel crashes. What should i do ?",
    "525853": "Here for the level upgrade.",
    "525327": "This really helped in understanding the problem as the overview is not very lucid. Cheers !!",
    "522727": "Several have wondered why we are recording at ~4 MHz when google searches are not finding accelerometers that can measure such a high acoustic frequency. In fact, in the \"supporting information\" referenced, we see:\n&gt; The sampling rate of the acoustic data is 330kHz. We are band-limited by the accelerometer (the frequency response is poor above about 50kHz).\n\nThis suggests that state-of-the-art fails above 50 kHz. The spectrum we see in the competition data  goes much higher (see the attached image for a spectral measurement).\n\nThe question I have is: was there a significant upgrade in the acoustic sensor for the data we are working with? Should we trust all of that high frequency \"signal\". Are we looking at a set of spurious detector resonances?\n\n It is important information! \n\nI hope the organizers can help with this.",
    "504433": "Does it help to apply a low pass filter to the data? ",
    "472130": "Hi Bertrand,\n\n1) the acoustic sensor is attached to the central plunger, or to one of the 2 side plates that provide normal force?\n\n2) in the experiment corresponding to the kaggle data, is instantaneous pressure or instantaneous velocity constant? my gut feeling tells me that on longer timescales both are constant, but this can obviously not be the case instantaneously, is the vertical central plunger say actuated by a screw or hydraulically or ...?\n\nEDIT: regarding question 2, I am trying to estimate what the \"output resistance\" is of the driver, if we make an analogy of pressure = voltage, then when there is movement = current the pressure is lower, until friction has taken over and pressure is restored... in the electrical case this is equivalent to stating that in practice every power source has an intrinsic or output resistance...",
    "469866": "12ms gap does not seem be big to ttf, but it actually matters a lot for frequency (e.g., Fourier) analysis. ",
    "468712": "As I understand, time_to_failure is calculated somehow and the true remaining time for each bin is the value corresponding to it's last point. Is that correct?\n\nIs the acoustic emission data contiguous, is the last point of each bin (with  about 4096 points/rows in the file) just before the first of the next bin, are all the points recorded with the same rate?\nOr, there is 0.012 s interval with no acoustic emission data after each bin and the recorded bins are just concatenated at the end?",
    "464429": "Hi, many thanks for these details.\nI have a question, maybe obvious, but a confirmation would be great for the dummy I am :-) :\nthe reference point to time to failure is the end of the chunk, right ?",
    "494667": "",
    "466000": "",
    "460274": "",
    "526451": "Thank you.\n",
    "526102": "Thank you a lot!",
    "496097": "Thank you for this context!",
    "469123": "Thank you, this has been very helpful.",
    "458336": "Thank you!"
  }
}