{
  "id": 90838,
  "title": "LANL p4581 data is on Kaggle now",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90838",
  "author_name": "",
  "post_date": "2019-04-28T02:01:31.389825900Z",
  "votes": 34,
  "comment_count": 18,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/redstr/lanl-p4581\">https://www.kaggle.com/redstr/lanl-p4581</a></p>\n\n<p>Update: actually, better use Leigh's version: <a href=\"https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581\">https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581</a></p>\n\n<p>Update2: I have split the data into channels properly.</p>\n\n<p>I downloaded all data from <a href=\"https://sites.psu.edu/chasbolton/\">https://sites.psu.edu/chasbolton/</a> and put it on Kaggle as a dataset. It should be possible to use within kernels.</p>\n\n<p>The data is large: the size is over 20x the size of the training set in the competition data. As to how useful it is, you can decide for yourself. It's certainly not a plug-and-play training set extension. The quake information is not present, and it's unclear how to define quakes there. The basic statistics are also different:\n<img src=\"https://i.imgur.com/3sVuoXc.png\" alt=\"\"></p>\n\n<p>So models based on basic statistics like variance, are going to have a hard time using this data for training as it is. Maybe someone can find a way to normalize it.</p>",
  "messages": [
    {
      "id": "524125",
      "postDate": "04/28/2019 02:01:31",
      "content": "<p><a href=\"https://www.kaggle.com/redstr/lanl-p4581\">https://www.kaggle.com/redstr/lanl-p4581</a></p>\n\n<p>Update: actually, better use Leigh's version: <a href=\"https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581\">https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581</a></p>\n\n<p>Update2: I have split the data into channels properly.</p>\n\n<p>I downloaded all data from <a href=\"https://sites.psu.edu/chasbolton/\">https://sites.psu.edu/chasbolton/</a> and put it on Kaggle as a dataset. It should be possible to use within kernels.</p>\n\n<p>The data is large: the size is over 20x the size of the training set in the competition data. As to how useful it is, you can decide for yourself. It's certainly not a plug-and-play training set extension. The quake information is not present, and it's unclear how to define quakes there. The basic statistics are also different:\n<img src=\"https://i.imgur.com/3sVuoXc.png\" alt=\"\"></p>\n\n<p>So models based on basic statistics like variance, are going to have a hard time using this data for training as it is. Maybe someone can find a way to normalize it.</p>",
      "rawMarkdown": "https://www.kaggle.com/redstr/lanl-p4581\n\nUpdate: actually, better use Leigh's version: https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581\n\nUpdate2: I have split the data into channels properly.\n\nI downloaded all data from https://sites.psu.edu/chasbolton/ and put it on Kaggle as a dataset. It should be possible to use within kernels.\n\nThe data is large: the size is over 20x the size of the training set in the competition data. As to how useful it is, you can decide for yourself. It's certainly not a plug-and-play training set extension. The quake information is not present, and it's unclear how to define quakes there. The basic statistics are also different:\n![](https://i.imgur.com/3sVuoXc.png)\n\nSo models based on basic statistics like variance, are going to have a hard time using this data for training as it is. Maybe someone can find a way to normalize it.",
      "votes": null
    },
    {
      "id": "524132",
      "postDate": "04/28/2019 02:53:04",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "Thanks.",
      "votes": null
    },
    {
      "id": "524251",
      "postDate": "04/28/2019 10:03:23",
      "content": "<p>Thanks for sharing!</p>\n\n<p>If this data is the same as we are using, I would say that we are having a case of data leakage. Besides this, it seems that the training data provided by the problem owner is not really representative or even from a similar of to the whole experimental data.  These characteristic makes really hard to make a decent prediction.</p>",
      "rawMarkdown": "Thanks for sharing!\n\nIf this data is the same as we are using, I would say that we are having a case of data leakage. Besides this, it seems that the training data provided by the problem owner is not really representative or even from a similar of to the whole experimental data.  These characteristic makes really hard to make a decent prediction.",
      "votes": null
    },
    {
      "id": "524271",
      "postDate": "04/28/2019 11:13:57",
      "content": "<p>It's not the same</p>",
      "rawMarkdown": "It's not the same",
      "votes": null
    },
    {
      "id": "524296",
      "postDate": "04/28/2019 12:58:20",
      "content": "<p>I compresed single channel to 11GB (data have 2 channel: 33 and 34)</p>",
      "rawMarkdown": "I compresed single channel to 11GB (data have 2 channel: 33 and 34)",
      "votes": null
    },
    {
      "id": "524336",
      "postDate": "04/28/2019 14:40:17",
      "content": "<p>Can you clarify? What is the channel, and how to distinguish them? Maybe this is the source of difference in the statistics?</p>",
      "rawMarkdown": "Can you clarify? What is the channel, and how to distinguish them? Maybe this is the source of difference in the statistics?",
      "votes": null
    },
    {
      "id": "524339",
      "postDate": "04/28/2019 14:43:29",
      "content": "<p>Yes, the data comes from another experiment. The experiment setup is very similar, only here they varied the shear stress in time. In the picture above you can see the staircase pattern: these are the times of different shear stresses. The stress range at around 500 mark on the plot is similar to the reported shear stress during the experiment which we are trying to predict. However, the \"acoustic power\" is still quite different from what is observed in the train set.</p>",
      "rawMarkdown": "Yes, the data comes from another experiment. The experiment setup is very similar, only here they varied the shear stress in time. In the picture above you can see the staircase pattern: these are the times of different shear stresses. The stress range at around 500 mark on the plot is similar to the reported shear stress during the experiment which we are trying to predict. However, the \"acoustic power\" is still quite different from what is observed in the train set.",
      "votes": null
    },
    {
      "id": "524350",
      "postDate": "04/28/2019 15:02:02",
      "content": "<p>What is the channel, and how to distinguish them?\n- channels2save: 33, 34. Read data description. Channel have same info</p>\n\n<p>Maybe this is the source of difference in the statistics?\n- No. Difference in various normal stress. Look at <a href=\"https://folk.uio.no/karenmai/publications/mair_JGR_2002.pdf\">this paper</a> (image on page 2)</p>",
      "rawMarkdown": "What is the channel, and how to distinguish them?\n- channels2save: 33, 34. Read data description. Channel have same info\n\nMaybe this is the source of difference in the statistics?\n- No. Difference in various normal stress. Look at [this paper](https://folk.uio.no/karenmai/publications/mair_JGR_2002.pdf) (image on page 2)",
      "votes": null
    },
    {
      "id": "524357",
      "postDate": "04/28/2019 15:31:30",
      "content": "<p>You are right, thanks for the info! I didn't even look in that file. So how are the channels arranged? Are they interleaved, or separated into first and the second half of the file?</p>",
      "rawMarkdown": "You are right, thanks for the info! I didn't even look in that file. So how are the channels arranged? Are they interleaved, or separated into first and the second half of the file?",
      "votes": null
    },
    {
      "id": "524365",
      "postDate": "04/28/2019 15:48:50",
      "content": "<p>Algorithm in file plotacousticdata.m</p>\n\n<p><a href=\"https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581\">https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581</a></p>\n\n<p>Splited by events</p>",
      "rawMarkdown": "Algorithm in file plotacousticdata.m\n\nhttps://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581\n\nSplited by events",
      "votes": null
    },
    {
      "id": "524394",
      "postDate": "04/28/2019 17:27:06",
      "content": "<p>Thanks, your data is much better organized than mine!\nAlso, the mystery of size 4095 bursts is resolved: it's probably an off-by-one mistake when joining the files. They likely missed the first or the last tick of each file. Note that each original .ac file has exactly 1280 size-4096 bursts.\nIn your opinion, is their event markup trustworthy? How similar are their events to our training quakes?</p>",
      "rawMarkdown": "Thanks, your data is much better organized than mine!\nAlso, the mystery of size 4095 bursts is resolved: it's probably an off-by-one mistake when joining the files. They likely missed the first or the last tick of each file. Note that each original .ac file has exactly 1280 size-4096 bursts.\nIn your opinion, is their event markup trustworthy? How similar are their events to our training quakes?",
      "votes": null
    },
    {
      "id": "524398",
      "postDate": "04/28/2019 17:31:45",
      "content": "<p>Bertrand said:\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.</p>\n\n<p>They splitted data by shear info. It's more accuracy. I splitted by picks of variation (not so accuracy).</p>",
      "rawMarkdown": "Bertrand said:\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\n\nThey splitted data by shear info. It's more accuracy. I splitted by picks of variation (not so accuracy).",
      "votes": null
    },
    {
      "id": "524402",
      "postDate": "04/28/2019 17:40:04",
      "content": "<p>Ah, so you split the data yourself? I thought you split by their \"manual event picks\" which are together with the data. I just looked at them and it's some weird stuff.\nCould you specify your criterion for quake detection?\nI know what Bertrand said. But he didn't say why every 1280 bins there is a bin of size 4095 in the train set. Now we know almost for sure. Because the very first bin is of size 4095, they probably missed the first tick in each file.</p>",
      "rawMarkdown": "Ah, so you split the data yourself? I thought you split by their \"manual event picks\" which are together with the data. I just looked at them and it's some weird stuff.\nCould you specify your criterion for quake detection?\nI know what Bertrand said. But he didn't say why every 1280 bins there is a bin of size 4095 in the train set. Now we know almost for sure. Because the very first bin is of size 4095, they probably missed the first tick in each file.",
      "votes": null
    },
    {
      "id": "524409",
      "postDate": "04/28/2019 17:46:26",
      "content": "<p>Automated splitted and fix some mistake. Splitted by friction (function of variation) as recommend Bertrand. Split where sudden dropped friction. Its easy where normal stress small, but if its large - detect splitted harder</p>",
      "rawMarkdown": "Automated splitted and fix some mistake. Splitted by friction (function of variation) as recommend Bertrand. Split where sudden dropped friction. Its easy where normal stress small, but if its large - detect splitted harder",
      "votes": null
    },
    {
      "id": "524419",
      "postDate": "04/28/2019 18:01:47",
      "content": "<p>Thanks. One more thing, which I think is important: the time between ticks within one 4096-tick bin in this data is specified as 2.5202e-07, while the time between bin starts is 0.001044. This means that there is a very little gap between bins.\nIn the training set, according to \"ttf\", the time between ticks within a bin is 1.1029617e-09 (~200 times smaller), while the time between bin starts is 0.001, i.e. almost the same. In your opinion, is TTF field in our train set wrong, or is the data indeed recorded at over 200x frequency compared to p4581?</p>",
      "rawMarkdown": "Thanks. One more thing, which I think is important: the time between ticks within one 4096-tick bin in this data is specified as 2.5202e-07, while the time between bin starts is 0.001044. This means that there is a very little gap between bins.\nIn the training set, according to \"ttf\", the time between ticks within a bin is 1.1029617e-09 (~200 times smaller), while the time between bin starts is 0.001, i.e. almost the same. In your opinion, is TTF field in our train set wrong, or is the data indeed recorded at over 200x frequency compared to p4581?",
      "votes": null
    },
    {
      "id": "524422",
      "postDate": "04/28/2019 18:05:26",
      "content": "<p>TTF fields are equal in our set and p4581. Answer here: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526</a></p>",
      "rawMarkdown": "TTF fields are equal in our set and p4581. Answer here: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526",
      "votes": null
    },
    {
      "id": "524428",
      "postDate": "04/28/2019 18:12:06",
      "content": "<p>Yep, thanks again for the clarification. Indeed it matches completely, the frequency and the 12us gap. So the TTF field in the train data must be wrong within bins. Seeing this other dataset is really convincing.</p>",
      "rawMarkdown": "Yep, thanks again for the clarification. Indeed it matches completely, the frequency and the 12us gap. So the TTF field in the train data must be wrong within bins. Seeing this other dataset is really convincing.",
      "votes": null
    },
    {
      "id": "535875",
      "postDate": "05/23/2019 15:05:41",
      "content": "<p>Hi <a href=\"/leighplt\">@leighplt</a>. Thanks for uploading the data from exp p4581. When you pulled the data from matlab, was there shearing force? If there was, can you please upload that as well? I would like to try to use shearing force. Thanks!</p>",
      "rawMarkdown": "Hi @leighplt. Thanks for uploading the data from exp p4581. When you pulled the data from matlab, was there shearing force? If there was, can you please upload that as well? I would like to try to use shearing force. Thanks!",
      "votes": null
    },
    {
      "id": "536337",
      "postDate": "05/24/2019 09:30:16",
      "content": "<p>Unfortunately shearing force not in data</p>",
      "rawMarkdown": "Unfortunately shearing force not in data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 524132,
      "author_name": "timmmmmms",
      "author_url": "",
      "post_date": "04/28/2019 02:53:04",
      "content": "<p>Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 524251,
      "author_name": "santiagoruiz",
      "author_url": "",
      "post_date": "04/28/2019 10:03:23",
      "content": "<p>Thanks for sharing!</p>\n\n<p>If this data is the same as we are using, I would say that we are having a case of data leakage. Besides this, it seems that the training data provided by the problem owner is not really representative or even from a similar of to the whole experimental data.  These characteristic makes really hard to make a decent prediction.</p>",
      "votes": null,
      "replies": [
        {
          "id": 524271,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/28/2019 11:13:57",
          "content": "<p>It's not the same</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524339,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/28/2019 14:43:29",
          "content": "<p>Yes, the data comes from another experiment. The experiment setup is very similar, only here they varied the shear stress in time. In the picture above you can see the staircase pattern: these are the times of different shear stresses. The stress range at around 500 mark on the plot is similar to the reported shear stress during the experiment which we are trying to predict. However, the \"acoustic power\" is still quite different from what is observed in the train set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 524296,
      "author_name": "leighplt",
      "author_url": "",
      "post_date": "04/28/2019 12:58:20",
      "content": "<p>I compresed single channel to 11GB (data have 2 channel: 33 and 34)</p>",
      "votes": null,
      "replies": [
        {
          "id": 524336,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/28/2019 14:40:17",
          "content": "<p>Can you clarify? What is the channel, and how to distinguish them? Maybe this is the source of difference in the statistics?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524350,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "04/28/2019 15:02:02",
          "content": "<p>What is the channel, and how to distinguish them?\n- channels2save: 33, 34. Read data description. Channel have same info</p>\n\n<p>Maybe this is the source of difference in the statistics?\n- No. Difference in various normal stress. Look at <a href=\"https://folk.uio.no/karenmai/publications/mair_JGR_2002.pdf\">this paper</a> (image on page 2)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524357,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/28/2019 15:31:30",
          "content": "<p>You are right, thanks for the info! I didn't even look in that file. So how are the channels arranged? Are they interleaved, or separated into first and the second half of the file?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524365,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "04/28/2019 15:48:50",
          "content": "<p>Algorithm in file plotacousticdata.m</p>\n\n<p><a href=\"https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581\">https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581</a></p>\n\n<p>Splited by events</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524394,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/28/2019 17:27:06",
          "content": "<p>Thanks, your data is much better organized than mine!\nAlso, the mystery of size 4095 bursts is resolved: it's probably an off-by-one mistake when joining the files. They likely missed the first or the last tick of each file. Note that each original .ac file has exactly 1280 size-4096 bursts.\nIn your opinion, is their event markup trustworthy? How similar are their events to our training quakes?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524398,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "04/28/2019 17:31:45",
          "content": "<p>Bertrand said:\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.</p>\n\n<p>They splitted data by shear info. It's more accuracy. I splitted by picks of variation (not so accuracy).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524402,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/28/2019 17:40:04",
          "content": "<p>Ah, so you split the data yourself? I thought you split by their \"manual event picks\" which are together with the data. I just looked at them and it's some weird stuff.\nCould you specify your criterion for quake detection?\nI know what Bertrand said. But he didn't say why every 1280 bins there is a bin of size 4095 in the train set. Now we know almost for sure. Because the very first bin is of size 4095, they probably missed the first tick in each file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524409,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "04/28/2019 17:46:26",
          "content": "<p>Automated splitted and fix some mistake. Splitted by friction (function of variation) as recommend Bertrand. Split where sudden dropped friction. Its easy where normal stress small, but if its large - detect splitted harder</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524419,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/28/2019 18:01:47",
          "content": "<p>Thanks. One more thing, which I think is important: the time between ticks within one 4096-tick bin in this data is specified as 2.5202e-07, while the time between bin starts is 0.001044. This means that there is a very little gap between bins.\nIn the training set, according to \"ttf\", the time between ticks within a bin is 1.1029617e-09 (~200 times smaller), while the time between bin starts is 0.001, i.e. almost the same. In your opinion, is TTF field in our train set wrong, or is the data indeed recorded at over 200x frequency compared to p4581?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524422,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "04/28/2019 18:05:26",
          "content": "<p>TTF fields are equal in our set and p4581. Answer here: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524428,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/28/2019 18:12:06",
          "content": "<p>Yep, thanks again for the clarification. Indeed it matches completely, the frequency and the 12us gap. So the TTF field in the train data must be wrong within bins. Seeing this other dataset is really convincing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 535875,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "05/23/2019 15:05:41",
      "content": "<p>Hi <a href=\"/leighplt\">@leighplt</a>. Thanks for uploading the data from exp p4581. When you pulled the data from matlab, was there shearing force? If there was, can you please upload that as well? I would like to try to use shearing force. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 536337,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "05/24/2019 09:30:16",
          "content": "<p>Unfortunately shearing force not in data</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "524125": "https://www.kaggle.com/redstr/lanl-p4581\n\nUpdate: actually, better use Leigh's version: https://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581\n\nUpdate2: I have split the data into channels properly.\n\nI downloaded all data from https://sites.psu.edu/chasbolton/ and put it on Kaggle as a dataset. It should be possible to use within kernels.\n\nThe data is large: the size is over 20x the size of the training set in the competition data. As to how useful it is, you can decide for yourself. It's certainly not a plug-and-play training set extension. The quake information is not present, and it's unclear how to define quakes there. The basic statistics are also different:\n![](https://i.imgur.com/3sVuoXc.png)\n\nSo models based on basic statistics like variance, are going to have a hard time using this data for training as it is. Maybe someone can find a way to normalize it.",
    "524132": "Thanks.",
    "524251": "Thanks for sharing!\n\nIf this data is the same as we are using, I would say that we are having a case of data leakage. Besides this, it seems that the training data provided by the problem owner is not really representative or even from a similar of to the whole experimental data.  These characteristic makes really hard to make a decent prediction.",
    "524271": "It's not the same",
    "524296": "I compresed single channel to 11GB (data have 2 channel: 33 and 34)",
    "524336": "Can you clarify? What is the channel, and how to distinguish them? Maybe this is the source of difference in the statistics?",
    "524339": "Yes, the data comes from another experiment. The experiment setup is very similar, only here they varied the shear stress in time. In the picture above you can see the staircase pattern: these are the times of different shear stresses. The stress range at around 500 mark on the plot is similar to the reported shear stress during the experiment which we are trying to predict. However, the \"acoustic power\" is still quite different from what is observed in the train set.",
    "524350": "What is the channel, and how to distinguish them?\n- channels2save: 33, 34. Read data description. Channel have same info\n\nMaybe this is the source of difference in the statistics?\n- No. Difference in various normal stress. Look at [this paper](https://folk.uio.no/karenmai/publications/mair_JGR_2002.pdf) (image on page 2)",
    "524357": "You are right, thanks for the info! I didn't even look in that file. So how are the channels arranged? Are they interleaved, or separated into first and the second half of the file?",
    "524365": "Algorithm in file plotacousticdata.m\n\nhttps://www.kaggle.com/leighplt/laboratory-acoustic-data-exp4581\n\nSplited by events",
    "524394": "Thanks, your data is much better organized than mine!\nAlso, the mystery of size 4095 bursts is resolved: it's probably an off-by-one mistake when joining the files. They likely missed the first or the last tick of each file. Note that each original .ac file has exactly 1280 size-4096 bursts.\nIn your opinion, is their event markup trustworthy? How similar are their events to our training quakes?",
    "524398": "Bertrand said:\nThe data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\n\nThey splitted data by shear info. It's more accuracy. I splitted by picks of variation (not so accuracy).",
    "524402": "Ah, so you split the data yourself? I thought you split by their \"manual event picks\" which are together with the data. I just looked at them and it's some weird stuff.\nCould you specify your criterion for quake detection?\nI know what Bertrand said. But he didn't say why every 1280 bins there is a bin of size 4095 in the train set. Now we know almost for sure. Because the very first bin is of size 4095, they probably missed the first tick in each file.",
    "524409": "Automated splitted and fix some mistake. Splitted by friction (function of variation) as recommend Bertrand. Split where sudden dropped friction. Its easy where normal stress small, but if its large - detect splitted harder",
    "524419": "Thanks. One more thing, which I think is important: the time between ticks within one 4096-tick bin in this data is specified as 2.5202e-07, while the time between bin starts is 0.001044. This means that there is a very little gap between bins.\nIn the training set, according to \"ttf\", the time between ticks within a bin is 1.1029617e-09 (~200 times smaller), while the time between bin starts is 0.001, i.e. almost the same. In your opinion, is TTF field in our train set wrong, or is the data indeed recorded at over 200x frequency compared to p4581?",
    "524422": "TTF fields are equal in our set and p4581. Answer here: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526",
    "524428": "Yep, thanks again for the clarification. Indeed it matches completely, the frequency and the 12us gap. So the TTF field in the train data must be wrong within bins. Seeing this other dataset is really convincing.",
    "535875": "Hi @leighplt. Thanks for uploading the data from exp p4581. When you pulled the data from matlab, was there shearing force? If there was, can you please upload that as well? I would like to try to use shearing force. Thanks!",
    "536337": "Unfortunately shearing force not in data"
  },
  "source": "meta"
}