{
  "id": 90664,
  "title": "Are data from p4677 ?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90664",
  "author_name": "",
  "post_date": "2019-04-25T18:14:24.302942100Z",
  "votes": 57,
  "comment_count": 59,
  "views": 0,
  "content": "<p>Based on:\n\"Earthquake Catalog‐Based Machine Learning Identification of Laboratory Fault States and the Effects of Magnitude of Completeness\"\n<a href=\"https://doi.org/10.1029/2018GL079712\">Only abstract, but has similar information under \"Supporting Information\"</a>\n<a href=\"https://arxiv.org/pdf/1810.11539.pdf\">arXiv Version</a></p>\n\n<p><strong>Assumption:</strong>\nFailure event has somewhat arbitrary definition.</p>\n\n<blockquote>\n  <p>We define large failure events as times for which stress drop exceeds 0.05 MPa within 1 ms.</p>\n</blockquote>\n\n<p>If we compare train part of <strong>Figure 3</strong> (from arXiv version) with plot of time to failure of our data: sequence looks similar with exception of cycle (~2125 s and ~1500). Shear stress at this point drops and under the assumption could be seen as failure with different criteria.\nttf plot by @mks2192\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/523143/13096/download.png\" alt=\"ttf plot\"></p>\n\n<p>P.S. This paper contains a lot of useful information</p>",
  "messages": [
    {
      "id": "523197",
      "postDate": "04/25/2019 18:14:24",
      "content": "<p>Based on:\n\"Earthquake Catalog‐Based Machine Learning Identification of Laboratory Fault States and the Effects of Magnitude of Completeness\"\n<a href=\"https://doi.org/10.1029/2018GL079712\">Only abstract, but has similar information under \"Supporting Information\"</a>\n<a href=\"https://arxiv.org/pdf/1810.11539.pdf\">arXiv Version</a></p>\n\n<p><strong>Assumption:</strong>\nFailure event has somewhat arbitrary definition.</p>\n\n<blockquote>\n  <p>We define large failure events as times for which stress drop exceeds 0.05 MPa within 1 ms.</p>\n</blockquote>\n\n<p>If we compare train part of <strong>Figure 3</strong> (from arXiv version) with plot of time to failure of our data: sequence looks similar with exception of cycle (~2125 s and ~1500). Shear stress at this point drops and under the assumption could be seen as failure with different criteria.\nttf plot by @mks2192\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/523143/13096/download.png\" alt=\"ttf plot\"></p>\n\n<p>P.S. This paper contains a lot of useful information</p>",
      "rawMarkdown": "Based on:\n\"Earthquake Catalog‐Based Machine Learning Identification of Laboratory Fault States and the Effects of Magnitude of Completeness\"\n[Only abstract, but has similar information under \"Supporting Information\"](https://doi.org/10.1029/2018GL079712)\n[arXiv Version](https://arxiv.org/pdf/1810.11539.pdf)\n\n**Assumption:**\nFailure event has somewhat arbitrary definition.\n&gt; We define large failure events as times for which stress drop exceeds 0.05 MPa within 1 ms.\n\nIf we compare train part of **Figure 3** (from arXiv version) with plot of time to failure of our data: sequence looks similar with exception of cycle (~2125 s and ~1500). Shear stress at this point drops and under the assumption could be seen as failure with different criteria.\nttf plot by @mks2192\n![ttf plot](https://storage.googleapis.com/kaggle-forum-message-attachments/523143/13096/download.png)\n\nP.S. This paper contains a lot of useful information",
      "votes": null
    },
    {
      "id": "523207",
      "postDate": "04/25/2019 18:28:11",
      "content": "<p>I agree, it does look very similar (with that one exception at ~1500, which might just be because of another definition of failure).\nEven the time scale is surprisingly close to ours here. </p>\n\n<p>Great finding! Now we \"only\" need to find the test set ;)</p>\n\n<p>In the paper, the authors state that all data is available at <a href=\"http://www3.geosc.psu.edu/~cjm38/\">http://www3.geosc.psu.edu/~cjm38/</a>\nApparently data for experiment p4677 is missing / has been removed for the time of this competition. </p>\n\n<p>Experiment p4581 is available here: <a href=\"https://psu.app.box.com/s/6c2mkkb8s3qhb74urx18igbk5sg7iius/folder/53413362846\">https://psu.app.box.com/s/6c2mkkb8s3qhb74urx18igbk5sg7iius/folder/53413362846</a></p>\n\n<p>could be VERY interesting for some extra training data</p>\n\n<p>Here: <a href=\"https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL076708\">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL076708</a> we see both experiments being used by our friend Bertrand</p>",
      "rawMarkdown": "I agree, it does look very similar (with that one exception at ~1500, which might just be because of another definition of failure).\nEven the time scale is surprisingly close to ours here. \n\nGreat finding! Now we \"only\" need to find the test set ;)\n\nIn the paper, the authors state that all data is available at http://www3.geosc.psu.edu/~cjm38/\nApparently data for experiment p4677 is missing / has been removed for the time of this competition. \n\nExperiment p4581 is available here: https://psu.app.box.com/s/6c2mkkb8s3qhb74urx18igbk5sg7iius/folder/53413362846\n\ncould be VERY interesting for some extra training data\n\nHere: https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL076708 we see both experiments being used by our friend Bertrand",
      "votes": null
    },
    {
      "id": "523257",
      "postDate": "04/25/2019 20:53:49",
      "content": "<p>weird</p>",
      "rawMarkdown": "weird",
      "votes": null
    },
    {
      "id": "523283",
      "postDate": "04/25/2019 22:45:21",
      "content": "<p>Good find <a href=\"/mykper\">@mykper</a>. I think it is worth stepping back from it all and think through this problem a bit more.</p>",
      "rawMarkdown": "Good find @mykper. I think it is worth stepping back from it all and think through this problem a bit more.",
      "votes": null
    },
    {
      "id": "523369",
      "postDate": "04/26/2019 05:29:40",
      "content": "<p>If this is true, then it shows that test data isn't much different form train data, it does not have more 'small ttf' than train.</p>",
      "rawMarkdown": "If this is true, then it shows that test data isn't much different form train data, it does not have more 'small ttf' than train.",
      "votes": null
    },
    {
      "id": "523418",
      "postDate": "04/26/2019 08:16:20",
      "content": "<p>That's not necessarily true. \nThe test set could be from a much later/earlier period of time in that same experiment or it could be sampled in a specific way.</p>",
      "rawMarkdown": "That's not necessarily true. \nThe test set could be from a much later/earlier period of time in that same experiment or it could be sampled in a specific way.",
      "votes": null
    },
    {
      "id": "523453",
      "postDate": "04/26/2019 10:07:23",
      "content": "<p>Why on earth would the paper authors not use ALL the experiment data?  Not only the train part matches quite precisely our train part, but the length of their test part is similar to the length of our test part.  I bet they split test into 150k chunks then shuffled them.</p>\n\n<p>There is a general principle in science called Occam razor: favor the simplest explanation.</p>\n\n<p>It is very useful when analyzing data.  And it is also useful when modeling data.  Simpler mdoels are better.  For instance a model with, say 10 features, is likely to generalize better than a model with 300 features, even if they look the same on CV and public LB.</p>",
      "rawMarkdown": "Why on earth would the paper authors not use ALL the experiment data?  Not only the train part matches quite precisely our train part, but the length of their test part is similar to the length of our test part.  I bet they split test into 150k chunks then shuffled them.\n\nThere is a general principle in science called Occam razor: favor the simplest explanation.\n\nIt is very useful when analyzing data.  And it is also useful when modeling data.  Simpler mdoels are better.  For instance a model with, say 10 features, is likely to generalize better than a model with 300 features, even if they look the same on CV and public LB.",
      "votes": null
    },
    {
      "id": "523454",
      "postDate": "04/26/2019 10:16:30",
      "content": "<p>They did on bigger part, <strong>Figure S1</strong> and <strong>Figure S2</strong>. Shear stress there clearly has downward trend, and drastic change in behavior, maybe due to loss of material.</p>",
      "rawMarkdown": "They did on bigger part, **Figure S1** and **Figure S2**. Shear stress there clearly has downward trend, and drastic change in behavior, maybe due to loss of material.",
      "votes": null
    },
    {
      "id": "523472",
      "postDate": "04/26/2019 11:14:43",
      "content": "<p>This cant not be. They have to be exactly similar. The train data are continues. I can agree there are some similarities between them, but this can be explained by the fact they always use the same experiment setup which provide more or less similar behaviors for each measurement. </p>",
      "rawMarkdown": "This cant not be. They have to be exactly similar. The train data are continues. I can agree there are some similarities between them, but this can be explained by the fact they always use the same experiment setup which provide more or less similar behaviors for each measurement.",
      "votes": null
    },
    {
      "id": "523485",
      "postDate": "04/26/2019 11:52:51",
      "content": "<p>CPMP, they did NOT use the entire experminent but just a small fraction of it. Using a longer period of time the long term drift in the experiment would need to be taken into account. \nApparently the fraction they used was a small plateau inside the bigger sequence. </p>\n\n<p>Please have a closer look at the published papers again!</p>\n\n<p>Though, I agree with you in terms of the test data here. It very much looks like the same test data as used in the paper. I was only saying that you can not tell for sure. It's only a reasonable guess. </p>\n\n<p>You can clearly see that e.g. the frequency of quakes decreases with time and so does the shear stress.</p>\n\n<p><img src=\"https://i.imgur.com/Lq3PWux.png\" alt=\"fraction\"> </p>",
      "rawMarkdown": "CPMP, they did NOT use the entire experminent but just a small fraction of it. Using a longer period of time the long term drift in the experiment would need to be taken into account. \nApparently the fraction they used was a small plateau inside the bigger sequence. \n\nPlease have a closer look at the published papers again!\n\nThough, I agree with you in terms of the test data here. It very much looks like the same test data as used in the paper. I was only saying that you can not tell for sure. It's only a reasonable guess. \n\nYou can clearly see that e.g. the frequency of quakes decreases with time and so does the shear stress.\n\n![fraction](https://i.imgur.com/Lq3PWux.png)",
      "votes": null
    },
    {
      "id": "523487",
      "postDate": "04/26/2019 11:56:59",
      "content": "<p>Are you referring to the test data as a whole, or just the public LB data?</p>\n\n<p>Models generally work best around a TTF range of  2-8s. The LB scores are consistently better than overall CV and it's reasonable to conclude this range is overrepresented on the LB. Apart from that I have observed few differences in the train/test distribution of features. The most obvious are mean and one other statistical feature I'm trying to understand.</p>",
      "rawMarkdown": "Are you referring to the test data as a whole, or just the public LB data?\n\nModels generally work best around a TTF range of  2-8s. The LB scores are consistently better than overall CV and it's reasonable to conclude this range is overrepresented on the LB. Apart from that I have observed few differences in the train/test distribution of features. The most obvious are mean and one other statistical feature I'm trying to understand.",
      "votes": null
    },
    {
      "id": "523498",
      "postDate": "04/26/2019 12:14:02",
      "content": "<blockquote>\n  <p>it's reasonable to conclude this range is overrepresented on the LB</p>\n</blockquote>\n\n<p>Why is it reasonable assumption?</p>",
      "rawMarkdown": "&gt; it's reasonable to conclude this range is overrepresented on the LB\n\nWhy is it reasonable assumption?",
      "votes": null
    },
    {
      "id": "523508",
      "postDate": "04/26/2019 12:32:56",
      "content": "<p>Is it not generally unusual to obtain better scores on a test set than your own CV? It strongly implies that the portion of the test set on which we are scored contains examples which correspond to the range of lowest CV error. Every model I have produced consistently scores MAE &lt; 2 for the TTF range between 2-7/8s. Outside this window the error increases rapidly. This has been the case for every model architecture I've tried. Is my assumption unreasonable?</p>\n\n<p>I don't think it's a useful observation anyway since it tells us nothing about the test set as a whole, only the public fraction.  For reference, my MAE/TTF generally looks like this:</p>\n\n<p><img src=\"https://i.imgur.com/7JzXNeL.png\" alt=\"\"></p>",
      "rawMarkdown": "Is it not generally unusual to obtain better scores on a test set than your own CV? It strongly implies that the portion of the test set on which we are scored contains examples which correspond to the range of lowest CV error. Every model I have produced consistently scores MAE &lt; 2 for the TTF range between 2-7/8s. Outside this window the error increases rapidly. This has been the case for every model architecture I've tried. Is my assumption unreasonable?\n\nI don't think it's a useful observation anyway since it tells us nothing about the test set as a whole, only the public fraction.  For reference, my MAE/TTF generally looks like this:\n\n![](https://i.imgur.com/7JzXNeL.png)",
      "votes": null
    },
    {
      "id": "523565",
      "postDate": "04/26/2019 14:25:34",
      "content": "<p>I think we can safely say this is indeed the data from experiment p4677.\nSee the overlay of the two pictures (TTF from the experiment on kaggle, shear stress from the experiment p4677 in the paper):\n<img src=\"https://i.imgur.com/TTvkiWn.png\" alt=\"overlay\"></p>",
      "rawMarkdown": "I think we can safely say this is indeed the data from experiment p4677.\nSee the overlay of the two pictures (TTF from the experiment on kaggle, shear stress from the experiment p4677 in the paper):\n![overlay](https://i.imgur.com/TTvkiWn.png)",
      "votes": null
    },
    {
      "id": "523574",
      "postDate": "04/26/2019 14:35:04",
      "content": "<p>Securing academic data is like grabbing a fistful of water... I hope leakage doesn't spoil this competition like it has so many others.</p>",
      "rawMarkdown": "Securing academic data is like grabbing a fistful of water... I hope leakage doesn't spoil this competition like it has so many others.",
      "votes": null
    },
    {
      "id": "523576",
      "postDate": "04/26/2019 14:42:48",
      "content": "<p>Simple experiment.\nStarting from things we know: mean ttf of public set is 4.017 (score of sample submission) and for train ~5.68. It seems unlikely that random sample would have so much lower mean ttf, assuming train and test have similar distribution.\nSo lets assume that public set is continuous in time. :)\nI choose sample length of 350 pieces.\n<code>np.convolve(ttf, np.ones(350) / 350, mode='valid')</code>\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/523576/13108/ttf_plot.png\" alt=\"ttf_plot\">\nThen select points of intersection with LB value, that will be starting points for separate validation sets, rest - train.\n<img src=\"https://imgur.com/RnaaIby.png\" alt=\"pred_plot\">\nFixed an error.</p>",
      "rawMarkdown": "Simple experiment.\nStarting from things we know: mean ttf of public set is 4.017 (score of sample submission) and for train ~5.68. It seems unlikely that random sample would have so much lower mean ttf, assuming train and test have similar distribution.\nSo lets assume that public set is continuous in time. :)\nI choose sample length of 350 pieces.\n`np.convolve(ttf, np.ones(350) / 350, mode='valid')`\n![ttf_plot](https://storage.googleapis.com/kaggle-forum-message-attachments/523576/13108/ttf_plot.png)\nThen select points of intersection with LB value, that will be starting points for separate validation sets, rest - train.\n![pred_plot](https://imgur.com/RnaaIby.png)\nFixed an error.",
      "votes": null
    },
    {
      "id": "523601",
      "postDate": "04/26/2019 15:38:16",
      "content": "<p><a href=\"/mykper\">@mykper</a> Nice!</p>",
      "rawMarkdown": "mykper Nice!",
      "votes": null
    },
    {
      "id": "523605",
      "postDate": "04/26/2019 15:51:29",
      "content": "<p>If this is true then this is a huge leak IMHO. E.g. just looking at Fig. 1  in your last link: it shows the test data at quite high resolution, def. enough to map the segments onto it. It may be the death of this compo.</p>",
      "rawMarkdown": "If this is true then this is a huge leak IMHO. E.g. just looking at Fig. 1  in your last link: it shows the test data at quite high resolution, def. enough to map the segments onto it. It may be the death of this compo.",
      "votes": null
    },
    {
      "id": "523631",
      "postDate": "04/26/2019 16:50:43",
      "content": "<p>So p4677 cannot be downloaded anywhere, right? </p>",
      "rawMarkdown": "So p4677 cannot be downloaded anywhere, right?",
      "votes": null
    },
    {
      "id": "523647",
      "postDate": "04/26/2019 17:12:52",
      "content": "<p>I cold not find it, but it does not mean nobody can.  I hope nobody can.</p>",
      "rawMarkdown": "I cold not find it, but it does not mean nobody can.  I hope nobody can.",
      "votes": null
    },
    {
      "id": "523650",
      "postDate": "04/26/2019 17:15:42",
      "content": "<p>Well it would be a clear leak if someone could and only discloses (if at all) the external data in the end.</p>",
      "rawMarkdown": "Well it would be a clear leak if someone could and only discloses (if at all) the external data in the end.",
      "votes": null
    },
    {
      "id": "523653",
      "postDate": "04/26/2019 17:17:11",
      "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> I guess we are not speaking about the same here. I said test data, not public test data.  </p>",
      "rawMarkdown": "bigironsphere I guess we are not speaking about the same here. I said test data, not public test data.",
      "votes": null
    },
    {
      "id": "523659",
      "postDate": "04/26/2019 17:21:21",
      "content": "<p><a href=\"/ilu000\">@ilu000</a> </p>\n\n<blockquote>\n  <p>CPMP, they did NOT use the entire experminent but just a small fraction of it. </p>\n</blockquote>\n\n<p>Yes, my point is not that.  Sorry if my original wording was misleading. </p>\n\n<p>They selected a subset of the data for their paper, because that subset exhibit stationary behavior to some extent.  Why on earth would they select a different subset for the competition?  The training part is almost the same, why would the test part be different?</p>",
      "rawMarkdown": "ilu000 \n\n&gt; CPMP, they did NOT use the entire experminent but just a small fraction of it. \n\nYes, my point is not that.  Sorry if my original wording was misleading. \n\nThey selected a subset of the data for their paper, because that subset exhibit stationary behavior to some extent.  Why on earth would they select a different subset for the competition?  The training part is almost the same, why would the test part be different?",
      "votes": null
    },
    {
      "id": "523673",
      "postDate": "04/26/2019 17:59:26",
      "content": "<p>Unless you use it to tune your models, but not actually train on it , therefore you dont need to disclose using it</p>",
      "rawMarkdown": "Unless you use it to tune your models, but not actually train on it , therefore you dont need to disclose using it",
      "votes": null
    },
    {
      "id": "523677",
      "postDate": "04/26/2019 18:12:16",
      "content": "<p>My understanding of this is different. So you are using it? :)</p>",
      "rawMarkdown": "My understanding of this is different. So you are using it? :)",
      "votes": null
    },
    {
      "id": "523683",
      "postDate": "04/26/2019 18:34:54",
      "content": "<p>I suppose only the competition creators can tell if the data could be leaked. We don't know if it was ever up to download, or if they sent it privately to anyone.</p>",
      "rawMarkdown": "I suppose only the competition creators can tell if the data could be leaked. We don't know if it was ever up to download, or if they sent it privately to anyone.",
      "votes": null
    },
    {
      "id": "523736",
      "postDate": "04/26/2019 22:04:11",
      "content": "<p>All ok</p>",
      "rawMarkdown": "All ok",
      "votes": null
    },
    {
      "id": "523737",
      "postDate": "04/26/2019 22:10:36",
      "content": "<p>This is sad. </p>",
      "rawMarkdown": "This is sad.",
      "votes": null
    },
    {
      "id": "523756",
      "postDate": "04/26/2019 23:54:27",
      "content": "<p>Can anyone who pulled all the data on the experiments from p4581 and p2394 dataset please upload the files here onto kaggle? It was be helpful for everyone who doesn't have MATLAB.</p>",
      "rawMarkdown": "Can anyone who pulled all the data on the experiments from p4581 and p2394 dataset please upload the files here onto kaggle? It was be helpful for everyone who doesn't have MATLAB.",
      "votes": null
    },
    {
      "id": "523763",
      "postDate": "04/27/2019 00:38:56",
      "content": "<p>I took a look at the p4581 files. I didn't download all the data, but, from what I can tell, only acoustic data is available. While there could be some usefulness of the acoustic data by itself, there is no shear stress or TTF data to use as targets.</p>",
      "rawMarkdown": "I took a look at the p4581 files. I didn't download all the data, but, from what I can tell, only acoustic data is available. While there could be some usefulness of the acoustic data by itself, there is no shear stress or TTF data to use as targets.",
      "votes": null
    },
    {
      "id": "523777",
      "postDate": "04/27/2019 01:52:11",
      "content": "<p>Did anybody even figure out how to download all the data from this box thing? I even have a box account, but there is no way to copy or download it except by manually clicking \"Save link as...\" on thousands of files.</p>",
      "rawMarkdown": "Did anybody even figure out how to download all the data from this box thing? I even have a box account, but there is no way to copy or download it except by manually clicking \"Save link as...\" on thousands of files.",
      "votes": null
    },
    {
      "id": "523821",
      "postDate": "04/27/2019 05:54:06",
      "content": "<p>If I blow the figure up so that train is 680 pixels wide, over 629 145 480, and test is 150 000, means each test segment would be .16 pixels wide?</p>",
      "rawMarkdown": "If I blow the figure up so that train is 680 pixels wide, over 629 145 480, and test is 150 000, means each test segment would be .16 pixels wide?",
      "votes": null
    },
    {
      "id": "523872",
      "postDate": "04/27/2019 09:08:49",
      "content": "<p>I guess, if an acoustic data is not splitted into random chunks, it would be easy to determine the time of failure.</p>",
      "rawMarkdown": "I guess, if an acoustic data is not splitted into random chunks, it would be easy to determine the time of failure.",
      "votes": null
    },
    {
      "id": "523879",
      "postDate": "04/27/2019 09:35:35",
      "content": "<p>@CPMP I agree with you. And indeed the test part looks very similar. \nThankfully it seems that the authors of this challenge at least shuffled the test set. I'll create a post for that later the day. </p>",
      "rawMarkdown": "CPMP I agree with you. And indeed the test part looks very similar. \nThankfully it seems that the authors of this challenge at least shuffled the test set. I'll create a post for that later the day.",
      "votes": null
    },
    {
      "id": "523894",
      "postDate": "04/27/2019 09:59:41",
      "content": "<p>If they privately sent it to anyone, or anyone already had it, that person will surprisingly appear on the LB in the final days with just a few submissions. I think Kaggle/host must ensure that there will be no such things, otherwise Kaggle will get a massive trust blow from users, and that will be a huge scandal. </p>",
      "rawMarkdown": "If they privately sent it to anyone, or anyone already had it, that person will surprisingly appear on the LB in the final days with just a few submissions. I think Kaggle/host must ensure that there will be no such things, otherwise Kaggle will get a massive trust blow from users, and that will be a huge scandal.",
      "votes": null
    },
    {
      "id": "523922",
      "postDate": "04/27/2019 12:04:43",
      "content": "<p>For me, the most important insight is that test data is similar to training data as shown in Figure 1 in the paper. </p>\n\n<p>And I hope that no one finds any data that can cause a leakage 🙃</p>",
      "rawMarkdown": "For me, the most important insight is that test data is similar to training data as shown in Figure 1 in the paper. \n\nAnd I hope that no one finds any data that can cause a leakage 🙃",
      "votes": null
    },
    {
      "id": "523974",
      "postDate": "04/27/2019 14:43:52",
      "content": "<p>Yeah, seems like there is no way to map it actually. You can find the figure in higher resolution elsewhere, but you get at most 0.25pixels/segment. Which is a good thing. Still, a close call.</p>",
      "rawMarkdown": "Yeah, seems like there is no way to map it actually. You can find the figure in higher resolution elsewhere, but you get at most 0.25pixels/segment. Which is a good thing. Still, a close call.",
      "votes": null
    },
    {
      "id": "524020",
      "postDate": "04/27/2019 17:46:56",
      "content": "<p><a href=\"/mykper\">@mykper</a> How do we really know that mean ttf of public set is <strong>4.017</strong>?</p>",
      "rawMarkdown": "mykper How do we really know that mean ttf of public set is **4.017**?",
      "votes": null
    },
    {
      "id": "524023",
      "postDate": "04/27/2019 17:56:05",
      "content": "<p>The sample submission (all TTF==0) scores 4.017 on LB.</p>",
      "rawMarkdown": "The sample submission (all TTF==0) scores 4.017 on LB.",
      "votes": null
    },
    {
      "id": "524202",
      "postDate": "04/28/2019 07:50:12",
      "content": "<p>Lab made many experiments, some of it very similar:\n(look at p4***)\n<a href=\"https://ars.els-cdn.com/content/image/1-s2.0-S0040195119301325-gr2.jpg\">https://ars.els-cdn.com/content/image/1-s2.0-S0040195119301325-gr2.jpg</a></p>",
      "rawMarkdown": "Lab made many experiments, some of it very similar:\n(look at p4***)\nhttps://ars.els-cdn.com/content/image/1-s2.0-S0040195119301325-gr2.jpg",
      "votes": null
    },
    {
      "id": "524268",
      "postDate": "04/28/2019 11:05:21",
      "content": "<p><a href=\"https://doi.org/10.1016/j.tecto.2019.04.010\">https://doi.org/10.1016/j.tecto.2019.04.010</a> This paper?</p>",
      "rawMarkdown": "[https://doi.org/10.1016/j.tecto.2019.04.010](https://doi.org/10.1016/j.tecto.2019.04.010) This paper?",
      "votes": null
    },
    {
      "id": "524295",
      "postDate": "04/28/2019 12:50:51",
      "content": "<p>Yep. I think train data from another exp, but similar to p4677 (article with data from p4677 by ~2016y) Lab makes many experiments, so for kaggle can prepered new data.</p>",
      "rawMarkdown": "Yep. I think train data from another exp, but similar to p4677 (article with data from p4677 by ~2016y) Lab makes many experiments, so for kaggle can prepered new data.",
      "votes": null
    },
    {
      "id": "524321",
      "postDate": "04/28/2019 13:55:08",
      "content": "<p>Isn't it a requirement to show your methodology to the organisers if you're a prizewinner? Could be a nasty surprise for any aspiring cheaters... </p>",
      "rawMarkdown": "Isn't it a requirement to show your methodology to the organisers if you're a prizewinner? Could be a nasty surprise for any aspiring cheaters...",
      "votes": null
    },
    {
      "id": "524355",
      "postDate": "04/28/2019 15:23:33",
      "content": "<p><a href=\"http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf\">Characterization of Acoustic Emissions from Laboratory Stick-Slip events in Simulated Fault Gouge</a> They are using different part of p4677 there.</p>",
      "rawMarkdown": "[Characterization of Acoustic Emissions from Laboratory Stick-Slip events in Simulated Fault Gouge](http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf) They are using different part of p4677 there.",
      "votes": null
    },
    {
      "id": "524476",
      "postDate": "04/28/2019 21:17:14",
      "content": "<p>nice!</p>",
      "rawMarkdown": "nice!",
      "votes": null
    },
    {
      "id": "524637",
      "postDate": "04/29/2019 08:32:59",
      "content": "<p>Why do you say it would be cheating?  What specific rule would this violate?</p>",
      "rawMarkdown": "Why do you say it would be cheating?  What specific rule would this violate?",
      "votes": null
    },
    {
      "id": "524672",
      "postDate": "04/29/2019 10:16:27",
      "content": "<p>Allowing use of external data in this competition looks like a bad joke. :(</p>",
      "rawMarkdown": "Allowing use of external data in this competition looks like a bad joke. :(",
      "votes": null
    },
    {
      "id": "524696",
      "postDate": "04/29/2019 11:37:10",
      "content": "<blockquote>\n  <p>Allowing use of external data in this competition looks like a bad joke. :(</p>\n</blockquote>\n\n<p>Indeed, but it is allowed.  Let's pray that there is no leak of test data...  And if that happens let's pray that the person finding it will be sharing openly so that the competition is reset.</p>",
      "rawMarkdown": "&gt; Allowing use of external data in this competition looks like a bad joke. :(\n\nIndeed, but it is allowed.  Let's pray that there is no leak of test data...  And if that happens let's pray that the person finding it will be sharing openly so that the competition is reset.",
      "votes": null
    },
    {
      "id": "524718",
      "postDate": "04/29/2019 12:17:21",
      "content": "<p>I still believe there's no one having it.</p>",
      "rawMarkdown": "I still believe there's no one having it.",
      "votes": null
    },
    {
      "id": "524768",
      "postDate": "04/29/2019 13:40:17",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<blockquote>\n  <p>The following provision supersedes General Rules Section 7.C. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. <strong>The source of any external data must be posted to the official competition forum prior to the Entry Deadline.</strong>\"</p>\n</blockquote>\n\n<p>Emphasis mine. Using the test data covertly to your advantage would invalidate your submission. I should have clarified that by 'cheating' I meant not declaring your private use of the test data in order to gain an advantage.</p>\n\n<p>Hopefully even if the test data is leaked, the LANL researchers will have other data they can use for evaluation. </p>",
      "rawMarkdown": "cpmpml \n&gt; The following provision supersedes General Rules Section 7.C. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. **The source of any external data must be posted to the official competition forum prior to the Entry Deadline.**\"\n\nEmphasis mine. Using the test data covertly to your advantage would invalidate your submission. I should have clarified that by 'cheating' I meant not declaring your private use of the test data in order to gain an advantage.\n\nHopefully even if the test data is leaked, the LANL researchers will have other data they can use for evaluation.",
      "votes": null
    },
    {
      "id": "524786",
      "postDate": "04/29/2019 14:16:12",
      "content": "<p>Wow... if this is true and test dataset is same with the paper, we can obtain some information about test dataset from the figure in this paper...</p>",
      "rawMarkdown": "Wow... if this is true and test dataset is same with the paper, we can obtain some information about test dataset from the figure in this paper...",
      "votes": null
    },
    {
      "id": "524904",
      "postDate": "04/29/2019 18:38:52",
      "content": "<blockquote>\n  <p><strong>ELIGIBILITY</strong>\n  The following provision supersedes General Rules Section 2.B. below: “Only Employees, interns, contractors, officers and directors of the Competition Sponsor, Kaggle Inc., and any other parties with access to the original research or ground truth of this competition's dataset, and their parent companies, subsidiaries and affiliates, are not eligible to win the competition. Individuals without access to the original research or ground truth are allowed to both enter and win.\"</p>\n</blockquote>\n\n<p>So, can we really use it or not? I'm confused.</p>",
      "rawMarkdown": "&gt; **ELIGIBILITY**\nThe following provision supersedes General Rules Section 2.B. below: “Only Employees, interns, contractors, officers and directors of the Competition Sponsor, Kaggle Inc., and any other parties with access to the original research or ground truth of this competition's dataset, and their parent companies, subsidiaries and affiliates, are not eligible to win the competition. Individuals without access to the original research or ground truth are allowed to both enter and win.\"\n\nSo, can we really use it or not? I'm confused.",
      "votes": null
    },
    {
      "id": "524929",
      "postDate": "04/29/2019 19:12:43",
      "content": "<blockquote>\n  <p>So, can we really use it or not? I'm confused.</p>\n</blockquote>\n\n<p>The ground truth data? Of course you can't use it. If it will be revealed that it was leaked, the \ncompetition will most likely be reset or cancelled. There is no way they will award you a prize if you admit that you used the ground truth in your model.</p>\n\n<p>If you don't disclose it, then it all depends on how well you can cover it up. But that's a really miserable game to play.</p>",
      "rawMarkdown": "&gt; So, can we really use it or not? I'm confused.\n\nThe ground truth data? Of course you can't use it. If it will be revealed that it was leaked, the \ncompetition will most likely be reset or cancelled. There is no way they will award you a prize if you admit that you used the ground truth in your model.\n\nIf you don't disclose it, then it all depends on how well you can cover it up. But that's a really miserable game to play.",
      "votes": null
    },
    {
      "id": "524931",
      "postDate": "04/29/2019 19:24:12",
      "content": "<p>No.\n&gt; original research</p>\n\n<p>Data from p4677, if it is \"original research\".</p>",
      "rawMarkdown": "No.\n&gt; original research\n\nData from p4677, if it is \"original research\".",
      "votes": null
    },
    {
      "id": "524944",
      "postDate": "04/29/2019 20:21:51",
      "content": "<p>Late in the protein competition, some of the ground truth was found on a public website and disclosed by the (very straightforward) finders. It made a big difference to scores but there were scripts made to allow everyone to easily use it. The competition just proceeded.</p>\n\n<p>Two things:\n   - it was only a part of the ground truth with a score difference of around 100-200 places\n   - I am not absolutely sure how much the private leaderboard was affected, I believe it was similar to the public LB.</p>",
      "rawMarkdown": "Late in the protein competition, some of the ground truth was found on a public website and disclosed by the (very straightforward) finders. It made a big difference to scores but there were scripts made to allow everyone to easily use it. The competition just proceeded.\n\nTwo things:\n   - it was only a part of the ground truth with a score difference of around 100-200 places\n   - I am not absolutely sure how much the private leaderboard was affected, I believe it was similar to the public LB.",
      "votes": null
    },
    {
      "id": "524969",
      "postDate": "04/29/2019 21:41:50",
      "content": "<p>This is pretty difficult, if not impossible to do.\nEach chunk of 150000 data points corresponds to about a quarter of a pixel in that diagramm. I wasn't able to match the chunks and hopefully no one else is. </p>",
      "rawMarkdown": "This is pretty difficult, if not impossible to do.\nEach chunk of 150000 data points corresponds to about a quarter of a pixel in that diagramm. I wasn't able to match the chunks and hopefully no one else is.",
      "votes": null
    },
    {
      "id": "525218",
      "postDate": "04/30/2019 12:55:47",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice",
      "votes": null
    },
    {
      "id": "527302",
      "postDate": "05/05/2019 04:06:55",
      "content": "<p><a href=\"/ilu000\">@ilu000</a> What software did you use to overlay the ttf data on top of the training shearing force data?</p>",
      "rawMarkdown": "ilu000 What software did you use to overlay the ttf data on top of the training shearing force data?",
      "votes": null
    },
    {
      "id": "535844",
      "postDate": "05/23/2019 14:09:16",
      "content": "<p>PowerPoint :D\nQuick and dirty solution</p>",
      "rawMarkdown": "PowerPoint :D\nQuick and dirty solution",
      "votes": null
    },
    {
      "id": "543544",
      "postDate": "06/04/2019 15:29:26",
      "content": "<p>This is the source of the leakage.</p>",
      "rawMarkdown": "This is the source of the leakage.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 523207,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "04/25/2019 18:28:11",
      "content": "<p>I agree, it does look very similar (with that one exception at ~1500, which might just be because of another definition of failure).\nEven the time scale is surprisingly close to ours here. </p>\n\n<p>Great finding! Now we \"only\" need to find the test set ;)</p>\n\n<p>In the paper, the authors state that all data is available at <a href=\"http://www3.geosc.psu.edu/~cjm38/\">http://www3.geosc.psu.edu/~cjm38/</a>\nApparently data for experiment p4677 is missing / has been removed for the time of this competition. </p>\n\n<p>Experiment p4581 is available here: <a href=\"https://psu.app.box.com/s/6c2mkkb8s3qhb74urx18igbk5sg7iius/folder/53413362846\">https://psu.app.box.com/s/6c2mkkb8s3qhb74urx18igbk5sg7iius/folder/53413362846</a></p>\n\n<p>could be VERY interesting for some extra training data</p>\n\n<p>Here: <a href=\"https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL076708\">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL076708</a> we see both experiments being used by our friend Bertrand</p>",
      "votes": null,
      "replies": [
        {
          "id": 523257,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "04/25/2019 20:53:49",
          "content": "<p>weird</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523605,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/26/2019 15:51:29",
          "content": "<p>If this is true then this is a huge leak IMHO. E.g. just looking at Fig. 1  in your last link: it shows the test data at quite high resolution, def. enough to map the segments onto it. It may be the death of this compo.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523821,
          "author_name": "glimmung",
          "author_url": "",
          "post_date": "04/27/2019 05:54:06",
          "content": "<p>If I blow the figure up so that train is 680 pixels wide, over 629 145 480, and test is 150 000, means each test segment would be .16 pixels wide?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523974,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/27/2019 14:43:52",
          "content": "<p>Yeah, seems like there is no way to map it actually. You can find the figure in higher resolution elsewhere, but you get at most 0.25pixels/segment. Which is a good thing. Still, a close call.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 523283,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "04/25/2019 22:45:21",
      "content": "<p>Good find <a href=\"/mykper\">@mykper</a>. I think it is worth stepping back from it all and think through this problem a bit more.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 523369,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/26/2019 05:29:40",
      "content": "<p>If this is true, then it shows that test data isn't much different form train data, it does not have more 'small ttf' than train.</p>",
      "votes": null,
      "replies": [
        {
          "id": 523418,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/26/2019 08:16:20",
          "content": "<p>That's not necessarily true. \nThe test set could be from a much later/earlier period of time in that same experiment or it could be sampled in a specific way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523453,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/26/2019 10:07:23",
          "content": "<p>Why on earth would the paper authors not use ALL the experiment data?  Not only the train part matches quite precisely our train part, but the length of their test part is similar to the length of our test part.  I bet they split test into 150k chunks then shuffled them.</p>\n\n<p>There is a general principle in science called Occam razor: favor the simplest explanation.</p>\n\n<p>It is very useful when analyzing data.  And it is also useful when modeling data.  Simpler mdoels are better.  For instance a model with, say 10 features, is likely to generalize better than a model with 300 features, even if they look the same on CV and public LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523454,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "04/26/2019 10:16:30",
          "content": "<p>They did on bigger part, <strong>Figure S1</strong> and <strong>Figure S2</strong>. Shear stress there clearly has downward trend, and drastic change in behavior, maybe due to loss of material.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523485,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/26/2019 11:52:51",
          "content": "<p>CPMP, they did NOT use the entire experminent but just a small fraction of it. Using a longer period of time the long term drift in the experiment would need to be taken into account. \nApparently the fraction they used was a small plateau inside the bigger sequence. </p>\n\n<p>Please have a closer look at the published papers again!</p>\n\n<p>Though, I agree with you in terms of the test data here. It very much looks like the same test data as used in the paper. I was only saying that you can not tell for sure. It's only a reasonable guess. </p>\n\n<p>You can clearly see that e.g. the frequency of quakes decreases with time and so does the shear stress.</p>\n\n<p><img src=\"https://i.imgur.com/Lq3PWux.png\" alt=\"fraction\"> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523487,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/26/2019 11:56:59",
          "content": "<p>Are you referring to the test data as a whole, or just the public LB data?</p>\n\n<p>Models generally work best around a TTF range of  2-8s. The LB scores are consistently better than overall CV and it's reasonable to conclude this range is overrepresented on the LB. Apart from that I have observed few differences in the train/test distribution of features. The most obvious are mean and one other statistical feature I'm trying to understand.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523498,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/26/2019 12:14:02",
          "content": "<blockquote>\n  <p>it's reasonable to conclude this range is overrepresented on the LB</p>\n</blockquote>\n\n<p>Why is it reasonable assumption?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523508,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/26/2019 12:32:56",
          "content": "<p>Is it not generally unusual to obtain better scores on a test set than your own CV? It strongly implies that the portion of the test set on which we are scored contains examples which correspond to the range of lowest CV error. Every model I have produced consistently scores MAE &lt; 2 for the TTF range between 2-7/8s. Outside this window the error increases rapidly. This has been the case for every model architecture I've tried. Is my assumption unreasonable?</p>\n\n<p>I don't think it's a useful observation anyway since it tells us nothing about the test set as a whole, only the public fraction.  For reference, my MAE/TTF generally looks like this:</p>\n\n<p><img src=\"https://i.imgur.com/7JzXNeL.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523576,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "04/26/2019 14:42:48",
          "content": "<p>Simple experiment.\nStarting from things we know: mean ttf of public set is 4.017 (score of sample submission) and for train ~5.68. It seems unlikely that random sample would have so much lower mean ttf, assuming train and test have similar distribution.\nSo lets assume that public set is continuous in time. :)\nI choose sample length of 350 pieces.\n<code>np.convolve(ttf, np.ones(350) / 350, mode='valid')</code>\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/523576/13108/ttf_plot.png\" alt=\"ttf_plot\">\nThen select points of intersection with LB value, that will be starting points for separate validation sets, rest - train.\n<img src=\"https://imgur.com/RnaaIby.png\" alt=\"pred_plot\">\nFixed an error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523601,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/26/2019 15:38:16",
          "content": "<p><a href=\"/mykper\">@mykper</a> Nice!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523653,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/26/2019 17:17:11",
          "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> I guess we are not speaking about the same here. I said test data, not public test data.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523659,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/26/2019 17:21:21",
          "content": "<p><a href=\"/ilu000\">@ilu000</a> </p>\n\n<blockquote>\n  <p>CPMP, they did NOT use the entire experminent but just a small fraction of it. </p>\n</blockquote>\n\n<p>Yes, my point is not that.  Sorry if my original wording was misleading. </p>\n\n<p>They selected a subset of the data for their paper, because that subset exhibit stationary behavior to some extent.  Why on earth would they select a different subset for the competition?  The training part is almost the same, why would the test part be different?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523879,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/27/2019 09:35:35",
          "content": "<p>@CPMP I agree with you. And indeed the test part looks very similar. \nThankfully it seems that the authors of this challenge at least shuffled the test set. I'll create a post for that later the day. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524020,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/27/2019 17:46:56",
          "content": "<p><a href=\"/mykper\">@mykper</a> How do we really know that mean ttf of public set is <strong>4.017</strong>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524023,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/27/2019 17:56:05",
          "content": "<p>The sample submission (all TTF==0) scores 4.017 on LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 523472,
      "author_name": "greenwing1985",
      "author_url": "",
      "post_date": "04/26/2019 11:14:43",
      "content": "<p>This cant not be. They have to be exactly similar. The train data are continues. I can agree there are some similarities between them, but this can be explained by the fact they always use the same experiment setup which provide more or less similar behaviors for each measurement. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 523565,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "04/26/2019 14:25:34",
      "content": "<p>I think we can safely say this is indeed the data from experiment p4677.\nSee the overlay of the two pictures (TTF from the experiment on kaggle, shear stress from the experiment p4677 in the paper):\n<img src=\"https://i.imgur.com/TTvkiWn.png\" alt=\"overlay\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 523574,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/26/2019 14:35:04",
          "content": "<p>Securing academic data is like grabbing a fistful of water... I hope leakage doesn't spoil this competition like it has so many others.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524786,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "04/29/2019 14:16:12",
          "content": "<p>Wow... if this is true and test dataset is same with the paper, we can obtain some information about test dataset from the figure in this paper...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524969,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/29/2019 21:41:50",
          "content": "<p>This is pretty difficult, if not impossible to do.\nEach chunk of 150000 data points corresponds to about a quarter of a pixel in that diagramm. I wasn't able to match the chunks and hopefully no one else is. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527302,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "05/05/2019 04:06:55",
          "content": "<p><a href=\"/ilu000\">@ilu000</a> What software did you use to overlay the ttf data on top of the training shearing force data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535844,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/23/2019 14:09:16",
          "content": "<p>PowerPoint :D\nQuick and dirty solution</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 523631,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "04/26/2019 16:50:43",
      "content": "<p>So p4677 cannot be downloaded anywhere, right? </p>",
      "votes": null,
      "replies": [
        {
          "id": 523647,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/26/2019 17:12:52",
          "content": "<p>I cold not find it, but it does not mean nobody can.  I hope nobody can.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523650,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/26/2019 17:15:42",
          "content": "<p>Well it would be a clear leak if someone could and only discloses (if at all) the external data in the end.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523673,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "04/26/2019 17:59:26",
          "content": "<p>Unless you use it to tune your models, but not actually train on it , therefore you dont need to disclose using it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523677,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/26/2019 18:12:16",
          "content": "<p>My understanding of this is different. So you are using it? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523683,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/26/2019 18:34:54",
          "content": "<p>I suppose only the competition creators can tell if the data could be leaked. We don't know if it was ever up to download, or if they sent it privately to anyone.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523894,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "04/27/2019 09:59:41",
          "content": "<p>If they privately sent it to anyone, or anyone already had it, that person will surprisingly appear on the LB in the final days with just a few submissions. I think Kaggle/host must ensure that there will be no such things, otherwise Kaggle will get a massive trust blow from users, and that will be a huge scandal. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524321,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/28/2019 13:55:08",
          "content": "<p>Isn't it a requirement to show your methodology to the organisers if you're a prizewinner? Could be a nasty surprise for any aspiring cheaters... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524637,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/29/2019 08:32:59",
          "content": "<p>Why do you say it would be cheating?  What specific rule would this violate?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524672,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "04/29/2019 10:16:27",
          "content": "<p>Allowing use of external data in this competition looks like a bad joke. :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524696,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/29/2019 11:37:10",
          "content": "<blockquote>\n  <p>Allowing use of external data in this competition looks like a bad joke. :(</p>\n</blockquote>\n\n<p>Indeed, but it is allowed.  Let's pray that there is no leak of test data...  And if that happens let's pray that the person finding it will be sharing openly so that the competition is reset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524718,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "04/29/2019 12:17:21",
          "content": "<p>I still believe there's no one having it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524768,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/29/2019 13:40:17",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<blockquote>\n  <p>The following provision supersedes General Rules Section 7.C. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. <strong>The source of any external data must be posted to the official competition forum prior to the Entry Deadline.</strong>\"</p>\n</blockquote>\n\n<p>Emphasis mine. Using the test data covertly to your advantage would invalidate your submission. I should have clarified that by 'cheating' I meant not declaring your private use of the test data in order to gain an advantage.</p>\n\n<p>Hopefully even if the test data is leaked, the LANL researchers will have other data they can use for evaluation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524904,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "04/29/2019 18:38:52",
          "content": "<blockquote>\n  <p><strong>ELIGIBILITY</strong>\n  The following provision supersedes General Rules Section 2.B. below: “Only Employees, interns, contractors, officers and directors of the Competition Sponsor, Kaggle Inc., and any other parties with access to the original research or ground truth of this competition's dataset, and their parent companies, subsidiaries and affiliates, are not eligible to win the competition. Individuals without access to the original research or ground truth are allowed to both enter and win.\"</p>\n</blockquote>\n\n<p>So, can we really use it or not? I'm confused.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524929,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/29/2019 19:12:43",
          "content": "<blockquote>\n  <p>So, can we really use it or not? I'm confused.</p>\n</blockquote>\n\n<p>The ground truth data? Of course you can't use it. If it will be revealed that it was leaked, the \ncompetition will most likely be reset or cancelled. There is no way they will award you a prize if you admit that you used the ground truth in your model.</p>\n\n<p>If you don't disclose it, then it all depends on how well you can cover it up. But that's a really miserable game to play.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524931,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "04/29/2019 19:24:12",
          "content": "<p>No.\n&gt; original research</p>\n\n<p>Data from p4677, if it is \"original research\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 523736,
      "author_name": "alinayatsko",
      "author_url": "",
      "post_date": "04/26/2019 22:04:11",
      "content": "<p>All ok</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 523737,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "04/26/2019 22:10:36",
      "content": "<p>This is sad. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 523756,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "04/26/2019 23:54:27",
      "content": "<p>Can anyone who pulled all the data on the experiments from p4581 and p2394 dataset please upload the files here onto kaggle? It was be helpful for everyone who doesn't have MATLAB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 523763,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "04/27/2019 00:38:56",
          "content": "<p>I took a look at the p4581 files. I didn't download all the data, but, from what I can tell, only acoustic data is available. While there could be some usefulness of the acoustic data by itself, there is no shear stress or TTF data to use as targets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523777,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "04/27/2019 01:52:11",
          "content": "<p>Did anybody even figure out how to download all the data from this box thing? I even have a box account, but there is no way to copy or download it except by manually clicking \"Save link as...\" on thousands of files.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523872,
          "author_name": "sorokin",
          "author_url": "",
          "post_date": "04/27/2019 09:08:49",
          "content": "<p>I guess, if an acoustic data is not splitted into random chunks, it would be easy to determine the time of failure.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 523922,
      "author_name": "ammar111",
      "author_url": "",
      "post_date": "04/27/2019 12:04:43",
      "content": "<p>For me, the most important insight is that test data is similar to training data as shown in Figure 1 in the paper. </p>\n\n<p>And I hope that no one finds any data that can cause a leakage 🙃</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 524202,
      "author_name": "leighplt",
      "author_url": "",
      "post_date": "04/28/2019 07:50:12",
      "content": "<p>Lab made many experiments, some of it very similar:\n(look at p4***)\n<a href=\"https://ars.els-cdn.com/content/image/1-s2.0-S0040195119301325-gr2.jpg\">https://ars.els-cdn.com/content/image/1-s2.0-S0040195119301325-gr2.jpg</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 524268,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "04/28/2019 11:05:21",
          "content": "<p><a href=\"https://doi.org/10.1016/j.tecto.2019.04.010\">https://doi.org/10.1016/j.tecto.2019.04.010</a> This paper?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524295,
          "author_name": "leighplt",
          "author_url": "",
          "post_date": "04/28/2019 12:50:51",
          "content": "<p>Yep. I think train data from another exp, but similar to p4677 (article with data from p4677 by ~2016y) Lab makes many experiments, so for kaggle can prepered new data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524355,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "04/28/2019 15:23:33",
          "content": "<p><a href=\"http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf\">Characterization of Acoustic Emissions from Laboratory Stick-Slip events in Simulated Fault Gouge</a> They are using different part of p4677 there.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 524476,
      "author_name": "sahbasalarian",
      "author_url": "",
      "post_date": "04/28/2019 21:17:14",
      "content": "<p>nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 524944,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "04/29/2019 20:21:51",
      "content": "<p>Late in the protein competition, some of the ground truth was found on a public website and disclosed by the (very straightforward) finders. It made a big difference to scores but there were scripts made to allow everyone to easily use it. The competition just proceeded.</p>\n\n<p>Two things:\n   - it was only a part of the ground truth with a score difference of around 100-200 places\n   - I am not absolutely sure how much the private leaderboard was affected, I believe it was similar to the public LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 525218,
      "author_name": "ilikeevb",
      "author_url": "",
      "post_date": "04/30/2019 12:55:47",
      "content": "<p>Nice</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543544,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 15:29:26",
      "content": "<p>This is the source of the leakage.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "523197": "Based on:\n\"Earthquake Catalog‐Based Machine Learning Identification of Laboratory Fault States and the Effects of Magnitude of Completeness\"\n[Only abstract, but has similar information under \"Supporting Information\"](https://doi.org/10.1029/2018GL079712)\n[arXiv Version](https://arxiv.org/pdf/1810.11539.pdf)\n\n**Assumption:**\nFailure event has somewhat arbitrary definition.\n&gt; We define large failure events as times for which stress drop exceeds 0.05 MPa within 1 ms.\n\nIf we compare train part of **Figure 3** (from arXiv version) with plot of time to failure of our data: sequence looks similar with exception of cycle (~2125 s and ~1500). Shear stress at this point drops and under the assumption could be seen as failure with different criteria.\nttf plot by @mks2192\n![ttf plot](https://storage.googleapis.com/kaggle-forum-message-attachments/523143/13096/download.png)\n\nP.S. This paper contains a lot of useful information",
    "523207": "I agree, it does look very similar (with that one exception at ~1500, which might just be because of another definition of failure).\nEven the time scale is surprisingly close to ours here. \n\nGreat finding! Now we \"only\" need to find the test set ;)\n\nIn the paper, the authors state that all data is available at http://www3.geosc.psu.edu/~cjm38/\nApparently data for experiment p4677 is missing / has been removed for the time of this competition. \n\nExperiment p4581 is available here: https://psu.app.box.com/s/6c2mkkb8s3qhb74urx18igbk5sg7iius/folder/53413362846\n\ncould be VERY interesting for some extra training data\n\nHere: https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL076708 we see both experiments being used by our friend Bertrand",
    "523257": "weird",
    "523283": "Good find @mykper. I think it is worth stepping back from it all and think through this problem a bit more.",
    "523369": "If this is true, then it shows that test data isn't much different form train data, it does not have more 'small ttf' than train.",
    "523418": "That's not necessarily true. \nThe test set could be from a much later/earlier period of time in that same experiment or it could be sampled in a specific way.",
    "523453": "Why on earth would the paper authors not use ALL the experiment data?  Not only the train part matches quite precisely our train part, but the length of their test part is similar to the length of our test part.  I bet they split test into 150k chunks then shuffled them.\n\nThere is a general principle in science called Occam razor: favor the simplest explanation.\n\nIt is very useful when analyzing data.  And it is also useful when modeling data.  Simpler mdoels are better.  For instance a model with, say 10 features, is likely to generalize better than a model with 300 features, even if they look the same on CV and public LB.",
    "523454": "They did on bigger part, **Figure S1** and **Figure S2**. Shear stress there clearly has downward trend, and drastic change in behavior, maybe due to loss of material.",
    "523472": "This cant not be. They have to be exactly similar. The train data are continues. I can agree there are some similarities between them, but this can be explained by the fact they always use the same experiment setup which provide more or less similar behaviors for each measurement.",
    "523485": "CPMP, they did NOT use the entire experminent but just a small fraction of it. Using a longer period of time the long term drift in the experiment would need to be taken into account. \nApparently the fraction they used was a small plateau inside the bigger sequence. \n\nPlease have a closer look at the published papers again!\n\nThough, I agree with you in terms of the test data here. It very much looks like the same test data as used in the paper. I was only saying that you can not tell for sure. It's only a reasonable guess. \n\nYou can clearly see that e.g. the frequency of quakes decreases with time and so does the shear stress.\n\n![fraction](https://i.imgur.com/Lq3PWux.png)",
    "523487": "Are you referring to the test data as a whole, or just the public LB data?\n\nModels generally work best around a TTF range of  2-8s. The LB scores are consistently better than overall CV and it's reasonable to conclude this range is overrepresented on the LB. Apart from that I have observed few differences in the train/test distribution of features. The most obvious are mean and one other statistical feature I'm trying to understand.",
    "523498": "&gt; it's reasonable to conclude this range is overrepresented on the LB\n\nWhy is it reasonable assumption?",
    "523508": "Is it not generally unusual to obtain better scores on a test set than your own CV? It strongly implies that the portion of the test set on which we are scored contains examples which correspond to the range of lowest CV error. Every model I have produced consistently scores MAE &lt; 2 for the TTF range between 2-7/8s. Outside this window the error increases rapidly. This has been the case for every model architecture I've tried. Is my assumption unreasonable?\n\nI don't think it's a useful observation anyway since it tells us nothing about the test set as a whole, only the public fraction.  For reference, my MAE/TTF generally looks like this:\n\n![](https://i.imgur.com/7JzXNeL.png)",
    "523565": "I think we can safely say this is indeed the data from experiment p4677.\nSee the overlay of the two pictures (TTF from the experiment on kaggle, shear stress from the experiment p4677 in the paper):\n![overlay](https://i.imgur.com/TTvkiWn.png)",
    "523574": "Securing academic data is like grabbing a fistful of water... I hope leakage doesn't spoil this competition like it has so many others.",
    "523576": "Simple experiment.\nStarting from things we know: mean ttf of public set is 4.017 (score of sample submission) and for train ~5.68. It seems unlikely that random sample would have so much lower mean ttf, assuming train and test have similar distribution.\nSo lets assume that public set is continuous in time. :)\nI choose sample length of 350 pieces.\n`np.convolve(ttf, np.ones(350) / 350, mode='valid')`\n![ttf_plot](https://storage.googleapis.com/kaggle-forum-message-attachments/523576/13108/ttf_plot.png)\nThen select points of intersection with LB value, that will be starting points for separate validation sets, rest - train.\n![pred_plot](https://imgur.com/RnaaIby.png)\nFixed an error.",
    "523601": "mykper Nice!",
    "523605": "If this is true then this is a huge leak IMHO. E.g. just looking at Fig. 1  in your last link: it shows the test data at quite high resolution, def. enough to map the segments onto it. It may be the death of this compo.",
    "523631": "So p4677 cannot be downloaded anywhere, right?",
    "523647": "I cold not find it, but it does not mean nobody can.  I hope nobody can.",
    "523650": "Well it would be a clear leak if someone could and only discloses (if at all) the external data in the end.",
    "523653": "bigironsphere I guess we are not speaking about the same here. I said test data, not public test data.",
    "523659": "ilu000 \n\n&gt; CPMP, they did NOT use the entire experminent but just a small fraction of it. \n\nYes, my point is not that.  Sorry if my original wording was misleading. \n\nThey selected a subset of the data for their paper, because that subset exhibit stationary behavior to some extent.  Why on earth would they select a different subset for the competition?  The training part is almost the same, why would the test part be different?",
    "523673": "Unless you use it to tune your models, but not actually train on it , therefore you dont need to disclose using it",
    "523677": "My understanding of this is different. So you are using it? :)",
    "523683": "I suppose only the competition creators can tell if the data could be leaked. We don't know if it was ever up to download, or if they sent it privately to anyone.",
    "523736": "All ok",
    "523737": "This is sad.",
    "523756": "Can anyone who pulled all the data on the experiments from p4581 and p2394 dataset please upload the files here onto kaggle? It was be helpful for everyone who doesn't have MATLAB.",
    "523763": "I took a look at the p4581 files. I didn't download all the data, but, from what I can tell, only acoustic data is available. While there could be some usefulness of the acoustic data by itself, there is no shear stress or TTF data to use as targets.",
    "523777": "Did anybody even figure out how to download all the data from this box thing? I even have a box account, but there is no way to copy or download it except by manually clicking \"Save link as...\" on thousands of files.",
    "523821": "If I blow the figure up so that train is 680 pixels wide, over 629 145 480, and test is 150 000, means each test segment would be .16 pixels wide?",
    "523872": "I guess, if an acoustic data is not splitted into random chunks, it would be easy to determine the time of failure.",
    "523879": "CPMP I agree with you. And indeed the test part looks very similar. \nThankfully it seems that the authors of this challenge at least shuffled the test set. I'll create a post for that later the day.",
    "523894": "If they privately sent it to anyone, or anyone already had it, that person will surprisingly appear on the LB in the final days with just a few submissions. I think Kaggle/host must ensure that there will be no such things, otherwise Kaggle will get a massive trust blow from users, and that will be a huge scandal.",
    "523922": "For me, the most important insight is that test data is similar to training data as shown in Figure 1 in the paper. \n\nAnd I hope that no one finds any data that can cause a leakage 🙃",
    "523974": "Yeah, seems like there is no way to map it actually. You can find the figure in higher resolution elsewhere, but you get at most 0.25pixels/segment. Which is a good thing. Still, a close call.",
    "524020": "mykper How do we really know that mean ttf of public set is **4.017**?",
    "524023": "The sample submission (all TTF==0) scores 4.017 on LB.",
    "524202": "Lab made many experiments, some of it very similar:\n(look at p4***)\nhttps://ars.els-cdn.com/content/image/1-s2.0-S0040195119301325-gr2.jpg",
    "524268": "[https://doi.org/10.1016/j.tecto.2019.04.010](https://doi.org/10.1016/j.tecto.2019.04.010) This paper?",
    "524295": "Yep. I think train data from another exp, but similar to p4677 (article with data from p4677 by ~2016y) Lab makes many experiments, so for kaggle can prepered new data.",
    "524321": "Isn't it a requirement to show your methodology to the organisers if you're a prizewinner? Could be a nasty surprise for any aspiring cheaters...",
    "524355": "[Characterization of Acoustic Emissions from Laboratory Stick-Slip events in Simulated Fault Gouge](http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf) They are using different part of p4677 there.",
    "524476": "nice!",
    "524637": "Why do you say it would be cheating?  What specific rule would this violate?",
    "524672": "Allowing use of external data in this competition looks like a bad joke. :(",
    "524696": "&gt; Allowing use of external data in this competition looks like a bad joke. :(\n\nIndeed, but it is allowed.  Let's pray that there is no leak of test data...  And if that happens let's pray that the person finding it will be sharing openly so that the competition is reset.",
    "524718": "I still believe there's no one having it.",
    "524768": "cpmpml \n&gt; The following provision supersedes General Rules Section 7.C. below: “You may use data, other than the Competition Data, as allowed on the Competition Website to develop and test your models and Submissions; provided, you have the right and authority to use such external data for the purposes of the Competition, and to share such data with Sponsor and Kaggle as may be required. **The source of any external data must be posted to the official competition forum prior to the Entry Deadline.**\"\n\nEmphasis mine. Using the test data covertly to your advantage would invalidate your submission. I should have clarified that by 'cheating' I meant not declaring your private use of the test data in order to gain an advantage.\n\nHopefully even if the test data is leaked, the LANL researchers will have other data they can use for evaluation.",
    "524786": "Wow... if this is true and test dataset is same with the paper, we can obtain some information about test dataset from the figure in this paper...",
    "524904": "&gt; **ELIGIBILITY**\nThe following provision supersedes General Rules Section 2.B. below: “Only Employees, interns, contractors, officers and directors of the Competition Sponsor, Kaggle Inc., and any other parties with access to the original research or ground truth of this competition's dataset, and their parent companies, subsidiaries and affiliates, are not eligible to win the competition. Individuals without access to the original research or ground truth are allowed to both enter and win.\"\n\nSo, can we really use it or not? I'm confused.",
    "524929": "&gt; So, can we really use it or not? I'm confused.\n\nThe ground truth data? Of course you can't use it. If it will be revealed that it was leaked, the \ncompetition will most likely be reset or cancelled. There is no way they will award you a prize if you admit that you used the ground truth in your model.\n\nIf you don't disclose it, then it all depends on how well you can cover it up. But that's a really miserable game to play.",
    "524931": "No.\n&gt; original research\n\nData from p4677, if it is \"original research\".",
    "524944": "Late in the protein competition, some of the ground truth was found on a public website and disclosed by the (very straightforward) finders. It made a big difference to scores but there were scripts made to allow everyone to easily use it. The competition just proceeded.\n\nTwo things:\n   - it was only a part of the ground truth with a score difference of around 100-200 places\n   - I am not absolutely sure how much the private leaderboard was affected, I believe it was similar to the public LB.",
    "524969": "This is pretty difficult, if not impossible to do.\nEach chunk of 150000 data points corresponds to about a quarter of a pixel in that diagramm. I wasn't able to match the chunks and hopefully no one else is.",
    "525218": "Nice",
    "527302": "ilu000 What software did you use to overlay the ttf data on top of the training shearing force data?",
    "535844": "PowerPoint :D\nQuick and dirty solution",
    "543544": "This is the source of the leakage."
  },
  "source": "meta"
}