{
  "id": 91583,
  "title": "Public set time to failure distribution",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91583",
  "author_name": "",
  "post_date": "2019-05-06T18:16:34.369772100Z",
  "votes": 36,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Here I will try to estimate some values associated with public set.\nMean (as already been discussed, submission with 0 constant) - 4.017\nMaximum value can be estimated by finding minimum constant value, at which equality holds:\n<code>MAE(C) + MAE(0) == C</code>\nFor 11, 10, 9: 6.982, 5.982, 5.017. Maximum value is less than 10. We could get better estimate by reducing step, but since it is unlikely that public set is randomly sampled (given the mean and the maximum value), finding the best fit in train set could be an option.\nMAE scores for constant values from 0 to 10:\n<code>4.017, 3.163, 2.589, 2.295, 2.281, 2.532, 2.943, 3.494, 4.185, 5.017, 5.982</code>\nFor next step we will need to know the length of public set, 9 values with log variance greater than 9 were set to 10000, rest - ~0 (in train log_var &gt; 9 corresponds to segments with ttf close to 0). With score of 61.485, best guess is 348.\n```\n(4.017 * 348 + 2 * 10000) / 348\n61.48826436781609</p>\n\n<p>```\nI sampled series of 3 random cycles (EQ), merged them in continuous pieces and selected sets that meet conditions (3.5 &lt; mean &lt; 4.5 and 9 &lt;= max &lt;= 10).\nClosest to public set (euclidean distance between MAE scores):</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13173/best_fit.png\" alt=\"\">\nMAE scores.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13174/mae_comp.png\" alt=\"\">\nSorted time to failure (for public estimated by LGB)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13175/sort_comp.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "527970",
      "postDate": "05/06/2019 18:16:34",
      "content": "<p>Here I will try to estimate some values associated with public set.\nMean (as already been discussed, submission with 0 constant) - 4.017\nMaximum value can be estimated by finding minimum constant value, at which equality holds:\n<code>MAE(C) + MAE(0) == C</code>\nFor 11, 10, 9: 6.982, 5.982, 5.017. Maximum value is less than 10. We could get better estimate by reducing step, but since it is unlikely that public set is randomly sampled (given the mean and the maximum value), finding the best fit in train set could be an option.\nMAE scores for constant values from 0 to 10:\n<code>4.017, 3.163, 2.589, 2.295, 2.281, 2.532, 2.943, 3.494, 4.185, 5.017, 5.982</code>\nFor next step we will need to know the length of public set, 9 values with log variance greater than 9 were set to 10000, rest - ~0 (in train log_var &gt; 9 corresponds to segments with ttf close to 0). With score of 61.485, best guess is 348.\n```\n(4.017 * 348 + 2 * 10000) / 348\n61.48826436781609</p>\n\n<p>```\nI sampled series of 3 random cycles (EQ), merged them in continuous pieces and selected sets that meet conditions (3.5 &lt; mean &lt; 4.5 and 9 &lt;= max &lt;= 10).\nClosest to public set (euclidean distance between MAE scores):</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13173/best_fit.png\" alt=\"\">\nMAE scores.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13174/mae_comp.png\" alt=\"\">\nSorted time to failure (for public estimated by LGB)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13175/sort_comp.png\" alt=\"\"></p>",
      "rawMarkdown": "Here I will try to estimate some values associated with public set.\nMean (as already been discussed, submission with 0 constant) - 4.017\nMaximum value can be estimated by finding minimum constant value, at which equality holds:\n`MAE(C) + MAE(0) == C`\nFor 11, 10, 9: 6.982, 5.982, 5.017. Maximum value is less than 10. We could get better estimate by reducing step, but since it is unlikely that public set is randomly sampled (given the mean and the maximum value), finding the best fit in train set could be an option.\nMAE scores for constant values from 0 to 10:\n`4.017, 3.163, 2.589, 2.295, 2.281, 2.532, 2.943, 3.494, 4.185, 5.017, 5.982`\nFor next step we will need to know the length of public set, 9 values with log variance greater than 9 were set to 10000, rest - ~0 (in train log_var &gt; 9 corresponds to segments with ttf close to 0). With score of 61.485, best guess is 348.\n```\n(4.017 * 348 + 2 * 10000) / 348\n61.48826436781609\n\n```\nI sampled series of 3 random cycles (EQ), merged them in continuous pieces and selected sets that meet conditions (3.5 &lt; mean &lt; 4.5 and 9 &lt;= max &lt;= 10).\nClosest to public set (euclidean distance between MAE scores):\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13173/best_fit.png)\nMAE scores.\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13174/mae_comp.png)\nSorted time to failure (for public estimated by LGB)\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13175/sort_comp.png)",
      "votes": null
    },
    {
      "id": "527979",
      "postDate": "05/06/2019 18:35:20",
      "content": "<p>That look interesting... What are the last 2 graphs representing ? Do you see a very strong correlation betwen on CV score on this folder and your public LB ?</p>",
      "rawMarkdown": "That look interesting... What are the last 2 graphs representing ? Do you see a very strong correlation betwen on CV score on this folder and your public LB ?",
      "votes": null
    },
    {
      "id": "527985",
      "postDate": "05/06/2019 18:46:31",
      "content": "<p>MAE scores plot. x is constant value and y - mae score of that constant value.\nWith last graph I tried to show distribution of ttf values, length of the step corresponds to estimated count of values that lie between consecutive constant values.\nRegarding last question, I've not tested yet :)</p>",
      "rawMarkdown": "MAE scores plot. x is constant value and y - mae score of that constant value.\nWith last graph I tried to show distribution of ttf values, length of the step corresponds to estimated count of values that lie between consecutive constant values.\nRegarding last question, I've not tested yet :)",
      "votes": null
    },
    {
      "id": "527991",
      "postDate": "05/06/2019 18:59:37",
      "content": "<p>We know the length of public set, right? 13% should be 341 samples</p>\n\n<p>Your subset from training is a continuous segment, right? I think that's the only way to meet the criteria of public properties you came up with.</p>",
      "rawMarkdown": "We know the length of public set, right? 13% should be 341 samples\n\nYour subset from training is a continuous segment, right? I think that's the only way to meet the criteria of public properties you came up with.",
      "votes": null
    },
    {
      "id": "527992",
      "postDate": "05/06/2019 19:05:01",
      "content": "<p>Length of 341 contradicts with result of LB probing.\nSegment on the graph might be continuous in training data (one could select almost exact from training data that will be continuous), but I don't think that is necessary since I'm looking only at ttf values.</p>",
      "rawMarkdown": "Length of 341 contradicts with result of LB probing.\nSegment on the graph might be continuous in training data (one could select almost exact from training data that will be continuous), but I don't think that is necessary since I'm looking only at ttf values.",
      "votes": null
    },
    {
      "id": "528002",
      "postDate": "05/06/2019 20:14:32",
      "content": "<p>Thanks for the analysis. Much appreciated. I have a question on the second part. At some point you say that</p>\n\n<p>&gt; 9 values with log variance greater than 9 were set to 10000</p>\n\n<p>Shouldn't be a 9 in the equation that comes afterwards instead of a 2? Am I forgetting something?\nThanks!</p>",
      "rawMarkdown": "Thanks for the analysis. Much appreciated. I have a question on the second part. At some point you say that\n\n&gt; 9 values with log variance greater than 9 were set to 10000\n\nShouldn't be a 9 in the equation that comes afterwards instead of a 2? Am I forgetting something?\nThanks!",
      "votes": null
    },
    {
      "id": "528003",
      "postDate": "05/06/2019 20:18:15",
      "content": "<p>Only 2 appear in public set. 9 in total</p>",
      "rawMarkdown": "Only 2 appear in public set. 9 in total",
      "votes": null
    },
    {
      "id": "528004",
      "postDate": "05/06/2019 20:20:54",
      "content": "<p>Ok, I get it now. Thanks.</p>",
      "rawMarkdown": "Ok, I get it now. Thanks.",
      "votes": null
    },
    {
      "id": "528018",
      "postDate": "05/06/2019 21:28:36",
      "content": "<p>I didn't understand anything. Writing full sentences and labeling the axes of the plots can help a better communication. </p>",
      "rawMarkdown": "I didn't understand anything. Writing full sentences and labeling the axes of the plots can help a better communication.",
      "votes": null
    },
    {
      "id": "528063",
      "postDate": "05/07/2019 02:37:36",
      "content": "<p>This statement: \n\"it is unlikely that public set is randomly sampled (given the mean and the maximum value)\"</p>\n\n<p>I don't agree with that. Your assumption is subjective \n You can try a random picking 300 samples from the training ttf, what is the probability of picking a set with mean 4? If you can prove that, your topic is valid, otherwise it is a huge mistake.</p>",
      "rawMarkdown": "This statement: \n\"it is unlikely that public set is randomly sampled (given the mean and the maximum value)\"\n\nI don't agree with that. Your assumption is subjective \n You can try a random picking 300 samples from the training ttf, what is the probability of picking a set with mean 4? If you can prove that, your topic is valid, otherwise it is a huge mistake.",
      "votes": null
    },
    {
      "id": "528094",
      "postDate": "05/07/2019 04:28:07",
      "content": "<blockquote>\n  <p>You can try a random picking 300 samples from the training ttf, what is the probability of picking a set with mean 4? </p>\n</blockquote>\n\n<p>Extremely low.  Just do a random sampling and you'll see it.</p>",
      "rawMarkdown": "&gt; You can try a random picking 300 samples from the training ttf, what is the probability of picking a set with mean 4? \n\nExtremely low.  Just do a random sampling and you'll see it.",
      "votes": null
    },
    {
      "id": "528096",
      "postDate": "05/07/2019 04:31:21",
      "content": "<p>Well done!</p>\n\n<p>The only thing I have issue with is your first plot.  There is only one quake in it while you assume there are two in public set.  What am I missing?</p>",
      "rawMarkdown": "Well done!\n\nThe only thing I have issue with is your first plot.  There is only one quake in it while you assume there are two in public set.  What am I missing?",
      "votes": null
    },
    {
      "id": "528102",
      "postDate": "05/07/2019 04:49:53",
      "content": "<p>Let me try to rephrase it.  </p>\n\n<p>First part is about finding the max value in public set.  Let MAE(c)be the score of a constant submission of value c.  Then if c1 and c2 are two values larger than the max of public set we have:</p>\n\n<pre><code>MAE(c1) - MAE(c2) = c1 - c2\n</code></pre>\n\n<p>Then, you fix c1 and look for the smallest c2 where the above equality hold.  This smallest value is the max value in public test set.</p>\n\n<p>Second part is about finding the number of samples in public test set.  Some assumptions are made.  First, high acoustic peaks before quakes have acoustic data std above exp(9).  This assumption is true on train data.   A submission with all 0 except for the 9 high peak segments set to 10000 should have a public LB score of</p>\n\n<pre><code>( 4.017 * n + k * 10000 ) / n\n</code></pre>\n\n<p>where n is the number samples in public test data and k the number of high acoustic peaks.  The score of that submission is 61.485, and the pair n,k that gives the closest value to it is 348, 2.  Proportion of public data in test set is then 348 / 2624 = 13.26 %, which is consistent with what we were told.</p>\n\n<p>Last part is where I have an issue with, see my other comment.  <a href=\"/mykper\">@mykper</a> made the assumption (which I agree with) that public test data is from a continuous portion of experience data.  Then he looked for a continuous subset of 3 quake cycles that look like public test data\n- Mean close to 4.017\n- 348 samples\n- Max value between 9 and 10\n- Contains 2 quakes \nWhat is not clear is why the one he says is best contains only one quake in it.</p>",
      "rawMarkdown": "Let me try to rephrase it.  \n\nFirst part is about finding the max value in public set.  Let MAE(c)be the score of a constant submission of value c.  Then if c1 and c2 are two values larger than the max of public set we have:\n\n    MAE(c1) - MAE(c2) = c1 - c2\n\nThen, you fix c1 and look for the smallest c2 where the above equality hold.  This smallest value is the max value in public test set.\n\nSecond part is about finding the number of samples in public test set.  Some assumptions are made.  First, high acoustic peaks before quakes have acoustic data std above exp(9).  This assumption is true on train data.   A submission with all 0 except for the 9 high peak segments set to 10000 should have a public LB score of\n\n    ( 4.017 * n + k * 10000 ) / n\n\nwhere n is the number samples in public test data and k the number of high acoustic peaks.  The score of that submission is 61.485, and the pair n,k that gives the closest value to it is 348, 2.  Proportion of public data in test set is then 348 / 2624 = 13.26 %, which is consistent with what we were told.\n\nLast part is where I have an issue with, see my other comment.  @mykper made the assumption (which I agree with) that public test data is from a continuous portion of experience data.  Then he looked for a continuous subset of 3 quake cycles that look like public test data\n- Mean close to 4.017\n- 348 samples\n- Max value between 9 and 10\n- Contains 2 quakes \nWhat is not clear is why the one he says is best contains only one quake in it.",
      "votes": null
    },
    {
      "id": "528103",
      "postDate": "05/07/2019 04:51:48",
      "content": "<p>Sorry for the empty messages.</p>",
      "rawMarkdown": "Sorry for the empty messages.",
      "votes": null
    },
    {
      "id": "528104",
      "postDate": "05/07/2019 04:52:27",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "528105",
      "postDate": "05/07/2019 04:53:13",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "528108",
      "postDate": "05/07/2019 04:57:27",
      "content": "<p>Genius detective. Thanks!</p>",
      "rawMarkdown": "Genius detective. Thanks!",
      "votes": null
    },
    {
      "id": "528115",
      "postDate": "05/07/2019 05:14:17",
      "content": "<p>Two (1 full short, 1 incomplete short/long) or even one (incomplete long quake), as long as the continuous segment is 348 samples long. </p>",
      "rawMarkdown": "Two (1 full short, 1 incomplete short/long) or even one (incomplete long quake), as long as the continuous segment is 348 samples long.",
      "votes": null
    },
    {
      "id": "528150",
      "postDate": "05/07/2019 06:49:34",
      "content": "<p>Thank you captain. Great explanation.</p>",
      "rawMarkdown": "Thank you captain. Great explanation.",
      "votes": null
    },
    {
      "id": "528151",
      "postDate": "05/07/2019 06:50:59",
      "content": "<p>We did that, the chance is very low, specifically with the properties posted above.</p>",
      "rawMarkdown": "We did that, the chance is very low, specifically with the properties posted above.",
      "votes": null
    },
    {
      "id": "528155",
      "postDate": "05/07/2019 06:55:26",
      "content": "<p>Did he say that it only contains one EQ? I think it's impossible. It can only be the end of 1 EQ and the start of another (or another full one).</p>",
      "rawMarkdown": "Did he say that it only contains one EQ? I think it's impossible. It can only be the end of 1 EQ and the start of another (or another full one).",
      "votes": null
    },
    {
      "id": "528156",
      "postDate": "05/07/2019 06:56:48",
      "content": "<p>I tried to exploit this observation a while ago but frustratingly could not get anything reasonable out of it. So either I was doing something wrong (very likely) or something is strange. Has anyone been able to utilize this to overfit to public LB?</p>",
      "rawMarkdown": "I tried to exploit this observation a while ago but frustratingly could not get anything reasonable out of it. So either I was doing something wrong (very likely) or something is strange. Has anyone been able to utilize this to overfit to public LB?",
      "votes": null
    },
    {
      "id": "528215",
      "postDate": "05/07/2019 09:14:59",
      "content": "<p><a href=\"/amjad85\">@amjad85</a> Sorry for that.\n<a href=\"/cpmpml\">@cpmpml</a> Thank you for clear explanation :)\n&gt; What is not clear is why the one he says is best contains only one quake in it.</p>\n\n<p>My bad, I created same MAE scores vector for each sample, and calculated distance between train MAE scores vectors and one I got for public set. This one has smallest distance to public \"MAE vector\". Sadly it gives almost no additional information about location of similar EQ in full test set.</p>\n\n<blockquote>\n  <p>248 samples</p>\n</blockquote>\n\n<p>348.</p>\n\n<blockquote>\n  <p>Contains 2 quakes </p>\n</blockquote>\n\n<p>Contains two precursors (high energy acoustic burst).</p>",
      "rawMarkdown": "amjad85 Sorry for that.\n@cpmpml Thank you for clear explanation :)\n&gt; What is not clear is why the one he says is best contains only one quake in it.\n\nMy bad, I created same MAE scores vector for each sample, and calculated distance between train MAE scores vectors and one I got for public set. This one has smallest distance to public \"MAE vector\". Sadly it gives almost no additional information about location of similar EQ in full test set.\n&gt; 248 samples\n\n348.\n&gt; Contains 2 quakes \n\nContains two precursors (high energy acoustic burst).",
      "votes": null
    },
    {
      "id": "528241",
      "postDate": "05/07/2019 10:32:29",
      "content": "<p>Thanks, fixed the typo.</p>",
      "rawMarkdown": "Thanks, fixed the typo.",
      "votes": null
    },
    {
      "id": "528266",
      "postDate": "05/07/2019 11:41:29",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Segments with log var &gt; 9 have ttf of ~0.32 in train, last point on plot has ttf ~0.137, few segments away from reset of ttf (start of new cycle). If we ignore missing segments and just roll array, plot could look like this (has exactly same properties):\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/528266/13178/roll.png\" alt=\"roll.png\"></p>\n\n<p>In presence of EQ with ~4 sec cycle that might be possible.\nApproach I choose is limited by variety of EQ in train set, i.e. assuming that train/test have similar EQ distribution, might not be the case, of course.</p>",
      "rawMarkdown": "cpmpml Segments with log var &gt; 9 have ttf of ~0.32 in train, last point on plot has ttf ~0.137, few segments away from reset of ttf (start of new cycle). If we ignore missing segments and just roll array, plot could look like this (has exactly same properties):\n![roll.png](https://storage.googleapis.com/kaggle-forum-message-attachments/528266/13178/roll.png)\n\nIn presence of EQ with ~4 sec cycle that might be possible.\nApproach I choose is limited by variety of EQ in train set, i.e. assuming that train/test have similar EQ distribution, might not be the case, of course.",
      "votes": null
    },
    {
      "id": "528313",
      "postDate": "05/07/2019 13:24:08",
      "content": "<p>Thanks, I like this one better, even if it is mathematically equivalent to the first one!</p>",
      "rawMarkdown": "Thanks, I like this one better, even if it is mathematically equivalent to the first one!",
      "votes": null
    },
    {
      "id": "528326",
      "postDate": "05/07/2019 13:46:13",
      "content": "<p>Something as simple as adjusting the mean value of the predictions by an offset or a multiplicative factor gave me worse LB results. So... no luck.</p>",
      "rawMarkdown": "Something as simple as adjusting the mean value of the predictions by an offset or a multiplicative factor gave me worse LB results. So... no luck.",
      "votes": null
    },
    {
      "id": "528359",
      "postDate": "05/07/2019 15:25:56",
      "content": "<blockquote>\n  <p>Did he say that it only contains one EQ? </p>\n</blockquote>\n\n<p>He didn't but his last figure in the post was a bit misleading as the second quake is the last point.  He then published a variant with two explicit peaks .</p>",
      "rawMarkdown": "&gt; Did he say that it only contains one EQ? \n\nHe didn't but his last figure in the post was a bit misleading as the second quake is the last point.  He then published a variant with two explicit peaks .",
      "votes": null
    },
    {
      "id": "528457",
      "postDate": "05/07/2019 21:56:43",
      "content": "<p>These are great observations. However, I think the private LB is much different than the public LB from all the aspects you investigated herein. </p>",
      "rawMarkdown": "These are great observations. However, I think the private LB is much different than the public LB from all the aspects you investigated herein.",
      "votes": null
    },
    {
      "id": "528650",
      "postDate": "05/08/2019 09:18:58",
      "content": "<p>-- deleted -- (I was not adding a new method)</p>",
      "rawMarkdown": "deleted -- (I was not adding a new method)",
      "votes": null
    },
    {
      "id": "529039",
      "postDate": "05/09/2019 04:09:11",
      "content": "<p>Great Kernel! </p>",
      "rawMarkdown": "Great Kernel!",
      "votes": null
    },
    {
      "id": "529858",
      "postDate": "05/11/2019 02:11:06",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "535215",
      "postDate": "05/22/2019 14:03:27",
      "content": "<p>By a little math we can actually approximate the cumulative distribution function (cdf) and probability density function (pdf) from this:\nDifferentiate <code>mae</code> for a constant model (which is, what was submitted here) with respect to the model value and you get the difference of the fraction of observations smaller and those larger the model value. Solve for the fraction and we have a cdf. Our <code>mae</code> function is quite coarse, but with a finite differences we can (quite roughly) approximate the derivative. This is what we get:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/535215/13267/test_ttf_cdf.png\" alt=\"cdf\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/535215/13268/test_ttf_pdf.png\" alt=\"pdf\"></p>",
      "rawMarkdown": "By a little math we can actually approximate the cumulative distribution function (cdf) and probability density function (pdf) from this:\nDifferentiate `mae` for a constant model (which is, what was submitted here) with respect to the model value and you get the difference of the fraction of observations smaller and those larger the model value. Solve for the fraction and we have a cdf. Our `mae` function is quite coarse, but with a finite differences we can (quite roughly) approximate the derivative. This is what we get:\n\n![cdf](https://storage.googleapis.com/kaggle-forum-message-attachments/535215/13267/test_ttf_cdf.png)\n\n![pdf](https://storage.googleapis.com/kaggle-forum-message-attachments/535215/13268/test_ttf_pdf.png)",
      "votes": null
    },
    {
      "id": "535218",
      "postDate": "05/22/2019 14:11:23",
      "content": "<p>With enough points, we could reconstruct cdf and pdf much more accurately. This is only limited by the number of submissions. If we would help together... Maybe another reason to like or dislike MAE.</p>\n\n<p>However, as ´pdf´ is exactly equal for TTF = 1,2,3 and 6,7,8 respectively, <a href=\"/mykper\">@mykper</a> is probably right about his assumption about the 2 segments of continuous data that was shuffled for the public test set.</p>",
      "rawMarkdown": "With enough points, we could reconstruct cdf and pdf much more accurately. This is only limited by the number of submissions. If we would help together... Maybe another reason to like or dislike MAE.\n\nHowever, as ´pdf´ is exactly equal for TTF = 1,2,3 and 6,7,8 respectively, @mykper is probably right about his assumption about the 2 segments of continuous data that was shuffled for the public test set.",
      "votes": null
    },
    {
      "id": "535262",
      "postDate": "05/22/2019 15:25:07",
      "content": "<p>Interesting, but this is for public test data only unfortunately.</p>",
      "rawMarkdown": "Interesting, but this is for public test data only unfortunately.",
      "votes": null
    },
    {
      "id": "535268",
      "postDate": "05/22/2019 15:34:39",
      "content": "<p>True. Actually not that helpful for this competition. I find it interesting in general, though.</p>",
      "rawMarkdown": "True. Actually not that helpful for this competition. I find it interesting in general, though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 527979,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "05/06/2019 18:35:20",
      "content": "<p>That look interesting... What are the last 2 graphs representing ? Do you see a very strong correlation betwen on CV score on this folder and your public LB ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 527985,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "05/06/2019 18:46:31",
          "content": "<p>MAE scores plot. x is constant value and y - mae score of that constant value.\nWith last graph I tried to show distribution of ttf values, length of the step corresponds to estimated count of values that lie between consecutive constant values.\nRegarding last question, I've not tested yet :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527991,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "05/06/2019 18:59:37",
      "content": "<p>We know the length of public set, right? 13% should be 341 samples</p>\n\n<p>Your subset from training is a continuous segment, right? I think that's the only way to meet the criteria of public properties you came up with.</p>",
      "votes": null,
      "replies": [
        {
          "id": 527992,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "05/06/2019 19:05:01",
          "content": "<p>Length of 341 contradicts with result of LB probing.\nSegment on the graph might be continuous in training data (one could select almost exact from training data that will be continuous), but I don't think that is necessary since I'm looking only at ttf values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528650,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "05/08/2019 09:18:58",
          "content": "<p>-- deleted -- (I was not adding a new method)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528002,
      "author_name": "ricarddelgado",
      "author_url": "",
      "post_date": "05/06/2019 20:14:32",
      "content": "<p>Thanks for the analysis. Much appreciated. I have a question on the second part. At some point you say that</p>\n\n<p>&gt; 9 values with log variance greater than 9 were set to 10000</p>\n\n<p>Shouldn't be a 9 in the equation that comes afterwards instead of a 2? Am I forgetting something?\nThanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 528003,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "05/06/2019 20:18:15",
          "content": "<p>Only 2 appear in public set. 9 in total</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528004,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "05/06/2019 20:20:54",
          "content": "<p>Ok, I get it now. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528018,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "05/06/2019 21:28:36",
      "content": "<p>I didn't understand anything. Writing full sentences and labeling the axes of the plots can help a better communication. </p>",
      "votes": null,
      "replies": [
        {
          "id": 528102,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 04:49:53",
          "content": "<p>Let me try to rephrase it.  </p>\n\n<p>First part is about finding the max value in public set.  Let MAE(c)be the score of a constant submission of value c.  Then if c1 and c2 are two values larger than the max of public set we have:</p>\n\n<pre><code>MAE(c1) - MAE(c2) = c1 - c2\n</code></pre>\n\n<p>Then, you fix c1 and look for the smallest c2 where the above equality hold.  This smallest value is the max value in public test set.</p>\n\n<p>Second part is about finding the number of samples in public test set.  Some assumptions are made.  First, high acoustic peaks before quakes have acoustic data std above exp(9).  This assumption is true on train data.   A submission with all 0 except for the 9 high peak segments set to 10000 should have a public LB score of</p>\n\n<pre><code>( 4.017 * n + k * 10000 ) / n\n</code></pre>\n\n<p>where n is the number samples in public test data and k the number of high acoustic peaks.  The score of that submission is 61.485, and the pair n,k that gives the closest value to it is 348, 2.  Proportion of public data in test set is then 348 / 2624 = 13.26 %, which is consistent with what we were told.</p>\n\n<p>Last part is where I have an issue with, see my other comment.  <a href=\"/mykper\">@mykper</a> made the assumption (which I agree with) that public test data is from a continuous portion of experience data.  Then he looked for a continuous subset of 3 quake cycles that look like public test data\n- Mean close to 4.017\n- 348 samples\n- Max value between 9 and 10\n- Contains 2 quakes \nWhat is not clear is why the one he says is best contains only one quake in it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528150,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/07/2019 06:49:34",
          "content": "<p>Thank you captain. Great explanation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528155,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/07/2019 06:55:26",
          "content": "<p>Did he say that it only contains one EQ? I think it's impossible. It can only be the end of 1 EQ and the start of another (or another full one).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528215,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "05/07/2019 09:14:59",
          "content": "<p><a href=\"/amjad85\">@amjad85</a> Sorry for that.\n<a href=\"/cpmpml\">@cpmpml</a> Thank you for clear explanation :)\n&gt; What is not clear is why the one he says is best contains only one quake in it.</p>\n\n<p>My bad, I created same MAE scores vector for each sample, and calculated distance between train MAE scores vectors and one I got for public set. This one has smallest distance to public \"MAE vector\". Sadly it gives almost no additional information about location of similar EQ in full test set.</p>\n\n<blockquote>\n  <p>248 samples</p>\n</blockquote>\n\n<p>348.</p>\n\n<blockquote>\n  <p>Contains 2 quakes </p>\n</blockquote>\n\n<p>Contains two precursors (high energy acoustic burst).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528241,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 10:32:29",
          "content": "<p>Thanks, fixed the typo.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528359,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 15:25:56",
          "content": "<blockquote>\n  <p>Did he say that it only contains one EQ? </p>\n</blockquote>\n\n<p>He didn't but his last figure in the post was a bit misleading as the second quake is the last point.  He then published a variant with two explicit peaks .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528063,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "05/07/2019 02:37:36",
      "content": "<p>This statement: \n\"it is unlikely that public set is randomly sampled (given the mean and the maximum value)\"</p>\n\n<p>I don't agree with that. Your assumption is subjective \n You can try a random picking 300 samples from the training ttf, what is the probability of picking a set with mean 4? If you can prove that, your topic is valid, otherwise it is a huge mistake.</p>",
      "votes": null,
      "replies": [
        {
          "id": 528094,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 04:28:07",
          "content": "<blockquote>\n  <p>You can try a random picking 300 samples from the training ttf, what is the probability of picking a set with mean 4? </p>\n</blockquote>\n\n<p>Extremely low.  Just do a random sampling and you'll see it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528151,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/07/2019 06:50:59",
          "content": "<p>We did that, the chance is very low, specifically with the properties posted above.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528096,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/07/2019 04:31:21",
      "content": "<p>Well done!</p>\n\n<p>The only thing I have issue with is your first plot.  There is only one quake in it while you assume there are two in public set.  What am I missing?</p>",
      "votes": null,
      "replies": [
        {
          "id": 528115,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/07/2019 05:14:17",
          "content": "<p>Two (1 full short, 1 incomplete short/long) or even one (incomplete long quake), as long as the continuous segment is 348 samples long. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528266,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "05/07/2019 11:41:29",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Segments with log var &gt; 9 have ttf of ~0.32 in train, last point on plot has ttf ~0.137, few segments away from reset of ttf (start of new cycle). If we ignore missing segments and just roll array, plot could look like this (has exactly same properties):\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/528266/13178/roll.png\" alt=\"roll.png\"></p>\n\n<p>In presence of EQ with ~4 sec cycle that might be possible.\nApproach I choose is limited by variety of EQ in train set, i.e. assuming that train/test have similar EQ distribution, might not be the case, of course.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528313,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 13:24:08",
          "content": "<p>Thanks, I like this one better, even if it is mathematically equivalent to the first one!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528103,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/07/2019 04:51:48",
      "content": "<p>Sorry for the empty messages.</p>",
      "votes": null,
      "replies": [
        {
          "id": 528104,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 04:52:27",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 528105,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 04:53:13",
          "content": "",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528108,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "05/07/2019 04:57:27",
      "content": "<p>Genius detective. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 528156,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "05/07/2019 06:56:48",
      "content": "<p>I tried to exploit this observation a while ago but frustratingly could not get anything reasonable out of it. So either I was doing something wrong (very likely) or something is strange. Has anyone been able to utilize this to overfit to public LB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 528326,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "05/07/2019 13:46:13",
          "content": "<p>Something as simple as adjusting the mean value of the predictions by an offset or a multiplicative factor gave me worse LB results. So... no luck.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528457,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "05/07/2019 21:56:43",
      "content": "<p>These are great observations. However, I think the private LB is much different than the public LB from all the aspects you investigated herein. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529039,
      "author_name": "akhileshrai",
      "author_url": "",
      "post_date": "05/09/2019 04:09:11",
      "content": "<p>Great Kernel! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529858,
      "author_name": "dzabalar",
      "author_url": "",
      "post_date": "05/11/2019 02:11:06",
      "content": "<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 535215,
      "author_name": "bernir",
      "author_url": "",
      "post_date": "05/22/2019 14:03:27",
      "content": "<p>By a little math we can actually approximate the cumulative distribution function (cdf) and probability density function (pdf) from this:\nDifferentiate <code>mae</code> for a constant model (which is, what was submitted here) with respect to the model value and you get the difference of the fraction of observations smaller and those larger the model value. Solve for the fraction and we have a cdf. Our <code>mae</code> function is quite coarse, but with a finite differences we can (quite roughly) approximate the derivative. This is what we get:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/535215/13267/test_ttf_cdf.png\" alt=\"cdf\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/535215/13268/test_ttf_pdf.png\" alt=\"pdf\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 535218,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "05/22/2019 14:11:23",
          "content": "<p>With enough points, we could reconstruct cdf and pdf much more accurately. This is only limited by the number of submissions. If we would help together... Maybe another reason to like or dislike MAE.</p>\n\n<p>However, as ´pdf´ is exactly equal for TTF = 1,2,3 and 6,7,8 respectively, <a href=\"/mykper\">@mykper</a> is probably right about his assumption about the 2 segments of continuous data that was shuffled for the public test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535262,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/22/2019 15:25:07",
          "content": "<p>Interesting, but this is for public test data only unfortunately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535268,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "05/22/2019 15:34:39",
          "content": "<p>True. Actually not that helpful for this competition. I find it interesting in general, though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "527970": "Here I will try to estimate some values associated with public set.\nMean (as already been discussed, submission with 0 constant) - 4.017\nMaximum value can be estimated by finding minimum constant value, at which equality holds:\n`MAE(C) + MAE(0) == C`\nFor 11, 10, 9: 6.982, 5.982, 5.017. Maximum value is less than 10. We could get better estimate by reducing step, but since it is unlikely that public set is randomly sampled (given the mean and the maximum value), finding the best fit in train set could be an option.\nMAE scores for constant values from 0 to 10:\n`4.017, 3.163, 2.589, 2.295, 2.281, 2.532, 2.943, 3.494, 4.185, 5.017, 5.982`\nFor next step we will need to know the length of public set, 9 values with log variance greater than 9 were set to 10000, rest - ~0 (in train log_var &gt; 9 corresponds to segments with ttf close to 0). With score of 61.485, best guess is 348.\n```\n(4.017 * 348 + 2 * 10000) / 348\n61.48826436781609\n\n```\nI sampled series of 3 random cycles (EQ), merged them in continuous pieces and selected sets that meet conditions (3.5 &lt; mean &lt; 4.5 and 9 &lt;= max &lt;= 10).\nClosest to public set (euclidean distance between MAE scores):\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13173/best_fit.png)\nMAE scores.\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13174/mae_comp.png)\nSorted time to failure (for public estimated by LGB)\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/527970/13175/sort_comp.png)",
    "527979": "That look interesting... What are the last 2 graphs representing ? Do you see a very strong correlation betwen on CV score on this folder and your public LB ?",
    "527985": "MAE scores plot. x is constant value and y - mae score of that constant value.\nWith last graph I tried to show distribution of ttf values, length of the step corresponds to estimated count of values that lie between consecutive constant values.\nRegarding last question, I've not tested yet :)",
    "527991": "We know the length of public set, right? 13% should be 341 samples\n\nYour subset from training is a continuous segment, right? I think that's the only way to meet the criteria of public properties you came up with.",
    "527992": "Length of 341 contradicts with result of LB probing.\nSegment on the graph might be continuous in training data (one could select almost exact from training data that will be continuous), but I don't think that is necessary since I'm looking only at ttf values.",
    "528002": "Thanks for the analysis. Much appreciated. I have a question on the second part. At some point you say that\n\n&gt; 9 values with log variance greater than 9 were set to 10000\n\nShouldn't be a 9 in the equation that comes afterwards instead of a 2? Am I forgetting something?\nThanks!",
    "528003": "Only 2 appear in public set. 9 in total",
    "528004": "Ok, I get it now. Thanks.",
    "528018": "I didn't understand anything. Writing full sentences and labeling the axes of the plots can help a better communication.",
    "528063": "This statement: \n\"it is unlikely that public set is randomly sampled (given the mean and the maximum value)\"\n\nI don't agree with that. Your assumption is subjective \n You can try a random picking 300 samples from the training ttf, what is the probability of picking a set with mean 4? If you can prove that, your topic is valid, otherwise it is a huge mistake.",
    "528094": "&gt; You can try a random picking 300 samples from the training ttf, what is the probability of picking a set with mean 4? \n\nExtremely low.  Just do a random sampling and you'll see it.",
    "528096": "Well done!\n\nThe only thing I have issue with is your first plot.  There is only one quake in it while you assume there are two in public set.  What am I missing?",
    "528102": "Let me try to rephrase it.  \n\nFirst part is about finding the max value in public set.  Let MAE(c)be the score of a constant submission of value c.  Then if c1 and c2 are two values larger than the max of public set we have:\n\n    MAE(c1) - MAE(c2) = c1 - c2\n\nThen, you fix c1 and look for the smallest c2 where the above equality hold.  This smallest value is the max value in public test set.\n\nSecond part is about finding the number of samples in public test set.  Some assumptions are made.  First, high acoustic peaks before quakes have acoustic data std above exp(9).  This assumption is true on train data.   A submission with all 0 except for the 9 high peak segments set to 10000 should have a public LB score of\n\n    ( 4.017 * n + k * 10000 ) / n\n\nwhere n is the number samples in public test data and k the number of high acoustic peaks.  The score of that submission is 61.485, and the pair n,k that gives the closest value to it is 348, 2.  Proportion of public data in test set is then 348 / 2624 = 13.26 %, which is consistent with what we were told.\n\nLast part is where I have an issue with, see my other comment.  @mykper made the assumption (which I agree with) that public test data is from a continuous portion of experience data.  Then he looked for a continuous subset of 3 quake cycles that look like public test data\n- Mean close to 4.017\n- 348 samples\n- Max value between 9 and 10\n- Contains 2 quakes \nWhat is not clear is why the one he says is best contains only one quake in it.",
    "528103": "Sorry for the empty messages.",
    "528104": "",
    "528105": "",
    "528108": "Genius detective. Thanks!",
    "528115": "Two (1 full short, 1 incomplete short/long) or even one (incomplete long quake), as long as the continuous segment is 348 samples long.",
    "528150": "Thank you captain. Great explanation.",
    "528151": "We did that, the chance is very low, specifically with the properties posted above.",
    "528155": "Did he say that it only contains one EQ? I think it's impossible. It can only be the end of 1 EQ and the start of another (or another full one).",
    "528156": "I tried to exploit this observation a while ago but frustratingly could not get anything reasonable out of it. So either I was doing something wrong (very likely) or something is strange. Has anyone been able to utilize this to overfit to public LB?",
    "528215": "amjad85 Sorry for that.\n@cpmpml Thank you for clear explanation :)\n&gt; What is not clear is why the one he says is best contains only one quake in it.\n\nMy bad, I created same MAE scores vector for each sample, and calculated distance between train MAE scores vectors and one I got for public set. This one has smallest distance to public \"MAE vector\". Sadly it gives almost no additional information about location of similar EQ in full test set.\n&gt; 248 samples\n\n348.\n&gt; Contains 2 quakes \n\nContains two precursors (high energy acoustic burst).",
    "528241": "Thanks, fixed the typo.",
    "528266": "cpmpml Segments with log var &gt; 9 have ttf of ~0.32 in train, last point on plot has ttf ~0.137, few segments away from reset of ttf (start of new cycle). If we ignore missing segments and just roll array, plot could look like this (has exactly same properties):\n![roll.png](https://storage.googleapis.com/kaggle-forum-message-attachments/528266/13178/roll.png)\n\nIn presence of EQ with ~4 sec cycle that might be possible.\nApproach I choose is limited by variety of EQ in train set, i.e. assuming that train/test have similar EQ distribution, might not be the case, of course.",
    "528313": "Thanks, I like this one better, even if it is mathematically equivalent to the first one!",
    "528326": "Something as simple as adjusting the mean value of the predictions by an offset or a multiplicative factor gave me worse LB results. So... no luck.",
    "528359": "&gt; Did he say that it only contains one EQ? \n\nHe didn't but his last figure in the post was a bit misleading as the second quake is the last point.  He then published a variant with two explicit peaks .",
    "528457": "These are great observations. However, I think the private LB is much different than the public LB from all the aspects you investigated herein.",
    "528650": "deleted -- (I was not adding a new method)",
    "529039": "Great Kernel!",
    "529858": "Thanks!",
    "535215": "By a little math we can actually approximate the cumulative distribution function (cdf) and probability density function (pdf) from this:\nDifferentiate `mae` for a constant model (which is, what was submitted here) with respect to the model value and you get the difference of the fraction of observations smaller and those larger the model value. Solve for the fraction and we have a cdf. Our `mae` function is quite coarse, but with a finite differences we can (quite roughly) approximate the derivative. This is what we get:\n\n![cdf](https://storage.googleapis.com/kaggle-forum-message-attachments/535215/13267/test_ttf_cdf.png)\n\n![pdf](https://storage.googleapis.com/kaggle-forum-message-attachments/535215/13268/test_ttf_pdf.png)",
    "535218": "With enough points, we could reconstruct cdf and pdf much more accurately. This is only limited by the number of submissions. If we would help together... Maybe another reason to like or dislike MAE.\n\nHowever, as ´pdf´ is exactly equal for TTF = 1,2,3 and 6,7,8 respectively, @mykper is probably right about his assumption about the 2 segments of continuous data that was shuffled for the public test set.",
    "535262": "Interesting, but this is for public test data only unfortunately.",
    "535268": "True. Actually not that helpful for this competition. I find it interesting in general, though."
  },
  "source": "meta"
}