{
  "id": 94359,
  "title": "Some elements of 7th place Solution",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94359",
  "author_name": "Antoine",
  "post_date": "2019-06-04T04:31:20.050000",
  "votes": 26,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Thanks Kaggle for this interesting competition ! </p>\n\n<p>Here are a few points from our solution, at least my part, (I'll let my teammate describes their interesting ideas):\n-1 EQ out validation scheme\n-Data augmentation with 30k chunk size\n-Bases Models are simple lgb/xgb model with fixed number of trees, huber or gamma objective and sample weights\n-Calculated around 200 features based mainly on STFT (see image attached), there was a correlation between the energy on some frequency band, and the ttf. I created those features based on hann window with size 1000 and 5000 for different frequency bands (44khz -  60khz, 60khz - 136khz,  140khz - 212khz, 216khz - 356 khz)\n-Used different indicators on on those bands : mean / std / quantiles + some some simple features on raw signal like Numbers of peaks above threshold or quantiles.\n-One way to improve our CV score was to tweak the original to lower the ttf or very long EQ. I basically did a parallel shift of ttf between the start of long EQ to the mini-EQ.</p>\n\n<p>This solution was in top 15 on public leaderboard.\nThen the difficult part began when we realized that Private Distribution might be very different from Public LB and also from Train Data.\nSo we developed different measures based on different possible test sets. We realized that the distribution of EQ type (short or long) and distribution of ttf had a significant impact on our CV score. We optimized 2 of those measures, with the hypothesis that distribution could be the one in the published paper. Public LB of those 2 subs were : 1.39 and 1.29. The best on private LB was the first one ;) I'm glad we were able to survive the shake-up, this was really not an easy task. </p>\n\n<p>Congrats to all of you who shared great content in the Kernels and in the Forum. Congratulation to the winning team ! </p>\n\n<p>And a special thanks to my teammate @cpmpml and @bluetrain who brought so many very interesting ideas that I hope they will share with you soon ! Great team work !</p>",
  "messages": [
    {
      "id": 542751,
      "postDate": "2019-06-04T04:31:20.050Z",
      "content": "<p>Thanks Kaggle for this interesting competition ! </p>\n\n<p>Here are a few points from our solution, at least my part, (I'll let my teammate describes their interesting ideas):\n-1 EQ out validation scheme\n-Data augmentation with 30k chunk size\n-Bases Models are simple lgb/xgb model with fixed number of trees, huber or gamma objective and sample weights\n-Calculated around 200 features based mainly on STFT (see image attached), there was a correlation between the energy on some frequency band, and the ttf. I created those features based on hann window with size 1000 and 5000 for different frequency bands (44khz -  60khz, 60khz - 136khz,  140khz - 212khz, 216khz - 356 khz)\n-Used different indicators on on those bands : mean / std / quantiles + some some simple features on raw signal like Numbers of peaks above threshold or quantiles.\n-One way to improve our CV score was to tweak the original to lower the ttf or very long EQ. I basically did a parallel shift of ttf between the start of long EQ to the mini-EQ.</p>\n\n<p>This solution was in top 15 on public leaderboard.\nThen the difficult part began when we realized that Private Distribution might be very different from Public LB and also from Train Data.\nSo we developed different measures based on different possible test sets. We realized that the distribution of EQ type (short or long) and distribution of ttf had a significant impact on our CV score. We optimized 2 of those measures, with the hypothesis that distribution could be the one in the published paper. Public LB of those 2 subs were : 1.39 and 1.29. The best on private LB was the first one ;) I'm glad we were able to survive the shake-up, this was really not an easy task. </p>\n\n<p>Congrats to all of you who shared great content in the Kernels and in the Forum. Congratulation to the winning team ! </p>\n\n<p>And a special thanks to my teammate @cpmpml and @bluetrain who brought so many very interesting ideas that I hope they will share with you soon ! Great team work !</p>",
      "rawMarkdown": "Thanks Kaggle for this interesting competition ! \n\nHere are a few points from our solution, at least my part, (I'll let my teammate describes their interesting ideas):\n-1 EQ out validation scheme\n-Data augmentation with 30k chunk size\n-Bases Models are simple lgb/xgb model with fixed number of trees, huber or gamma objective and sample weights\n-Calculated around 200 features based mainly on STFT (see image attached), there was a correlation between the energy on some frequency band, and the ttf. I created those features based on hann window with size 1000 and 5000 for different frequency bands (44khz -  60khz, 60khz - 136khz,  140khz - 212khz, 216khz - 356 khz)\n-Used different indicators on on those bands : mean / std / quantiles + some some simple features on raw signal like Numbers of peaks above threshold or quantiles.\n-One way to improve our CV score was to tweak the original to lower the ttf or very long EQ. I basically did a parallel shift of ttf between the start of long EQ to the mini-EQ.\n\nThis solution was in top 15 on public leaderboard.\nThen the difficult part began when we realized that Private Distribution might be very different from Public LB and also from Train Data.\nSo we developed different measures based on different possible test sets. We realized that the distribution of EQ type (short or long) and distribution of ttf had a significant impact on our CV score. We optimized 2 of those measures, with the hypothesis that distribution could be the one in the published paper. Public LB of those 2 subs were : 1.39 and 1.29. The best on private LB was the first one ;) I'm glad we were able to survive the shake-up, this was really not an easy task. \n\nCongrats to all of you who shared great content in the Kernels and in the Forum. Congratulation to the winning team ! \n\nAnd a special thanks to my teammate @cpmpml and @bluetrain who brought so many very interesting ideas that I hope they will share with you soon ! Great team work !",
      "votes": 26
    },
    {
      "id": 543150,
      "postDate": "2019-06-04T10:56:55.133Z",
      "content": "<p>Congrats on the gold. Thanks for sharing. </p>",
      "rawMarkdown": "Congrats on the gold. Thanks for sharing. ",
      "votes": 1
    },
    {
      "id": 543118,
      "postDate": "2019-06-04T10:39:14.663Z",
      "content": "<p>It was a pleasure to team with you!</p>",
      "rawMarkdown": "It was a pleasure to team with you!",
      "votes": 1
    },
    {
      "id": 542789,
      "postDate": "2019-06-04T05:02:19.877Z",
      "content": "<p>Congrats <a href=\"/areveillon\">@areveillon</a>  and thanks for sharing!</p>",
      "rawMarkdown": "Congrats @areveillon  and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 542761,
      "postDate": "2019-06-04T04:39:13.383Z",
      "content": "<p>Thanks. It's good to know your solution also exploits the private test leak, which is similar to my solution. One out of your two submissions is a gamble, which finally gave you the good private score. That's contradicting what <a href=\"/cpmpml\">@cpmpml</a>  has said in some topic: \"trust your CV, don't gamble\". To me, \"gamble\" here means using any kind of information that is not available in the pure training set, and local CV is only dealing with the training set only.</p>",
      "rawMarkdown": "Thanks. It's good to know your solution also exploits the private test leak, which is similar to my solution. One out of your two submissions is a gamble, which finally gave you the good private score. That's contradicting what @cpmpml  has said in some topic: \"trust your CV, don't gamble\". To me, \"gamble\" here means using any kind of information that is not available in the pure training set, and local CV is only dealing with the training set only.",
      "votes": 1,
      "replies": [
        {
          "id": 542791,
          "postDate": "2019-06-04T05:05:20.587Z",
          "content": "<p>Trust your CV is the key.  I trusted my CV and it worked. The key is build a validation strategy that matches the differences between Train and Private Test distribution. Knowing apriori the testset mean  is the same as knowing apriori your validation fold mean . Think about that ;)</p>",
          "rawMarkdown": "Trust your CV is the key.  I trusted my CV and it worked. The key is build a validation strategy that matches the differences between Train and Private Test distribution. Knowing apriori the testset mean  is the same as knowing apriori your validation fold mean . Think about that ;)",
          "votes": 2
        },
        {
          "id": 542800,
          "postDate": "2019-06-04T05:10:04.073Z",
          "content": "<p>Yes I understand what you mean. But if some newbies ask and get the answer \"trust your CV\", that's totally useless for them, or even harmful. That's too vague. What if a guy did not trust CV and exploit the leak? He will also end up well. So \"trust your CV\" is really a useless phrase. \"Doing CV without leak\" and \"doing CV with leak\" are 2 completely different stories.</p>",
          "rawMarkdown": "Yes I understand what you mean. But if some newbies ask and get the answer \"trust your CV\", that's totally useless for them, or even harmful. That's too vague. What if a guy did not trust CV and exploit the leak? He will also end up well. So \"trust your CV\" is really a useless phrase. \"Doing CV without leak\" and \"doing CV with leak\" are 2 completely different stories.",
          "votes": 4
        },
        {
          "id": 543073,
          "postDate": "2019-06-04T10:11:55.133Z",
          "content": "<blockquote>\n  <p>That's contradicting what <a href=\"/cpmpml\">@cpmpml</a> has said in some topic: \"trust your CV, don't gamble\". T</p>\n</blockquote>\n\n<p>It does not contradict, we selected our best CV submissions, see my write up.</p>\n\n<p>Not sure why you want to prove I was lying in the forum.  Seems I always have one person trying to do this in every competition now.</p>",
          "rawMarkdown": "&gt; That's contradicting what @cpmpml has said in some topic: \"trust your CV, don't gamble\". T\n\nIt does not contradict, we selected our best CV submissions, see my write up.\n\nNot sure why you want to prove I was lying in the forum.  Seems I always have one person trying to do this in every competition now."
        },
        {
          "id": 543119,
          "postDate": "2019-06-04T10:39:28.080Z",
          "content": "<p>I did not say you lie. All I mean is that if we cannot reveal our method, it’s better to keep silent than stating the so called phrase “Trust your CV”, which is totally useless for the people who asked. I already said it: “doing CV” in my opinion does not take into account any information outside of the training set. Furthermore, if I “don’t gamble”, I would end up rank 300+, and I bet your team would to. </p>\n\n<p>Please don’t be sensitive. I never accused you of anything. You even helped me to find better features and ended up with a small good feature set. I may be too sensitive on what people share, because I was a lecturer, so anything useless for newbies can easily turn out harmful for them. </p>",
          "rawMarkdown": "I did not say you lie. All I mean is that if we cannot reveal our method, it’s better to keep silent than stating the so called phrase “Trust your CV”, which is totally useless for the people who asked. I already said it: “doing CV” in my opinion does not take into account any information outside of the training set. Furthermore, if I “don’t gamble”, I would end up rank 300+, and I bet your team would to. \n\n Please don’t be sensitive. I never accused you of anything. You even helped me to find better features and ended up with a small good feature set. I may be too sensitive on what people share, because I was a lecturer, so anything useless for newbies can easily turn out harmful for them. "
        },
        {
          "id": 543127,
          "postDate": "2019-06-04T10:44:05.970Z",
          "content": "<p>You wrote this:</p>\n\n<blockquote>\n  <p>One out of your two submissions is a gamble, which finally gave you the good private score.  That's contradicting what <a href=\"/cpmpml\">@cpmpml</a> has said in some topic: \"trust your CV, don't gamble\". </p>\n</blockquote>\n\n<p>It is not  a gamble and our behavior did not contradict our writing:  we selected the best CV score models.  We used two different ways to compute CV score (different sample weights, see my writeup or reread Antoine's).  Please stop trolling about what I wrote.  </p>",
          "rawMarkdown": "You wrote this:\n\n&gt; One out of your two submissions is a gamble, which finally gave you the good private score.  That's contradicting what @cpmpml has said in some topic: \"trust your CV, don't gamble\". \n\nIt is not  a gamble and our behavior did not contradict our writing:  we selected the best CV score models.  We used two different ways to compute CV score (different sample weights, see my writeup or reread Antoine's).  Please stop trolling about what I wrote.  "
        },
        {
          "id": 543139,
          "postDate": "2019-06-04T10:50:09.177Z",
          "content": "<p>My and your opinions on “gamble”  and CV just are different. My sincere apology for any inconvenience to you caused by my comment. </p>",
          "rawMarkdown": "My and your opinions on “gamble”  and CV just are different. My sincere apology for any inconvenience to you caused by my comment. "
        },
        {
          "id": 543154,
          "postDate": "2019-06-04T11:01:05.437Z",
          "content": "<p>Apologies are not needed.  Simply admit that we did select best CV score for our two final submissions.  Not sure why you find it hard to believe.</p>",
          "rawMarkdown": "Apologies are not needed.  Simply admit that we did select best CV score for our two final submissions.  Not sure why you find it hard to believe."
        },
        {
          "id": 543160,
          "postDate": "2019-06-04T11:04:03.803Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 543164,
          "postDate": "2019-06-04T11:05:29.303Z",
          "content": "<p>Indeed, we use test knowledge.  Like all top teams I guess.  And both our final subs use that knowledge.</p>",
          "rawMarkdown": "Indeed, we use test knowledge.  Like all top teams I guess.  And both our final subs use that knowledge."
        },
        {
          "id": 543169,
          "postDate": "2019-06-04T11:07:30.203Z",
          "content": "<p>I think my apology is the good way to end this discussion here. We all did that, and I might have said something inappropriate. </p>",
          "rawMarkdown": "I think my apology is the good way to end this discussion here. We all did that, and I might have said something inappropriate. "
        },
        {
          "id": 543172,
          "postDate": "2019-06-04T11:11:45.997Z",
          "content": "<p>No need for apology, I reacted to the fact that you say I contradicted y own advice to others.  I really try to never tell people something I know is false and misleading. I was trolled about that in a recent competition which explains why I react so firmly.  Thanks for apologizing, but this was not necessary, I am not hurt ;)</p>\n\n<p>To your other points, telling people to trust their CV is vague.  What matters is to set a good CV indeed.  I always explain my CV setting after competition end.  People can read my previous writeups if they want to learn a bit about how to create good CV settings.  And many top kagglers also explain what they do after every competition.   </p>",
          "rawMarkdown": "No need for apology, I reacted to the fact that you say I contradicted y own advice to others.  I really try to never tell people something I know is false and misleading. I was trolled about that in a recent competition which explains why I react so firmly.  Thanks for apologizing, but this was not necessary, I am not hurt ;)\n\nTo your other points, telling people to trust their CV is vague.  What matters is to set a good CV indeed.  I always explain my CV setting after competition end.  People can read my previous writeups if they want to learn a bit about how to create good CV settings.  And many top kagglers also explain what they do after every competition.   \n\n"
        },
        {
          "id": 543205,
          "postDate": "2019-06-04T11:28:16.440Z",
          "content": "<p>You are a popular figure on Kaggle. A lot of new players joined and follow your advice. If you share vague things it would massively affect a lot of people. To me, any term with 2 letters “CV” means that we “cross validate” the train set itself by any way. But with outside test leak, it should not be called CV anymore. People can use test leak to do a 1-time fit and got to top 10. So with the leak, one even does not need any CV method. You did exactly what you said, but just not clear. Can you imagine how many people out there spend time trying for better features and models to improve their CV, which is robust I assume, trust it, and get the disappointment at the end? The test leak is the key to success, not a robust CV or even any kind of feature. </p>",
          "rawMarkdown": "You are a popular figure on Kaggle. A lot of new players joined and follow your advice. If you share vague things it would massively affect a lot of people. To me, any term with 2 letters “CV” means that we “cross validate” the train set itself by any way. But with outside test leak, it should not be called CV anymore. People can use test leak to do a 1-time fit and got to top 10. So with the leak, one even does not need any CV method. You did exactly what you said, but just not clear. Can you imagine how many people out there spend time trying for better features and models to improve their CV, which is robust I assume, trust it, and get the disappointment at the end? The test leak is the key to success, not a robust CV or even any kind of feature. "
        },
        {
          "id": 543234,
          "postDate": "2019-06-04T12:01:09.243Z",
          "content": "<blockquote>\n  <p>You are a popular figure on Kaggle. A lot of new players joined and follow your advice. If you share vague things it would massively affect a lot of people. </p>\n</blockquote>\n\n<p>Yes, I'm popular, and it is precisely because of my sharing and the way I share.  I force no one to read or follow what I write, and I won't stop sharing vague things like I did so far.  If you or anyone disagree then just don't read my posts.  </p>\n\n<blockquote>\n  <p>People can use test leak to do a 1-time fit and got to top 10. </p>\n</blockquote>\n\n<p>First of all, who did that before competition end?  Doing it after competition end is not relevant.</p>\n\n<p>Second, and most important, CV is not a way to train models.  You can get best ever model without CV, by training on all data and being lucky with your choice of model, features, and parameters.  CV is a validation procedure which let you evaluate models, features, and parameters instead of relying on luck.</p>\n\n<p>Third, I can assure you that I only used test inspired weights in the last week or so, and that models that were trained before were among my best.  I have two models worth a gold and my team mates got some as well that were trained and validated without any use of test data.  </p>\n\n<p>Fourth, using trest estimate while stacking got us further up for sure.  I would not have used stacking without proper CV assessment.</p>",
          "rawMarkdown": "&gt; You are a popular figure on Kaggle. A lot of new players joined and follow your advice. If you share vague things it would massively affect a lot of people. \n\nYes, I'm popular, and it is precisely because of my sharing and the way I share.  I force no one to read or follow what I write, and I won't stop sharing vague things like I did so far.  If you or anyone disagree then just don't read my posts.  \n\n&gt; People can use test leak to do a 1-time fit and got to top 10. \n\nFirst of all, who did that before competition end?  Doing it after competition end is not relevant.\n\nSecond, and most important, CV is not a way to train models.  You can get best ever model without CV, by training on all data and being lucky with your choice of model, features, and parameters.  CV is a validation procedure which let you evaluate models, features, and parameters instead of relying on luck.\n\nThird, I can assure you that I only used test inspired weights in the last week or so, and that models that were trained before were among my best.  I have two models worth a gold and my team mates got some as well that were trained and validated without any use of test data.  \n\nFourth, using trest estimate while stacking got us further up for sure.  I would not have used stacking without proper CV assessment.\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 543177,
      "postDate": "2019-06-04T11:14:50.503Z",
      "content": "<p>Bravo <a href=\"/areveillon\">@areveillon</a> !\nIt's great to see you orange finally ;) Well deserved!</p>",
      "rawMarkdown": "Bravo @areveillon !\nIt's great to see you orange finally ;) Well deserved!",
      "votes": 2,
      "replies": [
        {
          "id": 543202,
          "postDate": "2019-06-04T11:24:36.577Z",
          "content": "<p>Rightfully so, well deserved Antoine!</p>",
          "rawMarkdown": "Rightfully so, well deserved Antoine!",
          "votes": 1
        }
      ]
    },
    {
      "id": 542849,
      "postDate": "2019-06-04T06:31:21.720Z",
      "content": "<p>I learned about 'Trust your CV' from the last competition, 'Microsoft Malware Prediction'. I learned how to build a good cv system at this competition. Thanks for sharing !</p>",
      "rawMarkdown": "I learned about 'Trust your CV' from the last competition, 'Microsoft Malware Prediction'. I learned how to build a good cv system at this competition. Thanks for sharing !",
      "votes": 2
    },
    {
      "id": 543131,
      "postDate": "2019-06-04T10:46:49.903Z",
      "content": "<p>For 1.29, was it generated using Leave 1 Earthquake Out? Before merging on my team, we were using 1 EQ out, then we went to Leave2EQOut, using <code>LeavePGroupsOut</code> function, but decided to switch because we thought each model had a good chance to overfit, since 1 or 2 EQ's didn't leave a lot of segments to predict on. Basically, what did you do to change the 1.39 to 1.29 submission? Thank you, and congratulations on gold medal</p>",
      "rawMarkdown": "For 1.29, was it generated using Leave 1 Earthquake Out? Before merging on my team, we were using 1 EQ out, then we went to Leave2EQOut, using `LeavePGroupsOut` function, but decided to switch because we thought each model had a good chance to overfit, since 1 or 2 EQ's didn't leave a lot of segments to predict on. Basically, what did you do to change the 1.39 to 1.29 submission? Thank you, and congratulations on gold medal",
      "replies": [
        {
          "id": 543137,
          "postDate": "2019-06-04T10:49:02.180Z",
          "content": "<p>We used same CV setting from 1.29 and 1.39 stacks, but we used different sample weights.  In both cases we selected the best Cv score stack.</p>",
          "rawMarkdown": "We used same CV setting from 1.29 and 1.39 stacks, but we used different sample weights.  In both cases we selected the best Cv score stack.",
          "votes": 1
        }
      ]
    },
    {
      "id": 543892,
      "postDate": "2019-06-04T23:13:54.943Z",
      "content": "<p>Congrats, thank you for sharing!</p>",
      "rawMarkdown": "Congrats, thank you for sharing!",
      "votes": 1
    },
    {
      "id": 543047,
      "postDate": "2019-06-04T09:49:01.577Z",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "rawMarkdown": "Congrats! Thanks for sharing :)",
      "votes": 1
    },
    {
      "id": 542974,
      "postDate": "2019-06-04T08:46:10.973Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 543150,
      "author_name": "Karan Jakhar",
      "author_url": "",
      "post_date": "2019-06-04T10:56:55.133000",
      "content": "<p>Congrats on the gold. Thanks for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543118,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-06-04T10:39:14.663000",
      "content": "<p>It was a pleasure to team with you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542789,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-06-04T05:02:19.877000",
      "content": "<p>Congrats <a href=\"/areveillon\">@areveillon</a>  and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542761,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-06-04T04:39:13.383000",
      "content": "<p>Thanks. It's good to know your solution also exploits the private test leak, which is similar to my solution. One out of your two submissions is a gamble, which finally gave you the good private score. That's contradicting what <a href=\"/cpmpml\">@cpmpml</a>  has said in some topic: \"trust your CV, don't gamble\". To me, \"gamble\" here means using any kind of information that is not available in the pure training set, and local CV is only dealing with the training set only.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 542791,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-06-04T05:05:20.587000",
          "content": "<p>Trust your CV is the key.  I trusted my CV and it worked. The key is build a validation strategy that matches the differences between Train and Private Test distribution. Knowing apriori the testset mean  is the same as knowing apriori your validation fold mean . Think about that ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 542800,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-06-04T05:10:04.073000",
          "content": "<p>Yes I understand what you mean. But if some newbies ask and get the answer \"trust your CV\", that's totally useless for them, or even harmful. That's too vague. What if a guy did not trust CV and exploit the leak? He will also end up well. So \"trust your CV\" is really a useless phrase. \"Doing CV without leak\" and \"doing CV with leak\" are 2 completely different stories.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 543073,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T10:11:55.133000",
          "content": "<blockquote>\n  <p>That's contradicting what <a href=\"/cpmpml\">@cpmpml</a> has said in some topic: \"trust your CV, don't gamble\". T</p>\n</blockquote>\n\n<p>It does not contradict, we selected our best CV submissions, see my write up.</p>\n\n<p>Not sure why you want to prove I was lying in the forum.  Seems I always have one person trying to do this in every competition now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543119,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-06-04T10:39:28.080000",
          "content": "<p>I did not say you lie. All I mean is that if we cannot reveal our method, it’s better to keep silent than stating the so called phrase “Trust your CV”, which is totally useless for the people who asked. I already said it: “doing CV” in my opinion does not take into account any information outside of the training set. Furthermore, if I “don’t gamble”, I would end up rank 300+, and I bet your team would to. </p>\n\n<p>Please don’t be sensitive. I never accused you of anything. You even helped me to find better features and ended up with a small good feature set. I may be too sensitive on what people share, because I was a lecturer, so anything useless for newbies can easily turn out harmful for them. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543127,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T10:44:05.970000",
          "content": "<p>You wrote this:</p>\n\n<blockquote>\n  <p>One out of your two submissions is a gamble, which finally gave you the good private score.  That's contradicting what <a href=\"/cpmpml\">@cpmpml</a> has said in some topic: \"trust your CV, don't gamble\". </p>\n</blockquote>\n\n<p>It is not  a gamble and our behavior did not contradict our writing:  we selected the best CV score models.  We used two different ways to compute CV score (different sample weights, see my writeup or reread Antoine's).  Please stop trolling about what I wrote.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543139,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-06-04T10:50:09.177000",
          "content": "<p>My and your opinions on “gamble”  and CV just are different. My sincere apology for any inconvenience to you caused by my comment. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543154,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T11:01:05.437000",
          "content": "<p>Apologies are not needed.  Simply admit that we did select best CV score for our two final submissions.  Not sure why you find it hard to believe.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543160,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T11:04:03.803000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543164,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T11:05:29.303000",
          "content": "<p>Indeed, we use test knowledge.  Like all top teams I guess.  And both our final subs use that knowledge.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543169,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-06-04T11:07:30.203000",
          "content": "<p>I think my apology is the good way to end this discussion here. We all did that, and I might have said something inappropriate. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543172,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T11:11:45.997000",
          "content": "<p>No need for apology, I reacted to the fact that you say I contradicted y own advice to others.  I really try to never tell people something I know is false and misleading. I was trolled about that in a recent competition which explains why I react so firmly.  Thanks for apologizing, but this was not necessary, I am not hurt ;)</p>\n\n<p>To your other points, telling people to trust their CV is vague.  What matters is to set a good CV indeed.  I always explain my CV setting after competition end.  People can read my previous writeups if they want to learn a bit about how to create good CV settings.  And many top kagglers also explain what they do after every competition.   </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543205,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-06-04T11:28:16.440000",
          "content": "<p>You are a popular figure on Kaggle. A lot of new players joined and follow your advice. If you share vague things it would massively affect a lot of people. To me, any term with 2 letters “CV” means that we “cross validate” the train set itself by any way. But with outside test leak, it should not be called CV anymore. People can use test leak to do a 1-time fit and got to top 10. So with the leak, one even does not need any CV method. You did exactly what you said, but just not clear. Can you imagine how many people out there spend time trying for better features and models to improve their CV, which is robust I assume, trust it, and get the disappointment at the end? The test leak is the key to success, not a robust CV or even any kind of feature. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543234,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T12:01:09.243000",
          "content": "<blockquote>\n  <p>You are a popular figure on Kaggle. A lot of new players joined and follow your advice. If you share vague things it would massively affect a lot of people. </p>\n</blockquote>\n\n<p>Yes, I'm popular, and it is precisely because of my sharing and the way I share.  I force no one to read or follow what I write, and I won't stop sharing vague things like I did so far.  If you or anyone disagree then just don't read my posts.  </p>\n\n<blockquote>\n  <p>People can use test leak to do a 1-time fit and got to top 10. </p>\n</blockquote>\n\n<p>First of all, who did that before competition end?  Doing it after competition end is not relevant.</p>\n\n<p>Second, and most important, CV is not a way to train models.  You can get best ever model without CV, by training on all data and being lucky with your choice of model, features, and parameters.  CV is a validation procedure which let you evaluate models, features, and parameters instead of relying on luck.</p>\n\n<p>Third, I can assure you that I only used test inspired weights in the last week or so, and that models that were trained before were among my best.  I have two models worth a gold and my team mates got some as well that were trained and validated without any use of test data.  </p>\n\n<p>Fourth, using trest estimate while stacking got us further up for sure.  I would not have used stacking without proper CV assessment.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543177,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2019-06-04T11:14:50.503000",
      "content": "<p>Bravo <a href=\"/areveillon\">@areveillon</a> !\nIt's great to see you orange finally ;) Well deserved!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 543202,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T11:24:36.577000",
          "content": "<p>Rightfully so, well deserved Antoine!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 542849,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2019-06-04T06:31:21.720000",
      "content": "<p>I learned about 'Trust your CV' from the last competition, 'Microsoft Malware Prediction'. I learned how to build a good cv system at this competition. Thanks for sharing !</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 543131,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2019-06-04T10:46:49.903000",
      "content": "<p>For 1.29, was it generated using Leave 1 Earthquake Out? Before merging on my team, we were using 1 EQ out, then we went to Leave2EQOut, using <code>LeavePGroupsOut</code> function, but decided to switch because we thought each model had a good chance to overfit, since 1 or 2 EQ's didn't leave a lot of segments to predict on. Basically, what did you do to change the 1.39 to 1.29 submission? Thank you, and congratulations on gold medal</p>",
      "votes": 0,
      "replies": [
        {
          "id": 543137,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T10:49:02.180000",
          "content": "<p>We used same CV setting from 1.29 and 1.39 stacks, but we used different sample weights.  In both cases we selected the best Cv score stack.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543892,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-06-04T23:13:54.943000",
      "content": "<p>Congrats, thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543047,
      "author_name": "Prashanth Thangavel",
      "author_url": "",
      "post_date": "2019-06-04T09:49:01.577000",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542974,
      "author_name": "Stanislav Blinov",
      "author_url": "",
      "post_date": "2019-06-04T08:46:10.973000",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542751": "Thanks Kaggle for this interesting competition ! \n\nHere are a few points from our solution, at least my part, (I'll let my teammate describes their interesting ideas):\n-1 EQ out validation scheme\n-Data augmentation with 30k chunk size\n-Bases Models are simple lgb/xgb model with fixed number of trees, huber or gamma objective and sample weights\n-Calculated around 200 features based mainly on STFT (see image attached), there was a correlation between the energy on some frequency band, and the ttf. I created those features based on hann window with size 1000 and 5000 for different frequency bands (44khz -  60khz, 60khz - 136khz,  140khz - 212khz, 216khz - 356 khz)\n-Used different indicators on on those bands : mean / std / quantiles + some some simple features on raw signal like Numbers of peaks above threshold or quantiles.\n-One way to improve our CV score was to tweak the original to lower the ttf or very long EQ. I basically did a parallel shift of ttf between the start of long EQ to the mini-EQ.\n\nThis solution was in top 15 on public leaderboard.\nThen the difficult part began when we realized that Private Distribution might be very different from Public LB and also from Train Data.\nSo we developed different measures based on different possible test sets. We realized that the distribution of EQ type (short or long) and distribution of ttf had a significant impact on our CV score. We optimized 2 of those measures, with the hypothesis that distribution could be the one in the published paper. Public LB of those 2 subs were : 1.39 and 1.29. The best on private LB was the first one ;) I'm glad we were able to survive the shake-up, this was really not an easy task. \n\nCongrats to all of you who shared great content in the Kernels and in the Forum. Congratulation to the winning team ! \n\nAnd a special thanks to my teammate @cpmpml and @bluetrain who brought so many very interesting ideas that I hope they will share with you soon ! Great team work !",
    "543150": "Congrats on the gold. Thanks for sharing. ",
    "543118": "It was a pleasure to team with you!",
    "542789": "Congrats @areveillon  and thanks for sharing!",
    "542761": "Thanks. It's good to know your solution also exploits the private test leak, which is similar to my solution. One out of your two submissions is a gamble, which finally gave you the good private score. That's contradicting what @cpmpml  has said in some topic: \"trust your CV, don't gamble\". To me, \"gamble\" here means using any kind of information that is not available in the pure training set, and local CV is only dealing with the training set only.",
    "543177": "Bravo @areveillon !\nIt's great to see you orange finally ;) Well deserved!",
    "542849": "I learned about 'Trust your CV' from the last competition, 'Microsoft Malware Prediction'. I learned how to build a good cv system at this competition. Thanks for sharing !",
    "543131": "For 1.29, was it generated using Leave 1 Earthquake Out? Before merging on my team, we were using 1 EQ out, then we went to Leave2EQOut, using `LeavePGroupsOut` function, but decided to switch because we thought each model had a good chance to overfit, since 1 or 2 EQ's didn't leave a lot of segments to predict on. Basically, what did you do to change the 1.39 to 1.29 submission? Thank you, and congratulations on gold medal",
    "543892": "Congrats, thank you for sharing!",
    "543047": "Congrats! Thanks for sharing :)",
    "542974": "Congrats and thanks for sharing!"
  }
}