{
  "id": 69737,
  "title": "What is the secret?",
  "url": "/competitions/PLAsTiCC-2018/discussion/69737",
  "author_name": "",
  "post_date": "2018-10-26T16:10:33.216130900Z",
  "votes": 14,
  "comment_count": 54,
  "views": 0,
  "content": "<p>There is a significant gap between 4th and 5th on the LB.  Leaders have found something...</p>",
  "messages": [
    {
      "id": "410760",
      "postDate": "10/26/2018 16:10:33",
      "content": "<p>There is a significant gap between 4th and 5th on the LB.  Leaders have found something...</p>",
      "rawMarkdown": "There is a significant gap between 4th and 5th on the LB.  Leaders have found something...",
      "votes": null
    },
    {
      "id": "410787",
      "postDate": "10/26/2018 16:55:31",
      "content": "<p>Maybe they found a way to deal with minority classes. I can't see there is any way that gradient boosting machine can handle class 53 &amp; 62</p>",
      "rawMarkdown": "Maybe they found a way to deal with minority classes. I can't see there is any way that gradient boosting machine can handle class 53 &amp; 62",
      "votes": null
    },
    {
      "id": "410789",
      "postDate": "10/26/2018 16:59:45",
      "content": "<p>My guess is that they had some success with a feature extraction library such as <a href=\"http://cesium-ml.org/\">cesium</a>.</p>",
      "rawMarkdown": "My guess is that they had some success with a feature extraction library such as [cesium](http://cesium-ml.org/).",
      "votes": null
    },
    {
      "id": "410803",
      "postDate": "10/26/2018 17:33:23",
      "content": "<p>I think they've found the way to handle class 99 problem</p>",
      "rawMarkdown": "I think they've found the way to handle class 99 problem",
      "votes": null
    },
    {
      "id": "410808",
      "postDate": "10/26/2018 17:41:03",
      "content": "<p>I don't think the gap is significant. It is even small. I am sure that if I stop today, one week later I will be between 10th and 20th which is even optimistic. And personally, I don't have anything like <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268#325119\">this</a> yet.</p>",
      "rawMarkdown": "I don't think the gap is significant. It is even small. I am sure that if I stop today, one week later I will be between 10th and 20th which is even optimistic. And personally, I don't have anything like [this][1] yet.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268#325119",
      "votes": null
    },
    {
      "id": "410812",
      "postDate": "10/26/2018 17:45:27",
      "content": "<p>It was a real nightmare (for me)... I hope it won't happen again!</p>",
      "rawMarkdown": "It was a real nightmare (for me)... I hope it won't happen again!",
      "votes": null
    },
    {
      "id": "410819",
      "postDate": "10/26/2018 17:57:06",
      "content": "<p>You are probably right, but I don't know what. </p>\n\n<p>Until today I wanted to change my group's name to: \"I really don't know what I'm doing\"</p>\n\n<p>I don't do anything fancy, just simple NN, with some common practice training and handling.</p>\n\n<p>I also have a few tiny tricks - but nothing big (every fancy thing I tried didn't seem to work)</p>\n\n<p>And most important, I didn't waste my time  on class 99 (for the time being) </p>\n\n<p>(of course I will have to do something with it sometime because this is the difference between my current score - 0.958, and my estimate of the best achievable score &lt;0.7) </p>",
      "rawMarkdown": "You are probably right, but I don't know what. \n\nUntil today I wanted to change my group's name to: \"I really don't know what I'm doing\"\n\nI don't do anything fancy, just simple NN, with some common practice training and handling.\n\nI also have a few tiny tricks - but nothing big (every fancy thing I tried didn't seem to work)\n\nAnd most important, I didn't waste my time  on class 99 (for the time being) \n\n(of course I will have to do something with it sometime because this is the difference between my current score - 0.958, and my estimate of the best achievable score &lt;0.7)",
      "votes": null
    },
    {
      "id": "410822",
      "postDate": "10/26/2018 17:58:45",
      "content": "<blockquote>\n  <p>It was a real nightmare (for me)... I hope it won't happen again!</p>\n</blockquote>\n\n<p>Ah, i knew your name was familiar!</p>",
      "rawMarkdown": "&gt; It was a real nightmare (for me)... I hope it won't happen again!\n\nAh, i knew your name was familiar!",
      "votes": null
    },
    {
      "id": "410825",
      "postDate": "10/26/2018 18:01:13",
      "content": "<p>I agree. The gap is insignificant, it can be breached with small changes.\nThe breakthrough will be when someone will score less then 0.8 </p>",
      "rawMarkdown": "I agree. The gap is insignificant, it can be breached with small changes.\nThe breakthrough will be when someone will score less then 0.8",
      "votes": null
    },
    {
      "id": "410831",
      "postDate": "10/26/2018 18:09:02",
      "content": "<p>I don't want to talk about my method, but I can only say I'm 100% sure my method and your method are completely different, the score is really close, though.</p>",
      "rawMarkdown": "I don't want to talk about my method, but I can only say I'm 100% sure my method and your method are completely different, the score is really close, though.",
      "votes": null
    },
    {
      "id": "410862",
      "postDate": "10/26/2018 19:18:05",
      "content": "<p>Hey Yuval,</p>\n\n<p>By NN do you mean a RNN or simply NN?</p>",
      "rawMarkdown": "Hey Yuval,\n\nBy NN do you mean a RNN or simply NN?",
      "votes": null
    },
    {
      "id": "410877",
      "postDate": "10/26/2018 19:55:40",
      "content": "<p>neither</p>",
      "rawMarkdown": "neither",
      "votes": null
    },
    {
      "id": "411154",
      "postDate": "10/27/2018 14:17:53",
      "content": "<blockquote>\n  <p>my estimate of the best achievable score &lt;0.7</p>\n</blockquote>\n\n<p>How do you estimate that?</p>",
      "rawMarkdown": "&gt;  my estimate of the best achievable score &lt;0.7\n\nHow do you estimate that?",
      "votes": null
    },
    {
      "id": "411213",
      "postDate": "10/27/2018 16:21:44",
      "content": "<p>I believe a large portion of the difference between local scores and LB scores is related to class 99. I can achieve about 0.6 local score. If anyone can find a reliable way to predict class 99, he should get close to the local score (there will still be a difference because of the huge test set). If class 99 is a real type of event then it is possible to find features to classify it.    </p>",
      "rawMarkdown": "I believe a large portion of the difference between local scores and LB scores is related to class 99. I can achieve about 0.6 local score. If anyone can find a reliable way to predict class 99, he should get close to the local score (there will still be a difference because of the huge test set). If class 99 is a real type of event then it is possible to find features to classify it.",
      "votes": null
    },
    {
      "id": "411256",
      "postDate": "10/27/2018 17:56:38",
      "content": "<p>And more accurately:</p>\n\n<p>If we give class 99 a constant value - V then the score penalty is:  </p>\n\n<pre><code>8/9*log(1-v)+1/9*log(V)\n</code></pre>\n\n<p>Assuming the weight for class 99,64,15 is indeed 2</p>\n\n<p>The optimum is about V=0.1 where we get penalty of about 0.35</p>\n\n<p>This is consistent with the current difference between my LB and local scores (I use a slightly better selection but not much).</p>\n\n<p>If we can really discover class 99, we can probably decrease the penalty to around 0.15 and improve the model by 0.05, an LB score is achievable (especially when considering the early stage of this competition) </p>",
      "rawMarkdown": "And more accurately:\n\nIf we give class 99 a constant value - V then the score penalty is:  \n\n    8/9*log(1-v)+1/9*log(V)\n\nAssuming the weight for class 99,64,15 is indeed 2\n\nThe optimum is about V=0.1 where we get penalty of about 0.35\n\nThis is consistent with the current difference between my LB and local scores (I use a slightly better selection but not much).\n\nIf we can really discover class 99, we can probably decrease the penalty to around 0.15 and improve the model by 0.05, an LB score is achievable (especially when considering the early stage of this competition)",
      "votes": null
    },
    {
      "id": "411698",
      "postDate": "10/28/2018 20:13:46",
      "content": "<p>I approve of your proposed team name</p>",
      "rawMarkdown": "I approve of your proposed team name",
      "votes": null
    },
    {
      "id": "434658",
      "postDate": "12/06/2018 18:35:25",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a></p>\n\n<p>this question now comes back to you. There is a significant gap between 0.7x and 0.8x in LB... leader have found something... </p>",
      "rawMarkdown": "cpmpml\n\n this question now comes back to you. There is a significant gap between 0.7x and 0.8x in LB... leader have found something...",
      "votes": null
    },
    {
      "id": "434713",
      "postDate": "12/06/2018 20:44:54",
      "content": "<p>the answer will probably come in 12 days or so</p>",
      "rawMarkdown": "the answer will probably come in 12 days or so",
      "votes": null
    },
    {
      "id": "434762",
      "postDate": "12/06/2018 22:43:02",
      "content": "<p>We found a hole in space time texture of the universe...</p>\n\n<p>Just kidding, we'll share what we did at the end of the competition.  Some of what we did may surprise others ;)</p>",
      "rawMarkdown": "We found a hole in space time texture of the universe...\n\nJust kidding, we'll share what we did at the end of the competition.  Some of what we did may surprise others ;)",
      "votes": null
    },
    {
      "id": "434791",
      "postDate": "12/07/2018 00:31:54",
      "content": "<p>I wonder if we're dealing with a bunch of features with e.g. 0.02 improvement or if there are is some radically different way to look at things which gets the top people a big boost...</p>",
      "rawMarkdown": "I wonder if we're dealing with a bunch of features with e.g. 0.02 improvement or if there are is some radically different way to look at things which gets the top people a big boost...",
      "votes": null
    },
    {
      "id": "434828",
      "postDate": "12/07/2018 02:33:32",
      "content": "<p>There are holes in the space-time, physicists call it black holes... but if you've found it you would not return... nothing can escape the black hole :-)</p>",
      "rawMarkdown": "There are holes in the space-time, physicists call it black holes... but if you've found it you would not return... nothing can escape the black hole :-)",
      "votes": null
    },
    {
      "id": "435205",
      "postDate": "12/07/2018 17:07:59",
      "content": "<p>I don't think the gap is significant. It is even small. I am sure that if I stop today, one week later I will be between 10th and 20th which is even optimistic.</p>",
      "rawMarkdown": "I don't think the gap is significant. It is even small. I am sure that if I stop today, one week later I will be between 10th and 20th which is even optimistic.",
      "votes": null
    },
    {
      "id": "435246",
      "postDate": "12/07/2018 18:49:21",
      "content": "<p>@mamas I bet that a week from now 10th rank LB score will be higher than 0.755.</p>",
      "rawMarkdown": "mamas I bet that a week from now 10th rank LB score will be higher than 0.755.",
      "votes": null
    },
    {
      "id": "435759",
      "postDate": "12/08/2018 17:59:58",
      "content": "<blockquote>\n  <p>leader have found something… </p>\n</blockquote>\n\n<p>Others seem to find something too, and the gap is shrinking fast.</p>",
      "rawMarkdown": "&gt; leader have found something… \n\nOthers seem to find something too, and the gap is shrinking fast.",
      "votes": null
    },
    {
      "id": "435900",
      "postDate": "12/09/2018 03:24:53",
      "content": "<p>aaaaa, but not me... I am still missing something... where to look for it ? :-( </p>\n\n<p>PS This kaggle thing is addictive, apparently. They should put a warning sign on the website: KAGGLE IS ADDICTIVE. ENTER AT YOUR OWN RISK ! </p>",
      "rawMarkdown": "aaaaa, but not me... I am still missing something... where to look for it ? :-( \n\nPS This kaggle thing is addictive, apparently. They should put a warning sign on the website: KAGGLE IS ADDICTIVE. ENTER AT YOUR OWN RISK !",
      "votes": null
    },
    {
      "id": "436027",
      "postDate": "12/09/2018 11:04:59",
      "content": "<p><a href=\"/blondinka\">@blondinka</a> I can relate</p>",
      "rawMarkdown": "blondinka I can relate",
      "votes": null
    },
    {
      "id": "436042",
      "postDate": "12/09/2018 12:08:21",
      "content": "<p>Kaggle responsibly! ;)</p>",
      "rawMarkdown": "Kaggle responsibly! ;)",
      "votes": null
    },
    {
      "id": "437782",
      "postDate": "12/12/2018 13:40:20",
      "content": "<p>Seems now top 2 have found something others haven't.  The gap is large.</p>\n\n<p>I am writing this now because each time I did I closed the gap in the next few days ;)  Hope it can work again...</p>\n\n<p>Just kidding of course, but trying to catch up still.</p>",
      "rawMarkdown": "Seems now top 2 have found something others haven't.  The gap is large.\n\nI am writing this now because each time I did I closed the gap in the next few days ;)  Hope it can work again...\n\nJust kidding of course, but trying to catch up still.",
      "votes": null
    },
    {
      "id": "437787",
      "postDate": "12/12/2018 13:44:51",
      "content": "<p>where you able to calculate the hessian properly?</p>",
      "rawMarkdown": "where you able to calculate the hessian properly?",
      "votes": null
    },
    {
      "id": "437790",
      "postDate": "12/12/2018 13:47:23",
      "content": "<p>What for?  I don't see the need for it.</p>",
      "rawMarkdown": "What for?  I don't see the need for it.",
      "votes": null
    },
    {
      "id": "437793",
      "postDate": "12/12/2018 13:53:47",
      "content": "<p>I can confirm that having the Hessian doesn't help. I've managed to calculate it analytically using <code>sympy</code> and it helps when using it with the LGBM loss function used in the kernels. In other words it is better than simply outputting a vector of 1s. But it is worse than properly using the <code>sample_weight</code> parameter. </p>",
      "rawMarkdown": "I can confirm that having the Hessian doesn't help. I've managed to calculate it analytically using `sympy` and it helps when using it with the LGBM loss function used in the kernels. In other words it is better than simply outputting a vector of 1s. But it is worse than properly using the `sample_weight` parameter.",
      "votes": null
    },
    {
      "id": "437801",
      "postDate": "12/12/2018 14:17:02",
      "content": "<p>I think they have found the magic feature object_id%27</p>",
      "rawMarkdown": "I think they have found the magic feature object_id%27",
      "votes": null
    },
    {
      "id": "437803",
      "postDate": "12/12/2018 14:23:22",
      "content": "<p>Are you sure it's 27? 28 is working better for me.</p>",
      "rawMarkdown": "Are you sure it's 27? 28 is working better for me.",
      "votes": null
    },
    {
      "id": "437804",
      "postDate": "12/12/2018 14:26:13",
      "content": "<p>Good to know, I am trying values one by one and I am only at object_id%22</p>",
      "rawMarkdown": "Good to know, I am trying values one by one and I am only at object_id%22",
      "votes": null
    },
    {
      "id": "437876",
      "postDate": "12/12/2018 17:16:59",
      "content": "<blockquote>\n  <p>I think they have found the magic feature object_id%27</p>\n</blockquote>\n\n<p>Oh damn! so that was it. But, now I have only 6 days to go. </p>",
      "rawMarkdown": "&gt; I think they have found the magic feature object_id%27\n\nOh damn! so that was it. But, now I have only 6 days to go.",
      "votes": null
    },
    {
      "id": "437955",
      "postDate": "12/12/2018 20:48:08",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a></p>\n\n<p>I am wondering about it for the last week, we keep trying and cannot find... It is not iterative. Kyle is doing really well. If we assume you know ML better than Kyle, then it has to lie in the astronomy domain... that's what I think</p>\n\n<p>PS object_id%13 is my favorite :)</p>",
      "rawMarkdown": "cpmpml\n\nI am wondering about it for the last week, we keep trying and cannot find... It is not iterative. Kyle is doing really well. If we assume you know ML better than Kyle, then it has to lie in the astronomy domain... that's what I think\n\nPS object_id%13 is my favorite :)",
      "votes": null
    },
    {
      "id": "437977",
      "postDate": "12/12/2018 21:57:39",
      "content": "<p>we have managed to find the answer to the universe, it's class 42.</p>",
      "rawMarkdown": "we have managed to find the answer to the universe, it's class 42.",
      "votes": null
    },
    {
      "id": "438007",
      "postDate": "12/12/2018 23:38:06",
      "content": "<p>@cpmp and <a href=\"/maxhalford\">@maxhalford</a> thanks for confirmation. now i can stop trying to do it. haha</p>",
      "rawMarkdown": "cpmp and @maxhalford thanks for confirmation. now i can stop trying to do it. haha",
      "votes": null
    },
    {
      "id": "438065",
      "postDate": "12/13/2018 03:40:15",
      "content": "<p>@Blonde, you use object_id%13 as a feature?</p>",
      "rawMarkdown": "Blonde, you use object_id%13 as a feature?",
      "votes": null
    },
    {
      "id": "438625",
      "postDate": "12/14/2018 00:38:05",
      "content": "<p><a href=\"/niclasdoce\">@niclasdoce</a>  </p>\n\n<p>nop, I just stare at it... I managed to find it on the night sky when kaggling at night :-)</p>\n\n<p>PS use proper nick names, you see them in down-left corner when you navigate in the use name, my is <a href=\"/blondinka\">@blondinka</a></p>",
      "rawMarkdown": "niclasdoce  \n\nnop, I just stare at it... I managed to find it on the night sky when kaggling at night :-)\n\nPS use proper nick names, you see them in down-left corner when you navigate in the use name, my is @blondinka",
      "votes": null
    },
    {
      "id": "438629",
      "postDate": "12/14/2018 00:45:13",
      "content": "<p>i don't understand.</p>",
      "rawMarkdown": "i don't understand.",
      "votes": null
    },
    {
      "id": "438698",
      "postDate": "12/14/2018 03:29:26",
      "content": "<p>I think that there are quite a few non-obvious ways to increase your score. My first model used basically the same features/methods that I am using now, and it came in around 1.03. My CV really hasn't changed very much from that model. Most of the improvement has come from trying not to overfit the data to get the \"gap\" down. There are a lot of different tricks that I have found, and I think that I can guess who has figured out what from the leaderboard scores...</p>",
      "rawMarkdown": "I think that there are quite a few non-obvious ways to increase your score. My first model used basically the same features/methods that I am using now, and it came in around 1.03. My CV really hasn't changed very much from that model. Most of the improvement has come from trying not to overfit the data to get the \"gap\" down. There are a lot of different tricks that I have found, and I think that I can guess who has figured out what from the leaderboard scores...",
      "votes": null
    },
    {
      "id": "438705",
      "postDate": "12/14/2018 03:43:09",
      "content": "<p>i see. that's great to hear.</p>",
      "rawMarkdown": "i see. that's great to hear.",
      "votes": null
    },
    {
      "id": "438748",
      "postDate": "12/14/2018 05:57:17",
      "content": "<p>Thanks for sharing Kyle. In another discussion you have written that you have 0.4 CV and here you say your CV didn't change from LB score 1.03. it is really great that you obtained such CV score that early in the competition. But it means you had serious overfitting issues which causes 0.6 gap as far as I understand.</p>",
      "rawMarkdown": "Thanks for sharing Kyle. In another discussion you have written that you have 0.4 CV and here you say your CV didn't change from LB score 1.03. it is really great that you obtained such CV score that early in the competition. But it means you had serious overfitting issues which causes 0.6 gap as far as I understand.",
      "votes": null
    },
    {
      "id": "438761",
      "postDate": "12/14/2018 06:24:29",
      "content": "<p>Kyle,  you are sharing more and more, are you feeling alone at the top?</p>\n\n<p>I agree with you, lots of our progress as well was along the lines of what you just disclosed.   I guess we had to spend more time on feature engineering than you as we don't have your background.</p>",
      "rawMarkdown": "Kyle,  you are sharing more and more, are you feeling alone at the top?\n\nI agree with you, lots of our progress as well was along the lines of what you just disclosed.   I guess we had to spend more time on feature engineering than you as we don't have your background.",
      "votes": null
    },
    {
      "id": "438783",
      "postDate": "12/14/2018 07:19:23",
      "content": "<p>I had roughly a CV of 0.5 with a LB of 1.0 at the start, so my gap was large. I dropped the gap down to ~0.4 fairly quickly, and lowering the CV has seemed to also lower the gap with the methods that I'm using. I posted on the CV thread several times, so you can see my progress there if you want. This is also my first kaggle competition and I've been learning a lot of ML as I go. For everything that I've done with ML in the past, tuning the model to get an extra 5% out didn't really matter. Bad tuning was hurting me a lot at first.</p>\n\n<p>I have a weird relationship with helping people out for this competition. On one hand I want to win, but on the other I am likely going to be using LSST data in my future research. I want the final models to be the best that they can be!</p>",
      "rawMarkdown": "I had roughly a CV of 0.5 with a LB of 1.0 at the start, so my gap was large. I dropped the gap down to ~0.4 fairly quickly, and lowering the CV has seemed to also lower the gap with the methods that I'm using. I posted on the CV thread several times, so you can see my progress there if you want. This is also my first kaggle competition and I've been learning a lot of ML as I go. For everything that I've done with ML in the past, tuning the model to get an extra 5% out didn't really matter. Bad tuning was hurting me a lot at first.\n\nI have a weird relationship with helping people out for this competition. On one hand I want to win, but on the other I am likely going to be using LSST data in my future research. I want the final models to be the best that they can be!",
      "votes": null
    },
    {
      "id": "438798",
      "postDate": "12/14/2018 07:49:27",
      "content": "<blockquote>\n  <p>I want the final models to be the best that they can be!</p>\n</blockquote>\n\n<p>There should be a follow up competition next year ;)</p>",
      "rawMarkdown": "&gt; I want the final models to be the best that they can be!\n\nThere should be a follow up competition next year ;)",
      "votes": null
    },
    {
      "id": "438819",
      "postDate": "12/14/2018 08:39:02",
      "content": "<p>&gt; @mamas I bet that a week from now 10th rank LB score will be higher than 0.755.</p>\n\n<p>I was right, only 5 teams are below 0.755 now...  10th is 0.823 as I write.</p>\n\n<p>Crossing 0.8 is hard</p>",
      "rawMarkdown": "&gt; @mamas I bet that a week from now 10th rank LB score will be higher than 0.755.\n\nI was right, only 5 teams are below 0.755 now...  10th is 0.823 as I write.\n\nCrossing 0.8 is hard",
      "votes": null
    },
    {
      "id": "438846",
      "postDate": "12/14/2018 09:21:30",
      "content": "<p>hmm, seems they couldn't find 'secret'...</p>",
      "rawMarkdown": "hmm, seems they couldn't find 'secret'...",
      "votes": null
    },
    {
      "id": "438897",
      "postDate": "12/14/2018 10:57:34",
      "content": "<p>yea. can't. ive been stuck for a while haha. I'm excited to see what you guys did.</p>",
      "rawMarkdown": "yea. can't. ive been stuck for a while haha. I'm excited to see what you guys did.",
      "votes": null
    },
    {
      "id": "439430",
      "postDate": "12/15/2018 12:48:40",
      "content": "<p><a href=\"/kyleboone\">@kyleboone</a> </p>\n\n<p>Kyle, as you are again in a good mood for sharing, I have a problem with Doppler effect. It's clear how to take it into account for individual spectroscopic lines, but I still do not get it how to recalculate redshift coefficients for integrated spectrum we have here. I've seen they use z in TypeIa templates, but it's still not clear how they recalculate brightness in different bands taking z into account for such an intergrated spectrum... I've seen many papers and people do not write clearly how they recalculate different bands, is it something obvious for astronomers?  </p>",
      "rawMarkdown": "kyleboone \n\nKyle, as you are again in a good mood for sharing, I have a problem with Doppler effect. It's clear how to take it into account for individual spectroscopic lines, but I still do not get it how to recalculate redshift coefficients for integrated spectrum we have here. I've seen they use z in TypeIa templates, but it's still not clear how they recalculate brightness in different bands taking z into account for such an intergrated spectrum... I've seen many papers and people do not write clearly how they recalculate different bands, is it something obvious for astronomers?",
      "votes": null
    },
    {
      "id": "439463",
      "postDate": "12/15/2018 14:25:21",
      "content": "<p>I don't want to provide too much information here because we're in the last few days of the competition and it could be unfair. This is a general challenge in astronomy, and we use the term \"k-correction\" to refer to it. Ask me again after the competition is over and I'll provide more details.</p>",
      "rawMarkdown": "I don't want to provide too much information here because we're in the last few days of the competition and it could be unfair. This is a general challenge in astronomy, and we use the term \"k-correction\" to refer to it. Ask me again after the competition is over and I'll provide more details.",
      "votes": null
    },
    {
      "id": "439471",
      "postDate": "12/15/2018 14:45:26",
      "content": "<p>Since we're in the last days, let us just reiterate that we designed this contest so you would not have to have too much domain-specific knowledge in order to compete and do well. No one is expecting you to derive and apply k-corrections, and remember that a lot of classification thus far has been visual. We're definitely not doing multiple integrals in our head when we look at data to decide what an object might be!</p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
      "rawMarkdown": "Since we're in the last days, let us just reiterate that we designed this contest so you would not have to have too much domain-specific knowledge in order to compete and do well. No one is expecting you to derive and apply k-corrections, and remember that a lot of classification thus far has been visual. We're definitely not doing multiple integrals in our head when we look at data to decide what an object might be!\n\nCheers,\n\n-Gautham for the PLAsTiCC team",
      "votes": null
    },
    {
      "id": "439474",
      "postDate": "12/15/2018 15:02:31",
      "content": "<p>I was thinking of it for a long time, not just these few last days... I just could not get how you people do it, even though I looked through many papers... and when I saw templates I still could not understand how it was done. I am still curios to know how to do it. Now I can at least find some info, thank you</p>",
      "rawMarkdown": "I was thinking of it for a long time, not just these few last days... I just could not get how you people do it, even though I looked through many papers... and when I saw templates I still could not understand how it was done. I am still curios to know how to do it. Now I can at least find some info, thank you",
      "votes": null
    },
    {
      "id": "440137",
      "postDate": "12/17/2018 05:42:08",
      "content": "<p><a href=\"/gautam\">@gautam</a>\nbtw , if someone calculates can they have added advantage ?</p>",
      "rawMarkdown": "gautam\nbtw , if someone calculates can they have added advantage ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 410787,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "10/26/2018 16:55:31",
      "content": "<p>Maybe they found a way to deal with minority classes. I can't see there is any way that gradient boosting machine can handle class 53 &amp; 62</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 410789,
      "author_name": "maxhalford",
      "author_url": "",
      "post_date": "10/26/2018 16:59:45",
      "content": "<p>My guess is that they had some success with a feature extraction library such as <a href=\"http://cesium-ml.org/\">cesium</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 410803,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "10/26/2018 17:33:23",
      "content": "<p>I think they've found the way to handle class 99 problem</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 410808,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "10/26/2018 17:41:03",
      "content": "<p>I don't think the gap is significant. It is even small. I am sure that if I stop today, one week later I will be between 10th and 20th which is even optimistic. And personally, I don't have anything like <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268#325119\">this</a> yet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 410812,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "10/26/2018 17:45:27",
          "content": "<p>It was a real nightmare (for me)... I hope it won't happen again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410822,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/26/2018 17:58:45",
          "content": "<blockquote>\n  <p>It was a real nightmare (for me)... I hope it won't happen again!</p>\n</blockquote>\n\n<p>Ah, i knew your name was familiar!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410825,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "10/26/2018 18:01:13",
          "content": "<p>I agree. The gap is insignificant, it can be breached with small changes.\nThe breakthrough will be when someone will score less then 0.8 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410819,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "10/26/2018 17:57:06",
      "content": "<p>You are probably right, but I don't know what. </p>\n\n<p>Until today I wanted to change my group's name to: \"I really don't know what I'm doing\"</p>\n\n<p>I don't do anything fancy, just simple NN, with some common practice training and handling.</p>\n\n<p>I also have a few tiny tricks - but nothing big (every fancy thing I tried didn't seem to work)</p>\n\n<p>And most important, I didn't waste my time  on class 99 (for the time being) </p>\n\n<p>(of course I will have to do something with it sometime because this is the difference between my current score - 0.958, and my estimate of the best achievable score &lt;0.7) </p>",
      "votes": null,
      "replies": [
        {
          "id": 410831,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "10/26/2018 18:09:02",
          "content": "<p>I don't want to talk about my method, but I can only say I'm 100% sure my method and your method are completely different, the score is really close, though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410862,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "10/26/2018 19:18:05",
          "content": "<p>Hey Yuval,</p>\n\n<p>By NN do you mean a RNN or simply NN?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410877,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "10/26/2018 19:55:40",
          "content": "<p>neither</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411154,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/27/2018 14:17:53",
          "content": "<blockquote>\n  <p>my estimate of the best achievable score &lt;0.7</p>\n</blockquote>\n\n<p>How do you estimate that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411213,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "10/27/2018 16:21:44",
          "content": "<p>I believe a large portion of the difference between local scores and LB scores is related to class 99. I can achieve about 0.6 local score. If anyone can find a reliable way to predict class 99, he should get close to the local score (there will still be a difference because of the huge test set). If class 99 is a real type of event then it is possible to find features to classify it.    </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411256,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "10/27/2018 17:56:38",
          "content": "<p>And more accurately:</p>\n\n<p>If we give class 99 a constant value - V then the score penalty is:  </p>\n\n<pre><code>8/9*log(1-v)+1/9*log(V)\n</code></pre>\n\n<p>Assuming the weight for class 99,64,15 is indeed 2</p>\n\n<p>The optimum is about V=0.1 where we get penalty of about 0.35</p>\n\n<p>This is consistent with the current difference between my LB and local scores (I use a slightly better selection but not much).</p>\n\n<p>If we can really discover class 99, we can probably decrease the penalty to around 0.15 and improve the model by 0.05, an LB score is achievable (especially when considering the early stage of this competition) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411698,
          "author_name": "abednadir",
          "author_url": "",
          "post_date": "10/28/2018 20:13:46",
          "content": "<p>I approve of your proposed team name</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434658,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "12/06/2018 18:35:25",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a></p>\n\n<p>this question now comes back to you. There is a significant gap between 0.7x and 0.8x in LB... leader have found something... </p>",
      "votes": null,
      "replies": [
        {
          "id": 434713,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "12/06/2018 20:44:54",
          "content": "<p>the answer will probably come in 12 days or so</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434762,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/06/2018 22:43:02",
          "content": "<p>We found a hole in space time texture of the universe...</p>\n\n<p>Just kidding, we'll share what we did at the end of the competition.  Some of what we did may surprise others ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434828,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/07/2018 02:33:32",
          "content": "<p>There are holes in the space-time, physicists call it black holes... but if you've found it you would not return... nothing can escape the black hole :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 435205,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/07/2018 17:07:59",
          "content": "<p>I don't think the gap is significant. It is even small. I am sure that if I stop today, one week later I will be between 10th and 20th which is even optimistic.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 435246,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/07/2018 18:49:21",
          "content": "<p>@mamas I bet that a week from now 10th rank LB score will be higher than 0.755.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 435759,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/08/2018 17:59:58",
          "content": "<blockquote>\n  <p>leader have found something… </p>\n</blockquote>\n\n<p>Others seem to find something too, and the gap is shrinking fast.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 435900,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/09/2018 03:24:53",
          "content": "<p>aaaaa, but not me... I am still missing something... where to look for it ? :-( </p>\n\n<p>PS This kaggle thing is addictive, apparently. They should put a warning sign on the website: KAGGLE IS ADDICTIVE. ENTER AT YOUR OWN RISK ! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436027,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "12/09/2018 11:04:59",
          "content": "<p><a href=\"/blondinka\">@blondinka</a> I can relate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436042,
          "author_name": "kainsama",
          "author_url": "",
          "post_date": "12/09/2018 12:08:21",
          "content": "<p>Kaggle responsibly! ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438819,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/14/2018 08:39:02",
          "content": "<p>&gt; @mamas I bet that a week from now 10th rank LB score will be higher than 0.755.</p>\n\n<p>I was right, only 5 teams are below 0.755 now...  10th is 0.823 as I write.</p>\n\n<p>Crossing 0.8 is hard</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438846,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/14/2018 09:21:30",
          "content": "<p>hmm, seems they couldn't find 'secret'...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438897,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "12/14/2018 10:57:34",
          "content": "<p>yea. can't. ive been stuck for a while haha. I'm excited to see what you guys did.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434791,
      "author_name": "sdoria",
      "author_url": "",
      "post_date": "12/07/2018 00:31:54",
      "content": "<p>I wonder if we're dealing with a bunch of features with e.g. 0.02 improvement or if there are is some radically different way to look at things which gets the top people a big boost...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 437782,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/12/2018 13:40:20",
      "content": "<p>Seems now top 2 have found something others haven't.  The gap is large.</p>\n\n<p>I am writing this now because each time I did I closed the gap in the next few days ;)  Hope it can work again...</p>\n\n<p>Just kidding of course, but trying to catch up still.</p>",
      "votes": null,
      "replies": [
        {
          "id": 437787,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "12/12/2018 13:44:51",
          "content": "<p>where you able to calculate the hessian properly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437790,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/12/2018 13:47:23",
          "content": "<p>What for?  I don't see the need for it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437793,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "12/12/2018 13:53:47",
          "content": "<p>I can confirm that having the Hessian doesn't help. I've managed to calculate it analytically using <code>sympy</code> and it helps when using it with the LGBM loss function used in the kernels. In other words it is better than simply outputting a vector of 1s. But it is worse than properly using the <code>sample_weight</code> parameter. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437801,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "12/12/2018 14:17:02",
          "content": "<p>I think they have found the magic feature object_id%27</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437803,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "12/12/2018 14:23:22",
          "content": "<p>Are you sure it's 27? 28 is working better for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437804,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/12/2018 14:26:13",
          "content": "<p>Good to know, I am trying values one by one and I am only at object_id%22</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437876,
          "author_name": "vignam",
          "author_url": "",
          "post_date": "12/12/2018 17:16:59",
          "content": "<blockquote>\n  <p>I think they have found the magic feature object_id%27</p>\n</blockquote>\n\n<p>Oh damn! so that was it. But, now I have only 6 days to go. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437955,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/12/2018 20:48:08",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a></p>\n\n<p>I am wondering about it for the last week, we keep trying and cannot find... It is not iterative. Kyle is doing really well. If we assume you know ML better than Kyle, then it has to lie in the astronomy domain... that's what I think</p>\n\n<p>PS object_id%13 is my favorite :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437977,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "12/12/2018 21:57:39",
          "content": "<p>we have managed to find the answer to the universe, it's class 42.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438007,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "12/12/2018 23:38:06",
          "content": "<p>@cpmp and <a href=\"/maxhalford\">@maxhalford</a> thanks for confirmation. now i can stop trying to do it. haha</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438065,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "12/13/2018 03:40:15",
          "content": "<p>@Blonde, you use object_id%13 as a feature?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438625,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/14/2018 00:38:05",
          "content": "<p><a href=\"/niclasdoce\">@niclasdoce</a>  </p>\n\n<p>nop, I just stare at it... I managed to find it on the night sky when kaggling at night :-)</p>\n\n<p>PS use proper nick names, you see them in down-left corner when you navigate in the use name, my is <a href=\"/blondinka\">@blondinka</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438629,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "12/14/2018 00:45:13",
          "content": "<p>i don't understand.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 438698,
      "author_name": "kyleboone",
      "author_url": "",
      "post_date": "12/14/2018 03:29:26",
      "content": "<p>I think that there are quite a few non-obvious ways to increase your score. My first model used basically the same features/methods that I am using now, and it came in around 1.03. My CV really hasn't changed very much from that model. Most of the improvement has come from trying not to overfit the data to get the \"gap\" down. There are a lot of different tricks that I have found, and I think that I can guess who has figured out what from the leaderboard scores...</p>",
      "votes": null,
      "replies": [
        {
          "id": 438705,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "12/14/2018 03:43:09",
          "content": "<p>i see. that's great to hear.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438748,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "12/14/2018 05:57:17",
          "content": "<p>Thanks for sharing Kyle. In another discussion you have written that you have 0.4 CV and here you say your CV didn't change from LB score 1.03. it is really great that you obtained such CV score that early in the competition. But it means you had serious overfitting issues which causes 0.6 gap as far as I understand.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438761,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/14/2018 06:24:29",
          "content": "<p>Kyle,  you are sharing more and more, are you feeling alone at the top?</p>\n\n<p>I agree with you, lots of our progress as well was along the lines of what you just disclosed.   I guess we had to spend more time on feature engineering than you as we don't have your background.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438783,
          "author_name": "kyleboone",
          "author_url": "",
          "post_date": "12/14/2018 07:19:23",
          "content": "<p>I had roughly a CV of 0.5 with a LB of 1.0 at the start, so my gap was large. I dropped the gap down to ~0.4 fairly quickly, and lowering the CV has seemed to also lower the gap with the methods that I'm using. I posted on the CV thread several times, so you can see my progress there if you want. This is also my first kaggle competition and I've been learning a lot of ML as I go. For everything that I've done with ML in the past, tuning the model to get an extra 5% out didn't really matter. Bad tuning was hurting me a lot at first.</p>\n\n<p>I have a weird relationship with helping people out for this competition. On one hand I want to win, but on the other I am likely going to be using LSST data in my future research. I want the final models to be the best that they can be!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438798,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/14/2018 07:49:27",
          "content": "<blockquote>\n  <p>I want the final models to be the best that they can be!</p>\n</blockquote>\n\n<p>There should be a follow up competition next year ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439430,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/15/2018 12:48:40",
          "content": "<p><a href=\"/kyleboone\">@kyleboone</a> </p>\n\n<p>Kyle, as you are again in a good mood for sharing, I have a problem with Doppler effect. It's clear how to take it into account for individual spectroscopic lines, but I still do not get it how to recalculate redshift coefficients for integrated spectrum we have here. I've seen they use z in TypeIa templates, but it's still not clear how they recalculate brightness in different bands taking z into account for such an intergrated spectrum... I've seen many papers and people do not write clearly how they recalculate different bands, is it something obvious for astronomers?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439463,
          "author_name": "kyleboone",
          "author_url": "",
          "post_date": "12/15/2018 14:25:21",
          "content": "<p>I don't want to provide too much information here because we're in the last few days of the competition and it could be unfair. This is a general challenge in astronomy, and we use the term \"k-correction\" to refer to it. Ask me again after the competition is over and I'll provide more details.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439471,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "12/15/2018 14:45:26",
          "content": "<p>Since we're in the last days, let us just reiterate that we designed this contest so you would not have to have too much domain-specific knowledge in order to compete and do well. No one is expecting you to derive and apply k-corrections, and remember that a lot of classification thus far has been visual. We're definitely not doing multiple integrals in our head when we look at data to decide what an object might be!</p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439474,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/15/2018 15:02:31",
          "content": "<p>I was thinking of it for a long time, not just these few last days... I just could not get how you people do it, even though I looked through many papers... and when I saw templates I still could not understand how it was done. I am still curios to know how to do it. Now I can at least find some info, thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440137,
          "author_name": "mks2192",
          "author_url": "",
          "post_date": "12/17/2018 05:42:08",
          "content": "<p><a href=\"/gautam\">@gautam</a>\nbtw , if someone calculates can they have added advantage ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "410760": "There is a significant gap between 4th and 5th on the LB.  Leaders have found something...",
    "410787": "Maybe they found a way to deal with minority classes. I can't see there is any way that gradient boosting machine can handle class 53 &amp; 62",
    "410789": "My guess is that they had some success with a feature extraction library such as [cesium](http://cesium-ml.org/).",
    "410803": "I think they've found the way to handle class 99 problem",
    "410808": "I don't think the gap is significant. It is even small. I am sure that if I stop today, one week later I will be between 10th and 20th which is even optimistic. And personally, I don't have anything like [this][1] yet.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268#325119",
    "410812": "It was a real nightmare (for me)... I hope it won't happen again!",
    "410819": "You are probably right, but I don't know what. \n\nUntil today I wanted to change my group's name to: \"I really don't know what I'm doing\"\n\nI don't do anything fancy, just simple NN, with some common practice training and handling.\n\nI also have a few tiny tricks - but nothing big (every fancy thing I tried didn't seem to work)\n\nAnd most important, I didn't waste my time  on class 99 (for the time being) \n\n(of course I will have to do something with it sometime because this is the difference between my current score - 0.958, and my estimate of the best achievable score &lt;0.7)",
    "410822": "&gt; It was a real nightmare (for me)... I hope it won't happen again!\n\nAh, i knew your name was familiar!",
    "410825": "I agree. The gap is insignificant, it can be breached with small changes.\nThe breakthrough will be when someone will score less then 0.8",
    "410831": "I don't want to talk about my method, but I can only say I'm 100% sure my method and your method are completely different, the score is really close, though.",
    "410862": "Hey Yuval,\n\nBy NN do you mean a RNN or simply NN?",
    "410877": "neither",
    "411154": "&gt;  my estimate of the best achievable score &lt;0.7\n\nHow do you estimate that?",
    "411213": "I believe a large portion of the difference between local scores and LB scores is related to class 99. I can achieve about 0.6 local score. If anyone can find a reliable way to predict class 99, he should get close to the local score (there will still be a difference because of the huge test set). If class 99 is a real type of event then it is possible to find features to classify it.",
    "411256": "And more accurately:\n\nIf we give class 99 a constant value - V then the score penalty is:  \n\n    8/9*log(1-v)+1/9*log(V)\n\nAssuming the weight for class 99,64,15 is indeed 2\n\nThe optimum is about V=0.1 where we get penalty of about 0.35\n\nThis is consistent with the current difference between my LB and local scores (I use a slightly better selection but not much).\n\nIf we can really discover class 99, we can probably decrease the penalty to around 0.15 and improve the model by 0.05, an LB score is achievable (especially when considering the early stage of this competition)",
    "411698": "I approve of your proposed team name",
    "434658": "cpmpml\n\n this question now comes back to you. There is a significant gap between 0.7x and 0.8x in LB... leader have found something...",
    "434713": "the answer will probably come in 12 days or so",
    "434762": "We found a hole in space time texture of the universe...\n\nJust kidding, we'll share what we did at the end of the competition.  Some of what we did may surprise others ;)",
    "434791": "I wonder if we're dealing with a bunch of features with e.g. 0.02 improvement or if there are is some radically different way to look at things which gets the top people a big boost...",
    "434828": "There are holes in the space-time, physicists call it black holes... but if you've found it you would not return... nothing can escape the black hole :-)",
    "435205": "I don't think the gap is significant. It is even small. I am sure that if I stop today, one week later I will be between 10th and 20th which is even optimistic.",
    "435246": "mamas I bet that a week from now 10th rank LB score will be higher than 0.755.",
    "435759": "&gt; leader have found something… \n\nOthers seem to find something too, and the gap is shrinking fast.",
    "435900": "aaaaa, but not me... I am still missing something... where to look for it ? :-( \n\nPS This kaggle thing is addictive, apparently. They should put a warning sign on the website: KAGGLE IS ADDICTIVE. ENTER AT YOUR OWN RISK !",
    "436027": "blondinka I can relate",
    "436042": "Kaggle responsibly! ;)",
    "437782": "Seems now top 2 have found something others haven't.  The gap is large.\n\nI am writing this now because each time I did I closed the gap in the next few days ;)  Hope it can work again...\n\nJust kidding of course, but trying to catch up still.",
    "437787": "where you able to calculate the hessian properly?",
    "437790": "What for?  I don't see the need for it.",
    "437793": "I can confirm that having the Hessian doesn't help. I've managed to calculate it analytically using `sympy` and it helps when using it with the LGBM loss function used in the kernels. In other words it is better than simply outputting a vector of 1s. But it is worse than properly using the `sample_weight` parameter.",
    "437801": "I think they have found the magic feature object_id%27",
    "437803": "Are you sure it's 27? 28 is working better for me.",
    "437804": "Good to know, I am trying values one by one and I am only at object_id%22",
    "437876": "&gt; I think they have found the magic feature object_id%27\n\nOh damn! so that was it. But, now I have only 6 days to go.",
    "437955": "cpmpml\n\nI am wondering about it for the last week, we keep trying and cannot find... It is not iterative. Kyle is doing really well. If we assume you know ML better than Kyle, then it has to lie in the astronomy domain... that's what I think\n\nPS object_id%13 is my favorite :)",
    "437977": "we have managed to find the answer to the universe, it's class 42.",
    "438007": "cpmp and @maxhalford thanks for confirmation. now i can stop trying to do it. haha",
    "438065": "Blonde, you use object_id%13 as a feature?",
    "438625": "niclasdoce  \n\nnop, I just stare at it... I managed to find it on the night sky when kaggling at night :-)\n\nPS use proper nick names, you see them in down-left corner when you navigate in the use name, my is @blondinka",
    "438629": "i don't understand.",
    "438698": "I think that there are quite a few non-obvious ways to increase your score. My first model used basically the same features/methods that I am using now, and it came in around 1.03. My CV really hasn't changed very much from that model. Most of the improvement has come from trying not to overfit the data to get the \"gap\" down. There are a lot of different tricks that I have found, and I think that I can guess who has figured out what from the leaderboard scores...",
    "438705": "i see. that's great to hear.",
    "438748": "Thanks for sharing Kyle. In another discussion you have written that you have 0.4 CV and here you say your CV didn't change from LB score 1.03. it is really great that you obtained such CV score that early in the competition. But it means you had serious overfitting issues which causes 0.6 gap as far as I understand.",
    "438761": "Kyle,  you are sharing more and more, are you feeling alone at the top?\n\nI agree with you, lots of our progress as well was along the lines of what you just disclosed.   I guess we had to spend more time on feature engineering than you as we don't have your background.",
    "438783": "I had roughly a CV of 0.5 with a LB of 1.0 at the start, so my gap was large. I dropped the gap down to ~0.4 fairly quickly, and lowering the CV has seemed to also lower the gap with the methods that I'm using. I posted on the CV thread several times, so you can see my progress there if you want. This is also my first kaggle competition and I've been learning a lot of ML as I go. For everything that I've done with ML in the past, tuning the model to get an extra 5% out didn't really matter. Bad tuning was hurting me a lot at first.\n\nI have a weird relationship with helping people out for this competition. On one hand I want to win, but on the other I am likely going to be using LSST data in my future research. I want the final models to be the best that they can be!",
    "438798": "&gt; I want the final models to be the best that they can be!\n\nThere should be a follow up competition next year ;)",
    "438819": "&gt; @mamas I bet that a week from now 10th rank LB score will be higher than 0.755.\n\nI was right, only 5 teams are below 0.755 now...  10th is 0.823 as I write.\n\nCrossing 0.8 is hard",
    "438846": "hmm, seems they couldn't find 'secret'...",
    "438897": "yea. can't. ive been stuck for a while haha. I'm excited to see what you guys did.",
    "439430": "kyleboone \n\nKyle, as you are again in a good mood for sharing, I have a problem with Doppler effect. It's clear how to take it into account for individual spectroscopic lines, but I still do not get it how to recalculate redshift coefficients for integrated spectrum we have here. I've seen they use z in TypeIa templates, but it's still not clear how they recalculate brightness in different bands taking z into account for such an intergrated spectrum... I've seen many papers and people do not write clearly how they recalculate different bands, is it something obvious for astronomers?",
    "439463": "I don't want to provide too much information here because we're in the last few days of the competition and it could be unfair. This is a general challenge in astronomy, and we use the term \"k-correction\" to refer to it. Ask me again after the competition is over and I'll provide more details.",
    "439471": "Since we're in the last days, let us just reiterate that we designed this contest so you would not have to have too much domain-specific knowledge in order to compete and do well. No one is expecting you to derive and apply k-corrections, and remember that a lot of classification thus far has been visual. We're definitely not doing multiple integrals in our head when we look at data to decide what an object might be!\n\nCheers,\n\n-Gautham for the PLAsTiCC team",
    "439474": "I was thinking of it for a long time, not just these few last days... I just could not get how you people do it, even though I looked through many papers... and when I saw templates I still could not understand how it was done. I am still curios to know how to do it. Now I can at least find some info, thank you",
    "440137": "gautam\nbtw , if someone calculates can they have added advantage ?"
  },
  "source": "meta"
}