{
  "id": 75156,
  "title": "21st Solution ~super tough road~",
  "url": "/competitions/PLAsTiCC-2018/discussion/75156",
  "author_name": "",
  "post_date": "2018-12-19T01:07:50.279598500Z",
  "votes": 24,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi, dear kagglers.\nFirst of all, big thanks to Kaggle and LSST teams for holding such a flawless and leakage-free competition. And congratulations to the winners and all Kagglers.</p>\n\n<p>It’s really painful and tough during last few weeks since our score didn’t improve well. Although our team were unable to get gold medal. but we still write down our solution here in order to give some feedback to this wonderful community :))</p>\n\n<p>Our final model is a blending of 4 LGBMs. CV=0.41653, Public=0.85191, Private=0.86777, We’ll write what we’ve tried detailedly below :</p>\n\n<h3>Feature engineering</h3>\n\n<ul>\n<li><p>Aggregation <br>\nagg: {mean, median, max, skew, kurt, percentile(10,25,75,90,99), iqr, max-min} <br>\napply agg on :\n{flux, flux_err, flux_by_flux_ratio_sq, flux_ratio_sq, shift flux, phase shift flux, rolling flux, phase rolling flux, detected, normalized flux, flux*photoz, flux*photoz^2, flux / photoz, flux / photoz^2} <br>\ndifferent group : {object_id}, {object_id,passband}, {object_id,detected}, {object_id,passband,detected}, {kmeans of hostgal_photoz}</p></li>\n<li><p>Color features <br>\ndifference between the {max,mean,median} flux of different passbands (like <a href=\"https://arxiv.org/pdf/1701.05689.pdf\">this</a>, see 3.1, but not only computed with adjacent bands), these features give us really big boost.   </p></li>\n<li><p>Features from paper and different library <br>\n-- <a href=\"https://github.com/cesium-ml/cesium\">cesium library</a>, we computed all the features in it. <br>\n-- <a href=\"https://github.com/blue-yonder/tsfresh\">tsfresh</a>, just tried some of it. <br>\n-- <a href=\"http://isadoranun.github.io/tsfeat/FeaturesDocumentation.html\">FATS library</a>, we've tried some features based on the importance plot in these paper : <a href=\"https://arxiv.org/pdf/1506.00010.pdf\">paper_1</a>, <a href=\"https://arxiv.org/ftp/arxiv/papers/1710/1710.06804.pdf\">paper_2</a>, <a href=\"https://arxiv.org/pdf/1809.00763.pdf\">paper_3</a> <br>\n-- Some other features from <a href=\"https://arxiv.org/pdf/1801.07323.pdf\">this paper</a> and <a href=\"https://arxiv.org/pdf/1511.03456.pdf\">how to compute</a> : 4 passband independent\ntime scale features, Autocorrelation Integral, Shannon entropy, HL Ratio, Median Absolute Deviation, Von-Neumann Ratio and some Statistic like Shapiro-Wilk or Jarque-Bera. but most of them ending up with no improvement.</p></li>\n<li><p>some other useful features <br>\n-- max(mjd) - min(mjd) when detected == 1 <br>\n-- max(mjd) - min(mjd) when detected == 1 and only select flux that greater than mean(flux) <br>\n-- Some diff features like \"max(flux) - mean(flux)\" or ratio features <br>\n-- Percentage of passband at max flux per time, use mjd and phase, see below : <br>\n<img src=\"https://imgur.com/9kK71x2.png\" alt=\"class\">\n-- PCA on whole dataset(train + test) <br>\n-- Making a lgbm to predict hostgal_specz, using oof as a feature <br>\n-- Number of \"going up\" and \"going down\"</p></li>\n<li><p>something not worked for us this time <br>\n-- Measuring rising time and declining time of the light curve in different passband. <br>\n-- Decay speed of flux after peak in different time window <br>\n-- Conv1D, RNN for time series features <br>\n-- kmeans of some important features then doing WOE, target encoding\nex. kmeans of hostgal_photoz cluster=15, then target encoding or WOE per cluster.\n-- Stacking, still can't figure out why <br>\n-- Gaussian Process with RBF Kernel and Spline kernel in \"kernlab\" package in R, and doing same aggregation like above didn't improve our CV. <br>\n-- Training a model with objective = \"ova\" in lgbm, this one is only worse about 0.03-0.05 than gbdt one, but it didn't give any positive feedback when blending.</p></li>\n</ul>\n\n<h3>Special Sauce</h3>\n\n<p>There are some secret sauce that improved our score a lot, one is we use GPyOpt to optimized the sample weight in lightgbm, it suprisingly gave us another 0.04 boost in LB. We use following sample weight in our best single model :</p>\n\n<p>&gt; labels2weight = {6: 4.343454,\n 15: 3.283419,\n 16: 2.074140,\n 42: 0.2,\n 52: 3.671259,\n 53: 17.475224,\n 62: 0.837049,\n 64: 9.155481,\n 65: 0.728893,\n 67: 3.800462,\n 88: 4.003533,\n 90: 0.1,\n 92: 2.984069,\n 95: 4.337909\n}</p>\n\n<p>And another thicks is to doing some post-processing, which <a href=\"/fatihozturk\">@fatihozturk</a> will share details with validation in another write-up, we have done some optimization with following formula : pred_class[X] * param[X] and X = 1~14, we optimized the parameter with Nelder-Mead solver, and it also gain me another boost of 0.03. Also, applying </p>\n\n<p>&gt; pred[i] = pred[i].apply(lambda x: 1.1 if x&gt;0.76 else x )   </p>\n\n<p>can improve lb by 0.005 ! </p>\n\n<p>And that’s all, thanks for your reading!</p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75140\">fatih’s post processing thread</a></p>",
  "messages": [
    {
      "id": "441727",
      "postDate": "12/19/2018 01:07:50",
      "content": "<p>Hi, dear kagglers.\nFirst of all, big thanks to Kaggle and LSST teams for holding such a flawless and leakage-free competition. And congratulations to the winners and all Kagglers.</p>\n\n<p>It’s really painful and tough during last few weeks since our score didn’t improve well. Although our team were unable to get gold medal. but we still write down our solution here in order to give some feedback to this wonderful community :))</p>\n\n<p>Our final model is a blending of 4 LGBMs. CV=0.41653, Public=0.85191, Private=0.86777, We’ll write what we’ve tried detailedly below :</p>\n\n<h3>Feature engineering</h3>\n\n<ul>\n<li><p>Aggregation <br>\nagg: {mean, median, max, skew, kurt, percentile(10,25,75,90,99), iqr, max-min} <br>\napply agg on :\n{flux, flux_err, flux_by_flux_ratio_sq, flux_ratio_sq, shift flux, phase shift flux, rolling flux, phase rolling flux, detected, normalized flux, flux*photoz, flux*photoz^2, flux / photoz, flux / photoz^2} <br>\ndifferent group : {object_id}, {object_id,passband}, {object_id,detected}, {object_id,passband,detected}, {kmeans of hostgal_photoz}</p></li>\n<li><p>Color features <br>\ndifference between the {max,mean,median} flux of different passbands (like <a href=\"https://arxiv.org/pdf/1701.05689.pdf\">this</a>, see 3.1, but not only computed with adjacent bands), these features give us really big boost.   </p></li>\n<li><p>Features from paper and different library <br>\n-- <a href=\"https://github.com/cesium-ml/cesium\">cesium library</a>, we computed all the features in it. <br>\n-- <a href=\"https://github.com/blue-yonder/tsfresh\">tsfresh</a>, just tried some of it. <br>\n-- <a href=\"http://isadoranun.github.io/tsfeat/FeaturesDocumentation.html\">FATS library</a>, we've tried some features based on the importance plot in these paper : <a href=\"https://arxiv.org/pdf/1506.00010.pdf\">paper_1</a>, <a href=\"https://arxiv.org/ftp/arxiv/papers/1710/1710.06804.pdf\">paper_2</a>, <a href=\"https://arxiv.org/pdf/1809.00763.pdf\">paper_3</a> <br>\n-- Some other features from <a href=\"https://arxiv.org/pdf/1801.07323.pdf\">this paper</a> and <a href=\"https://arxiv.org/pdf/1511.03456.pdf\">how to compute</a> : 4 passband independent\ntime scale features, Autocorrelation Integral, Shannon entropy, HL Ratio, Median Absolute Deviation, Von-Neumann Ratio and some Statistic like Shapiro-Wilk or Jarque-Bera. but most of them ending up with no improvement.</p></li>\n<li><p>some other useful features <br>\n-- max(mjd) - min(mjd) when detected == 1 <br>\n-- max(mjd) - min(mjd) when detected == 1 and only select flux that greater than mean(flux) <br>\n-- Some diff features like \"max(flux) - mean(flux)\" or ratio features <br>\n-- Percentage of passband at max flux per time, use mjd and phase, see below : <br>\n<img src=\"https://imgur.com/9kK71x2.png\" alt=\"class\">\n-- PCA on whole dataset(train + test) <br>\n-- Making a lgbm to predict hostgal_specz, using oof as a feature <br>\n-- Number of \"going up\" and \"going down\"</p></li>\n<li><p>something not worked for us this time <br>\n-- Measuring rising time and declining time of the light curve in different passband. <br>\n-- Decay speed of flux after peak in different time window <br>\n-- Conv1D, RNN for time series features <br>\n-- kmeans of some important features then doing WOE, target encoding\nex. kmeans of hostgal_photoz cluster=15, then target encoding or WOE per cluster.\n-- Stacking, still can't figure out why <br>\n-- Gaussian Process with RBF Kernel and Spline kernel in \"kernlab\" package in R, and doing same aggregation like above didn't improve our CV. <br>\n-- Training a model with objective = \"ova\" in lgbm, this one is only worse about 0.03-0.05 than gbdt one, but it didn't give any positive feedback when blending.</p></li>\n</ul>\n\n<h3>Special Sauce</h3>\n\n<p>There are some secret sauce that improved our score a lot, one is we use GPyOpt to optimized the sample weight in lightgbm, it suprisingly gave us another 0.04 boost in LB. We use following sample weight in our best single model :</p>\n\n<p>&gt; labels2weight = {6: 4.343454,\n 15: 3.283419,\n 16: 2.074140,\n 42: 0.2,\n 52: 3.671259,\n 53: 17.475224,\n 62: 0.837049,\n 64: 9.155481,\n 65: 0.728893,\n 67: 3.800462,\n 88: 4.003533,\n 90: 0.1,\n 92: 2.984069,\n 95: 4.337909\n}</p>\n\n<p>And another thicks is to doing some post-processing, which <a href=\"/fatihozturk\">@fatihozturk</a> will share details with validation in another write-up, we have done some optimization with following formula : pred_class[X] * param[X] and X = 1~14, we optimized the parameter with Nelder-Mead solver, and it also gain me another boost of 0.03. Also, applying </p>\n\n<p>&gt; pred[i] = pred[i].apply(lambda x: 1.1 if x&gt;0.76 else x )   </p>\n\n<p>can improve lb by 0.005 ! </p>\n\n<p>And that’s all, thanks for your reading!</p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75140\">fatih’s post processing thread</a></p>",
      "rawMarkdown": "Hi, dear kagglers.\nFirst of all, big thanks to Kaggle and LSST teams for holding such a flawless and leakage-free competition. And congratulations to the winners and all Kagglers.\n\nIt’s really painful and tough during last few weeks since our score didn’t improve well. Although our team were unable to get gold medal. but we still write down our solution here in order to give some feedback to this wonderful community :))\n\nOur final model is a blending of 4 LGBMs. CV=0.41653, Public=0.85191, Private=0.86777, We’ll write what we’ve tried detailedly below :\n\n### Feature engineering\n\n+ Aggregation   \nagg: {mean, median, max, skew, kurt, percentile(10,25,75,90,99), iqr, max-min}  \napply agg on :\n{flux, flux_err, flux_by_flux_ratio_sq, flux_ratio_sq, shift flux, phase shift flux, rolling flux, phase rolling flux, detected, normalized flux, flux*photoz, flux*photoz^2, flux / photoz, flux / photoz^2}  \ndifferent group : {object_id}, {object_id,passband}, {object_id,detected}, {object_id,passband,detected}, {kmeans of hostgal_photoz}\n\n\n+ Color features   \n  difference between the {max,mean,median} flux of different passbands (like [this](https://arxiv.org/pdf/1701.05689.pdf), see 3.1, but not only computed with adjacent bands), these features give us really big boost.   \n  \n+ Features from paper and different library  \n  -- [cesium library](https://github.com/cesium-ml/cesium), we computed all the features in it.  \n  -- [tsfresh](https://github.com/blue-yonder/tsfresh), just tried some of it.  \n  -- [FATS library](http://isadoranun.github.io/tsfeat/FeaturesDocumentation.html), we've tried some features based on the importance plot in these paper : [paper_1](https://arxiv.org/pdf/1506.00010.pdf), [paper_2](https://arxiv.org/ftp/arxiv/papers/1710/1710.06804.pdf), [paper_3](https://arxiv.org/pdf/1809.00763.pdf)  \n  -- Some other features from [this paper](https://arxiv.org/pdf/1801.07323.pdf) and [how to compute](https://arxiv.org/pdf/1511.03456.pdf) : 4 passband independent\ntime scale features, Autocorrelation Integral, Shannon entropy, HL Ratio, Median Absolute Deviation, Von-Neumann Ratio and some Statistic like Shapiro-Wilk or Jarque-Bera. but most of them ending up with no improvement.\n  \n+ some other useful features   \n  -- max(mjd) - min(mjd) when detected == 1  \n  -- max(mjd) - min(mjd) when detected == 1 and only select flux that greater than mean(flux)  \n  -- Some diff features like \"max(flux) - mean(flux)\" or ratio features  \n  -- Percentage of passband at max flux per time, use mjd and phase, see below :  \n![class][1]\n  -- PCA on whole dataset(train + test)  \n  -- Making a lgbm to predict hostgal_specz, using oof as a feature  \n  -- Number of \"going up\" and \"going down\"\n  \n+ something not worked for us this time  \n  -- Measuring rising time and declining time of the light curve in different passband.  \n  -- Decay speed of flux after peak in different time window   \n  -- Conv1D, RNN for time series features  \n  -- kmeans of some important features then doing WOE, target encoding\nex. kmeans of hostgal_photoz cluster=15, then target encoding or WOE per cluster.\n  -- Stacking, still can't figure out why  \n  -- Gaussian Process with RBF Kernel and Spline kernel in \"kernlab\" package in R, and doing same aggregation like above didn't improve our CV.  \n  -- Training a model with objective = \"ova\" in lgbm, this one is only worse about 0.03-0.05 than gbdt one, but it didn't give any positive feedback when blending.\n\n\n### Special Sauce\nThere are some secret sauce that improved our score a lot, one is we use GPyOpt to optimized the sample weight in lightgbm, it suprisingly gave us another 0.04 boost in LB. We use following sample weight in our best single model :\n\n&gt; labels2weight = {6: 4.343454,\n 15: 3.283419,\n 16: 2.074140,\n 42: 0.2,\n 52: 3.671259,\n 53: 17.475224,\n 62: 0.837049,\n 64: 9.155481,\n 65: 0.728893,\n 67: 3.800462,\n 88: 4.003533,\n 90: 0.1,\n 92: 2.984069,\n 95: 4.337909\n}\n\nAnd another thicks is to doing some post-processing, which @fatihozturk will share details with validation in another write-up, we have done some optimization with following formula : pred_class[X] * param[X] and X = 1~14, we optimized the parameter with Nelder-Mead solver, and it also gain me another boost of 0.03. Also, applying \n\n&gt; pred[i] = pred[i].apply(lambda x: 1.1 if x&gt;0.76 else x )   \n\ncan improve lb by 0.005 ! \n\n\nAnd that’s all, thanks for your reading!\n\n[fatih’s post processing thread](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75140)\n\n[1]:https://imgur.com/9kK71x2.png",
      "votes": null
    },
    {
      "id": "441731",
      "postDate": "12/19/2018 01:17:15",
      "content": "<p>Thanks for sharing, and congrats on the result.  Your weights are close to the optimal ones obtained by class weights divided by class frequency, but some are quite different.  I incldue below the ones recommended by theory. Did you found that your weights were actually better than the theoretically optimal weights?  And isn't your postprocessing trick needed because you don't use optimal weights?</p>\n\n<pre><code>    weight\n 6  3.248344\n15  1.981818\n16  0.530844\n42  0.411148\n52  2.680328\n53  16.350000\n62  1.013430\n64  9.617647\n65  0.500000\n67  2.358173\n88  1.325676\n90  0.212062\n92  2.052301\n95  2.802857\n</code></pre>\n\n<p>​Edited.</p>",
      "rawMarkdown": "Thanks for sharing, and congrats on the result.  Your weights are close to the optimal ones obtained by class weights divided by class frequency, but some are quite different.  I incldue below the ones recommended by theory. Did you found that your weights were actually better than the theoretically optimal weights?  And isn't your postprocessing trick needed because you don't use optimal weights?\n\n     \tweight\n     6 \t3.248344\n    15 \t1.981818\n    16 \t0.530844\n    42 \t0.411148\n    52 \t2.680328\n    53 \t16.350000\n    62 \t1.013430\n    64 \t9.617647\n    65 \t0.500000\n    67 \t2.358173\n    88 \t1.325676\n    90 \t0.212062\n    92 \t2.052301\n    95 \t2.802857\n\n​Edited.",
      "votes": null
    },
    {
      "id": "441735",
      "postDate": "12/19/2018 01:33:50",
      "content": "<p>&gt;Did you found that your weights were actually better than the theoretically optimal weights?</p>\n\n<p>Yes. First, I used optimal weights. After changing weight, it improved score 0.04.\nI optimized weight with CV that calculated with theoretically  optimal weights. And best weight CV was better than theoretically one.</p>\n\n<p>&gt;And isn't your postprocessing trick needed because you don't use optimal weights?</p>\n\n<p>I don't think so. I changed only weight (not used post processing), it improved score 0.04.</p>\n\n<p>But this process didn't boost fatih's score. His improvement was below 0.01. I don't know the reason. But in my model it boosted score well.</p>",
      "rawMarkdown": "&gt;Did you found that your weights were actually better than the theoretically optimal weights?\n\nYes. First, I used optimal weights. After changing weight, it improved score 0.04.\nI optimized weight with CV that calculated with theoretically  optimal weights. And best weight CV was better than theoretically one.\n\n&gt;And isn't your postprocessing trick needed because you don't use optimal weights?\n\nI don't think so. I changed only weight (not used post processing), it improved score 0.04.\n\nBut this process didn't boost fatih's score. His improvement was below 0.01. I don't know the reason. But in my model it boosted score well.",
      "votes": null
    },
    {
      "id": "441736",
      "postDate": "12/19/2018 01:41:39",
      "content": "<p>I think this issue really depends on each model. I also used very different weights than these ones and mine was always better for me. <a href=\"/cpmpml\">@cpmpml</a> you should give a try postprocessing in a way that I explained in the thread and let us know if it also improves your model or not even with the 'optimal' weights.</p>",
      "rawMarkdown": "I think this issue really depends on each model. I also used very different weights than these ones and mine was always better for me. @cpmpml you should give a try postprocessing in a way that I explained in the thread and let us know if it also improves your model or not even with the 'optimal' weights.",
      "votes": null
    },
    {
      "id": "441737",
      "postDate": "12/19/2018 01:49:04",
      "content": "<p>Sorry, did not include the right weights.  I am surprised by this I must say.  I'll try if I have the energy ;)</p>",
      "rawMarkdown": "Sorry, did not include the right weights.  I am surprised by this I must say.  I'll try if I have the energy ;)",
      "votes": null
    },
    {
      "id": "441748",
      "postDate": "12/19/2018 02:30:23",
      "content": "<p>Congrats and thanks for sharing your solution. After noticing difficult to classify objects kept getting dumped into class 90 (others had mentioned this in the discussions) I had played around with giving class 90 a smaller weight and found 0.1 gave my model the best improvement. Looking at your solution I regret not having pursued that idea further. </p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution. After noticing difficult to classify objects kept getting dumped into class 90 (others had mentioned this in the discussions) I had played around with giving class 90 a smaller weight and found 0.1 gave my model the best improvement. Looking at your solution I regret not having pursued that idea further.",
      "votes": null
    },
    {
      "id": "441851",
      "postDate": "12/19/2018 06:54:59",
      "content": "<p>Congrats and thanks for sharing. I had the same sort of experience with stacking. All the tools I have been using during the last 3 years of Kaggle competitions did not work here. I'm still puzzled.</p>",
      "rawMarkdown": "Congrats and thanks for sharing. I had the same sort of experience with stacking. All the tools I have been using during the last 3 years of Kaggle competitions did not work here. I'm still puzzled.",
      "votes": null
    },
    {
      "id": "441856",
      "postDate": "12/19/2018 07:06:12",
      "content": "<p>Congrats takuoko, It looks like a very good idea to tune sample weight, I should have tried it.</p>",
      "rawMarkdown": "Congrats takuoko, It looks like a very good idea to tune sample weight, I should have tried it.",
      "votes": null
    },
    {
      "id": "441860",
      "postDate": "12/19/2018 07:09:28",
      "content": "<p>yes, in this competition, almost all the tools we got from kaggle didn't work. what was really necessary was fresh and logical ideas. Actually, my teammate, yuval didn't know what stacking means when I teamed-up with him, despite his super performance.</p>",
      "rawMarkdown": "yes, in this competition, almost all the tools we got from kaggle didn't work. what was really necessary was fresh and logical ideas. Actually, my teammate, yuval didn't know what stacking means when I teamed-up with him, despite his super performance.",
      "votes": null
    },
    {
      "id": "441888",
      "postDate": "12/19/2018 07:54:24",
      "content": "<p>Yes, let's. It may boost your team for 2nd or 1st place:)</p>",
      "rawMarkdown": "Yes, let's. It may boost your team for 2nd or 1st place:)",
      "votes": null
    },
    {
      "id": "441890",
      "postDate": "12/19/2018 08:00:18",
      "content": "<p>Thank you! \nActually, optimizing weight is not easy task. The result changed a little every time. So it takes much optimization. I used 300 epochs optimization with GPyOpt, but it may be better that using more epochs.</p>",
      "rawMarkdown": "Thank you! \nActually, optimizing weight is not easy task. The result changed a little every time. So it takes much optimization. I used 300 epochs optimization with GPyOpt, but it may be better that using more epochs.",
      "votes": null
    },
    {
      "id": "441892",
      "postDate": "12/19/2018 08:03:56",
      "content": "<p>Yes. Stacking caused overfitting terribly. One reason may be difference between distribution of train and test. Stacking tends to extract train result, IMHO.\n<a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75012\">This post</a> shows good stacking solution. </p>",
      "rawMarkdown": "Yes. Stacking caused overfitting terribly. One reason may be difference between distribution of train and test. Stacking tends to extract train result, IMHO.\n[This post](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75012) shows good stacking solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 441731,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/19/2018 01:17:15",
      "content": "<p>Thanks for sharing, and congrats on the result.  Your weights are close to the optimal ones obtained by class weights divided by class frequency, but some are quite different.  I incldue below the ones recommended by theory. Did you found that your weights were actually better than the theoretically optimal weights?  And isn't your postprocessing trick needed because you don't use optimal weights?</p>\n\n<pre><code>    weight\n 6  3.248344\n15  1.981818\n16  0.530844\n42  0.411148\n52  2.680328\n53  16.350000\n62  1.013430\n64  9.617647\n65  0.500000\n67  2.358173\n88  1.325676\n90  0.212062\n92  2.052301\n95  2.802857\n</code></pre>\n\n<p>​Edited.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441735,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "12/19/2018 01:33:50",
          "content": "<p>&gt;Did you found that your weights were actually better than the theoretically optimal weights?</p>\n\n<p>Yes. First, I used optimal weights. After changing weight, it improved score 0.04.\nI optimized weight with CV that calculated with theoretically  optimal weights. And best weight CV was better than theoretically one.</p>\n\n<p>&gt;And isn't your postprocessing trick needed because you don't use optimal weights?</p>\n\n<p>I don't think so. I changed only weight (not used post processing), it improved score 0.04.</p>\n\n<p>But this process didn't boost fatih's score. His improvement was below 0.01. I don't know the reason. But in my model it boosted score well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441736,
          "author_name": "fatihozturk",
          "author_url": "",
          "post_date": "12/19/2018 01:41:39",
          "content": "<p>I think this issue really depends on each model. I also used very different weights than these ones and mine was always better for me. <a href=\"/cpmpml\">@cpmpml</a> you should give a try postprocessing in a way that I explained in the thread and let us know if it also improves your model or not even with the 'optimal' weights.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441737,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/19/2018 01:49:04",
          "content": "<p>Sorry, did not include the right weights.  I am surprised by this I must say.  I'll try if I have the energy ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441748,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/19/2018 02:30:23",
      "content": "<p>Congrats and thanks for sharing your solution. After noticing difficult to classify objects kept getting dumped into class 90 (others had mentioned this in the discussions) I had played around with giving class 90 a smaller weight and found 0.1 gave my model the best improvement. Looking at your solution I regret not having pursued that idea further. </p>",
      "votes": null,
      "replies": [
        {
          "id": 441890,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "12/19/2018 08:00:18",
          "content": "<p>Thank you! \nActually, optimizing weight is not easy task. The result changed a little every time. So it takes much optimization. I used 300 epochs optimization with GPyOpt, but it may be better that using more epochs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441851,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "12/19/2018 06:54:59",
      "content": "<p>Congrats and thanks for sharing. I had the same sort of experience with stacking. All the tools I have been using during the last 3 years of Kaggle competitions did not work here. I'm still puzzled.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441860,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/19/2018 07:09:28",
          "content": "<p>yes, in this competition, almost all the tools we got from kaggle didn't work. what was really necessary was fresh and logical ideas. Actually, my teammate, yuval didn't know what stacking means when I teamed-up with him, despite his super performance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441892,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "12/19/2018 08:03:56",
          "content": "<p>Yes. Stacking caused overfitting terribly. One reason may be difference between distribution of train and test. Stacking tends to extract train result, IMHO.\n<a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75012\">This post</a> shows good stacking solution. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441856,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/19/2018 07:06:12",
      "content": "<p>Congrats takuoko, It looks like a very good idea to tune sample weight, I should have tried it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441888,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "12/19/2018 07:54:24",
          "content": "<p>Yes, let's. It may boost your team for 2nd or 1st place:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "441727": "Hi, dear kagglers.\nFirst of all, big thanks to Kaggle and LSST teams for holding such a flawless and leakage-free competition. And congratulations to the winners and all Kagglers.\n\nIt’s really painful and tough during last few weeks since our score didn’t improve well. Although our team were unable to get gold medal. but we still write down our solution here in order to give some feedback to this wonderful community :))\n\nOur final model is a blending of 4 LGBMs. CV=0.41653, Public=0.85191, Private=0.86777, We’ll write what we’ve tried detailedly below :\n\n### Feature engineering\n\n+ Aggregation   \nagg: {mean, median, max, skew, kurt, percentile(10,25,75,90,99), iqr, max-min}  \napply agg on :\n{flux, flux_err, flux_by_flux_ratio_sq, flux_ratio_sq, shift flux, phase shift flux, rolling flux, phase rolling flux, detected, normalized flux, flux*photoz, flux*photoz^2, flux / photoz, flux / photoz^2}  \ndifferent group : {object_id}, {object_id,passband}, {object_id,detected}, {object_id,passband,detected}, {kmeans of hostgal_photoz}\n\n\n+ Color features   \n  difference between the {max,mean,median} flux of different passbands (like [this](https://arxiv.org/pdf/1701.05689.pdf), see 3.1, but not only computed with adjacent bands), these features give us really big boost.   \n  \n+ Features from paper and different library  \n  -- [cesium library](https://github.com/cesium-ml/cesium), we computed all the features in it.  \n  -- [tsfresh](https://github.com/blue-yonder/tsfresh), just tried some of it.  \n  -- [FATS library](http://isadoranun.github.io/tsfeat/FeaturesDocumentation.html), we've tried some features based on the importance plot in these paper : [paper_1](https://arxiv.org/pdf/1506.00010.pdf), [paper_2](https://arxiv.org/ftp/arxiv/papers/1710/1710.06804.pdf), [paper_3](https://arxiv.org/pdf/1809.00763.pdf)  \n  -- Some other features from [this paper](https://arxiv.org/pdf/1801.07323.pdf) and [how to compute](https://arxiv.org/pdf/1511.03456.pdf) : 4 passband independent\ntime scale features, Autocorrelation Integral, Shannon entropy, HL Ratio, Median Absolute Deviation, Von-Neumann Ratio and some Statistic like Shapiro-Wilk or Jarque-Bera. but most of them ending up with no improvement.\n  \n+ some other useful features   \n  -- max(mjd) - min(mjd) when detected == 1  \n  -- max(mjd) - min(mjd) when detected == 1 and only select flux that greater than mean(flux)  \n  -- Some diff features like \"max(flux) - mean(flux)\" or ratio features  \n  -- Percentage of passband at max flux per time, use mjd and phase, see below :  \n![class][1]\n  -- PCA on whole dataset(train + test)  \n  -- Making a lgbm to predict hostgal_specz, using oof as a feature  \n  -- Number of \"going up\" and \"going down\"\n  \n+ something not worked for us this time  \n  -- Measuring rising time and declining time of the light curve in different passband.  \n  -- Decay speed of flux after peak in different time window   \n  -- Conv1D, RNN for time series features  \n  -- kmeans of some important features then doing WOE, target encoding\nex. kmeans of hostgal_photoz cluster=15, then target encoding or WOE per cluster.\n  -- Stacking, still can't figure out why  \n  -- Gaussian Process with RBF Kernel and Spline kernel in \"kernlab\" package in R, and doing same aggregation like above didn't improve our CV.  \n  -- Training a model with objective = \"ova\" in lgbm, this one is only worse about 0.03-0.05 than gbdt one, but it didn't give any positive feedback when blending.\n\n\n### Special Sauce\nThere are some secret sauce that improved our score a lot, one is we use GPyOpt to optimized the sample weight in lightgbm, it suprisingly gave us another 0.04 boost in LB. We use following sample weight in our best single model :\n\n&gt; labels2weight = {6: 4.343454,\n 15: 3.283419,\n 16: 2.074140,\n 42: 0.2,\n 52: 3.671259,\n 53: 17.475224,\n 62: 0.837049,\n 64: 9.155481,\n 65: 0.728893,\n 67: 3.800462,\n 88: 4.003533,\n 90: 0.1,\n 92: 2.984069,\n 95: 4.337909\n}\n\nAnd another thicks is to doing some post-processing, which @fatihozturk will share details with validation in another write-up, we have done some optimization with following formula : pred_class[X] * param[X] and X = 1~14, we optimized the parameter with Nelder-Mead solver, and it also gain me another boost of 0.03. Also, applying \n\n&gt; pred[i] = pred[i].apply(lambda x: 1.1 if x&gt;0.76 else x )   \n\ncan improve lb by 0.005 ! \n\n\nAnd that’s all, thanks for your reading!\n\n[fatih’s post processing thread](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75140)\n\n[1]:https://imgur.com/9kK71x2.png",
    "441731": "Thanks for sharing, and congrats on the result.  Your weights are close to the optimal ones obtained by class weights divided by class frequency, but some are quite different.  I incldue below the ones recommended by theory. Did you found that your weights were actually better than the theoretically optimal weights?  And isn't your postprocessing trick needed because you don't use optimal weights?\n\n     \tweight\n     6 \t3.248344\n    15 \t1.981818\n    16 \t0.530844\n    42 \t0.411148\n    52 \t2.680328\n    53 \t16.350000\n    62 \t1.013430\n    64 \t9.617647\n    65 \t0.500000\n    67 \t2.358173\n    88 \t1.325676\n    90 \t0.212062\n    92 \t2.052301\n    95 \t2.802857\n\n​Edited.",
    "441735": "&gt;Did you found that your weights were actually better than the theoretically optimal weights?\n\nYes. First, I used optimal weights. After changing weight, it improved score 0.04.\nI optimized weight with CV that calculated with theoretically  optimal weights. And best weight CV was better than theoretically one.\n\n&gt;And isn't your postprocessing trick needed because you don't use optimal weights?\n\nI don't think so. I changed only weight (not used post processing), it improved score 0.04.\n\nBut this process didn't boost fatih's score. His improvement was below 0.01. I don't know the reason. But in my model it boosted score well.",
    "441736": "I think this issue really depends on each model. I also used very different weights than these ones and mine was always better for me. @cpmpml you should give a try postprocessing in a way that I explained in the thread and let us know if it also improves your model or not even with the 'optimal' weights.",
    "441737": "Sorry, did not include the right weights.  I am surprised by this I must say.  I'll try if I have the energy ;)",
    "441748": "Congrats and thanks for sharing your solution. After noticing difficult to classify objects kept getting dumped into class 90 (others had mentioned this in the discussions) I had played around with giving class 90 a smaller weight and found 0.1 gave my model the best improvement. Looking at your solution I regret not having pursued that idea further.",
    "441851": "Congrats and thanks for sharing. I had the same sort of experience with stacking. All the tools I have been using during the last 3 years of Kaggle competitions did not work here. I'm still puzzled.",
    "441856": "Congrats takuoko, It looks like a very good idea to tune sample weight, I should have tried it.",
    "441860": "yes, in this competition, almost all the tools we got from kaggle didn't work. what was really necessary was fresh and logical ideas. Actually, my teammate, yuval didn't know what stacking means when I teamed-up with him, despite his super performance.",
    "441888": "Yes, let's. It may boost your team for 2nd or 1st place:)",
    "441890": "Thank you! \nActually, optimizing weight is not easy task. The result changed a little every time. So it takes much optimization. I used 300 epochs optimization with GPyOpt, but it may be better that using more epochs.",
    "441892": "Yes. Stacking caused overfitting terribly. One reason may be difference between distribution of train and test. Stacking tends to extract train result, IMHO.\n[This post](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75012) shows good stacking solution."
  },
  "source": "meta"
}