{
  "id": 75262,
  "title": "20th Place Solution",
  "url": "/competitions/PLAsTiCC-2018/writeups/macho-20th-place-solution",
  "author_name": "",
  "post_date": "2018-12-20T01:27:06.963779800Z",
  "votes": 36,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all, thanks Kaggle and the sponsors for this great, clear and leak free competition :-)</p>\n\n<p>This competition I started working hard in the beggining and practically stopped competing during last 10 days. My final solution is a stack ensemble and a blend of 4x LGB and 2x SVC models trained using different subsets of features. </p>\n\n<h1>Features</h1>\n\n<p>I used some subsets of around 250 features for the final models, but I'm sure I tried more than 1000 during the competition and dropped useless features. My features are built using all kind of aggregations and statistics, feets lib, coeficients of linear regressions for each passband, coeficient  of linear regression after a peak and custom made features base on observation of the light curves, for example: raise rate of peaks and valeys, peak ratios of 'time to raive' vs 'time to decrease to X%', etc...</p>\n\n<h1>Models</h1>\n\n<p>For each model fit I augmented the train set concatenating noise versions of the train set. That simple augmentation improved my score by about 0.04 and decreased the CV x LB gap.</p>\n\n<p>Model 1:  multi-class LGB trained using all 250 features and all whole set.</p>\n\n<p>Model 2:  multi-class LGB trained using a subset of the 250 features, selected via feature selection, using whole train set.</p>\n\n<p>Model 3:  multi-class LGB trained using another subset of the 250 features, selected via other feature selection algorithm, using whole train set.</p>\n\n<p>Model 4:  multi-label LGB trained using a subset of the 250 features, selected via feature selection for each target individually, using whole train set.</p>\n\n<p>Model 5:  multi-label SVC using rbf kernel trained using all 250 features and whole train set.</p>\n\n<p>Model 6:  multi-class LGB trained using all 250 features. Fit one model for intra galaxy and other for extra galaxy.</p>\n\n<p>Model 7:  multi-label LGB trained using a subset of the 250 features, selected via feature selection for each target individually.  Fit one model for ddf==0 and other for ddf==1.</p>\n\n<p>Stack ensemble 1, multi-class, models 1,2,3,4,5.</p>\n\n<p>Stack ensemble 2, multi-class, model 6.</p>\n\n<p>Stack ensemble 3, multi-class, model 7.</p>\n\n<p>Blend Stack Ensemble 1,2,3 using simple geometric weighted average.</p>\n\n<p>For class 99 I ended using Scirpus(thanks @scirpus) equation plus some minor LB probings. </p>\n\n<h1>What didn't worked for me:</h1>\n\n<ul>\n<li><p>List item</p></li>\n<li><p>FFT, wavelet and periodogram features.</p></li>\n<li><p>DAE</p></li>\n<li><p>Curve reconstruction</p></li>\n<li><p>Fit a model for hoztgal_specz using train+test</p></li>\n<li><p>CNN models ( metric &gt; 1.3 )</p></li>\n<li><p>Target encoding</p></li>\n<li><p>Flux correction by redshift</p></li>\n<li><p>Flux passband correction using redshift</p></li>\n<li><p>Semi-supervised Learning (improved just 0.0001)</p></li>\n<li><p>SMOTE</p></li>\n</ul>\n\n<p>Giba</p>",
  "messages": [
    {
      "id": "442453",
      "postDate": "12/20/2018 01:27:06",
      "content": "<p>First of all, thanks Kaggle and the sponsors for this great, clear and leak free competition :-)</p>\n\n<p>This competition I started working hard in the beggining and practically stopped competing during last 10 days. My final solution is a stack ensemble and a blend of 4x LGB and 2x SVC models trained using different subsets of features. </p>\n\n<h1>Features</h1>\n\n<p>I used some subsets of around 250 features for the final models, but I'm sure I tried more than 1000 during the competition and dropped useless features. My features are built using all kind of aggregations and statistics, feets lib, coeficients of linear regressions for each passband, coeficient  of linear regression after a peak and custom made features base on observation of the light curves, for example: raise rate of peaks and valeys, peak ratios of 'time to raive' vs 'time to decrease to X%', etc...</p>\n\n<h1>Models</h1>\n\n<p>For each model fit I augmented the train set concatenating noise versions of the train set. That simple augmentation improved my score by about 0.04 and decreased the CV x LB gap.</p>\n\n<p>Model 1:  multi-class LGB trained using all 250 features and all whole set.</p>\n\n<p>Model 2:  multi-class LGB trained using a subset of the 250 features, selected via feature selection, using whole train set.</p>\n\n<p>Model 3:  multi-class LGB trained using another subset of the 250 features, selected via other feature selection algorithm, using whole train set.</p>\n\n<p>Model 4:  multi-label LGB trained using a subset of the 250 features, selected via feature selection for each target individually, using whole train set.</p>\n\n<p>Model 5:  multi-label SVC using rbf kernel trained using all 250 features and whole train set.</p>\n\n<p>Model 6:  multi-class LGB trained using all 250 features. Fit one model for intra galaxy and other for extra galaxy.</p>\n\n<p>Model 7:  multi-label LGB trained using a subset of the 250 features, selected via feature selection for each target individually.  Fit one model for ddf==0 and other for ddf==1.</p>\n\n<p>Stack ensemble 1, multi-class, models 1,2,3,4,5.</p>\n\n<p>Stack ensemble 2, multi-class, model 6.</p>\n\n<p>Stack ensemble 3, multi-class, model 7.</p>\n\n<p>Blend Stack Ensemble 1,2,3 using simple geometric weighted average.</p>\n\n<p>For class 99 I ended using Scirpus(thanks @scirpus) equation plus some minor LB probings. </p>\n\n<h1>What didn't worked for me:</h1>\n\n<ul>\n<li><p>List item</p></li>\n<li><p>FFT, wavelet and periodogram features.</p></li>\n<li><p>DAE</p></li>\n<li><p>Curve reconstruction</p></li>\n<li><p>Fit a model for hoztgal_specz using train+test</p></li>\n<li><p>CNN models ( metric &gt; 1.3 )</p></li>\n<li><p>Target encoding</p></li>\n<li><p>Flux correction by redshift</p></li>\n<li><p>Flux passband correction using redshift</p></li>\n<li><p>Semi-supervised Learning (improved just 0.0001)</p></li>\n<li><p>SMOTE</p></li>\n</ul>\n\n<p>Giba</p>",
      "rawMarkdown": "First of all, thanks Kaggle and the sponsors for this great, clear and leak free competition :-)\n\nThis competition I started working hard in the beggining and practically stopped competing during last 10 days. My final solution is a stack ensemble and a blend of 4x LGB and 2x SVC models trained using different subsets of features. \n\n#Features\nI used some subsets of around 250 features for the final models, but I'm sure I tried more than 1000 during the competition and dropped useless features. My features are built using all kind of aggregations and statistics, feets lib, coeficients of linear regressions for each passband, coeficient  of linear regression after a peak and custom made features base on observation of the light curves, for example: raise rate of peaks and valeys, peak ratios of 'time to raive' vs 'time to decrease to X%', etc...\n\n#Models\nFor each model fit I augmented the train set concatenating noise versions of the train set. That simple augmentation improved my score by about 0.04 and decreased the CV x LB gap.\n\nModel 1:  multi-class LGB trained using all 250 features and all whole set.\n\nModel 2:  multi-class LGB trained using a subset of the 250 features, selected via feature selection, using whole train set.\n\nModel 3:  multi-class LGB trained using another subset of the 250 features, selected via other feature selection algorithm, using whole train set.\n\nModel 4:  multi-label LGB trained using a subset of the 250 features, selected via feature selection for each target individually, using whole train set.\n\nModel 5:  multi-label SVC using rbf kernel trained using all 250 features and whole train set.\n\nModel 6:  multi-class LGB trained using all 250 features. Fit one model for intra galaxy and other for extra galaxy.\n\nModel 7:  multi-label LGB trained using a subset of the 250 features, selected via feature selection for each target individually.  Fit one model for ddf==0 and other for ddf==1.\n\nStack ensemble 1, multi-class, models 1,2,3,4,5.\n\nStack ensemble 2, multi-class, model 6.\n\nStack ensemble 3, multi-class, model 7.\n\nBlend Stack Ensemble 1,2,3 using simple geometric weighted average.\n\nFor class 99 I ended using Scirpus(thanks @scirpus) equation plus some minor LB probings. \n\n\n\n\n#What didn't worked for me:\n\n - List item\n\n - FFT, wavelet and periodogram features.\n\n - DAE\n\n - Curve reconstruction\n\n - Fit a model for hoztgal_specz using train+test\n\n - CNN models ( metric &gt; 1.3 )\n\n - Target encoding\n\n - Flux correction by redshift\n\n - Flux passband correction using redshift\n\n - Semi-supervised Learning (improved just 0.0001)\n\n - SMOTE\n\n\nGiba",
      "votes": null
    },
    {
      "id": "442470",
      "postDate": "12/20/2018 02:16:55",
      "content": "<p>Thanks for sharing, and congrats!   The list of things you tried is quite impressive. I guess you would have easily been in gold with some NN added to your ensemble.</p>",
      "rawMarkdown": "Thanks for sharing, and congrats!   The list of things you tried is quite impressive. I guess you would have easily been in gold with some NN added to your ensemble.",
      "votes": null
    },
    {
      "id": "442480",
      "postDate": "12/20/2018 02:27:30",
      "content": "<p>Great as usual!</p>",
      "rawMarkdown": "Great as usual!",
      "votes": null
    },
    {
      "id": "442493",
      "postDate": "12/20/2018 02:40:12",
      "content": "<p>You're right, I missed a good NN model.</p>",
      "rawMarkdown": "You're right, I missed a good NN model.",
      "votes": null
    },
    {
      "id": "442577",
      "postDate": "12/20/2018 06:38:02",
      "content": "<p>This is cool. What kind of feature selection algorithms did you use?</p>",
      "rawMarkdown": "This is cool. What kind of feature selection algorithms did you use?",
      "votes": null
    },
    {
      "id": "442690",
      "postDate": "12/20/2018 10:44:48",
      "content": "<p>Thanks for sharing! What was the score for SVC model? Is it comparable with LGB?</p>",
      "rawMarkdown": "Thanks for sharing! What was the score for SVC model? Is it comparable with LGB?",
      "votes": null
    },
    {
      "id": "442745",
      "postDate": "12/20/2018 12:32:27",
      "content": "<p>I used RFECV and other I coded by myself.</p>",
      "rawMarkdown": "I used RFECV and other I coded by myself.",
      "votes": null
    },
    {
      "id": "442748",
      "postDate": "12/20/2018 12:36:07",
      "content": "<p>I trained the SVC as a multi-label task so score isn't possible to compare directly. But the rbf kernel performed very well in terms of AUC and ACC and SVMs are very robust against outliers and added around 0.01 to the final blend.</p>",
      "rawMarkdown": "I trained the SVC as a multi-label task so score isn't possible to compare directly. But the rbf kernel performed very well in terms of AUC and ACC and SVMs are very robust against outliers and added around 0.01 to the final blend.",
      "votes": null
    },
    {
      "id": "527064",
      "postDate": "05/04/2019 13:37:21",
      "content": "<p>Amazing Work! Thanks for sharing! I loved the stack ensemble part of your work. </p>",
      "rawMarkdown": "Amazing Work! Thanks for sharing! I loved the stack ensemble part of your work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442470,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/20/2018 02:16:55",
      "content": "<p>Thanks for sharing, and congrats!   The list of things you tried is quite impressive. I guess you would have easily been in gold with some NN added to your ensemble.</p>",
      "votes": null,
      "replies": [
        {
          "id": 442493,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "12/20/2018 02:40:12",
          "content": "<p>You're right, I missed a good NN model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442480,
      "author_name": "longyin2",
      "author_url": "",
      "post_date": "12/20/2018 02:27:30",
      "content": "<p>Great as usual!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442577,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "12/20/2018 06:38:02",
      "content": "<p>This is cool. What kind of feature selection algorithms did you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 442745,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "12/20/2018 12:32:27",
          "content": "<p>I used RFECV and other I coded by myself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442690,
      "author_name": "antklen",
      "author_url": "",
      "post_date": "12/20/2018 10:44:48",
      "content": "<p>Thanks for sharing! What was the score for SVC model? Is it comparable with LGB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 442748,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "12/20/2018 12:36:07",
          "content": "<p>I trained the SVC as a multi-label task so score isn't possible to compare directly. But the rbf kernel performed very well in terms of AUC and ACC and SVMs are very robust against outliers and added around 0.01 to the final blend.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527064,
      "author_name": "akhileshrai",
      "author_url": "",
      "post_date": "05/04/2019 13:37:21",
      "content": "<p>Amazing Work! Thanks for sharing! I loved the stack ensemble part of your work. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "442453": "First of all, thanks Kaggle and the sponsors for this great, clear and leak free competition :-)\n\nThis competition I started working hard in the beggining and practically stopped competing during last 10 days. My final solution is a stack ensemble and a blend of 4x LGB and 2x SVC models trained using different subsets of features. \n\n#Features\nI used some subsets of around 250 features for the final models, but I'm sure I tried more than 1000 during the competition and dropped useless features. My features are built using all kind of aggregations and statistics, feets lib, coeficients of linear regressions for each passband, coeficient  of linear regression after a peak and custom made features base on observation of the light curves, for example: raise rate of peaks and valeys, peak ratios of 'time to raive' vs 'time to decrease to X%', etc...\n\n#Models\nFor each model fit I augmented the train set concatenating noise versions of the train set. That simple augmentation improved my score by about 0.04 and decreased the CV x LB gap.\n\nModel 1:  multi-class LGB trained using all 250 features and all whole set.\n\nModel 2:  multi-class LGB trained using a subset of the 250 features, selected via feature selection, using whole train set.\n\nModel 3:  multi-class LGB trained using another subset of the 250 features, selected via other feature selection algorithm, using whole train set.\n\nModel 4:  multi-label LGB trained using a subset of the 250 features, selected via feature selection for each target individually, using whole train set.\n\nModel 5:  multi-label SVC using rbf kernel trained using all 250 features and whole train set.\n\nModel 6:  multi-class LGB trained using all 250 features. Fit one model for intra galaxy and other for extra galaxy.\n\nModel 7:  multi-label LGB trained using a subset of the 250 features, selected via feature selection for each target individually.  Fit one model for ddf==0 and other for ddf==1.\n\nStack ensemble 1, multi-class, models 1,2,3,4,5.\n\nStack ensemble 2, multi-class, model 6.\n\nStack ensemble 3, multi-class, model 7.\n\nBlend Stack Ensemble 1,2,3 using simple geometric weighted average.\n\nFor class 99 I ended using Scirpus(thanks @scirpus) equation plus some minor LB probings. \n\n\n\n\n#What didn't worked for me:\n\n - List item\n\n - FFT, wavelet and periodogram features.\n\n - DAE\n\n - Curve reconstruction\n\n - Fit a model for hoztgal_specz using train+test\n\n - CNN models ( metric &gt; 1.3 )\n\n - Target encoding\n\n - Flux correction by redshift\n\n - Flux passband correction using redshift\n\n - Semi-supervised Learning (improved just 0.0001)\n\n - SMOTE\n\n\nGiba",
    "442470": "Thanks for sharing, and congrats!   The list of things you tried is quite impressive. I guess you would have easily been in gold with some NN added to your ensemble.",
    "442480": "Great as usual!",
    "442493": "You're right, I missed a good NN model.",
    "442577": "This is cool. What kind of feature selection algorithms did you use?",
    "442690": "Thanks for sharing! What was the score for SVC model? Is it comparable with LGB?",
    "442745": "I used RFECV and other I coded by myself.",
    "442748": "I trained the SVC as a multi-label task so score isn't possible to compare directly. But the rbf kernel performed very well in terms of AUC and ACC and SVMs are very robust against outliers and added around 0.01 to the final blend.",
    "527064": "Amazing Work! Thanks for sharing! I loved the stack ensemble part of your work."
  },
  "source": "meta"
}