{
  "id": 70669,
  "title": "Hard classes",
  "url": "/competitions/PLAsTiCC-2018/discussion/70669",
  "author_name": "CPMP",
  "post_date": "2018-11-06T12:54:01.433000",
  "votes": 30,
  "comment_count": 26,
  "views": 0,
  "content": "<p>My best model has hard time separating classes 42, 52, 62, 67, and 90.  Separating these classes is the key to winning the competition, more than class 99 IMHO.  See the confusion matrix:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416235/10623/confmatrix.png\" alt=\"confusion matrix\"></p>",
  "messages": [
    {
      "id": 416235,
      "postDate": "2018-11-06T12:54:01.433Z",
      "content": "<p>My best model has hard time separating classes 42, 52, 62, 67, and 90.  Separating these classes is the key to winning the competition, more than class 99 IMHO.  See the confusion matrix:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416235/10623/confmatrix.png\" alt=\"confusion matrix\"></p>",
      "rawMarkdown": "My best model has hard time separating classes 42, 52, 62, 67, and 90.  Separating these classes is the key to winning the competition, more than class 99 IMHO.  See the confusion matrix:\n\n![confusion matrix][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416235/10623/confmatrix.png",
      "votes": 30
    },
    {
      "id": 416477,
      "postDate": "2018-11-06T17:39:36.820Z",
      "content": "<p>I also made the same observations. The classes 42, 52, 62, 67, and 90 seem to be different types of supernovae. (This seems to be an argument for not focusing on the calculation of periods.) Most publications I have seen at actually handle those.</p>",
      "rawMarkdown": "I also made the same observations. The classes 42, 52, 62, 67, and 90 seem to be different types of supernovae. (This seems to be an argument for not focusing on the calculation of periods.) Most publications I have seen at actually handle those.",
      "votes": 5,
      "replies": [
        {
          "id": 416493,
          "postDate": "2018-11-06T18:35:27.600Z",
          "content": "<blockquote>\n  <p>This seems to be an argument for not focusing on the calculation of periods.</p>\n</blockquote>\n\n<p>Good point!</p>",
          "rawMarkdown": "&gt; This seems to be an argument for not focusing on the calculation of periods.\n\nGood point!",
          "votes": 1
        },
        {
          "id": 416516,
          "postDate": "2018-11-06T19:43:38.323Z",
          "content": "<p>I believe (haven't studied in deep yet), that supernovae can be in principle differentiated using lightcurve analysis as well (well, not period, of course, but speed of brightness decay, \"half-width\" and similar parameters). It could be helpful, if we managed to link these classes with specific supernovae types (i.e. 42 = Ia, just example).</p>",
          "rawMarkdown": "I believe (haven't studied in deep yet), that supernovae can be in principle differentiated using lightcurve analysis as well (well, not period, of course, but speed of brightness decay, \"half-width\" and similar parameters). It could be helpful, if we managed to link these classes with specific supernovae types (i.e. 42 = Ia, just example).",
          "votes": 2
        },
        {
          "id": 416817,
          "postDate": "2018-11-07T10:13:38.437Z",
          "content": "<p>Period features (via LSSA) is basically curve fitting -&gt; use fitted parameters as features. The same idea can definitely be used for other types of objects if you know the typical shape of the light curves and have a nice family of functions to approximate them. Actually if we know more about how supernovae differ, we might not need to know which label is which, as long as a fitted curve can describe such differences.</p>",
          "rawMarkdown": "Period features (via LSSA) is basically curve fitting -&gt; use fitted parameters as features. The same idea can definitely be used for other types of objects if you know the typical shape of the light curves and have a nice family of functions to approximate them. Actually if we know more about how supernovae differ, we might not need to know which label is which, as long as a fitted curve can describe such differences."
        },
        {
          "id": 416823,
          "postDate": "2018-11-07T10:23:54.230Z",
          "content": "<p>Template fitting is indeed a technique used by astronomers AFAIK.  I have not given it a try, and I am not even sure where to start on this.</p>",
          "rawMarkdown": "Template fitting is indeed a technique used by astronomers AFAIK.  I have not given it a try, and I am not even sure where to start on this.",
          "votes": 1
        },
        {
          "id": 417141,
          "postDate": "2018-11-07T20:23:23.407Z",
          "content": "<p>I am experiencing the same issue where class 90 is swallowing the other classifications of the aperiodic classes (supernova/nova?) since we have so many more training examples of that class as compared to the other \"spiky\" ones.  As far as I can tell, several models can use high-level observation statistics (avg flux, max flux, std dev) to differentiate between the periodic classes and the episodic ones and discriminate well within the periodic classes.  However when it comes to the episodic ones, my models give up and just say \"dump it in class 90.\"</p>\n\n<p>The last public light curve competition (SNPhotCC) from the LSST group focused exclusively on supernova classification. The paper below lists some of the key findings/techniques from that competition.  As Michal and Mithrillion pointed out, previous contestants heavily relied on curve fitting and using the model parameters as features in the classifiers.  On page 8, they list a specific curve that the was used to generate the samples in the prior competition.</p>\n\n<p><a href=\"http://astro.uchicago.edu/~frieman/Courses/A411-old/References/supernova-photometric-classification-1008.1024.pdf\">http://astro.uchicago.edu/~frieman/Courses/A411-old/References/supernova-photometric-classification-1008.1024.pdf</a></p>\n\n<p>Furthermore, these parameters were sensitive to redshift so closer SN will will have different parameters than ones of the same type that are far away. </p>\n\n<p>Just like many people are using branching models on the Milky Way/Extragalatic feature, I've been trying to test for periodicity in the light curves and then have separate models for the periodic vs cataclysmic classes.  One problem I have been facing is that there are some periodic training examples that also have transient spikes/weakening the strength of periodicity tests.</p>",
          "rawMarkdown": "I am experiencing the same issue where class 90 is swallowing the other classifications of the aperiodic classes (supernova/nova?) since we have so many more training examples of that class as compared to the other \"spiky\" ones.  As far as I can tell, several models can use high-level observation statistics (avg flux, max flux, std dev) to differentiate between the periodic classes and the episodic ones and discriminate well within the periodic classes.  However when it comes to the episodic ones, my models give up and just say \"dump it in class 90.\"\n\nThe last public light curve competition (SNPhotCC) from the LSST group focused exclusively on supernova classification. The paper below lists some of the key findings/techniques from that competition.  As Michal and Mithrillion pointed out, previous contestants heavily relied on curve fitting and using the model parameters as features in the classifiers.  On page 8, they list a specific curve that the was used to generate the samples in the prior competition.\n\nhttp://astro.uchicago.edu/~frieman/Courses/A411-old/References/supernova-photometric-classification-1008.1024.pdf\n\nFurthermore, these parameters were sensitive to redshift so closer SN will will have different parameters than ones of the same type that are far away. \n\nJust like many people are using branching models on the Milky Way/Extragalatic feature, I've been trying to test for periodicity in the light curves and then have separate models for the periodic vs cataclysmic classes.  One problem I have been facing is that there are some periodic training examples that also have transient spikes/weakening the strength of periodicity tests.",
          "votes": 11
        },
        {
          "id": 417174,
          "postDate": "2018-11-07T22:14:16.433Z",
          "content": "<p>Thanks a lot for sharing this.</p>",
          "rawMarkdown": "Thanks a lot for sharing this."
        },
        {
          "id": 417353,
          "postDate": "2018-11-08T06:22:09.700Z",
          "content": "<blockquote>\n  <p>I am experiencing the same issue where class 90 is swallowing the other classifications of the aperiodic classes (supernova/nova?) since we have so many more training examples of that class as compared to the other \"spiky\" ones.</p>\n</blockquote>\n\n<p>I used the neural network model from Kernels and it seems to be right for that model. The result is more than 2 millions (out of 3.5 millions) of 'class 90' samples. I suspect this is wrong.\n@CPMP How many samples of 'class 90' do you have in your result? I used the maximum value in a row to determine the class for a sample.</p>",
          "rawMarkdown": "&gt; I am experiencing the same issue where class 90 is swallowing the other classifications of the aperiodic classes (supernova/nova?) since we have so many more training examples of that class as compared to the other \"spiky\" ones.\n\nI used the neural network model from Kernels and it seems to be right for that model. The result is more than 2 millions (out of 3.5 millions) of 'class 90' samples. I suspect this is wrong.\n@CPMP How many samples of 'class 90' do you have in your result? I used the maximum value in a row to determine the class for a sample."
        },
        {
          "id": 417561,
          "postDate": "2018-11-08T13:02:59.463Z",
          "content": "<blockquote>\n  <p>How many samples of 'class 90' do you have in your result?</p>\n</blockquote>\n\n<p>The confusion matrix is computed on out of fold predictions for all the training data.  I therefore have as many examples in class 90 as there are in training data.</p>",
          "rawMarkdown": "&gt; How many samples of 'class 90' do you have in your result?\n\nThe confusion matrix is computed on out of fold predictions for all the training data.  I therefore have as many examples in class 90 as there are in training data.",
          "votes": -2
        },
        {
          "id": 417576,
          "postDate": "2018-11-08T13:32:23.523Z",
          "content": "<p>I mean the number of 'class 90' samples in the submission file. I suspect there shouldn't be so many samples with high 'class 90' probability. </p>",
          "rawMarkdown": "I mean the number of 'class 90' samples in the submission file. I suspect there shouldn't be so many samples with high 'class 90' probability. "
        },
        {
          "id": 417586,
          "postDate": "2018-11-08T13:46:40.737Z",
          "content": "<p>Here are the mean probabilities for the submission file corresponding to that confusion matrix.  I did not compute classes as you suggest, as this is not how submissions are evaluated ;)</p>\n\n<pre><code>class_6      1.776892e-03\nclass_15     1.072527e-01\nclass_16     2.913544e-02\nclass_42     1.477831e-01\nclass_52     9.749433e-02\nclass_53     5.449143e-04\nclass_62     7.669120e-02\nclass_64     1.632178e-02\nclass_65     2.614934e-02\nclass_67     6.618325e-02\nclass_88     3.137809e-02\nclass_90     3.096742e-01\nclass_92     5.297100e-02\nclass_95     2.522340e-02\nclass_99     1.800000e-01\n</code></pre>",
          "rawMarkdown": "Here are the mean probabilities for the submission file corresponding to that confusion matrix.  I did not compute classes as you suggest, as this is not how submissions are evaluated ;)\n\n    class_6      1.776892e-03\n    class_15     1.072527e-01\n    class_16     2.913544e-02\n    class_42     1.477831e-01\n    class_52     9.749433e-02\n    class_53     5.449143e-04\n    class_62     7.669120e-02\n    class_64     1.632178e-02\n    class_65     2.614934e-02\n    class_67     6.618325e-02\n    class_88     3.137809e-02\n    class_90     3.096742e-01\n    class_92     5.297100e-02\n    class_95     2.522340e-02\n    class_99     1.800000e-01",
          "votes": 1
        },
        {
          "id": 417667,
          "postDate": "2018-11-08T15:53:10.293Z",
          "content": "<p>Thanks!</p>\n\n<p>Hmm.. So I checked it the same way as you and the mean values in the Kernel is ok. At least for 'class 90': 0.275285. And in the train set it's about 30% samples with it.</p>\n\n<p>Maybe the fractions of the classes in the train set and in the test set represent the real frequency in our Universe.</p>",
          "rawMarkdown": "Thanks!\n\nHmm.. So I checked it the same way as you and the mean values in the Kernel is ok. At least for 'class 90': 0.275285. And in the train set it's about 30% samples with it.\n\nMaybe the fractions of the classes in the train set and in the test set represent the real frequency in our Universe.\n",
          "votes": 1
        },
        {
          "id": 432209,
          "postDate": "2018-12-03T14:57:37.460Z",
          "content": "<p>Thanks! I will spend too much time in thinking of class 99 if I did not see this. Is that true that you just set all class 99 to be 0.18 in your case?</p>",
          "rawMarkdown": "Thanks! I will spend too much time in thinking of class 99 if I did not see this. Is that true that you just set all class 99 to be 0.18 in your case?"
        },
        {
          "id": 432619,
          "postDate": "2018-12-04T05:26:43.383Z",
          "content": "<blockquote>\n  <p>Is that true that you just set all class 99 to be 0.18 in your case?</p>\n</blockquote>\n\n<p>No.  I started from Olivier's method, see his public kernel.</p>",
          "rawMarkdown": "&gt; Is that true that you just set all class 99 to be 0.18 in your case?\n\nNo.  I started from Olivier's method, see his public kernel."
        }
      ]
    },
    {
      "id": 416411,
      "postDate": "2018-11-06T16:12:55.993Z",
      "content": "<p>I fully agree. My model is not that accurate yet but I face similar issues distinguishing classes 42, 52, 62, 67, and 90. My confusion matrix looks quite similar. \n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416411/10624/confusion_matrix.png\" alt=\"confusion matrix\"></p>",
      "rawMarkdown": "I fully agree. My model is not that accurate yet but I face similar issues distinguishing classes 42, 52, 62, 67, and 90. My confusion matrix looks quite similar. \n![confusion matrix][1]\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416411/10624/confusion_matrix.png",
      "votes": 2,
      "replies": [
        {
          "id": 416498,
          "postDate": "2018-11-06T18:40:34.393Z",
          "content": "<p>Very similar indeed.  Although mine is better in general, yours is better for some classes.  There is room for improvement for both of us!</p>",
          "rawMarkdown": "Very similar indeed.  Although mine is better in general, yours is better for some classes.  There is room for improvement for both of us!"
        }
      ]
    },
    {
      "id": 416357,
      "postDate": "2018-11-06T15:34:19.153Z",
      "content": "<p>I saw something similar (with a much simpler model; just a straightforward RandomForest); are you using any kind of oversampling process, like SMOTE? Although SMOTE and the like help, I suspect these oversampling techniques may end misleading the model if over abused by adding noise to the training data set, thus degrading the \"signal\" of other classes.</p>\n\n<p>Maybe a kind of hard negative mining may be key to improve these particular classes. Also, stacking generalization may also help.</p>",
      "rawMarkdown": "I saw something similar (with a much simpler model; just a straightforward RandomForest); are you using any kind of oversampling process, like SMOTE? Although SMOTE and the like help, I suspect these oversampling techniques may end misleading the model if over abused by adding noise to the training data set, thus degrading the \"signal\" of other classes.\n\nMaybe a kind of hard negative mining may be key to improve these particular classes. Also, stacking generalization may also help.",
      "votes": 2,
      "replies": [
        {
          "id": 416362,
          "postDate": "2018-11-06T15:41:10.197Z",
          "content": "<p>I am not using SMOTE. I never was able to use SMOTE successfully actually.</p>",
          "rawMarkdown": "I am not using SMOTE. I never was able to use SMOTE successfully actually."
        }
      ]
    },
    {
      "id": 418761,
      "postDate": "2018-11-10T15:10:13.767Z",
      "content": "<p>most of the object of these hard classes are being mistakenly classified as class90 </p>",
      "rawMarkdown": "most of the object of these hard classes are being mistakenly classified as class90 ",
      "replies": [
        {
          "id": 418767,
          "postDate": "2018-11-10T15:19:26.350Z",
          "content": "<p>Some, not most ;)</p>",
          "rawMarkdown": "Some, not most ;)"
        },
        {
          "id": 419023,
          "postDate": "2018-11-11T05:20:14.700Z",
          "content": "<p>i mean on my single lgbm model. most of my mistakes land on class 90.</p>",
          "rawMarkdown": "i mean on my single lgbm model. most of my mistakes land on class 90."
        },
        {
          "id": 419051,
          "postDate": "2018-11-11T06:26:39.377Z",
          "content": "<p>I agree most mistakes land on class 90. It is true for my model as well.</p>",
          "rawMarkdown": "I agree most mistakes land on class 90. It is true for my model as well."
        },
        {
          "id": 419533,
          "postDate": "2018-11-12T06:14:43.773Z",
          "content": "<p>do you know if 52, 67, and 90 are types of supernovae?</p>",
          "rawMarkdown": "do you know if 52, 67, and 90 are types of supernovae?"
        },
        {
          "id": 419598,
          "postDate": "2018-11-12T09:26:39.530Z",
          "content": "<p>I am not an astronomer, but they all have some kind of brightness burst, therefore I think some of these classes are supernova.  But there are other sources that can have a brightness burst AFAIK.</p>",
          "rawMarkdown": "I am not an astronomer, but they all have some kind of brightness burst, therefore I think some of these classes are supernova.  But there are other sources that can have a brightness burst AFAIK."
        },
        {
          "id": 419613,
          "postDate": "2018-11-12T09:44:00.563Z",
          "content": "<p>these classes are very difficult for me. i tried manual calculation of time to decay or how much decay after 15 days(<a href=\"http://astronomy.swin.edu.au/cosmos/T/Type+Ia+Supernova+Light+Curves\">http://astronomy.swin.edu.au/cosmos/T/Type+Ia+Supernova+Light+Curves</a>) but nothing improved my prediction. im stuck for 2 weeks now haha</p>",
          "rawMarkdown": "these classes are very difficult for me. i tried manual calculation of time to decay or how much decay after 15 days(http://astronomy.swin.edu.au/cosmos/T/Type+Ia+Supernova+Light+Curves) but nothing improved my prediction. im stuck for 2 weeks now haha"
        }
      ]
    },
    {
      "id": 420584,
      "postDate": "2018-11-13T21:10:15.107Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 416477,
      "author_name": "Helgi",
      "author_url": "",
      "post_date": "2018-11-06T17:39:36.820000",
      "content": "<p>I also made the same observations. The classes 42, 52, 62, 67, and 90 seem to be different types of supernovae. (This seems to be an argument for not focusing on the calculation of periods.) Most publications I have seen at actually handle those.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 416493,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-06T18:35:27.600000",
          "content": "<blockquote>\n  <p>This seems to be an argument for not focusing on the calculation of periods.</p>\n</blockquote>\n\n<p>Good point!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416516,
          "author_name": "Michal Haltuf",
          "author_url": "",
          "post_date": "2018-11-06T19:43:38.323000",
          "content": "<p>I believe (haven't studied in deep yet), that supernovae can be in principle differentiated using lightcurve analysis as well (well, not period, of course, but speed of brightness decay, \"half-width\" and similar parameters). It could be helpful, if we managed to link these classes with specific supernovae types (i.e. 42 = Ia, just example).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 416817,
          "author_name": "Mithrillion",
          "author_url": "",
          "post_date": "2018-11-07T10:13:38.437000",
          "content": "<p>Period features (via LSSA) is basically curve fitting -&gt; use fitted parameters as features. The same idea can definitely be used for other types of objects if you know the typical shape of the light curves and have a nice family of functions to approximate them. Actually if we know more about how supernovae differ, we might not need to know which label is which, as long as a fitted curve can describe such differences.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416823,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-07T10:23:54.230000",
          "content": "<p>Template fitting is indeed a technique used by astronomers AFAIK.  I have not given it a try, and I am not even sure where to start on this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 417141,
          "author_name": "Mike Holcomb",
          "author_url": "",
          "post_date": "2018-11-07T20:23:23.407000",
          "content": "<p>I am experiencing the same issue where class 90 is swallowing the other classifications of the aperiodic classes (supernova/nova?) since we have so many more training examples of that class as compared to the other \"spiky\" ones.  As far as I can tell, several models can use high-level observation statistics (avg flux, max flux, std dev) to differentiate between the periodic classes and the episodic ones and discriminate well within the periodic classes.  However when it comes to the episodic ones, my models give up and just say \"dump it in class 90.\"</p>\n\n<p>The last public light curve competition (SNPhotCC) from the LSST group focused exclusively on supernova classification. The paper below lists some of the key findings/techniques from that competition.  As Michal and Mithrillion pointed out, previous contestants heavily relied on curve fitting and using the model parameters as features in the classifiers.  On page 8, they list a specific curve that the was used to generate the samples in the prior competition.</p>\n\n<p><a href=\"http://astro.uchicago.edu/~frieman/Courses/A411-old/References/supernova-photometric-classification-1008.1024.pdf\">http://astro.uchicago.edu/~frieman/Courses/A411-old/References/supernova-photometric-classification-1008.1024.pdf</a></p>\n\n<p>Furthermore, these parameters were sensitive to redshift so closer SN will will have different parameters than ones of the same type that are far away. </p>\n\n<p>Just like many people are using branching models on the Milky Way/Extragalatic feature, I've been trying to test for periodicity in the light curves and then have separate models for the periodic vs cataclysmic classes.  One problem I have been facing is that there are some periodic training examples that also have transient spikes/weakening the strength of periodicity tests.</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 417174,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-07T22:14:16.433000",
          "content": "<p>Thanks a lot for sharing this.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417353,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-11-08T06:22:09.700000",
          "content": "<blockquote>\n  <p>I am experiencing the same issue where class 90 is swallowing the other classifications of the aperiodic classes (supernova/nova?) since we have so many more training examples of that class as compared to the other \"spiky\" ones.</p>\n</blockquote>\n\n<p>I used the neural network model from Kernels and it seems to be right for that model. The result is more than 2 millions (out of 3.5 millions) of 'class 90' samples. I suspect this is wrong.\n@CPMP How many samples of 'class 90' do you have in your result? I used the maximum value in a row to determine the class for a sample.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417561,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-08T13:02:59.463000",
          "content": "<blockquote>\n  <p>How many samples of 'class 90' do you have in your result?</p>\n</blockquote>\n\n<p>The confusion matrix is computed on out of fold predictions for all the training data.  I therefore have as many examples in class 90 as there are in training data.</p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 417576,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-11-08T13:32:23.523000",
          "content": "<p>I mean the number of 'class 90' samples in the submission file. I suspect there shouldn't be so many samples with high 'class 90' probability. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417586,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-08T13:46:40.737000",
          "content": "<p>Here are the mean probabilities for the submission file corresponding to that confusion matrix.  I did not compute classes as you suggest, as this is not how submissions are evaluated ;)</p>\n\n<pre><code>class_6      1.776892e-03\nclass_15     1.072527e-01\nclass_16     2.913544e-02\nclass_42     1.477831e-01\nclass_52     9.749433e-02\nclass_53     5.449143e-04\nclass_62     7.669120e-02\nclass_64     1.632178e-02\nclass_65     2.614934e-02\nclass_67     6.618325e-02\nclass_88     3.137809e-02\nclass_90     3.096742e-01\nclass_92     5.297100e-02\nclass_95     2.522340e-02\nclass_99     1.800000e-01\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 417667,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-11-08T15:53:10.293000",
          "content": "<p>Thanks!</p>\n\n<p>Hmm.. So I checked it the same way as you and the mean values in the Kernel is ok. At least for 'class 90': 0.275285. And in the train set it's about 30% samples with it.</p>\n\n<p>Maybe the fractions of the classes in the train set and in the test set represent the real frequency in our Universe.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 432209,
          "author_name": "Ningxiao Zhang",
          "author_url": "",
          "post_date": "2018-12-03T14:57:37.460000",
          "content": "<p>Thanks! I will spend too much time in thinking of class 99 if I did not see this. Is that true that you just set all class 99 to be 0.18 in your case?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 432619,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-04T05:26:43.383000",
          "content": "<blockquote>\n  <p>Is that true that you just set all class 99 to be 0.18 in your case?</p>\n</blockquote>\n\n<p>No.  I started from Olivier's method, see his public kernel.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416411,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2018-11-06T16:12:55.993000",
      "content": "<p>I fully agree. My model is not that accurate yet but I face similar issues distinguishing classes 42, 52, 62, 67, and 90. My confusion matrix looks quite similar. \n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416411/10624/confusion_matrix.png\" alt=\"confusion matrix\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 416498,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-06T18:40:34.393000",
          "content": "<p>Very similar indeed.  Although mine is better in general, yours is better for some classes.  There is room for improvement for both of us!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416357,
      "author_name": "Andreu Sancho-Asensio",
      "author_url": "",
      "post_date": "2018-11-06T15:34:19.153000",
      "content": "<p>I saw something similar (with a much simpler model; just a straightforward RandomForest); are you using any kind of oversampling process, like SMOTE? Although SMOTE and the like help, I suspect these oversampling techniques may end misleading the model if over abused by adding noise to the training data set, thus degrading the \"signal\" of other classes.</p>\n\n<p>Maybe a kind of hard negative mining may be key to improve these particular classes. Also, stacking generalization may also help.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 416362,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-06T15:41:10.197000",
          "content": "<p>I am not using SMOTE. I never was able to use SMOTE successfully actually.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 418761,
      "author_name": "dylonLL",
      "author_url": "",
      "post_date": "2018-11-10T15:10:13.767000",
      "content": "<p>most of the object of these hard classes are being mistakenly classified as class90 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 418767,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-10T15:19:26.350000",
          "content": "<p>Some, not most ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419023,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-11-11T05:20:14.700000",
          "content": "<p>i mean on my single lgbm model. most of my mistakes land on class 90.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419051,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-11T06:26:39.377000",
          "content": "<p>I agree most mistakes land on class 90. It is true for my model as well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419533,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-11-12T06:14:43.773000",
          "content": "<p>do you know if 52, 67, and 90 are types of supernovae?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419598,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-12T09:26:39.530000",
          "content": "<p>I am not an astronomer, but they all have some kind of brightness burst, therefore I think some of these classes are supernova.  But there are other sources that can have a brightness burst AFAIK.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419613,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-11-12T09:44:00.563000",
          "content": "<p>these classes are very difficult for me. i tried manual calculation of time to decay or how much decay after 15 days(<a href=\"http://astronomy.swin.edu.au/cosmos/T/Type+Ia+Supernova+Light+Curves\">http://astronomy.swin.edu.au/cosmos/T/Type+Ia+Supernova+Light+Curves</a>) but nothing improved my prediction. im stuck for 2 weeks now haha</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 420584,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-13T21:10:15.107000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "416235": "My best model has hard time separating classes 42, 52, 62, 67, and 90.  Separating these classes is the key to winning the competition, more than class 99 IMHO.  See the confusion matrix:\n\n![confusion matrix][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416235/10623/confmatrix.png",
    "416477": "I also made the same observations. The classes 42, 52, 62, 67, and 90 seem to be different types of supernovae. (This seems to be an argument for not focusing on the calculation of periods.) Most publications I have seen at actually handle those.",
    "416411": "I fully agree. My model is not that accurate yet but I face similar issues distinguishing classes 42, 52, 62, 67, and 90. My confusion matrix looks quite similar. \n![confusion matrix][1]\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416411/10624/confusion_matrix.png",
    "416357": "I saw something similar (with a much simpler model; just a straightforward RandomForest); are you using any kind of oversampling process, like SMOTE? Although SMOTE and the like help, I suspect these oversampling techniques may end misleading the model if over abused by adding noise to the training data set, thus degrading the \"signal\" of other classes.\n\nMaybe a kind of hard negative mining may be key to improve these particular classes. Also, stacking generalization may also help.",
    "418761": "most of the object of these hard classes are being mistakenly classified as class90 ",
    "420584": ""
  }
}