{
  "id": 68943,
  "title": "Predicting class 99",
  "url": "/competitions/PLAsTiCC-2018/discussion/68943",
  "author_name": "olivier",
  "post_date": "2018-10-18T17:21:01.405000",
  "votes": 50,
  "comment_count": 38,
  "views": 0,
  "content": "<p>I've tried so many things here that I'm not sure I can list them all. I must admit I'm puzzled as to how to derive class 99 probability from the training set. It's the very first time I'm facing a similar problem.</p>\n\n<p>Here are the things I've tried: </p>\n\n<ul>\n<li>Use a constant</li>\n<li>Going to logits, derive individual probabilities from that, compute p99 as the product of opposite class probabilities, go back to logit and use softmax to normalize (probably the worst result)</li>\n<li>Computing the probability as the opposite of the max proba over all classes for each sample</li>\n<li>Calculate p99 as the product of the other classes opposite proba, going to logits and using softmax to normalize all probanilities</li>\n<li>Used <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67205\">IsolationForest</a> but still have issues merging that with the other classes proba. I've found anomalies are way to high in my opinion.</li>\n</ul>\n\n<p>The best I've found this far is to use the product of all opposite probabilities and adapting the mean but I'm not satisfied with that since it kills the other classes logloss.</p>\n\n<p>If anyone is willing to share his ideas I would be gratefull ;-)</p>",
  "messages": [
    {
      "id": 406121,
      "postDate": "2018-10-18T17:21:01.407Z",
      "content": "<p>I've tried so many things here that I'm not sure I can list them all. I must admit I'm puzzled as to how to derive class 99 probability from the training set. It's the very first time I'm facing a similar problem.</p>\n\n<p>Here are the things I've tried: </p>\n\n<ul>\n<li>Use a constant</li>\n<li>Going to logits, derive individual probabilities from that, compute p99 as the product of opposite class probabilities, go back to logit and use softmax to normalize (probably the worst result)</li>\n<li>Computing the probability as the opposite of the max proba over all classes for each sample</li>\n<li>Calculate p99 as the product of the other classes opposite proba, going to logits and using softmax to normalize all probanilities</li>\n<li>Used <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67205\">IsolationForest</a> but still have issues merging that with the other classes proba. I've found anomalies are way to high in my opinion.</li>\n</ul>\n\n<p>The best I've found this far is to use the product of all opposite probabilities and adapting the mean but I'm not satisfied with that since it kills the other classes logloss.</p>\n\n<p>If anyone is willing to share his ideas I would be gratefull ;-)</p>",
      "rawMarkdown": "I've tried so many things here that I'm not sure I can list them all. I must admit I'm puzzled as to how to derive class 99 probability from the training set. It's the very first time I'm facing a similar problem.\n\nHere are the things I've tried: \n\n - Use a constant\n - Going to logits, derive individual probabilities from that, compute p99 as the product of opposite class probabilities, go back to logit and use softmax to normalize (probably the worst result)\n - Computing the probability as the opposite of the max proba over all classes for each sample\n - Calculate p99 as the product of the other classes opposite proba, going to logits and using softmax to normalize all probanilities\n - Used [IsolationForest](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67205) but still have issues merging that with the other classes proba. I've found anomalies are way to high in my opinion.\n\nThe best I've found this far is to use the product of all opposite probabilities and adapting the mean but I'm not satisfied with that since it kills the other classes logloss.\n\nIf anyone is willing to share his ideas I would be gratefull ;-)\n  ",
      "votes": 50
    },
    {
      "id": 410767,
      "postDate": "2018-10-26T16:15:57.050Z",
      "content": "<p>I believe a prerequisite for predicting Class 99 is having a very good model for the other classes.\nThe good new is that a score of ~1.0 is achievable without really dealing with class 99. \nThe only thing I did up to this point with this prediction is:\nclass_99=np.where(other_classes.max&gt;0.9 , 0.01, 0.1)\n[it is slightly better then uniform, If I use 0.8 the score degrades, and also other values of probabilities degrade the score]</p>",
      "rawMarkdown": "I believe a prerequisite for predicting Class 99 is having a very good model for the other classes.\nThe good new is that a score of ~1.0 is achievable without really dealing with class 99. \nThe only thing I did up to this point with this prediction is:\nclass_99=np.where(other_classes.max&gt;0.9 , 0.01, 0.1)\n[it is slightly better then uniform, If I use 0.8 the score degrades, and also other values of probabilities degrade the score]",
      "votes": 11,
      "replies": [
        {
          "id": 411329,
          "postDate": "2018-10-27T21:27:40.297Z",
          "content": "<p>I tried your way, it is worse than Olivier's for my last sub, 1.069 vs 1.05.</p>\n\n<p>I'll stick to Olivier's way for now ;)</p>",
          "rawMarkdown": "I tried your way, it is worse than Olivier's for my last sub, 1.069 vs 1.05.\n\nI'll stick to Olivier's way for now ;)"
        },
        {
          "id": 411509,
          "postDate": "2018-10-28T10:46:43.497Z",
          "content": "<p>It is probably model defendant. A small tweak on this calculation, just improved my LB by another 0.07.</p>",
          "rawMarkdown": "It is probably model defendant. A small tweak on this calculation, just improved my LB by another 0.07."
        },
        {
          "id": 411512,
          "postDate": "2018-10-28T10:52:36.313Z",
          "content": "<blockquote>\n  <p>It is probably model defendant.</p>\n</blockquote>\n\n<p>I'll focus on that when I'll be out of ideas for predicting the known classes.  But yes, I agree this has to be tuned for each model/.</p>",
          "rawMarkdown": "&gt; It is probably model defendant.\n\nI'll focus on that when I'll be out of ideas for predicting the known classes.  But yes, I agree this has to be tuned for each model/.",
          "votes": 1
        }
      ]
    },
    {
      "id": 406122,
      "postDate": "2018-10-18T17:32:16.567Z",
      "content": "<p>Hello Olivier,</p>\n\n<p>By \"inverse\" I assumed you're referring to <code>1 - x</code> and not <code>1 / x</code>. If so I think it's clearer to use the term \"opposite\".</p>\n\n<p>Personally I'm using <code>1 - max(P)</code> where <code>P</code> is the distribution of probabilities for the other classes. I guess this corresponds to the third approach you mentioned. I find this intuitive since essentially it produces high values for \"flat\" distributions where the model isn't sure on any of the known classes. Maybe this idea could be pushed further by looking at the skew/kurtosis of the predicted class distribution.</p>\n\n<p>Anyway, just my 2 cents. </p>",
      "rawMarkdown": "Hello Olivier,\n\nBy \"inverse\" I assumed you're referring to `1 - x` and not `1 / x`. If so I think it's clearer to use the term \"opposite\".\n\nPersonally I'm using `1 - max(P)` where `P` is the distribution of probabilities for the other classes. I guess this corresponds to the third approach you mentioned. I find this intuitive since essentially it produces high values for \"flat\" distributions where the model isn't sure on any of the known classes. Maybe this idea could be pushed further by looking at the skew/kurtosis of the predicted class distribution.\n\nAnyway, just my 2 cents. ",
      "votes": 11,
      "replies": [
        {
          "id": 406138,
          "postDate": "2018-10-18T18:06:07.493Z",
          "content": "<p>Thanks for sharing Max. </p>\n\n<p>Yeah I thought 1 - max(P) would give better results as it was, as you said, very intuitive. Maybe I'm not good enough yet for other classes ;-)</p>",
          "rawMarkdown": "Thanks for sharing Max. \n\nYeah I thought 1 - max(P) would give better results as it was, as you said, very intuitive. Maybe I'm not good enough yet for other classes ;-)",
          "votes": 2
        },
        {
          "id": 406143,
          "postDate": "2018-10-18T18:16:20.473Z",
          "content": "<p>Indeed, the predictions for <code>class_99</code> will naturally get better with better predictions for the other classes.</p>",
          "rawMarkdown": "Indeed, the predictions for `class_99` will naturally get better with better predictions for the other classes.",
          "votes": 3
        },
        {
          "id": 407819,
          "postDate": "2018-10-21T20:38:34.753Z",
          "content": "<p>Thanks for the insightful comments. Agreed, 1-max(P) sounds very promising. However, at the moment, I find the same as Olivier. For me, 1-max(p) was worse than the very crude p(class_99)=0.2.</p>\n\n<p>I have found that using 1-max(p) when max(p) is high (say &gt;0.8) works well to reduce p(class_99) when the classifier is confident about another class, but when max(p) is low this seems to give too much weight to class_99. I think I'll have to revisit this when my base classifier is better!</p>",
          "rawMarkdown": "Thanks for the insightful comments. Agreed, 1-max(P) sounds very promising. However, at the moment, I find the same as Olivier. For me, 1-max(p) was worse than the very crude p(class_99)=0.2.\n\nI have found that using 1-max(p) when max(p) is high (say &gt;0.8) works well to reduce p(class_99) when the classifier is confident about another class, but when max(p) is low this seems to give too much weight to class_99. I think I'll have to revisit this when my base classifier is better!",
          "votes": 3
        },
        {
          "id": 408344,
          "postDate": "2018-10-22T17:57:38.090Z",
          "content": "<p>Thanks Andy. I have not tried <code>1 - np.max(probas)</code> lately with my new score. I may have a go at it and burn a submission to check if I can get anything better than my current process described in the public kernel.</p>",
          "rawMarkdown": "Thanks Andy. I have not tried `1 - np.max(probas)` lately with my new score. I may have a go at it and burn a submission to check if I can get anything better than my current process described in the public kernel."
        }
      ]
    },
    {
      "id": 411143,
      "postDate": "2018-10-27T13:42:28.737Z",
      "content": "<p>My LB probing experiments, up to now, showed me that distribution of class_99 Galactic is around 10x smaller than class_99 ExtraGalactic.  Hard coding class_99 of Galactic and ExtraGalactic to constants 0.017 and 0.17 improved my scores.</p>\n\n<p>Anyone did any other probing experiments?</p>",
      "rawMarkdown": "My LB probing experiments, up to now, showed me that distribution of class_99 Galactic is around 10x smaller than class_99 ExtraGalactic.  Hard coding class_99 of Galactic and ExtraGalactic to constants 0.017 and 0.17 improved my scores.\n\nAnyone did any other probing experiments?",
      "votes": 10,
      "replies": [
        {
          "id": 411260,
          "postDate": "2018-10-27T18:10:26.903Z",
          "content": "<p>I've got almost the same result. I took 0.015 and 0.15 - this gave me the best probe (2.081).</p>\n\n<p>I also probed </p>\n\n<p>0.020/0.147 -&gt; 2.081</p>\n\n<p>0.010/0.151 -&gt; 2.081</p>\n\n<p>0/0.154 -&gt; 2.191</p>\n\n<p>0.038/0.141 -&gt; 2.084</p>\n\n<p>0.106/0.113 - &gt; 2.105</p>\n\n<p>0.167/0.154 -&gt; 2.118</p>\n\n<p>I made a kernel about this here: <a href=\"https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081\">https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081</a></p>",
          "rawMarkdown": "I've got almost the same result. I took 0.015 and 0.15 - this gave me the best probe (2.081).\n\nI also probed \n\n0.020/0.147 -&gt; 2.081\n\n0.010/0.151 -&gt; 2.081\n\n0/0.154 -&gt; 2.191\n\n0.038/0.141 -&gt; 2.084\n\n0.106/0.113 - &gt; 2.105\n\n0.167/0.154 -&gt; 2.118\n\nI made a kernel about this here: https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081",
          "votes": 4
        },
        {
          "id": 411287,
          "postDate": "2018-10-27T19:32:36.867Z",
          "content": "<p>I tried your two constants, and it is a bit worse than Olivier's way for my latest sub, 1.057 vs 1.052.  Maybe the constant should be tuned for a given sub.</p>",
          "rawMarkdown": "I tried your two constants, and it is a bit worse than Olivier's way for my latest sub, 1.057 vs 1.052.  Maybe the constant should be tuned for a given sub.",
          "votes": 1
        }
      ]
    },
    {
      "id": 406406,
      "postDate": "2018-10-19T07:37:55.637Z",
      "content": "<p>Most of the class_99 objects (~95% - 97.5%) are placed in the Extragalactic group (out of the Milky Way).</p>\n\n<p>I tried public LB multiple times to detect this. Check the kernel: <a href=\"https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081\">https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081</a></p>",
      "rawMarkdown": "Most of the class_99 objects (~95% - 97.5%) are placed in the Extragalactic group (out of the Milky Way).\n\nI tried public LB multiple times to detect this. Check the kernel: https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081",
      "votes": 9
    },
    {
      "id": 406189,
      "postDate": "2018-10-18T20:06:00.520Z",
      "content": "<p>Hi,</p>\n\n<p>I've implemented my own version of one vs all classifier, that trains any scikit model of my choice (and probably of other pkgs with similar interface). Once I get probabilities for all classes, I set P(99)=1-sum(P(.)) and adjust it to zero when it gets negative. Though I haven't yet submitted results and don't know how am I performing. </p>",
      "rawMarkdown": "Hi,\n\nI've implemented my own version of one vs all classifier, that trains any scikit model of my choice (and probably of other pkgs with similar interface). Once I get probabilities for all classes, I set P(99)=1-sum(P(.)) and adjust it to zero when it gets negative. Though I haven't yet submitted results and don't know how am I performing. ",
      "votes": 5,
      "replies": [
        {
          "id": 406355,
          "postDate": "2018-10-19T05:38:18.890Z",
          "content": "<p>Thanks for sharing Linards Kalvāns. </p>\n\n<p>I had a look at that using lightgbm raw scores <code>lgb.predict(raw_score=True)</code> and sum of individual proba can go up to around 5. </p>\n\n<p>Hope you'll hav more chance with that !</p>",
          "rawMarkdown": "Thanks for sharing Linards Kalvāns. \n\nI had a look at that using lightgbm raw scores `lgb.predict(raw_score=True)` and sum of individual proba can go up to around 5. \n\nHope you'll hav more chance with that !\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 406165,
      "postDate": "2018-10-18T18:50:24.710Z",
      "content": "<p>One more:</p>\n\n<ul>\n<li>probe the Public LB :-)</li>\n</ul>",
      "rawMarkdown": "One more:\n\n - probe the Public LB :-)",
      "votes": 4,
      "replies": [
        {
          "id": 406177,
          "postDate": "2018-10-18T19:29:48.783Z",
          "content": "<p>Thanks Giba but I must be too silly for that :) I still don't get how to do that...</p>",
          "rawMarkdown": "Thanks Giba but I must be too silly for that :) I still don't get how to do that...",
          "votes": 2
        },
        {
          "id": 406410,
          "postDate": "2018-10-19T07:46:13.967Z",
          "content": "<p>This way becomes universal for that competition :)</p>",
          "rawMarkdown": "This way becomes universal for that competition :)",
          "votes": 2
        },
        {
          "id": 408586,
          "postDate": "2018-10-23T05:51:38.323Z",
          "content": "<p>It is consider as common practice now a days ; )</p>",
          "rawMarkdown": "It is consider as common practice now a days ; )",
          "votes": -1
        }
      ]
    },
    {
      "id": 409628,
      "postDate": "2018-10-24T16:06:13.033Z",
      "content": "<p>Just another idea: add noise (random) rows to training data labeled as class 99.</p>\n\n<p>I haven't had time to try it yet, so don't known if it would help or not. But as the aim of class 99 is to identify new objects not seen previously, it can be considered heterogeneous data, so random features from a correct distribution could help to emulate the heterogeneity.</p>",
      "rawMarkdown": "Just another idea: add noise (random) rows to training data labeled as class 99.\n\nI haven't had time to try it yet, so don't known if it would help or not. But as the aim of class 99 is to identify new objects not seen previously, it can be considered heterogeneous data, so random features from a correct distribution could help to emulate the heterogeneity.",
      "votes": 3,
      "replies": [
        {
          "id": 410659,
          "postDate": "2018-10-26T12:18:08.077Z",
          "content": "<p>Awesome idea! But it might be not as simple as it seems to be since we will probably need to guess what the noise really is and how the noise should look like... But the idea is brilliant and maybe there is a way to build an interesting model!</p>",
          "rawMarkdown": "Awesome idea! But it might be not as simple as it seems to be since we will probably need to guess what the noise really is and how the noise should look like... But the idea is brilliant and maybe there is a way to build an interesting model!",
          "votes": 2
        }
      ]
    },
    {
      "id": 408951,
      "postDate": "2018-10-23T16:32:14.813Z",
      "content": "<p>Clustering with DBScan automatically labels rejects as label == -1.  It may be worth using a lookup table to convert dbscan labels to probabilities</p>",
      "rawMarkdown": "Clustering with DBScan automatically labels rejects as label == -1.  It may be worth using a lookup table to convert dbscan labels to probabilities",
      "votes": 4,
      "replies": [
        {
          "id": 408977,
          "postDate": "2018-10-23T17:19:17.917Z",
          "content": "<p>Thanks Scirpus, that's on my todo list... </p>",
          "rawMarkdown": "Thanks Scirpus, that's on my todo list... ",
          "votes": 1
        },
        {
          "id": 409318,
          "postDate": "2018-10-24T05:34:50.310Z",
          "content": "<p>@Scirpus</p>\n\n<p>I did some search with DBSAN with no real success... I will give it a new try . Thanks.</p>",
          "rawMarkdown": "@Scirpus\n\nI did some search with DBSAN with no real success... I will give it a new try . Thanks.",
          "votes": 1
        }
      ]
    },
    {
      "id": 408374,
      "postDate": "2018-10-22T18:42:53.553Z",
      "content": "<p>To approach this unseen class issue, I think that I'm going to try to layer on top a <a href=\"http://scikit-learn.org/stable/auto_examples/svm/plot_oneclass.html#sphx-glr-auto-examples-svm-plot-oneclass-py\">One-Class SVM</a> on the training data.  When predicting, I'll first try to detect whether the test example is in the same \"class\" as the \"training class.\"  If not, label it as class 99.  If it is, use a full 14-class classifier trained on the seen classes in the training data.​</p>",
      "rawMarkdown": "To approach this unseen class issue, I think that I'm going to try to layer on top a [One-Class SVM][1] on the training data.  When predicting, I'll first try to detect whether the test example is in the same \"class\" as the \"training class.\"  If not, label it as class 99.  If it is, use a full 14-class classifier trained on the seen classes in the training data.​\n\n  [1]: http://scikit-learn.org/stable/auto_examples/svm/plot_oneclass.html#sphx-glr-auto-examples-svm-plot-oneclass-py \"One-Class SVM\"",
      "votes": 4,
      "replies": [
        {
          "id": 408380,
          "postDate": "2018-10-22T18:49:58.147Z",
          "content": "<p>Interesting step!</p>",
          "rawMarkdown": "Interesting step!"
        },
        {
          "id": 408636,
          "postDate": "2018-10-23T07:41:57.620Z",
          "content": "<p>I tried that with the <em>IsolationForest</em> and I found very weird stuff; moreover: a t-SNE of the training set and another using a large sample of the test set (roughly 6000 rows taken at random) suggest that these two groups come from distinct distributions... I also tried a t-SNE using a concatenation of both the training and the test sets, and this anomaly still appear in the data... In conclusion: it may happen that one-class SVM \"fails\" at finding the \"true\" class_99 objects due to this inconsistency; as if a preprocessing (e.g., maybe a kind of a <em>renormalization</em> in the training or test sets) was necessary prior to doing anything else.</p>\n\n<p>Note that both the train and test sets used in my analysis included the time series data, conveniently characterized using about 950 features per time series (for example: skewness per band, entropy, median, and so on).</p>",
          "rawMarkdown": "I tried that with the *IsolationForest* and I found very weird stuff; moreover: a t-SNE of the training set and another using a large sample of the test set (roughly 6000 rows taken at random) suggest that these two groups come from distinct distributions... I also tried a t-SNE using a concatenation of both the training and the test sets, and this anomaly still appear in the data... In conclusion: it may happen that one-class SVM \"fails\" at finding the \"true\" class_99 objects due to this inconsistency; as if a preprocessing (e.g., maybe a kind of a *renormalization* in the training or test sets) was necessary prior to doing anything else.\n\nNote that both the train and test sets used in my analysis included the time series data, conveniently characterized using about 950 features per time series (for example: skewness per band, entropy, median, and so on).",
          "votes": 4
        },
        {
          "id": 408793,
          "postDate": "2018-10-23T13:17:40.933Z",
          "content": "<p>Interesting find.  After you mentioned that, this part of the data note sticks out to me.</p>\n\n<blockquote>\n  <p><strong>3.2. Training and test data</strong></p>\n  \n  <p>Moreover, the training data properties are non-representative of distributions of the the test[sic] data set. The training data are mostly composed of nearby, low-redshift, brighter objects while the test data contain more distant (higher redshift) and fainter objects.</p>\n</blockquote>\n\n<p>Your findings and Ilya Khristoforov's kernel are consistent with this.  Maybe developing a One-Class SVM for each category (Milky Way vs. Extragalactic) can help...</p>",
          "rawMarkdown": "Interesting find.  After you mentioned that, this part of the data note sticks out to me.\n\n&gt;  **3.2. Training and test data**\n\n&gt; Moreover, the training data properties are non-representative of distributions of the the test[sic] data set. The training data are mostly composed of nearby, low-redshift, brighter objects while the test data contain more distant (higher redshift) and fainter objects.\n\nYour findings and Ilya Khristoforov's kernel are consistent with this.  Maybe developing a One-Class SVM for each category (Milky Way vs. Extragalactic) can help...",
          "votes": 5
        },
        {
          "id": 408807,
          "postDate": "2018-10-23T13:41:48.957Z",
          "content": "<p>That's a great idea actually; I'll try the same analysis with the Milky Way vs. Extragalactic and see what happens. I'll post here the findings.</p>",
          "rawMarkdown": "That's a great idea actually; I'll try the same analysis with the Milky Way vs. Extragalactic and see what happens. I'll post here the findings.",
          "votes": 4
        }
      ]
    },
    {
      "id": 410609,
      "postDate": "2018-10-26T10:43:29.417Z",
      "content": "<p>Maybe one class randomforest ? (<a href=\"https://github.com/ngoix/OCRF\">https://github.com/ngoix/OCRF</a>)\nI try  my own version (some kind of mixed lightgbm in this algorithm)\nbut it didn't work , LB is terrible :)</p>",
      "rawMarkdown": "Maybe one class randomforest ? (https://github.com/ngoix/OCRF)\nI try  my own version (some kind of mixed lightgbm in this algorithm)\nbut it didn't work , LB is terrible :)",
      "votes": 1
    },
    {
      "id": 410576,
      "postDate": "2018-10-26T09:39:38.080Z",
      "content": "<p>Olivier, </p>\n\n<p>In my limited experience, the way you compute it in your kernel is a bit better than using a constant 1/9 prediction for class 99.   LB gain is about 0.02 for me.</p>",
      "rawMarkdown": "Olivier, \n\nIn my limited experience, the way you compute it in your kernel is a bit better than using a constant 1/9 prediction for class 99.   LB gain is about 0.02 for me.",
      "votes": 1,
      "replies": [
        {
          "id": 410693,
          "postDate": "2018-10-26T13:32:33.493Z",
          "content": "<p>@CPMP, yes that's what I found. Using different averages for galactic and extragalactic and prod of opposite proba gives the best results for me.</p>",
          "rawMarkdown": "@CPMP, yes that's what I found. Using different averages for galactic and extragalactic and prod of opposite proba gives the best results for me.",
          "votes": 1
        }
      ]
    },
    {
      "id": 410092,
      "postDate": "2018-10-25T11:53:01.113Z",
      "content": "<p>How about connecting 99 to flux rows where detected==0\nIf a flux row is detected then it must be one of the training targets if it isn't then it must be 99 by definition.</p>",
      "rawMarkdown": "How about connecting 99 to flux rows where detected==0\nIf a flux row is detected then it must be one of the training targets if it isn't then it must be 99 by definition.",
      "votes": 1
    },
    {
      "id": 409344,
      "postDate": "2018-10-24T06:37:02.297Z",
      "content": "<p>Seems LB probing will play a role here.</p>",
      "rawMarkdown": "Seems LB probing will play a role here.",
      "votes": 1,
      "replies": [
        {
          "id": 409360,
          "postDate": "2018-10-24T07:17:28.470Z",
          "content": "<p>Unfortunately I wholeheartedly agree with you\nOne other gaming approach would be to make sure that the predictions are not too confident.  A 1 will be punished heavily so I am going to reduce the predictions between a min and a max before row normalizing</p>",
          "rawMarkdown": "Unfortunately I wholeheartedly agree with you\nOne other gaming approach would be to make sure that the predictions are not too confident.  A 1 will be punished heavily so I am going to reduce the predictions between a min and a max before row normalizing",
          "votes": 2
        }
      ]
    },
    {
      "id": 406439,
      "postDate": "2018-10-19T08:44:25.033Z",
      "content": "<p>I'm still very far from it, I have also tried auto encoders and distance error but it did not do significant improvement to my score.</p>",
      "rawMarkdown": "I'm still very far from it, I have also tried auto encoders and distance error but it did not do significant improvement to my score.",
      "votes": 1
    },
    {
      "id": 406187,
      "postDate": "2018-10-18T20:02:08.690Z",
      "content": "<p>@ogreiller</p>\n\n<p>Thanks for raising this issue I have too, and thanks to all for the insightfull replies.</p>",
      "rawMarkdown": "@ogreiller\n\nThanks for raising this issue I have too, and thanks to all for the insightfull replies.",
      "votes": 1
    },
    {
      "id": 408351,
      "postDate": "2018-10-22T18:07:41.940Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 410767,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2018-10-26T16:15:57.050000",
      "content": "<p>I believe a prerequisite for predicting Class 99 is having a very good model for the other classes.\nThe good new is that a score of ~1.0 is achievable without really dealing with class 99. \nThe only thing I did up to this point with this prediction is:\nclass_99=np.where(other_classes.max&gt;0.9 , 0.01, 0.1)\n[it is slightly better then uniform, If I use 0.8 the score degrades, and also other values of probabilities degrade the score]</p>",
      "votes": 11,
      "replies": [
        {
          "id": 411329,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-27T21:27:40.297000",
          "content": "<p>I tried your way, it is worse than Olivier's for my last sub, 1.069 vs 1.05.</p>\n\n<p>I'll stick to Olivier's way for now ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 411509,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-10-28T10:46:43.497000",
          "content": "<p>It is probably model defendant. A small tweak on this calculation, just improved my LB by another 0.07.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 411512,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-28T10:52:36.313000",
          "content": "<blockquote>\n  <p>It is probably model defendant.</p>\n</blockquote>\n\n<p>I'll focus on that when I'll be out of ideas for predicting the known classes.  But yes, I agree this has to be tuned for each model/.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 406122,
      "author_name": "Max Halford",
      "author_url": "",
      "post_date": "2018-10-18T17:32:16.567000",
      "content": "<p>Hello Olivier,</p>\n\n<p>By \"inverse\" I assumed you're referring to <code>1 - x</code> and not <code>1 / x</code>. If so I think it's clearer to use the term \"opposite\".</p>\n\n<p>Personally I'm using <code>1 - max(P)</code> where <code>P</code> is the distribution of probabilities for the other classes. I guess this corresponds to the third approach you mentioned. I find this intuitive since essentially it produces high values for \"flat\" distributions where the model isn't sure on any of the known classes. Maybe this idea could be pushed further by looking at the skew/kurtosis of the predicted class distribution.</p>\n\n<p>Anyway, just my 2 cents. </p>",
      "votes": 11,
      "replies": [
        {
          "id": 406138,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-18T18:06:07.493000",
          "content": "<p>Thanks for sharing Max. </p>\n\n<p>Yeah I thought 1 - max(P) would give better results as it was, as you said, very intuitive. Maybe I'm not good enough yet for other classes ;-)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 406143,
          "author_name": "Max Halford",
          "author_url": "",
          "post_date": "2018-10-18T18:16:20.473000",
          "content": "<p>Indeed, the predictions for <code>class_99</code> will naturally get better with better predictions for the other classes.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 407819,
          "author_name": "Andy Penrose",
          "author_url": "",
          "post_date": "2018-10-21T20:38:34.753000",
          "content": "<p>Thanks for the insightful comments. Agreed, 1-max(P) sounds very promising. However, at the moment, I find the same as Olivier. For me, 1-max(p) was worse than the very crude p(class_99)=0.2.</p>\n\n<p>I have found that using 1-max(p) when max(p) is high (say &gt;0.8) works well to reduce p(class_99) when the classifier is confident about another class, but when max(p) is low this seems to give too much weight to class_99. I think I'll have to revisit this when my base classifier is better!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 408344,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-22T17:57:38.090000",
          "content": "<p>Thanks Andy. I have not tried <code>1 - np.max(probas)</code> lately with my new score. I may have a go at it and burn a submission to check if I can get anything better than my current process described in the public kernel.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 411143,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-10-27T13:42:28.737000",
      "content": "<p>My LB probing experiments, up to now, showed me that distribution of class_99 Galactic is around 10x smaller than class_99 ExtraGalactic.  Hard coding class_99 of Galactic and ExtraGalactic to constants 0.017 and 0.17 improved my scores.</p>\n\n<p>Anyone did any other probing experiments?</p>",
      "votes": 10,
      "replies": [
        {
          "id": 411260,
          "author_name": "Ilya Khristoforov",
          "author_url": "",
          "post_date": "2018-10-27T18:10:26.903000",
          "content": "<p>I've got almost the same result. I took 0.015 and 0.15 - this gave me the best probe (2.081).</p>\n\n<p>I also probed </p>\n\n<p>0.020/0.147 -&gt; 2.081</p>\n\n<p>0.010/0.151 -&gt; 2.081</p>\n\n<p>0/0.154 -&gt; 2.191</p>\n\n<p>0.038/0.141 -&gt; 2.084</p>\n\n<p>0.106/0.113 - &gt; 2.105</p>\n\n<p>0.167/0.154 -&gt; 2.118</p>\n\n<p>I made a kernel about this here: <a href=\"https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081\">https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 411287,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-27T19:32:36.867000",
          "content": "<p>I tried your two constants, and it is a bit worse than Olivier's way for my latest sub, 1.057 vs 1.052.  Maybe the constant should be tuned for a given sub.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 406406,
      "author_name": "Ilya Khristoforov",
      "author_url": "",
      "post_date": "2018-10-19T07:37:55.637000",
      "content": "<p>Most of the class_99 objects (~95% - 97.5%) are placed in the Extragalactic group (out of the Milky Way).</p>\n\n<p>I tried public LB multiple times to detect this. Check the kernel: <a href=\"https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081\">https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081</a></p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 406189,
      "author_name": "Linards Kalvāns",
      "author_url": "",
      "post_date": "2018-10-18T20:06:00.520000",
      "content": "<p>Hi,</p>\n\n<p>I've implemented my own version of one vs all classifier, that trains any scikit model of my choice (and probably of other pkgs with similar interface). Once I get probabilities for all classes, I set P(99)=1-sum(P(.)) and adjust it to zero when it gets negative. Though I haven't yet submitted results and don't know how am I performing. </p>",
      "votes": 5,
      "replies": [
        {
          "id": 406355,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-19T05:38:18.890000",
          "content": "<p>Thanks for sharing Linards Kalvāns. </p>\n\n<p>I had a look at that using lightgbm raw scores <code>lgb.predict(raw_score=True)</code> and sum of individual proba can go up to around 5. </p>\n\n<p>Hope you'll hav more chance with that !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 406165,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-10-18T18:50:24.710000",
      "content": "<p>One more:</p>\n\n<ul>\n<li>probe the Public LB :-)</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 406177,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-18T19:29:48.783000",
          "content": "<p>Thanks Giba but I must be too silly for that :) I still don't get how to do that...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 406410,
          "author_name": "Ilya Khristoforov",
          "author_url": "",
          "post_date": "2018-10-19T07:46:13.967000",
          "content": "<p>This way becomes universal for that competition :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 408586,
          "author_name": "Harsh",
          "author_url": "",
          "post_date": "2018-10-23T05:51:38.323000",
          "content": "<p>It is consider as common practice now a days ; )</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 409628,
      "author_name": "JaimeF",
      "author_url": "",
      "post_date": "2018-10-24T16:06:13.033000",
      "content": "<p>Just another idea: add noise (random) rows to training data labeled as class 99.</p>\n\n<p>I haven't had time to try it yet, so don't known if it would help or not. But as the aim of class 99 is to identify new objects not seen previously, it can be considered heterogeneous data, so random features from a correct distribution could help to emulate the heterogeneity.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 410659,
          "author_name": "Ilya Khristoforov",
          "author_url": "",
          "post_date": "2018-10-26T12:18:08.077000",
          "content": "<p>Awesome idea! But it might be not as simple as it seems to be since we will probably need to guess what the noise really is and how the noise should look like... But the idea is brilliant and maybe there is a way to build an interesting model!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 408951,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-10-23T16:32:14.813000",
      "content": "<p>Clustering with DBScan automatically labels rejects as label == -1.  It may be worth using a lookup table to convert dbscan labels to probabilities</p>",
      "votes": 4,
      "replies": [
        {
          "id": 408977,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-23T17:19:17.917000",
          "content": "<p>Thanks Scirpus, that's on my todo list... </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 409318,
          "author_name": "mezoganet",
          "author_url": "",
          "post_date": "2018-10-24T05:34:50.310000",
          "content": "<p>@Scirpus</p>\n\n<p>I did some search with DBSAN with no real success... I will give it a new try . Thanks.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 408374,
      "author_name": "Mike Holcomb",
      "author_url": "",
      "post_date": "2018-10-22T18:42:53.553000",
      "content": "<p>To approach this unseen class issue, I think that I'm going to try to layer on top a <a href=\"http://scikit-learn.org/stable/auto_examples/svm/plot_oneclass.html#sphx-glr-auto-examples-svm-plot-oneclass-py\">One-Class SVM</a> on the training data.  When predicting, I'll first try to detect whether the test example is in the same \"class\" as the \"training class.\"  If not, label it as class 99.  If it is, use a full 14-class classifier trained on the seen classes in the training data.​</p>",
      "votes": 4,
      "replies": [
        {
          "id": 408380,
          "author_name": "Ilya Khristoforov",
          "author_url": "",
          "post_date": "2018-10-22T18:49:58.147000",
          "content": "<p>Interesting step!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 408636,
          "author_name": "Andreu Sancho-Asensio",
          "author_url": "",
          "post_date": "2018-10-23T07:41:57.620000",
          "content": "<p>I tried that with the <em>IsolationForest</em> and I found very weird stuff; moreover: a t-SNE of the training set and another using a large sample of the test set (roughly 6000 rows taken at random) suggest that these two groups come from distinct distributions... I also tried a t-SNE using a concatenation of both the training and the test sets, and this anomaly still appear in the data... In conclusion: it may happen that one-class SVM \"fails\" at finding the \"true\" class_99 objects due to this inconsistency; as if a preprocessing (e.g., maybe a kind of a <em>renormalization</em> in the training or test sets) was necessary prior to doing anything else.</p>\n\n<p>Note that both the train and test sets used in my analysis included the time series data, conveniently characterized using about 950 features per time series (for example: skewness per band, entropy, median, and so on).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 408793,
          "author_name": "Mike Holcomb",
          "author_url": "",
          "post_date": "2018-10-23T13:17:40.933000",
          "content": "<p>Interesting find.  After you mentioned that, this part of the data note sticks out to me.</p>\n\n<blockquote>\n  <p><strong>3.2. Training and test data</strong></p>\n  \n  <p>Moreover, the training data properties are non-representative of distributions of the the test[sic] data set. The training data are mostly composed of nearby, low-redshift, brighter objects while the test data contain more distant (higher redshift) and fainter objects.</p>\n</blockquote>\n\n<p>Your findings and Ilya Khristoforov's kernel are consistent with this.  Maybe developing a One-Class SVM for each category (Milky Way vs. Extragalactic) can help...</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 408807,
          "author_name": "Andreu Sancho-Asensio",
          "author_url": "",
          "post_date": "2018-10-23T13:41:48.957000",
          "content": "<p>That's a great idea actually; I'll try the same analysis with the Milky Way vs. Extragalactic and see what happens. I'll post here the findings.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 410609,
      "author_name": "yentianbao",
      "author_url": "",
      "post_date": "2018-10-26T10:43:29.417000",
      "content": "<p>Maybe one class randomforest ? (<a href=\"https://github.com/ngoix/OCRF\">https://github.com/ngoix/OCRF</a>)\nI try  my own version (some kind of mixed lightgbm in this algorithm)\nbut it didn't work , LB is terrible :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 410576,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-10-26T09:39:38.080000",
      "content": "<p>Olivier, </p>\n\n<p>In my limited experience, the way you compute it in your kernel is a bit better than using a constant 1/9 prediction for class 99.   LB gain is about 0.02 for me.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 410693,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-26T13:32:33.493000",
          "content": "<p>@CPMP, yes that's what I found. Using different averages for galactic and extragalactic and prod of opposite proba gives the best results for me.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 410092,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-10-25T11:53:01.113000",
      "content": "<p>How about connecting 99 to flux rows where detected==0\nIf a flux row is detected then it must be one of the training targets if it isn't then it must be 99 by definition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 409344,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-10-24T06:37:02.297000",
      "content": "<p>Seems LB probing will play a role here.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 409360,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-10-24T07:17:28.470000",
          "content": "<p>Unfortunately I wholeheartedly agree with you\nOne other gaming approach would be to make sure that the predictions are not too confident.  A 1 will be punished heavily so I am going to reduce the predictions between a min and a max before row normalizing</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 406439,
      "author_name": "dylonLL",
      "author_url": "",
      "post_date": "2018-10-19T08:44:25.033000",
      "content": "<p>I'm still very far from it, I have also tried auto encoders and distance error but it did not do significant improvement to my score.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 406187,
      "author_name": "mezoganet",
      "author_url": "",
      "post_date": "2018-10-18T20:02:08.690000",
      "content": "<p>@ogreiller</p>\n\n<p>Thanks for raising this issue I have too, and thanks to all for the insightfull replies.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 408351,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-22T18:07:41.940000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "406121": "I've tried so many things here that I'm not sure I can list them all. I must admit I'm puzzled as to how to derive class 99 probability from the training set. It's the very first time I'm facing a similar problem.\n\nHere are the things I've tried: \n\n - Use a constant\n - Going to logits, derive individual probabilities from that, compute p99 as the product of opposite class probabilities, go back to logit and use softmax to normalize (probably the worst result)\n - Computing the probability as the opposite of the max proba over all classes for each sample\n - Calculate p99 as the product of the other classes opposite proba, going to logits and using softmax to normalize all probanilities\n - Used [IsolationForest](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67205) but still have issues merging that with the other classes proba. I've found anomalies are way to high in my opinion.\n\nThe best I've found this far is to use the product of all opposite probabilities and adapting the mean but I'm not satisfied with that since it kills the other classes logloss.\n\nIf anyone is willing to share his ideas I would be gratefull ;-)\n  ",
    "410767": "I believe a prerequisite for predicting Class 99 is having a very good model for the other classes.\nThe good new is that a score of ~1.0 is achievable without really dealing with class 99. \nThe only thing I did up to this point with this prediction is:\nclass_99=np.where(other_classes.max&gt;0.9 , 0.01, 0.1)\n[it is slightly better then uniform, If I use 0.8 the score degrades, and also other values of probabilities degrade the score]",
    "406122": "Hello Olivier,\n\nBy \"inverse\" I assumed you're referring to `1 - x` and not `1 / x`. If so I think it's clearer to use the term \"opposite\".\n\nPersonally I'm using `1 - max(P)` where `P` is the distribution of probabilities for the other classes. I guess this corresponds to the third approach you mentioned. I find this intuitive since essentially it produces high values for \"flat\" distributions where the model isn't sure on any of the known classes. Maybe this idea could be pushed further by looking at the skew/kurtosis of the predicted class distribution.\n\nAnyway, just my 2 cents. ",
    "411143": "My LB probing experiments, up to now, showed me that distribution of class_99 Galactic is around 10x smaller than class_99 ExtraGalactic.  Hard coding class_99 of Galactic and ExtraGalactic to constants 0.017 and 0.17 improved my scores.\n\nAnyone did any other probing experiments?",
    "406406": "Most of the class_99 objects (~95% - 97.5%) are placed in the Extragalactic group (out of the Milky Way).\n\nI tried public LB multiple times to detect this. Check the kernel: https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081",
    "406189": "Hi,\n\nI've implemented my own version of one vs all classifier, that trains any scikit model of my choice (and probably of other pkgs with similar interface). Once I get probabilities for all classes, I set P(99)=1-sum(P(.)) and adjust it to zero when it gets negative. Though I haven't yet submitted results and don't know how am I performing. ",
    "406165": "One more:\n\n - probe the Public LB :-)",
    "409628": "Just another idea: add noise (random) rows to training data labeled as class 99.\n\nI haven't had time to try it yet, so don't known if it would help or not. But as the aim of class 99 is to identify new objects not seen previously, it can be considered heterogeneous data, so random features from a correct distribution could help to emulate the heterogeneity.",
    "408951": "Clustering with DBScan automatically labels rejects as label == -1.  It may be worth using a lookup table to convert dbscan labels to probabilities",
    "408374": "To approach this unseen class issue, I think that I'm going to try to layer on top a [One-Class SVM][1] on the training data.  When predicting, I'll first try to detect whether the test example is in the same \"class\" as the \"training class.\"  If not, label it as class 99.  If it is, use a full 14-class classifier trained on the seen classes in the training data.​\n\n  [1]: http://scikit-learn.org/stable/auto_examples/svm/plot_oneclass.html#sphx-glr-auto-examples-svm-plot-oneclass-py \"One-Class SVM\"",
    "410609": "Maybe one class randomforest ? (https://github.com/ngoix/OCRF)\nI try  my own version (some kind of mixed lightgbm in this algorithm)\nbut it didn't work , LB is terrible :)",
    "410576": "Olivier, \n\nIn my limited experience, the way you compute it in your kernel is a bit better than using a constant 1/9 prediction for class 99.   LB gain is about 0.02 for me.",
    "410092": "How about connecting 99 to flux rows where detected==0\nIf a flux row is detected then it must be one of the training targets if it isn't then it must be 99 by definition.",
    "409344": "Seems LB probing will play a role here.",
    "406439": "I'm still very far from it, I have also tried auto encoders and distance error but it did not do significant improvement to my score.",
    "406187": "@ogreiller\n\nThanks for raising this issue I have too, and thanks to all for the insightfull replies.",
    "408351": ""
  }
}