{
  "id": 72104,
  "title": "Another way to calculate Unknown",
  "url": "/competitions/PLAsTiCC-2018/discussion/72104",
  "author_name": "Scirpus",
  "post_date": "2018-11-20T10:58:33.327000",
  "votes": 37,
  "comment_count": 42,
  "views": 0,
  "content": "<p>This one makes unknown a bit bigger that Olivier's method but it dropped my score by .01</p>\n\n<pre><code>def GenUnknown(data):\n    return ((((((data[\"mymedian\"]) + (((data[\"mymean\"]) / 2.0)))/2.0)) + (((((1.0) - (((data[\"mymax\"]) * (((data[\"mymax\"]) * (data[\"mymax\"]))))))) / 2.0)))/2.0)\n\nfeats = ['class_6', 'class_15', 'class_16', 'class_42', 'class_52', 'class_53',\n         'class_62', 'class_64', 'class_65', 'class_67', 'class_88', 'class_90',\n         'class_92', 'class_95']\n\ny = pd.DataFrame()\ny['mymean'] = preds_df[feats].mean(axis=1)\ny['mymedian'] = preds_df[feats].median(axis=1)\ny['mymax'] = preds_df[feats].max(axis=1)\n\nx.class_99 = GenUnknown(y)\n</code></pre>",
  "messages": [
    {
      "id": 424579,
      "postDate": "2018-11-20T10:58:33.327Z",
      "content": "<p>This one makes unknown a bit bigger that Olivier's method but it dropped my score by .01</p>\n\n<pre><code>def GenUnknown(data):\n    return ((((((data[\"mymedian\"]) + (((data[\"mymean\"]) / 2.0)))/2.0)) + (((((1.0) - (((data[\"mymax\"]) * (((data[\"mymax\"]) * (data[\"mymax\"]))))))) / 2.0)))/2.0)\n\nfeats = ['class_6', 'class_15', 'class_16', 'class_42', 'class_52', 'class_53',\n         'class_62', 'class_64', 'class_65', 'class_67', 'class_88', 'class_90',\n         'class_92', 'class_95']\n\ny = pd.DataFrame()\ny['mymean'] = preds_df[feats].mean(axis=1)\ny['mymedian'] = preds_df[feats].median(axis=1)\ny['mymax'] = preds_df[feats].max(axis=1)\n\nx.class_99 = GenUnknown(y)\n</code></pre>",
      "rawMarkdown": "This one makes unknown a bit bigger that Olivier's method but it dropped my score by .01\n\n    def GenUnknown(data):\n        return ((((((data[\"mymedian\"]) + (((data[\"mymean\"]) / 2.0)))/2.0)) + (((((1.0) - (((data[\"mymax\"]) * (((data[\"mymax\"]) * (data[\"mymax\"]))))))) / 2.0)))/2.0)\n    \n    feats = ['class_6', 'class_15', 'class_16', 'class_42', 'class_52', 'class_53',\n             'class_62', 'class_64', 'class_65', 'class_67', 'class_88', 'class_90',\n             'class_92', 'class_95']\n    \n    y = pd.DataFrame()\n    y['mymean'] = preds_df[feats].mean(axis=1)\n    y['mymedian'] = preds_df[feats].median(axis=1)\n    y['mymax'] = preds_df[feats].max(axis=1)\n    \n    x.class_99 = GenUnknown(y)",
      "votes": 37
    },
    {
      "id": 437060,
      "postDate": "2018-12-11T10:04:36.120Z",
      "content": "<p>Improved my LB score from 1.048 to 1.006. Thanks for your great contribution.</p>",
      "rawMarkdown": "Improved my LB score from 1.048 to 1.006. Thanks for your great contribution.",
      "votes": 1
    },
    {
      "id": 429452,
      "postDate": "2018-11-28T21:48:40.427Z",
      "content": "<p>This approach improved our score by 0.001 compared to @CPMP's method</p>",
      "rawMarkdown": "This approach improved our score by 0.001 compared to @CPMP's method",
      "votes": 1
    },
    {
      "id": 428519,
      "postDate": "2018-11-27T12:16:06.153Z",
      "content": "<p>Another one just using the max for Sergei</p>\n\n<pre><code>def GenUnknownII(data):\na = data[\"mymax\"]\nreturn -.0625*a**8-0.125*a**4-.0625*a+.25\n</code></pre>",
      "rawMarkdown": "Another one just using the max for Sergei\n\n    def GenUnknownII(data):\n    a = data[\"mymax\"]\n    return -.0625*a**8-0.125*a**4-.0625*a+.25\n\n    \n    ",
      "votes": 1,
      "replies": [
        {
          "id": 428539,
          "postDate": "2018-11-27T12:50:37.347Z",
          "content": "<p>@Scirpus, may I ask you, how are you getting these functions ? </p>",
          "rawMarkdown": "@Scirpus, may I ask you, how are you getting these functions ? "
        },
        {
          "id": 428549,
          "postDate": "2018-11-27T13:10:01.577Z",
          "content": "<p>This time I used a Genetic Algorithm rather than programming (There is only one variable)</p>",
          "rawMarkdown": "This time I used a Genetic Algorithm rather than programming (There is only one variable)",
          "votes": 1
        },
        {
          "id": 428767,
          "postDate": "2018-11-27T21:20:10.147Z",
          "content": "<p>For my best submission:</p>\n\n<p>old (with 'mymean' and 'mymedian') variant gives 1.041</p>\n\n<p>new variant (only 'mymax') gives 1.043</p>",
          "rawMarkdown": "For my best submission:\n\nold (with 'mymean' and 'mymedian') variant gives 1.041\n\nnew variant (only 'mymax') gives 1.043",
          "votes": 1
        },
        {
          "id": 428778,
          "postDate": "2018-11-27T21:46:11.113Z",
          "content": "<p>Thanks for the update - not too bad considering just one parameter;)</p>",
          "rawMarkdown": "Thanks for the update - not too bad considering just one parameter;)"
        },
        {
          "id": 429654,
          "postDate": "2018-11-29T06:15:27.763Z",
          "content": "<p>Scirpus! You can try to include Olivier's method in your genenetic algorithm.\nThe new (4-th) variable like data[\"oliver_proba\"] = 0.18 * product(1-p_i).</p>",
          "rawMarkdown": "Scirpus! You can try to include Olivier's method in your genenetic algorithm.\nThe new (4-th) variable like data[\"oliver_proba\"] = 0.18 * product(1-p_i)."
        }
      ]
    },
    {
      "id": 428446,
      "postDate": "2018-11-27T09:41:15.933Z",
      "content": "<p>Thanks Scirpus.  Using your method improved score from 0.995 --&gt; 0.983.</p>",
      "rawMarkdown": "Thanks Scirpus.  Using your method improved score from 0.995 --&gt; 0.983.",
      "votes": 1
    },
    {
      "id": 426988,
      "postDate": "2018-11-24T10:02:41.653Z",
      "content": "<p>I've tried the more simple expression</p>\n\n<p>0.5 * (1 - data['mymax'])</p>\n\n<p>It gave me a boost too, but a little bit lower (1.058 -&gt; 1.053)</p>",
      "rawMarkdown": "I've tried the more simple expression\n\n 0.5 * (1 - data['mymax'])\n\nIt gave me a boost too, but a little bit lower (1.058 -&gt; 1.053)",
      "votes": 1,
      "replies": [
        {
          "id": 427099,
          "postDate": "2018-11-24T15:09:31.493Z",
          "content": "<p>The difference is tiny so could just be noise - I like it!</p>",
          "rawMarkdown": "The difference is tiny so could just be noise - I like it!"
        }
      ]
    },
    {
      "id": 426780,
      "postDate": "2018-11-23T20:36:23.753Z",
      "content": "<p>Nice. For me 1.058 -&gt; 1.051</p>",
      "rawMarkdown": "Nice. For me 1.058 -&gt; 1.051",
      "votes": 1
    },
    {
      "id": 426711,
      "postDate": "2018-11-23T17:31:26.153Z",
      "content": "<p>tHANKS</p>",
      "rawMarkdown": "tHANKS",
      "votes": 1,
      "replies": [
        {
          "id": 426736,
          "postDate": "2018-11-23T18:21:58.990Z",
          "content": "<p>No problem - if you have trouble logging into websites you may want to turn CAPS lock off ;)</p>",
          "rawMarkdown": "No problem - if you have trouble logging into websites you may want to turn CAPS lock off ;)",
          "votes": 12
        }
      ]
    },
    {
      "id": 425572,
      "postDate": "2018-11-21T19:48:50.420Z",
      "content": "<p>Thanks Scirpus.  I tried my own algorithm (based on a similar principle) and my score got worse.  Using your method improved from 1.052 --&gt; 1.039.</p>",
      "rawMarkdown": "Thanks Scirpus.  I tried my own algorithm (based on a similar principle) and my score got worse.  Using your method improved from 1.052 --&gt; 1.039.",
      "votes": 1
    },
    {
      "id": 424648,
      "postDate": "2018-11-20T13:34:28.860Z",
      "content": "<p>gave me a 0.001 boost - probably noise ;)</p>",
      "rawMarkdown": "gave me a 0.001 boost - probably noise ;)",
      "votes": 1,
      "replies": [
        {
          "id": 424721,
          "postDate": "2018-11-20T15:12:27.657Z",
          "content": "<p>Thanks for letting me know - at least it didn't go down - LOL - as your models are never terrible you probably don't need any tweaks ;)</p>",
          "rawMarkdown": "Thanks for letting me know - at least it didn't go down - LOL - as your models are never terrible you probably don't need any tweaks ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 437282,
      "postDate": "2018-12-11T17:10:06.960Z",
      "content": "<p>FYI, this degrades a bit our best model score.  However, the fact that it degrades only very little while changing class-99 predictions quite a bit makes me think class-99 is not that important.  It kind of confirms what I wrote elsewhere about hard classes, esp class 52 that is hard to distinguish from other classes.</p>",
      "rawMarkdown": "FYI, this degrades a bit our best model score.  However, the fact that it degrades only very little while changing class-99 predictions quite a bit makes me think class-99 is not that important.  It kind of confirms what I wrote elsewhere about hard classes, esp class 52 that is hard to distinguish from other classes.",
      "votes": 2,
      "replies": [
        {
          "id": 437296,
          "postDate": "2018-12-11T17:28:27.223Z",
          "content": "<p>class_99 probably helps bad/overconfident models by dampening overly confident predictions - yours looks too good for that ;)</p>",
          "rawMarkdown": "class_99 probably helps bad/overconfident models by dampening overly confident predictions - yours looks too good for that ;)",
          "votes": 1
        },
        {
          "id": 437303,
          "postDate": "2018-12-11T17:36:37.547Z",
          "content": "<p>Maybe Olivier's method gets more reliable as the base model gets better.</p>",
          "rawMarkdown": "Maybe Olivier's method gets more reliable as the base model gets better.",
          "votes": 1
        },
        {
          "id": 437330,
          "postDate": "2018-12-11T18:12:53.127Z",
          "content": "<p>Olivers is typically smaller than this method - so you could be correct - good luck for Monday!</p>",
          "rawMarkdown": "Olivers is typically smaller than this method - so you could be correct - good luck for Monday!",
          "votes": 1
        },
        {
          "id": 437432,
          "postDate": "2018-12-11T23:08:27.320Z",
          "content": "<p>Are using Oliver's method still? I thought you have something better.</p>",
          "rawMarkdown": "Are using Oliver's method still? I thought you have something better."
        },
        {
          "id": 437512,
          "postDate": "2018-12-12T04:04:57.740Z",
          "content": "<p>Variant of it.  Do you have something better?</p>",
          "rawMarkdown": "Variant of it.  Do you have something better?"
        },
        {
          "id": 437601,
          "postDate": "2018-12-12T07:07:57.833Z",
          "content": "<p>No, we're using Scirpus' method. I thought people on the top found a better way to calculate class 99.\nWe hunted for it last days but no luck. :(</p>",
          "rawMarkdown": "No, we're using Scirpus' method. I thought people on the top found a better way to calculate class 99.\nWe hunted for it last days but no luck. :("
        },
        {
          "id": 437712,
          "postDate": "2018-12-12T10:54:22.257Z",
          "content": "<p>Maybe top 3 have found a better way ;)</p>",
          "rawMarkdown": "Maybe top 3 have found a better way ;)"
        },
        {
          "id": 437734,
          "postDate": "2018-12-12T11:51:27.443Z",
          "content": "<p>We didn't find any, but I believe, this is why we are both  at ~0.75 and they are at ~0.7 </p>",
          "rawMarkdown": "We didn't find any, but I believe, this is why we are both  at ~0.75 and they are at ~0.7 "
        }
      ]
    },
    {
      "id": 437265,
      "postDate": "2018-12-11T16:29:10.470Z",
      "content": "<p>Solve for y, where x are your probabilities, then choose the smallest y as class 99 probability?</p>\n\n<p>And rescale x with (1 - y)... And N = 14...</p>",
      "rawMarkdown": "Solve for y, where x are your probabilities, then choose the smallest y as class 99 probability?\n\nAnd rescale x with (1 - y)... And N = 14..."
    },
    {
      "id": 429535,
      "postDate": "2018-11-29T01:50:36.293Z",
      "content": "<p>Can anyone offer some intuition of why this might work? I can sort of understand why the original Product(1-p_i) might not work very well (the product of compliments gives high class 99 estimate to the most confused classes, but we know pretty well that the most confused classes are likely one of the known supernova classes), but I cannot make out why this scire relates to class 99 despite apparent usefulness.</p>",
      "rawMarkdown": "Can anyone offer some intuition of why this might work? I can sort of understand why the original Product(1-p_i) might not work very well (the product of compliments gives high class 99 estimate to the most confused classes, but we know pretty well that the most confused classes are likely one of the known supernova classes), but I cannot make out why this scire relates to class 99 despite apparent usefulness."
    },
    {
      "id": 429361,
      "postDate": "2018-11-28T18:34:21.067Z",
      "content": "<p>Do you treat or weight galatic / exg differently with this, or straight cut?</p>",
      "rawMarkdown": "Do you treat or weight galatic / exg differently with this, or straight cut?",
      "replies": [
        {
          "id": 437887,
          "postDate": "2018-12-12T17:37:36.157Z",
          "content": "<p>I applied Scirpus method and got 0.013 relative to Olivier's original (0.14) method.  I then applied another method that halved the class 99 probability for intergalactic.  That gave me another 0.001.</p>",
          "rawMarkdown": "I applied Scirpus method and got 0.013 relative to Olivier's original (0.14) method.  I then applied another method that halved the class 99 probability for intergalactic.  That gave me another 0.001."
        }
      ]
    },
    {
      "id": 429317,
      "postDate": "2018-11-28T17:07:54.420Z",
      "content": "<p>I got 0.005 lift.</p>",
      "rawMarkdown": "I got 0.005 lift.",
      "replies": [
        {
          "id": 429320,
          "postDate": "2018-11-28T17:22:47Z",
          "content": "<p>Compared to what?  I wanted to ask the question to each of those reporting improvement...</p>\n\n<p>I am not saying this has no merit, but I'd like to know what is the comparison point.</p>",
          "rawMarkdown": "Compared to what?  I wanted to ask the question to each of those reporting improvement...\n\nI am not saying this has no merit, but I'd like to know what is the comparison point.",
          "votes": 1
        },
        {
          "id": 429328,
          "postDate": "2018-11-28T17:33:32.047Z",
          "content": "<p>Its just an lgb from 1.012 to 1.007, compared to olivier's method.</p>",
          "rawMarkdown": "Its just an lgb from 1.012 to 1.007, compared to olivier's method.",
          "votes": 1
        },
        {
          "id": 429334,
          "postDate": "2018-11-28T17:49:25.257Z",
          "content": "<p>My improvements are compared to olivier's algorithm for predicting the unknown.</p>",
          "rawMarkdown": "My improvements are compared to olivier's algorithm for predicting the unknown.",
          "votes": 1
        },
        {
          "id": 429364,
          "postDate": "2018-11-28T18:41:47.610Z",
          "content": "<p>Thanks.  Olivier's method can easily be improved by using a larger mean for class_99 as I disclosed a while ago.  I wonder which way is better, Scirpus way or that.</p>",
          "rawMarkdown": "Thanks.  Olivier's method can easily be improved by using a larger mean for class_99 as I disclosed a while ago.  I wonder which way is better, Scirpus way or that.",
          "votes": 1
        },
        {
          "id": 429371,
          "postDate": "2018-11-28T18:53:42.633Z",
          "content": "<p>Could you give me the link to your method - I am defintely willing to use a submission and I'll report back.  Funnily enough my algorithm also increase class_99 on average - I am not sure if this is due to direct classification of class_99 or a way of reducing the confidence in the other classes ;)</p>",
          "rawMarkdown": "Could you give me the link to your method - I am defintely willing to use a submission and I'll report back.  Funnily enough my algorithm also increase class_99 on average - I am not sure if this is due to direct classification of class_99 or a way of reducing the confidence in the other classes ;)"
        },
        {
          "id": 429391,
          "postDate": "2018-11-28T19:42:12.360Z",
          "content": "<p>IIRC, cpmp uses 0.18 instead of 0.14.</p>",
          "rawMarkdown": "IIRC, cpmp uses 0.18 instead of 0.14."
        },
        {
          "id": 429393,
          "postDate": "2018-11-28T19:52:01.110Z",
          "content": "<p><a href=\"/authman\">@authman</a> - thanks for this - i tried to look at cpmps profiles to find it but as he is number one in discussion it was like looking for a needle in a haystack LOL ;)</p>",
          "rawMarkdown": "@authman - thanks for this - i tried to look at cpmps profiles to find it but as he is number one in discussion it was like looking for a needle in a haystack LOL ;)",
          "votes": 2
        },
        {
          "id": 429402,
          "postDate": "2018-11-28T20:02:32.470Z",
          "content": "<p>I hear you. Like looking for 52 in 90.</p>",
          "rawMarkdown": "I hear you. Like looking for 52 in 90.",
          "votes": 1
        }
      ]
    },
    {
      "id": 426782,
      "postDate": "2018-11-23T20:56:57.177Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 437850,
          "postDate": "2018-12-12T16:04:21.303Z",
          "content": "<p>This calculation brought down my public kernel from 1.048 to 1.040 (0.008). Which is close to yours.</p>",
          "rawMarkdown": "This calculation brought down my public kernel from 1.048 to 1.040 (0.008). Which is close to yours.",
          "votes": 1
        }
      ]
    },
    {
      "id": 427463,
      "postDate": "2018-11-25T15:10:26.057Z",
      "content": "<p>Thanks! From 1.008 -&gt; to 1.003 :)</p>",
      "rawMarkdown": "Thanks! From 1.008 -&gt; to 1.003 :)",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 437060,
      "author_name": "MuhammedBuyukkinaci",
      "author_url": "",
      "post_date": "2018-12-11T10:04:36.120000",
      "content": "<p>Improved my LB score from 1.048 to 1.006. Thanks for your great contribution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 429452,
      "author_name": "agarreta",
      "author_url": "",
      "post_date": "2018-11-28T21:48:40.427000",
      "content": "<p>This approach improved our score by 0.001 compared to @CPMP's method</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 428519,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-11-27T12:16:06.153000",
      "content": "<p>Another one just using the max for Sergei</p>\n\n<pre><code>def GenUnknownII(data):\na = data[\"mymax\"]\nreturn -.0625*a**8-0.125*a**4-.0625*a+.25\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 428539,
          "author_name": "Indranil Bhattacharya",
          "author_url": "",
          "post_date": "2018-11-27T12:50:37.347000",
          "content": "<p>@Scirpus, may I ask you, how are you getting these functions ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 428549,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-11-27T13:10:01.577000",
          "content": "<p>This time I used a Genetic Algorithm rather than programming (There is only one variable)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 428767,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-11-27T21:20:10.147000",
          "content": "<p>For my best submission:</p>\n\n<p>old (with 'mymean' and 'mymedian') variant gives 1.041</p>\n\n<p>new variant (only 'mymax') gives 1.043</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 428778,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-11-27T21:46:11.113000",
          "content": "<p>Thanks for the update - not too bad considering just one parameter;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429654,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-11-29T06:15:27.763000",
          "content": "<p>Scirpus! You can try to include Olivier's method in your genenetic algorithm.\nThe new (4-th) variable like data[\"oliver_proba\"] = 0.18 * product(1-p_i).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 428446,
      "author_name": "mrxnew",
      "author_url": "",
      "post_date": "2018-11-27T09:41:15.933000",
      "content": "<p>Thanks Scirpus.  Using your method improved score from 0.995 --&gt; 0.983.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 426988,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-11-24T10:02:41.653000",
      "content": "<p>I've tried the more simple expression</p>\n\n<p>0.5 * (1 - data['mymax'])</p>\n\n<p>It gave me a boost too, but a little bit lower (1.058 -&gt; 1.053)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 427099,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-11-24T15:09:31.493000",
          "content": "<p>The difference is tiny so could just be noise - I like it!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 426780,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-11-23T20:36:23.753000",
      "content": "<p>Nice. For me 1.058 -&gt; 1.051</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 426711,
      "author_name": "El-dosuky",
      "author_url": "",
      "post_date": "2018-11-23T17:31:26.153000",
      "content": "<p>tHANKS</p>",
      "votes": 1,
      "replies": [
        {
          "id": 426736,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-11-23T18:21:58.990000",
          "content": "<p>No problem - if you have trouble logging into websites you may want to turn CAPS lock off ;)</p>",
          "votes": 12,
          "replies": []
        }
      ]
    },
    {
      "id": 425572,
      "author_name": "Jim Sullivan",
      "author_url": "",
      "post_date": "2018-11-21T19:48:50.420000",
      "content": "<p>Thanks Scirpus.  I tried my own algorithm (based on a similar principle) and my score got worse.  Using your method improved from 1.052 --&gt; 1.039.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 424648,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2018-11-20T13:34:28.860000",
      "content": "<p>gave me a 0.001 boost - probably noise ;)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 424721,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-11-20T15:12:27.657000",
          "content": "<p>Thanks for letting me know - at least it didn't go down - LOL - as your models are never terrible you probably don't need any tweaks ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 437282,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-12-11T17:10:06.960000",
      "content": "<p>FYI, this degrades a bit our best model score.  However, the fact that it degrades only very little while changing class-99 predictions quite a bit makes me think class-99 is not that important.  It kind of confirms what I wrote elsewhere about hard classes, esp class 52 that is hard to distinguish from other classes.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 437296,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-12-11T17:28:27.223000",
          "content": "<p>class_99 probably helps bad/overconfident models by dampening overly confident predictions - yours looks too good for that ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437303,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-11T17:36:37.547000",
          "content": "<p>Maybe Olivier's method gets more reliable as the base model gets better.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437330,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-12-11T18:12:53.127000",
          "content": "<p>Olivers is typically smaller than this method - so you could be correct - good luck for Monday!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437432,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-12-11T23:08:27.320000",
          "content": "<p>Are using Oliver's method still? I thought you have something better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437512,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-12T04:04:57.740000",
          "content": "<p>Variant of it.  Do you have something better?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437601,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-12-12T07:07:57.833000",
          "content": "<p>No, we're using Scirpus' method. I thought people on the top found a better way to calculate class 99.\nWe hunted for it last days but no luck. :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437712,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-12T10:54:22.257000",
          "content": "<p>Maybe top 3 have found a better way ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437734,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-12-12T11:51:27.443000",
          "content": "<p>We didn't find any, but I believe, this is why we are both  at ~0.75 and they are at ~0.7 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 437265,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2018-12-11T16:29:10.470000",
      "content": "<p>Solve for y, where x are your probabilities, then choose the smallest y as class 99 probability?</p>\n\n<p>And rescale x with (1 - y)... And N = 14...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 429535,
      "author_name": "Mithrillion",
      "author_url": "",
      "post_date": "2018-11-29T01:50:36.293000",
      "content": "<p>Can anyone offer some intuition of why this might work? I can sort of understand why the original Product(1-p_i) might not work very well (the product of compliments gives high class 99 estimate to the most confused classes, but we know pretty well that the most confused classes are likely one of the known supernova classes), but I cannot make out why this scire relates to class 99 despite apparent usefulness.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 429361,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2018-11-28T18:34:21.067000",
      "content": "<p>Do you treat or weight galatic / exg differently with this, or straight cut?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 437887,
          "author_name": "Jim Sullivan",
          "author_url": "",
          "post_date": "2018-12-12T17:37:36.157000",
          "content": "<p>I applied Scirpus method and got 0.013 relative to Olivier's original (0.14) method.  I then applied another method that halved the class 99 probability for intergalactic.  That gave me another 0.001.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 429317,
      "author_name": "Chinta",
      "author_url": "",
      "post_date": "2018-11-28T17:07:54.420000",
      "content": "<p>I got 0.005 lift.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 429320,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-28T17:22:47",
          "content": "<p>Compared to what?  I wanted to ask the question to each of those reporting improvement...</p>\n\n<p>I am not saying this has no merit, but I'd like to know what is the comparison point.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429328,
          "author_name": "Chinta",
          "author_url": "",
          "post_date": "2018-11-28T17:33:32.047000",
          "content": "<p>Its just an lgb from 1.012 to 1.007, compared to olivier's method.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429334,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-11-28T17:49:25.257000",
          "content": "<p>My improvements are compared to olivier's algorithm for predicting the unknown.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429364,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-28T18:41:47.610000",
          "content": "<p>Thanks.  Olivier's method can easily be improved by using a larger mean for class_99 as I disclosed a while ago.  I wonder which way is better, Scirpus way or that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 429371,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-11-28T18:53:42.633000",
          "content": "<p>Could you give me the link to your method - I am defintely willing to use a submission and I'll report back.  Funnily enough my algorithm also increase class_99 on average - I am not sure if this is due to direct classification of class_99 or a way of reducing the confidence in the other classes ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429391,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-11-28T19:42:12.360000",
          "content": "<p>IIRC, cpmp uses 0.18 instead of 0.14.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429393,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-11-28T19:52:01.110000",
          "content": "<p><a href=\"/authman\">@authman</a> - thanks for this - i tried to look at cpmps profiles to find it but as he is number one in discussion it was like looking for a needle in a haystack LOL ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 429402,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-11-28T20:02:32.470000",
          "content": "<p>I hear you. Like looking for 52 in 90.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 426782,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-23T20:56:57.177000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 437850,
          "author_name": "Chia-Ta Tsai",
          "author_url": "",
          "post_date": "2018-12-12T16:04:21.303000",
          "content": "<p>This calculation brought down my public kernel from 1.048 to 1.040 (0.008). Which is close to yours.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 427463,
      "author_name": "Vadym",
      "author_url": "",
      "post_date": "2018-11-25T15:10:26.057000",
      "content": "<p>Thanks! From 1.008 -&gt; to 1.003 :)</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "424579": "This one makes unknown a bit bigger that Olivier's method but it dropped my score by .01\n\n    def GenUnknown(data):\n        return ((((((data[\"mymedian\"]) + (((data[\"mymean\"]) / 2.0)))/2.0)) + (((((1.0) - (((data[\"mymax\"]) * (((data[\"mymax\"]) * (data[\"mymax\"]))))))) / 2.0)))/2.0)\n    \n    feats = ['class_6', 'class_15', 'class_16', 'class_42', 'class_52', 'class_53',\n             'class_62', 'class_64', 'class_65', 'class_67', 'class_88', 'class_90',\n             'class_92', 'class_95']\n    \n    y = pd.DataFrame()\n    y['mymean'] = preds_df[feats].mean(axis=1)\n    y['mymedian'] = preds_df[feats].median(axis=1)\n    y['mymax'] = preds_df[feats].max(axis=1)\n    \n    x.class_99 = GenUnknown(y)",
    "437060": "Improved my LB score from 1.048 to 1.006. Thanks for your great contribution.",
    "429452": "This approach improved our score by 0.001 compared to @CPMP's method",
    "428519": "Another one just using the max for Sergei\n\n    def GenUnknownII(data):\n    a = data[\"mymax\"]\n    return -.0625*a**8-0.125*a**4-.0625*a+.25\n\n    \n    ",
    "428446": "Thanks Scirpus.  Using your method improved score from 0.995 --&gt; 0.983.",
    "426988": "I've tried the more simple expression\n\n 0.5 * (1 - data['mymax'])\n\nIt gave me a boost too, but a little bit lower (1.058 -&gt; 1.053)",
    "426780": "Nice. For me 1.058 -&gt; 1.051",
    "426711": "tHANKS",
    "425572": "Thanks Scirpus.  I tried my own algorithm (based on a similar principle) and my score got worse.  Using your method improved from 1.052 --&gt; 1.039.",
    "424648": "gave me a 0.001 boost - probably noise ;)",
    "437282": "FYI, this degrades a bit our best model score.  However, the fact that it degrades only very little while changing class-99 predictions quite a bit makes me think class-99 is not that important.  It kind of confirms what I wrote elsewhere about hard classes, esp class 52 that is hard to distinguish from other classes.",
    "437265": "Solve for y, where x are your probabilities, then choose the smallest y as class 99 probability?\n\nAnd rescale x with (1 - y)... And N = 14...",
    "429535": "Can anyone offer some intuition of why this might work? I can sort of understand why the original Product(1-p_i) might not work very well (the product of compliments gives high class 99 estimate to the most confused classes, but we know pretty well that the most confused classes are likely one of the known supernova classes), but I cannot make out why this scire relates to class 99 despite apparent usefulness.",
    "429361": "Do you treat or weight galatic / exg differently with this, or straight cut?",
    "429317": "I got 0.005 lift.",
    "426782": "",
    "427463": "Thanks! From 1.008 -&gt; to 1.003 :)"
  }
}