{
  "id": 67356,
  "title": "lgbm on raw scores 6.x",
  "url": "/competitions/PLAsTiCC-2018/discussion/67356",
  "author_name": "",
  "post_date": "2018-10-01T21:03:34.388356Z",
  "votes": 7,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I checked it a few times, maybe there is still an error in my processing code. A standard lgbm trained on the training_set_metadata.csv (10 numeric features) gives about 1.4 mlogloss on 5CV, but scores only 6.x on leaderboard ?? What is going on here?\nI tried the same with nn, got 1.35 on 5CV but only 5.9 on leaderboard ?? This is not a classic supervised setup. Anyone similar results?</p>",
  "messages": [
    {
      "id": "397071",
      "postDate": "10/01/2018 21:03:34",
      "content": "<p>I checked it a few times, maybe there is still an error in my processing code. A standard lgbm trained on the training_set_metadata.csv (10 numeric features) gives about 1.4 mlogloss on 5CV, but scores only 6.x on leaderboard ?? What is going on here?\nI tried the same with nn, got 1.35 on 5CV but only 5.9 on leaderboard ?? This is not a classic supervised setup. Anyone similar results?</p>",
      "rawMarkdown": "I checked it a few times, maybe there is still an error in my processing code. A standard lgbm trained on the training_set_metadata.csv (10 numeric features) gives about 1.4 mlogloss on 5CV, but scores only 6.x on leaderboard ?? What is going on here?\nI tried the same with nn, got 1.35 on 5CV but only 5.9 on leaderboard ?? This is not a classic supervised setup. Anyone similar results?",
      "votes": null
    },
    {
      "id": "397089",
      "postDate": "10/01/2018 21:57:42",
      "content": "<p><a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit</a></p>",
      "rawMarkdown": "https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit",
      "votes": null
    },
    {
      "id": "397095",
      "postDate": "10/01/2018 22:10:53",
      "content": "<p>Probably you made a mistake in the metric calculation. This is not a regular multiclass logloss. </p>",
      "rawMarkdown": "Probably you made a mistake in the metric calculation. This is not a regular multiclass logloss.",
      "votes": null
    },
    {
      "id": "397096",
      "postDate": "10/01/2018 22:13:46",
      "content": "<p>looks like Ni in the metric in not the size of the trainset, but the sum of positives cases for each class </p>",
      "rawMarkdown": "looks like Ni in the metric in not the size of the trainset, but the sum of positives cases for each class",
      "votes": null
    },
    {
      "id": "397111",
      "postDate": "10/01/2018 22:31:42",
      "content": "<p>from eval page: \"where N is the number of objects (positives only?) in the class set, M is the number of classes...\"</p>",
      "rawMarkdown": "from eval page: \"where N is the number of objects (positives only?) in the class set, M is the number of classes...\"",
      "votes": null
    },
    {
      "id": "397112",
      "postDate": "10/01/2018 22:32:37",
      "content": "<p>yep probably the Ni division.. dude, every time this special error functions for all the physic tasks :)\nIs there a reference code somewhere?</p>",
      "rawMarkdown": "yep probably the Ni division.. dude, every time this special error functions for all the physic tasks :)\nIs there a reference code somewhere?",
      "votes": null
    },
    {
      "id": "397118",
      "postDate": "10/01/2018 22:46:27",
      "content": "<p>I have done this for R, but the idea is simple. just sum all log(prediction) for all 1 in the OHE target matrix. Don't forget to normalize rows to unit and limit zeros and ones.</p>\n\n<p>score &lt;- function( tgt, ypred, W=rep(1/14,14) ){</p>\n\n<pre><code>pred &lt;- ypred / rowSums(ypred)  \npred[ pred &lt; 1e-15 ] &amp;lt;- 1e-15\npred[ pred &gt; (1-1e-15) ] &amp;lt;- (1-1e-15)\nres=0\nfor( i in 1:ncol(tgt) )  res &amp;lt;- res + W[i] * sum( tgt[,i]*log(pred[,i] ) ) / sum(tgt[,i])\nreturn( -res/sum(W) )\n</code></pre>\n\n<p>}</p>",
      "rawMarkdown": "I have done this for R, but the idea is simple. just sum all log(prediction) for all 1 in the OHE target matrix. Don't forget to normalize rows to unit and limit zeros and ones.\n\nscore &lt;- function( tgt, ypred, W=rep(1/14,14) ){\n\n    pred &lt;- ypred / rowSums(ypred)  \n    pred[ pred &lt; 1e-15 ] &lt;- 1e-15\n    pred[ pred &gt; (1-1e-15) ] &lt;- (1-1e-15)\n    res=0\n    for( i in 1:ncol(tgt) )  res &lt;- res + W[i] * sum( tgt[,i]*log(pred[,i] ) ) / sum(tgt[,i])\n    return( -res/sum(W) )\n\n}",
      "votes": null
    },
    {
      "id": "397273",
      "postDate": "10/02/2018 08:16:07",
      "content": "<p>I think I calculated this \"weighted multi-class logarithmic loss\" (wmcll) correctly (assume all w=1). I got wmcll=2.63906 for all classes equal likely on 5CV. With my last lgbm setting I got wmcll=2.18578 on 5CV, but why da heck scores this 6.x on leaderboard? Then I was thinking of the ghost \"class_99\", which I set to 0.0 for creating submission file, which is totally wrong. I tried setting the last 14th-class to 0.0 in my wmcll and I got 5.032 for all classes equal likely. So this drops a lot when set class_99=0. Who invented a 15th class without having any training samples? Maybe we can probe the base probability from class_99 with leaderboard feedback. Anyone done this?</p>",
      "rawMarkdown": "I think I calculated this \"weighted multi-class logarithmic loss\" (wmcll) correctly (assume all w=1). I got wmcll=2.63906 for all classes equal likely on 5CV. With my last lgbm setting I got wmcll=2.18578 on 5CV, but why da heck scores this 6.x on leaderboard? Then I was thinking of the ghost \"class_99\", which I set to 0.0 for creating submission file, which is totally wrong. I tried setting the last 14th-class to 0.0 in my wmcll and I got 5.032 for all classes equal likely. So this drops a lot when set class_99=0. Who invented a 15th class without having any training samples? Maybe we can probe the base probability from class_99 with leaderboard feedback. Anyone done this?",
      "votes": null
    },
    {
      "id": "397702",
      "postDate": "10/03/2018 00:37:00",
      "content": "<p>This is logloss, errors are hugely penalized. Someone must probe Class_99 distribution in Public LB. If I calculated my mlogloss correctly and assumed the right distribution for Class99 I got CV around 1.40  locally using equal weights and LB 1.54</p>",
      "rawMarkdown": "This is logloss, errors are hugely penalized. Someone must probe Class_99 distribution in Public LB. If I calculated my mlogloss correctly and assumed the right distribution for Class99 I got CV around 1.40  locally using equal weights and LB 1.54",
      "votes": null
    },
    {
      "id": "398775",
      "postDate": "10/04/2018 15:11:52",
      "content": "<p>got any tips on \"probing\"? feel like I'm missing out on a party</p>",
      "rawMarkdown": "got any tips on \"probing\"? feel like I'm missing out on a party",
      "votes": null
    },
    {
      "id": "399280",
      "postDate": "10/05/2018 14:35:39",
      "content": "<p>Hey Giba, are you also using a custom gradient and a custom hessian functions inside LightGBM? (assuming you are indeed using LightGBM).</p>",
      "rawMarkdown": "Hey Giba, are you also using a custom gradient and a custom hessian functions inside LightGBM? (assuming you are indeed using LightGBM).",
      "votes": null
    },
    {
      "id": "399649",
      "postDate": "10/06/2018 11:35:39",
      "content": "<p>Regular Lightgbm multiclass loss</p>",
      "rawMarkdown": "Regular Lightgbm multiclass loss",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 397089,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "10/01/2018 21:57:42",
      "content": "<p><a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 397095,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "10/01/2018 22:10:53",
      "content": "<p>Probably you made a mistake in the metric calculation. This is not a regular multiclass logloss. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 397096,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "10/01/2018 22:13:46",
      "content": "<p>looks like Ni in the metric in not the size of the trainset, but the sum of positives cases for each class </p>",
      "votes": null,
      "replies": [
        {
          "id": 397112,
          "author_name": "mjahrer",
          "author_url": "",
          "post_date": "10/01/2018 22:32:37",
          "content": "<p>yep probably the Ni division.. dude, every time this special error functions for all the physic tasks :)\nIs there a reference code somewhere?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 397118,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "10/01/2018 22:46:27",
          "content": "<p>I have done this for R, but the idea is simple. just sum all log(prediction) for all 1 in the OHE target matrix. Don't forget to normalize rows to unit and limit zeros and ones.</p>\n\n<p>score &lt;- function( tgt, ypred, W=rep(1/14,14) ){</p>\n\n<pre><code>pred &lt;- ypred / rowSums(ypred)  \npred[ pred &lt; 1e-15 ] &amp;lt;- 1e-15\npred[ pred &gt; (1-1e-15) ] &amp;lt;- (1-1e-15)\nres=0\nfor( i in 1:ncol(tgt) )  res &amp;lt;- res + W[i] * sum( tgt[,i]*log(pred[,i] ) ) / sum(tgt[,i])\nreturn( -res/sum(W) )\n</code></pre>\n\n<p>}</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 397273,
          "author_name": "mjahrer",
          "author_url": "",
          "post_date": "10/02/2018 08:16:07",
          "content": "<p>I think I calculated this \"weighted multi-class logarithmic loss\" (wmcll) correctly (assume all w=1). I got wmcll=2.63906 for all classes equal likely on 5CV. With my last lgbm setting I got wmcll=2.18578 on 5CV, but why da heck scores this 6.x on leaderboard? Then I was thinking of the ghost \"class_99\", which I set to 0.0 for creating submission file, which is totally wrong. I tried setting the last 14th-class to 0.0 in my wmcll and I got 5.032 for all classes equal likely. So this drops a lot when set class_99=0. Who invented a 15th class without having any training samples? Maybe we can probe the base probability from class_99 with leaderboard feedback. Anyone done this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 397702,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "10/03/2018 00:37:00",
          "content": "<p>This is logloss, errors are hugely penalized. Someone must probe Class_99 distribution in Public LB. If I calculated my mlogloss correctly and assumed the right distribution for Class99 I got CV around 1.40  locally using equal weights and LB 1.54</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 398775,
          "author_name": "abednadir",
          "author_url": "",
          "post_date": "10/04/2018 15:11:52",
          "content": "<p>got any tips on \"probing\"? feel like I'm missing out on a party</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 399280,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "10/05/2018 14:35:39",
          "content": "<p>Hey Giba, are you also using a custom gradient and a custom hessian functions inside LightGBM? (assuming you are indeed using LightGBM).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 399649,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "10/06/2018 11:35:39",
          "content": "<p>Regular Lightgbm multiclass loss</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 397111,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "10/01/2018 22:31:42",
      "content": "<p>from eval page: \"where N is the number of objects (positives only?) in the class set, M is the number of classes...\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "397071": "I checked it a few times, maybe there is still an error in my processing code. A standard lgbm trained on the training_set_metadata.csv (10 numeric features) gives about 1.4 mlogloss on 5CV, but scores only 6.x on leaderboard ?? What is going on here?\nI tried the same with nn, got 1.35 on 5CV but only 5.9 on leaderboard ?? This is not a classic supervised setup. Anyone similar results?",
    "397089": "https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit",
    "397095": "Probably you made a mistake in the metric calculation. This is not a regular multiclass logloss.",
    "397096": "looks like Ni in the metric in not the size of the trainset, but the sum of positives cases for each class",
    "397111": "from eval page: \"where N is the number of objects (positives only?) in the class set, M is the number of classes...\"",
    "397112": "yep probably the Ni division.. dude, every time this special error functions for all the physic tasks :)\nIs there a reference code somewhere?",
    "397118": "I have done this for R, but the idea is simple. just sum all log(prediction) for all 1 in the OHE target matrix. Don't forget to normalize rows to unit and limit zeros and ones.\n\nscore &lt;- function( tgt, ypred, W=rep(1/14,14) ){\n\n    pred &lt;- ypred / rowSums(ypred)  \n    pred[ pred &lt; 1e-15 ] &lt;- 1e-15\n    pred[ pred &gt; (1-1e-15) ] &lt;- (1-1e-15)\n    res=0\n    for( i in 1:ncol(tgt) )  res &lt;- res + W[i] * sum( tgt[,i]*log(pred[,i] ) ) / sum(tgt[,i])\n    return( -res/sum(W) )\n\n}",
    "397273": "I think I calculated this \"weighted multi-class logarithmic loss\" (wmcll) correctly (assume all w=1). I got wmcll=2.63906 for all classes equal likely on 5CV. With my last lgbm setting I got wmcll=2.18578 on 5CV, but why da heck scores this 6.x on leaderboard? Then I was thinking of the ghost \"class_99\", which I set to 0.0 for creating submission file, which is totally wrong. I tried setting the last 14th-class to 0.0 in my wmcll and I got 5.032 for all classes equal likely. So this drops a lot when set class_99=0. Who invented a 15th class without having any training samples? Maybe we can probe the base probability from class_99 with leaderboard feedback. Anyone done this?",
    "397702": "This is logloss, errors are hugely penalized. Someone must probe Class_99 distribution in Public LB. If I calculated my mlogloss correctly and assumed the right distribution for Class99 I got CV around 1.40  locally using equal weights and LB 1.54",
    "398775": "got any tips on \"probing\"? feel like I'm missing out on a party",
    "399280": "Hey Giba, are you also using a custom gradient and a custom hessian functions inside LightGBM? (assuming you are indeed using LightGBM).",
    "399649": "Regular Lightgbm multiclass loss"
  },
  "source": "meta"
}