{
  "id": 69795,
  "title": "Multi-weight logloss for keras",
  "url": "/competitions/PLAsTiCC-2018/discussion/69795",
  "author_name": "",
  "post_date": "2018-10-27T12:07:04.912555900Z",
  "votes": 7,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Is there any way to create multi-weight logloss loss function for keras, I am trying to create one but I can't make it work. \nHas anyone able to use it as a loss function and if not what are you guys using?</p>",
  "messages": [
    {
      "id": "411096",
      "postDate": "10/27/2018 12:07:04",
      "content": "<p>Is there any way to create multi-weight logloss loss function for keras, I am trying to create one but I can't make it work. \nHas anyone able to use it as a loss function and if not what are you guys using?</p>",
      "rawMarkdown": "Is there any way to create multi-weight logloss loss function for keras, I am trying to create one but I can't make it work. \nHas anyone able to use it as a loss function and if not what are you guys using?",
      "votes": null
    },
    {
      "id": "411103",
      "postDate": "10/27/2018 12:25:28",
      "content": "<pre><code>def mywloss(y_true,y_pred):  \n     yc=tf.clip_by_value(y_pred,1e-15,1-1e-15)\n     loss=-(tf.reduce_mean(tf.reduce_mean(y_true*tf.log(yc),axis=0)/wtable))\n     return loss\n</code></pre>\n\n<p>Edit:\n wtable - is a numpy 1d array with (the number of times class y_true occur in the data set)/(size of data set)  </p>",
      "rawMarkdown": "def mywloss(y_true,y_pred):  \n         yc=tf.clip_by_value(y_pred,1e-15,1-1e-15)\n         loss=-(tf.reduce_mean(tf.reduce_mean(y_true*tf.log(yc),axis=0)/wtable))\n         return loss\n\nEdit:\n wtable - is a numpy 1d array with (the number of times class y_true occur in the data set)/(size of data set)",
      "votes": null
    },
    {
      "id": "411104",
      "postDate": "10/27/2018 12:28:41",
      "content": "<p>Thanks! that is way shorter than what I was trying to do.</p>",
      "rawMarkdown": "Thanks! that is way shorter than what I was trying to do.",
      "votes": null
    },
    {
      "id": "411118",
      "postDate": "10/27/2018 13:01:04",
      "content": "<p>You divide by <code>wtable,</code> hence these are not what people usually call weights.  They are rather their inverse.</p>",
      "rawMarkdown": "You divide by `wtable,` hence these are not what people usually call weights.  They are rather their inverse.",
      "votes": null
    },
    {
      "id": "411122",
      "postDate": "10/27/2018 13:20:58",
      "content": "<p>You can also just use sample weights.</p>",
      "rawMarkdown": "You can also just use sample weights.",
      "votes": null
    },
    {
      "id": "411132",
      "postDate": "10/27/2018 13:32:50",
      "content": "<p>That's what I use indeed.</p>",
      "rawMarkdown": "That's what I use indeed.",
      "votes": null
    },
    {
      "id": "411145",
      "postDate": "10/27/2018 13:46:55",
      "content": "<p>You are correct. thanks,\nedited.</p>",
      "rawMarkdown": "You are correct. thanks,\nedited.",
      "votes": null
    },
    {
      "id": "417672",
      "postDate": "11/08/2018 16:04:02",
      "content": "<p>&gt;(the number of times class y_true occur in the data set)/(size of data set)</p>\n\n<p>why do you divide by the 'size of data set', isn't that just a constant and 'N_j' for class j should be enough?</p>",
      "rawMarkdown": "&gt;(the number of times class y_true occur in the data set)/(size of data set)\n\nwhy do you divide by the 'size of data set', isn't that just a constant and 'N_j' for class j should be enough?",
      "votes": null
    },
    {
      "id": "417805",
      "postDate": "11/08/2018 19:53:48",
      "content": "<p>When you do the calculation as a score on a full set, you don't need it. you do need to replace the reducemean with reducesum and it will take care of this constant. But when you calculate on a partial batch you must use reducemean, otherwise the result will depend on the size of the batch, hance you need to introduce this constant.</p>",
      "rawMarkdown": "When you do the calculation as a score on a full set, you don't need it. you do need to replace the reducemean with reducesum and it will take care of this constant. But when you calculate on a partial batch you must use reducemean, otherwise the result will depend on the size of the batch, hance you need to introduce this constant.",
      "votes": null
    },
    {
      "id": "417952",
      "postDate": "11/09/2018 03:42:53",
      "content": "<p>Have you run into issues with clipping blocking gradients while training? I tried a few simple networks and found that I am better off without the clipping in training loss.</p>",
      "rawMarkdown": "Have you run into issues with clipping blocking gradients while training? I tried a few simple networks and found that I am better off without the clipping in training loss.",
      "votes": null
    },
    {
      "id": "417989",
      "postDate": "11/09/2018 04:47:47",
      "content": "<p>I didn't. The clipping is for extreme valus, I don't see why it should cause issues. Actually, in later version I even clipped to 1e-4, because I didn't want one extreme value to degrade the full batch.</p>",
      "rawMarkdown": "I didn't. The clipping is for extreme valus, I don't see why it should cause issues. Actually, in later version I even clipped to 1e-4, because I didn't want one extreme value to degrade the full batch.",
      "votes": null
    },
    {
      "id": "418369",
      "postDate": "11/09/2018 18:39:34",
      "content": "<p>and you can also just use class weights.</p>",
      "rawMarkdown": "and you can also just use class weights.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 411103,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "10/27/2018 12:25:28",
      "content": "<pre><code>def mywloss(y_true,y_pred):  \n     yc=tf.clip_by_value(y_pred,1e-15,1-1e-15)\n     loss=-(tf.reduce_mean(tf.reduce_mean(y_true*tf.log(yc),axis=0)/wtable))\n     return loss\n</code></pre>\n\n<p>Edit:\n wtable - is a numpy 1d array with (the number of times class y_true occur in the data set)/(size of data set)  </p>",
      "votes": null,
      "replies": [
        {
          "id": 411104,
          "author_name": "satian",
          "author_url": "",
          "post_date": "10/27/2018 12:28:41",
          "content": "<p>Thanks! that is way shorter than what I was trying to do.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411118,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/27/2018 13:01:04",
          "content": "<p>You divide by <code>wtable,</code> hence these are not what people usually call weights.  They are rather their inverse.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411145,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "10/27/2018 13:46:55",
          "content": "<p>You are correct. thanks,\nedited.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417672,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "11/08/2018 16:04:02",
          "content": "<p>&gt;(the number of times class y_true occur in the data set)/(size of data set)</p>\n\n<p>why do you divide by the 'size of data set', isn't that just a constant and 'N_j' for class j should be enough?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417805,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "11/08/2018 19:53:48",
          "content": "<p>When you do the calculation as a score on a full set, you don't need it. you do need to replace the reducemean with reducesum and it will take care of this constant. But when you calculate on a partial batch you must use reducemean, otherwise the result will depend on the size of the batch, hance you need to introduce this constant.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417952,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "11/09/2018 03:42:53",
          "content": "<p>Have you run into issues with clipping blocking gradients while training? I tried a few simple networks and found that I am better off without the clipping in training loss.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417989,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "11/09/2018 04:47:47",
          "content": "<p>I didn't. The clipping is for extreme valus, I don't see why it should cause issues. Actually, in later version I even clipped to 1e-4, because I didn't want one extreme value to degrade the full batch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 411122,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "10/27/2018 13:20:58",
      "content": "<p>You can also just use sample weights.</p>",
      "votes": null,
      "replies": [
        {
          "id": 411132,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/27/2018 13:32:50",
          "content": "<p>That's what I use indeed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418369,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "11/09/2018 18:39:34",
          "content": "<p>and you can also just use class weights.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "411096": "Is there any way to create multi-weight logloss loss function for keras, I am trying to create one but I can't make it work. \nHas anyone able to use it as a loss function and if not what are you guys using?",
    "411103": "def mywloss(y_true,y_pred):  \n         yc=tf.clip_by_value(y_pred,1e-15,1-1e-15)\n         loss=-(tf.reduce_mean(tf.reduce_mean(y_true*tf.log(yc),axis=0)/wtable))\n         return loss\n\nEdit:\n wtable - is a numpy 1d array with (the number of times class y_true occur in the data set)/(size of data set)",
    "411104": "Thanks! that is way shorter than what I was trying to do.",
    "411118": "You divide by `wtable,` hence these are not what people usually call weights.  They are rather their inverse.",
    "411122": "You can also just use sample weights.",
    "411132": "That's what I use indeed.",
    "411145": "You are correct. thanks,\nedited.",
    "417672": "&gt;(the number of times class y_true occur in the data set)/(size of data set)\n\nwhy do you divide by the 'size of data set', isn't that just a constant and 'N_j' for class j should be enough?",
    "417805": "When you do the calculation as a score on a full set, you don't need it. you do need to replace the reducemean with reducesum and it will take care of this constant. But when you calculate on a partial batch you must use reducemean, otherwise the result will depend on the size of the batch, hance you need to introduce this constant.",
    "417952": "Have you run into issues with clipping blocking gradients while training? I tried a few simple networks and found that I am better off without the clipping in training loss.",
    "417989": "I didn't. The clipping is for extreme valus, I don't see why it should cause issues. Actually, in later version I even clipped to 1e-4, because I didn't want one extreme value to degrade the full batch.",
    "418369": "and you can also just use class weights."
  },
  "source": "meta"
}