{
  "id": 51303,
  "title": "Knowledge Distillation in Keras",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/51303",
  "author_name": "Gionata Benelli",
  "post_date": "2018-03-07T15:34:56.235000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone, I know that i'm late, but I need some insight on knowledge distillation, by Hinton.\nI have my best model at 77%LB with just 15K parameters, so i tried to build an enseble of 5 models getting 84%LB, not much, but it's something.</p>\n\n<p>I tried to use knowledge distillation to give some increase to the small model using different small temperatures(3,5,7) and my LB has always reduced with respect to the starting values of the model. \nIt is normal?consider that I'm finetuning the last layer of my model.</p>\n\n<p>I'm using keras as a framework and a custom loss for training, in short I'm using the weighted sum of categorical crossentropy on hard and soft target.</p>",
  "messages": [
    {
      "id": 292184,
      "postDate": "2018-03-07T15:34:56.237Z",
      "content": "<p>Hi everyone, I know that i'm late, but I need some insight on knowledge distillation, by Hinton.\nI have my best model at 77%LB with just 15K parameters, so i tried to build an enseble of 5 models getting 84%LB, not much, but it's something.</p>\n\n<p>I tried to use knowledge distillation to give some increase to the small model using different small temperatures(3,5,7) and my LB has always reduced with respect to the starting values of the model. \nIt is normal?consider that I'm finetuning the last layer of my model.</p>\n\n<p>I'm using keras as a framework and a custom loss for training, in short I'm using the weighted sum of categorical crossentropy on hard and soft target.</p>",
      "rawMarkdown": "Hi everyone, I know that i'm late, but I need some insight on knowledge distillation, by Hinton.\nI have my best model at 77%LB with just 15K parameters, so i tried to build an enseble of 5 models getting 84%LB, not much, but it's something.\n\nI tried to use knowledge distillation to give some increase to the small model using different small temperatures(3,5,7) and my LB has always reduced with respect to the starting values of the model. \nIt is normal?consider that I'm finetuning the last layer of my model.\n\nI'm using keras as a framework and a custom loss for training, in short I'm using the weighted sum of categorical crossentropy on hard and soft target."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "292184": "Hi everyone, I know that i'm late, but I need some insight on knowledge distillation, by Hinton.\nI have my best model at 77%LB with just 15K parameters, so i tried to build an enseble of 5 models getting 84%LB, not much, but it's something.\n\nI tried to use knowledge distillation to give some increase to the small model using different small temperatures(3,5,7) and my LB has always reduced with respect to the starting values of the model. \nIt is normal?consider that I'm finetuning the last layer of my model.\n\nI'm using keras as a framework and a custom loss for training, in short I'm using the weighted sum of categorical crossentropy on hard and soft target."
  }
}