{
  "id": 47645,
  "title": "LB0.88171 -> Why 9 CNN-based network MEL+MFCC fully-connected ensembling didnt work?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47645",
  "author_name": "",
  "post_date": "2018-01-17T08:44:55.604248800Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, I joined this competition pretty late (just 7 days before deadline), so I used Heng's baseline (THANKS for all your discussions and making code available).</p>\n\n<p>I then made a network consisting of (initially 12, then 9) subnetworks:</p>\n\n<pre><code> def forward(self, x_mel, x_mfcc):\n    #x1 = self.seResNet_16.forward(x_mel)\n    #x2 = self.seResNet_32.forward(x_mel)\n    x3 = self.densenet121.forward(x_mel)\n    x4 = self.densenet169.forward(x_mel)\n    x5 = self.resnet34.forward(x_mel)\n    x6 = self.resnet152.forward(x_mel)\n\n    #x7  = self.seResNet_16_mfcc.forward(x_mfcc)\n    x8  = self.seResNet_32_mfcc.forward(x_mfcc)\n    x9  = self.densenet161_mfcc.forward(x_mfcc)\n    x10 = self.densenet201_mfcc.forward(x_mfcc)\n    x11 = self.resnet50_mfcc.forward(x_mfcc)\n    x12 = self.resnet101_mfcc.forward(x_mfcc)\n\n    #x = torch.cat((x1,x2,x3,x4,x5,x6, x7,x8,x9,x10,x11,x12), 1) # batch_size, n * num_classes\n    x = torch.cat((x3,x4,x5,x6, x8,x9,x10,x11,x12), 1) # batch_size, n * num_classes\n\n    x = F.relu(self.fc1(x))\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.fc2(x))\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.fc3(x))\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = self.fc4(x)\n\n    return x, x3,x4,x5,x6, x8,x9,x10,x11,x12\n</code></pre>\n\n<p>as you can see the network returns each subnetwork predictions <code>x3,x4,x5,x6, x8,x9,x10,x11,x12</code> and the final ensembled prediction <code>x</code>. </p>\n\n<p>The main idea is to let a fully connected networks of 4 layers do the ensembling for us instead of us doing \"feature engineering on the ensembling\" (i.e. manually picking bad probabilities, choosing whether to do linear, geometric averaging, etc.).</p>\n\n<p>The loss function is such that at the beginning of training it focuses on the subnetwork losses (so we can later interpret what the ensembling is doing), and later shifts its weight to the ensembled prediction:</p>\n\n<pre><code>        logits, l_3,l_4,l_5,l_6, l_8,l_9,l_10,l_11,l_12   = data_parallel(net, (tensors_mel, tensors_mfcc))\n        probs   = F.softmax(logits,dim=1)\n\n        labels = Variable(labels).cuda()\n        partial_loss = (\\\n            F.cross_entropy(l_3, labels) + \\\n            F.cross_entropy(l_4, labels) + \\\n            F.cross_entropy(l_5, labels) + \\\n            F.cross_entropy(l_6, labels) + \\\n            F.cross_entropy(l_8, labels) + \\\n            F.cross_entropy(l_9, labels) + \\\n            F.cross_entropy(l_10, labels) + \\\n            F.cross_entropy(l_11, labels) + \\\n            F.cross_entropy(l_12, labels) ) / 9.\n\n        iter_cutoff = 10000\n        partial_loss_factor = float(np.clip((iter_cutoff - i)/iter_cutoff, 0.1, 1.))\n        cross_entropy_loss_factor = 1. - partial_loss_factor\n\n        loss    = cross_entropy_loss_factor * F.cross_entropy(logits, labels) + partial_loss_factor * partial_loss \n        acc     = top_accuracy(probs, labels, top_k=(1,))\n</code></pre>\n\n<p>The weird thing is that after enough iters (50k) if I submit ensembled predictions (from <code>x</code>) I got a lower score than submitting simply averaged predictions from  <code>l_3,l_4,l_5,l_6, l_8,l_9,l_10,l_11,l_12</code>.</p>\n\n<p>Code is available @ <a href=\"https://github.com/antorsae/tensorflow-speech-recognition-challenge\">https://github.com/antorsae/tensorflow-speech-recognition-challenge</a></p>",
  "messages": [
    {
      "id": "269761",
      "postDate": "01/17/2018 08:44:55",
      "content": "<p>Hi, I joined this competition pretty late (just 7 days before deadline), so I used Heng's baseline (THANKS for all your discussions and making code available).</p>\n\n<p>I then made a network consisting of (initially 12, then 9) subnetworks:</p>\n\n<pre><code> def forward(self, x_mel, x_mfcc):\n    #x1 = self.seResNet_16.forward(x_mel)\n    #x2 = self.seResNet_32.forward(x_mel)\n    x3 = self.densenet121.forward(x_mel)\n    x4 = self.densenet169.forward(x_mel)\n    x5 = self.resnet34.forward(x_mel)\n    x6 = self.resnet152.forward(x_mel)\n\n    #x7  = self.seResNet_16_mfcc.forward(x_mfcc)\n    x8  = self.seResNet_32_mfcc.forward(x_mfcc)\n    x9  = self.densenet161_mfcc.forward(x_mfcc)\n    x10 = self.densenet201_mfcc.forward(x_mfcc)\n    x11 = self.resnet50_mfcc.forward(x_mfcc)\n    x12 = self.resnet101_mfcc.forward(x_mfcc)\n\n    #x = torch.cat((x1,x2,x3,x4,x5,x6, x7,x8,x9,x10,x11,x12), 1) # batch_size, n * num_classes\n    x = torch.cat((x3,x4,x5,x6, x8,x9,x10,x11,x12), 1) # batch_size, n * num_classes\n\n    x = F.relu(self.fc1(x))\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.fc2(x))\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.fc3(x))\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = self.fc4(x)\n\n    return x, x3,x4,x5,x6, x8,x9,x10,x11,x12\n</code></pre>\n\n<p>as you can see the network returns each subnetwork predictions <code>x3,x4,x5,x6, x8,x9,x10,x11,x12</code> and the final ensembled prediction <code>x</code>. </p>\n\n<p>The main idea is to let a fully connected networks of 4 layers do the ensembling for us instead of us doing \"feature engineering on the ensembling\" (i.e. manually picking bad probabilities, choosing whether to do linear, geometric averaging, etc.).</p>\n\n<p>The loss function is such that at the beginning of training it focuses on the subnetwork losses (so we can later interpret what the ensembling is doing), and later shifts its weight to the ensembled prediction:</p>\n\n<pre><code>        logits, l_3,l_4,l_5,l_6, l_8,l_9,l_10,l_11,l_12   = data_parallel(net, (tensors_mel, tensors_mfcc))\n        probs   = F.softmax(logits,dim=1)\n\n        labels = Variable(labels).cuda()\n        partial_loss = (\\\n            F.cross_entropy(l_3, labels) + \\\n            F.cross_entropy(l_4, labels) + \\\n            F.cross_entropy(l_5, labels) + \\\n            F.cross_entropy(l_6, labels) + \\\n            F.cross_entropy(l_8, labels) + \\\n            F.cross_entropy(l_9, labels) + \\\n            F.cross_entropy(l_10, labels) + \\\n            F.cross_entropy(l_11, labels) + \\\n            F.cross_entropy(l_12, labels) ) / 9.\n\n        iter_cutoff = 10000\n        partial_loss_factor = float(np.clip((iter_cutoff - i)/iter_cutoff, 0.1, 1.))\n        cross_entropy_loss_factor = 1. - partial_loss_factor\n\n        loss    = cross_entropy_loss_factor * F.cross_entropy(logits, labels) + partial_loss_factor * partial_loss \n        acc     = top_accuracy(probs, labels, top_k=(1,))\n</code></pre>\n\n<p>The weird thing is that after enough iters (50k) if I submit ensembled predictions (from <code>x</code>) I got a lower score than submitting simply averaged predictions from  <code>l_3,l_4,l_5,l_6, l_8,l_9,l_10,l_11,l_12</code>.</p>\n\n<p>Code is available @ <a href=\"https://github.com/antorsae/tensorflow-speech-recognition-challenge\">https://github.com/antorsae/tensorflow-speech-recognition-challenge</a></p>",
      "rawMarkdown": "Hi, I joined this competition pretty late (just 7 days before deadline), so I used Heng's baseline (THANKS for all your discussions and making code available).\n\nI then made a network consisting of (initially 12, then 9) subnetworks:\n\n     def forward(self, x_mel, x_mfcc):\n        #x1 = self.seResNet_16.forward(x_mel)\n        #x2 = self.seResNet_32.forward(x_mel)\n        x3 = self.densenet121.forward(x_mel)\n        x4 = self.densenet169.forward(x_mel)\n        x5 = self.resnet34.forward(x_mel)\n        x6 = self.resnet152.forward(x_mel)\n\n        #x7  = self.seResNet_16_mfcc.forward(x_mfcc)\n        x8  = self.seResNet_32_mfcc.forward(x_mfcc)\n        x9  = self.densenet161_mfcc.forward(x_mfcc)\n        x10 = self.densenet201_mfcc.forward(x_mfcc)\n        x11 = self.resnet50_mfcc.forward(x_mfcc)\n        x12 = self.resnet101_mfcc.forward(x_mfcc)\n\n        #x = torch.cat((x1,x2,x3,x4,x5,x6, x7,x8,x9,x10,x11,x12), 1) # batch_size, n * num_classes\n        x = torch.cat((x3,x4,x5,x6, x8,x9,x10,x11,x12), 1) # batch_size, n * num_classes\n\n        x = F.relu(self.fc1(x))\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.fc2(x))\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.fc3(x))\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = self.fc4(x)\n\n        return x, x3,x4,x5,x6, x8,x9,x10,x11,x12\n\nas you can see the network returns each subnetwork predictions `x3,x4,x5,x6, x8,x9,x10,x11,x12` and the final ensembled prediction `x`. \n\nThe main idea is to let a fully connected networks of 4 layers do the ensembling for us instead of us doing \"feature engineering on the ensembling\" (i.e. manually picking bad probabilities, choosing whether to do linear, geometric averaging, etc.).\n\nThe loss function is such that at the beginning of training it focuses on the subnetwork losses (so we can later interpret what the ensembling is doing), and later shifts its weight to the ensembled prediction:\n\n            logits, l_3,l_4,l_5,l_6, l_8,l_9,l_10,l_11,l_12   = data_parallel(net, (tensors_mel, tensors_mfcc))\n            probs   = F.softmax(logits,dim=1)\n\n            labels = Variable(labels).cuda()\n            partial_loss = (\\\n                F.cross_entropy(l_3, labels) + \\\n                F.cross_entropy(l_4, labels) + \\\n                F.cross_entropy(l_5, labels) + \\\n                F.cross_entropy(l_6, labels) + \\\n                F.cross_entropy(l_8, labels) + \\\n                F.cross_entropy(l_9, labels) + \\\n                F.cross_entropy(l_10, labels) + \\\n                F.cross_entropy(l_11, labels) + \\\n                F.cross_entropy(l_12, labels) ) / 9.\n\n            iter_cutoff = 10000\n            partial_loss_factor = float(np.clip((iter_cutoff - i)/iter_cutoff, 0.1, 1.))\n            cross_entropy_loss_factor = 1. - partial_loss_factor\n\n            loss    = cross_entropy_loss_factor * F.cross_entropy(logits, labels) + partial_loss_factor * partial_loss \n            acc     = top_accuracy(probs, labels, top_k=(1,))\n\nThe weird thing is that after enough iters (50k) if I submit ensembled predictions (from `x`) I got a lower score than submitting simply averaged predictions from  `l_3,l_4,l_5,l_6, l_8,l_9,l_10,l_11,l_12`.\n\nCode is available @ https://github.com/antorsae/tensorflow-speech-recognition-challenge",
      "votes": null
    },
    {
      "id": "270393",
      "postDate": "01/18/2018 07:17:23",
      "content": "<p>for each of the models, post the csv file and raw probability and LB score (if you have).\npost also blend weights, the resulting csv and LB score.</p>\n\n<p>I see if i can combine better</p>",
      "rawMarkdown": "for each of the models, post the csv file and raw probability and LB score (if you have).\npost also blend weights, the resulting csv and LB score.\n\nI see if i can combine better",
      "votes": null
    },
    {
      "id": "270395",
      "postDate": "01/18/2018 07:25:44",
      "content": "<p>Hi Heng,</p>\n\n<p><code>probs.uint8.memmap</code> is the <code>x</code> ensembled by the 4 FC layers.\n<code>probs_ll.uint8.memmap</code> are the 9 individual logits (shape (-1, 12,9) iirc. </p>\n\n<p>I did a simple averaging on the 9 logits and the LB came out better than the <code>x</code> ensembling, which I don't understand.</p>",
      "rawMarkdown": "Hi Heng,\n\n`probs.uint8.memmap` is the `x` ensembled by the 4 FC layers.\n`probs_ll.uint8.memmap` are the 9 individual logits (shape (-1, 12,9) iirc. \n\nI did a simple averaging on the 9 logits and the LB came out better than the `x` ensembling, which I don't understand.",
      "votes": null
    },
    {
      "id": "270980",
      "postDate": "01/19/2018 11:15:10",
      "content": "<p>Here's a late test I just with a CTC approach (Baidu's Deep Speech): <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827</a></p>",
      "rawMarkdown": "Here's a late test I just with a CTC approach (Baidu's Deep Speech): https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 270393,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/18/2018 07:17:23",
      "content": "<p>for each of the models, post the csv file and raw probability and LB score (if you have).\npost also blend weights, the resulting csv and LB score.</p>\n\n<p>I see if i can combine better</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270395,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/18/2018 07:25:44",
      "content": "<p>Hi Heng,</p>\n\n<p><code>probs.uint8.memmap</code> is the <code>x</code> ensembled by the 4 FC layers.\n<code>probs_ll.uint8.memmap</code> are the 9 individual logits (shape (-1, 12,9) iirc. </p>\n\n<p>I did a simple averaging on the 9 logits and the LB came out better than the <code>x</code> ensembling, which I don't understand.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270980,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/19/2018 11:15:10",
      "content": "<p>Here's a late test I just with a CTC approach (Baidu's Deep Speech): <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "269761": "Hi, I joined this competition pretty late (just 7 days before deadline), so I used Heng's baseline (THANKS for all your discussions and making code available).\n\nI then made a network consisting of (initially 12, then 9) subnetworks:\n\n     def forward(self, x_mel, x_mfcc):\n        #x1 = self.seResNet_16.forward(x_mel)\n        #x2 = self.seResNet_32.forward(x_mel)\n        x3 = self.densenet121.forward(x_mel)\n        x4 = self.densenet169.forward(x_mel)\n        x5 = self.resnet34.forward(x_mel)\n        x6 = self.resnet152.forward(x_mel)\n\n        #x7  = self.seResNet_16_mfcc.forward(x_mfcc)\n        x8  = self.seResNet_32_mfcc.forward(x_mfcc)\n        x9  = self.densenet161_mfcc.forward(x_mfcc)\n        x10 = self.densenet201_mfcc.forward(x_mfcc)\n        x11 = self.resnet50_mfcc.forward(x_mfcc)\n        x12 = self.resnet101_mfcc.forward(x_mfcc)\n\n        #x = torch.cat((x1,x2,x3,x4,x5,x6, x7,x8,x9,x10,x11,x12), 1) # batch_size, n * num_classes\n        x = torch.cat((x3,x4,x5,x6, x8,x9,x10,x11,x12), 1) # batch_size, n * num_classes\n\n        x = F.relu(self.fc1(x))\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.fc2(x))\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.fc3(x))\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = self.fc4(x)\n\n        return x, x3,x4,x5,x6, x8,x9,x10,x11,x12\n\nas you can see the network returns each subnetwork predictions `x3,x4,x5,x6, x8,x9,x10,x11,x12` and the final ensembled prediction `x`. \n\nThe main idea is to let a fully connected networks of 4 layers do the ensembling for us instead of us doing \"feature engineering on the ensembling\" (i.e. manually picking bad probabilities, choosing whether to do linear, geometric averaging, etc.).\n\nThe loss function is such that at the beginning of training it focuses on the subnetwork losses (so we can later interpret what the ensembling is doing), and later shifts its weight to the ensembled prediction:\n\n            logits, l_3,l_4,l_5,l_6, l_8,l_9,l_10,l_11,l_12   = data_parallel(net, (tensors_mel, tensors_mfcc))\n            probs   = F.softmax(logits,dim=1)\n\n            labels = Variable(labels).cuda()\n            partial_loss = (\\\n                F.cross_entropy(l_3, labels) + \\\n                F.cross_entropy(l_4, labels) + \\\n                F.cross_entropy(l_5, labels) + \\\n                F.cross_entropy(l_6, labels) + \\\n                F.cross_entropy(l_8, labels) + \\\n                F.cross_entropy(l_9, labels) + \\\n                F.cross_entropy(l_10, labels) + \\\n                F.cross_entropy(l_11, labels) + \\\n                F.cross_entropy(l_12, labels) ) / 9.\n\n            iter_cutoff = 10000\n            partial_loss_factor = float(np.clip((iter_cutoff - i)/iter_cutoff, 0.1, 1.))\n            cross_entropy_loss_factor = 1. - partial_loss_factor\n\n            loss    = cross_entropy_loss_factor * F.cross_entropy(logits, labels) + partial_loss_factor * partial_loss \n            acc     = top_accuracy(probs, labels, top_k=(1,))\n\nThe weird thing is that after enough iters (50k) if I submit ensembled predictions (from `x`) I got a lower score than submitting simply averaged predictions from  `l_3,l_4,l_5,l_6, l_8,l_9,l_10,l_11,l_12`.\n\nCode is available @ https://github.com/antorsae/tensorflow-speech-recognition-challenge",
    "270393": "for each of the models, post the csv file and raw probability and LB score (if you have).\npost also blend weights, the resulting csv and LB score.\n\nI see if i can combine better",
    "270395": "Hi Heng,\n\n`probs.uint8.memmap` is the `x` ensembled by the 4 FC layers.\n`probs_ll.uint8.memmap` are the 9 individual logits (shape (-1, 12,9) iirc. \n\nI did a simple averaging on the 9 logits and the LB came out better than the `x` ensembling, which I don't understand.",
    "270980": "Here's a late test I just with a CTC approach (Baidu's Deep Speech): https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47827"
  },
  "source": "meta"
}