{
  "id": 47660,
  "title": "single models with high scores",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47660",
  "author_name": "",
  "post_date": "2018-01-17T12:50:26.119210600Z",
  "votes": 14,
  "comment_count": 14,
  "views": 0,
  "content": "<p>We had some good single models but failed to ensemble. The difference between our best model and the ensemble with 30 models is only 0.00646 (private score). On the other hand, the winner Heng CherKeng could get 89% from 86% models. So I am really interested how others are doing.</p>\n\n<p>Here are our private scores of the best single models:</p>\n\n<ul>\n<li>ResNext trained with all data and with crop: 0.89886 (same as the 43-th place in LB!)</li>\n<li>RexNext trained on train+valid with crop: 0.89686</li>\n<li>RexNext trained on train+valid without crop: 0.89533</li>\n<li>Densenet190 with all data and with crop: 0.89486</li>\n<li>Densenet190 on train+valid with crop: 0.89416</li>\n<li>WideResNet52-10 on train+test with crop: 0.89345</li>\n</ul>\n\n<p>What is your best single model score and the difference to the ensemble?</p>",
  "messages": [
    {
      "id": "269874",
      "postDate": "01/17/2018 12:50:26",
      "content": "<p>We had some good single models but failed to ensemble. The difference between our best model and the ensemble with 30 models is only 0.00646 (private score). On the other hand, the winner Heng CherKeng could get 89% from 86% models. So I am really interested how others are doing.</p>\n\n<p>Here are our private scores of the best single models:</p>\n\n<ul>\n<li>ResNext trained with all data and with crop: 0.89886 (same as the 43-th place in LB!)</li>\n<li>RexNext trained on train+valid with crop: 0.89686</li>\n<li>RexNext trained on train+valid without crop: 0.89533</li>\n<li>Densenet190 with all data and with crop: 0.89486</li>\n<li>Densenet190 on train+valid with crop: 0.89416</li>\n<li>WideResNet52-10 on train+test with crop: 0.89345</li>\n</ul>\n\n<p>What is your best single model score and the difference to the ensemble?</p>",
      "rawMarkdown": "We had some good single models but failed to ensemble. The difference between our best model and the ensemble with 30 models is only 0.00646 (private score). On the other hand, the winner Heng CherKeng could get 89% from 86% models. So I am really interested how others are doing.\n\nHere are our private scores of the best single models:\n\n - ResNext trained with all data and with crop: 0.89886 (same as the 43-th place in LB!)\n - RexNext trained on train+valid with crop: 0.89686\n - RexNext trained on train+valid without crop: 0.89533\n - Densenet190 with all data and with crop: 0.89486\n - Densenet190 on train+valid with crop: 0.89416\n - WideResNet52-10 on train+test with crop: 0.89345\n\nWhat is your best single model score and the difference to the ensemble?",
      "votes": null
    },
    {
      "id": "269909",
      "postDate": "01/17/2018 14:11:32",
      "content": "<p>My best single model (without bagging) is only 0.87 on public LB and 0.88 on private LB. Bagging it 5 times could lead to 0.88/0.89 but yours look very impressive! You mind sharing a bit more details/code?</p>",
      "rawMarkdown": "My best single model (without bagging) is only 0.87 on public LB and 0.88 on private LB. Bagging it 5 times could lead to 0.88/0.89 but yours look very impressive! You mind sharing a bit more details/code?",
      "votes": null
    },
    {
      "id": "269913",
      "postDate": "01/17/2018 14:21:52",
      "content": "<p>All our good models use 40x32 mel-spectrum which is quite similar to CIFAR 10 problems. So we took best performing CIFAR10 networks and trained them. So far, only PreAct-ResNets and DPN92 are failed. They got only 87% public LB.</p>",
      "rawMarkdown": "All our good models use 40x32 mel-spectrum which is quite similar to CIFAR 10 problems. So we took best performing CIFAR10 networks and trained them. So far, only PreAct-ResNets and DPN92 are failed. They got only 87% public LB.",
      "votes": null
    },
    {
      "id": "269917",
      "postDate": "01/17/2018 14:26:30",
      "content": "<p>That sounds fascinating! I used 40x110 mel-spectrum by default and didn't tune it. Was 40x32 the result of tuning from your side? You were doing conv2d on it correct? (not conv1d over the time axis?)</p>",
      "rawMarkdown": "That sounds fascinating! I used 40x110 mel-spectrum by default and didn't tune it. Was 40x32 the result of tuning from your side? You were doing conv2d on it correct? (not conv1d over the time axis?)",
      "votes": null
    },
    {
      "id": "269923",
      "postDate": "01/17/2018 14:31:48",
      "content": "<p>We use normal networks with conv2d as is :) Even normal VGG-19BN can get 87% that is why I am saying that PreAct-ResNets and DPN92 are failed.</p>",
      "rawMarkdown": "We use normal networks with conv2d as is :) Even normal VGG-19BN can get 87% that is why I am saying that PreAct-ResNets and DPN92 are failed.",
      "votes": null
    },
    {
      "id": "269949",
      "postDate": "01/17/2018 15:03:57",
      "content": "<p>got it. Thanks a lot for sharing!</p>",
      "rawMarkdown": "got it. Thanks a lot for sharing!",
      "votes": null
    },
    {
      "id": "269988",
      "postDate": "01/17/2018 16:09:03",
      "content": "<p>I got very little improvement from ensembling either: best single model was 0.89874, best ensemble 0.90191. </p>\n\n<p>I did try different ensembling techniques:\n- average log(probs)\n- voting\n- weighted log(probs), with weights estimated via logit model on validation set\n- weighted log(probs) with weights dependent on the word predicted (again estimated on the validation set))... </p>\n\n<p>I had reasonably uncorrelated inputs: \nmelspectrogram + VGG11 (0.89x)\n1-d wav + deep conv network (0.88x)\nwav + melspectrogram (2 input model) (0.88x)\nmfcc + wav (2 input model) (0.88x).</p>\n\n<p>The best single model was 40x110 melspectrogram, VGG 11, background noise / time-shift augmentation, training data included pseudo labelled data from the test set.</p>",
      "rawMarkdown": "I got very little improvement from ensembling either: best single model was 0.89874, best ensemble 0.90191. \n\nI did try different ensembling techniques:\n- average log(probs)\n- voting\n- weighted log(probs), with weights estimated via logit model on validation set\n- weighted log(probs) with weights dependent on the word predicted (again estimated on the validation set))... \n\nI had reasonably uncorrelated inputs: \nmelspectrogram + VGG11 (0.89x)\n1-d wav + deep conv network (0.88x)\nwav + melspectrogram (2 input model) (0.88x)\nmfcc + wav (2 input model) (0.88x).\n\nThe best single model was 40x110 melspectrogram, VGG 11, background noise / time-shift augmentation, training data included pseudo labelled data from the test set.",
      "votes": null
    },
    {
      "id": "270026",
      "postDate": "01/17/2018 17:04:47",
      "content": "<p>Thanks for reporting your results, could you tell us what input feature you used? </p>",
      "rawMarkdown": "Thanks for reporting your results, could you tell us what input feature you used?",
      "votes": null
    },
    {
      "id": "270067",
      "postDate": "01/17/2018 18:02:04",
      "content": "<p>40x32 mel-spectrogram</p>",
      "rawMarkdown": "40x32 mel-spectrogram",
      "votes": null
    },
    {
      "id": "270094",
      "postDate": "01/17/2018 18:42:23",
      "content": "<p>We fed a 295x295 mel spectrogram from <a href=\"https://github.com/keunwoochoi/kapre\">https://github.com/keunwoochoi/kapre</a> into the Keras InceptionResNetv2 model (with random initial weights) with 3 fully connected layers at the end.  Training that with all of the train data through 45 epochs (each epoch was the whole train dataset), got a 0.9005 on the private LB (0.89035 on public).  </p>\n\n<p>In training, each wav file was dilated in time by a random value between 60% and 100%.  </p>\n\n<p>In prediction, we ran the model across 5 different dilations for each file and then averaged the outputs across the 5 dilations to make the submission.</p>\n\n<p>Our final score involved an ensemble of many similarly good models.</p>",
      "rawMarkdown": "We fed a 295x295 mel spectrogram from https://github.com/keunwoochoi/kapre into the Keras InceptionResNetv2 model (with random initial weights) with 3 fully connected layers at the end.  Training that with all of the train data through 45 epochs (each epoch was the whole train dataset), got a 0.9005 on the private LB (0.89035 on public).  \n\nIn training, each wav file was dilated in time by a random value between 60% and 100%.  \n\nIn prediction, we ran the model across 5 different dilations for each file and then averaged the outputs across the 5 dilations to make the submission.\n\nOur final score involved an ensemble of many similarly good models.",
      "votes": null
    },
    {
      "id": "270109",
      "postDate": "01/17/2018 18:59:22",
      "content": "<p>My best single model was 90.9% on the private LB and 90.2% on public LB. I didn't think much about ensembling until it was too late to get much out of it, just varied some parameters here or there and averaged the sqrt of the probabilities, but honestly never thought I'd get far enough where it mattered haha. Just honored to compete here with some really smart and great people</p>",
      "rawMarkdown": "My best single model was 90.9% on the private LB and 90.2% on public LB. I didn't think much about ensembling until it was too late to get much out of it, just varied some parameters here or there and averaged the sqrt of the probabilities, but honestly never thought I'd get far enough where it mattered haha. Just honored to compete here with some really smart and great people",
      "votes": null
    },
    {
      "id": "270115",
      "postDate": "01/17/2018 19:04:36",
      "content": "<p>Crazy... Could you share little more details about your architecture?</p>",
      "rawMarkdown": "Crazy... Could you share little more details about your architecture?",
      "votes": null
    },
    {
      "id": "270206",
      "postDate": "01/17/2018 22:45:59",
      "content": "<p>Sure, it was getting long so I made a thread on it <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47715\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47715</a></p>",
      "rawMarkdown": "Sure, it was getting long so I made a thread on it https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47715",
      "votes": null
    },
    {
      "id": "270211",
      "postDate": "01/17/2018 23:03:24",
      "content": "<p>Did you try different mel spectrogram sizes? how much of a difference did that make?  I was wondering how far to go to see if any extra detail could be pulled out of the spectrogram or I would just be slowing down the model.</p>",
      "rawMarkdown": "Did you try different mel spectrogram sizes? how much of a difference did that make?  I was wondering how far to go to see if any extra detail could be pulled out of the spectrogram or I would just be slowing down the model.",
      "votes": null
    },
    {
      "id": "270260",
      "postDate": "01/18/2018 01:00:09",
      "content": "<p>Smaller sizes worked less well. This was the biggest we tried.</p>",
      "rawMarkdown": "Smaller sizes worked less well. This was the biggest we tried.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 269909,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "01/17/2018 14:11:32",
      "content": "<p>My best single model (without bagging) is only 0.87 on public LB and 0.88 on private LB. Bagging it 5 times could lead to 0.88/0.89 but yours look very impressive! You mind sharing a bit more details/code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 269913,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "01/17/2018 14:21:52",
          "content": "<p>All our good models use 40x32 mel-spectrum which is quite similar to CIFAR 10 problems. So we took best performing CIFAR10 networks and trained them. So far, only PreAct-ResNets and DPN92 are failed. They got only 87% public LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269917,
          "author_name": "xiaozhouwang",
          "author_url": "",
          "post_date": "01/17/2018 14:26:30",
          "content": "<p>That sounds fascinating! I used 40x110 mel-spectrum by default and didn't tune it. Was 40x32 the result of tuning from your side? You were doing conv2d on it correct? (not conv1d over the time axis?)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269923,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "01/17/2018 14:31:48",
          "content": "<p>We use normal networks with conv2d as is :) Even normal VGG-19BN can get 87% that is why I am saying that PreAct-ResNets and DPN92 are failed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269949,
          "author_name": "xiaozhouwang",
          "author_url": "",
          "post_date": "01/17/2018 15:03:57",
          "content": "<p>got it. Thanks a lot for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269988,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "01/17/2018 16:09:03",
      "content": "<p>I got very little improvement from ensembling either: best single model was 0.89874, best ensemble 0.90191. </p>\n\n<p>I did try different ensembling techniques:\n- average log(probs)\n- voting\n- weighted log(probs), with weights estimated via logit model on validation set\n- weighted log(probs) with weights dependent on the word predicted (again estimated on the validation set))... </p>\n\n<p>I had reasonably uncorrelated inputs: \nmelspectrogram + VGG11 (0.89x)\n1-d wav + deep conv network (0.88x)\nwav + melspectrogram (2 input model) (0.88x)\nmfcc + wav (2 input model) (0.88x).</p>\n\n<p>The best single model was 40x110 melspectrogram, VGG 11, background noise / time-shift augmentation, training data included pseudo labelled data from the test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270026,
      "author_name": "welldone2094",
      "author_url": "",
      "post_date": "01/17/2018 17:04:47",
      "content": "<p>Thanks for reporting your results, could you tell us what input feature you used? </p>",
      "votes": null,
      "replies": [
        {
          "id": 270067,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "01/17/2018 18:02:04",
          "content": "<p>40x32 mel-spectrogram</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270094,
      "author_name": "gte620v",
      "author_url": "",
      "post_date": "01/17/2018 18:42:23",
      "content": "<p>We fed a 295x295 mel spectrogram from <a href=\"https://github.com/keunwoochoi/kapre\">https://github.com/keunwoochoi/kapre</a> into the Keras InceptionResNetv2 model (with random initial weights) with 3 fully connected layers at the end.  Training that with all of the train data through 45 epochs (each epoch was the whole train dataset), got a 0.9005 on the private LB (0.89035 on public).  </p>\n\n<p>In training, each wav file was dilated in time by a random value between 60% and 100%.  </p>\n\n<p>In prediction, we ran the model across 5 different dilations for each file and then averaged the outputs across the 5 dilations to make the submission.</p>\n\n<p>Our final score involved an ensemble of many similarly good models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 270211,
          "author_name": "boltz0",
          "author_url": "",
          "post_date": "01/17/2018 23:03:24",
          "content": "<p>Did you try different mel spectrogram sizes? how much of a difference did that make?  I was wondering how far to go to see if any extra detail could be pulled out of the spectrogram or I would just be slowing down the model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270260,
          "author_name": "gte620v",
          "author_url": "",
          "post_date": "01/18/2018 01:00:09",
          "content": "<p>Smaller sizes worked less well. This was the biggest we tried.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270109,
      "author_name": "omalleyt",
      "author_url": "",
      "post_date": "01/17/2018 18:59:22",
      "content": "<p>My best single model was 90.9% on the private LB and 90.2% on public LB. I didn't think much about ensembling until it was too late to get much out of it, just varied some parameters here or there and averaged the sqrt of the probabilities, but honestly never thought I'd get far enough where it mattered haha. Just honored to compete here with some really smart and great people</p>",
      "votes": null,
      "replies": [
        {
          "id": 270115,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "01/17/2018 19:04:36",
          "content": "<p>Crazy... Could you share little more details about your architecture?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270206,
          "author_name": "omalleyt",
          "author_url": "",
          "post_date": "01/17/2018 22:45:59",
          "content": "<p>Sure, it was getting long so I made a thread on it <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47715\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47715</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "269874": "We had some good single models but failed to ensemble. The difference between our best model and the ensemble with 30 models is only 0.00646 (private score). On the other hand, the winner Heng CherKeng could get 89% from 86% models. So I am really interested how others are doing.\n\nHere are our private scores of the best single models:\n\n - ResNext trained with all data and with crop: 0.89886 (same as the 43-th place in LB!)\n - RexNext trained on train+valid with crop: 0.89686\n - RexNext trained on train+valid without crop: 0.89533\n - Densenet190 with all data and with crop: 0.89486\n - Densenet190 on train+valid with crop: 0.89416\n - WideResNet52-10 on train+test with crop: 0.89345\n\nWhat is your best single model score and the difference to the ensemble?",
    "269909": "My best single model (without bagging) is only 0.87 on public LB and 0.88 on private LB. Bagging it 5 times could lead to 0.88/0.89 but yours look very impressive! You mind sharing a bit more details/code?",
    "269913": "All our good models use 40x32 mel-spectrum which is quite similar to CIFAR 10 problems. So we took best performing CIFAR10 networks and trained them. So far, only PreAct-ResNets and DPN92 are failed. They got only 87% public LB.",
    "269917": "That sounds fascinating! I used 40x110 mel-spectrum by default and didn't tune it. Was 40x32 the result of tuning from your side? You were doing conv2d on it correct? (not conv1d over the time axis?)",
    "269923": "We use normal networks with conv2d as is :) Even normal VGG-19BN can get 87% that is why I am saying that PreAct-ResNets and DPN92 are failed.",
    "269949": "got it. Thanks a lot for sharing!",
    "269988": "I got very little improvement from ensembling either: best single model was 0.89874, best ensemble 0.90191. \n\nI did try different ensembling techniques:\n- average log(probs)\n- voting\n- weighted log(probs), with weights estimated via logit model on validation set\n- weighted log(probs) with weights dependent on the word predicted (again estimated on the validation set))... \n\nI had reasonably uncorrelated inputs: \nmelspectrogram + VGG11 (0.89x)\n1-d wav + deep conv network (0.88x)\nwav + melspectrogram (2 input model) (0.88x)\nmfcc + wav (2 input model) (0.88x).\n\nThe best single model was 40x110 melspectrogram, VGG 11, background noise / time-shift augmentation, training data included pseudo labelled data from the test set.",
    "270026": "Thanks for reporting your results, could you tell us what input feature you used?",
    "270067": "40x32 mel-spectrogram",
    "270094": "We fed a 295x295 mel spectrogram from https://github.com/keunwoochoi/kapre into the Keras InceptionResNetv2 model (with random initial weights) with 3 fully connected layers at the end.  Training that with all of the train data through 45 epochs (each epoch was the whole train dataset), got a 0.9005 on the private LB (0.89035 on public).  \n\nIn training, each wav file was dilated in time by a random value between 60% and 100%.  \n\nIn prediction, we ran the model across 5 different dilations for each file and then averaged the outputs across the 5 dilations to make the submission.\n\nOur final score involved an ensemble of many similarly good models.",
    "270109": "My best single model was 90.9% on the private LB and 90.2% on public LB. I didn't think much about ensembling until it was too late to get much out of it, just varied some parameters here or there and averaged the sqrt of the probabilities, but honestly never thought I'd get far enough where it mattered haha. Just honored to compete here with some really smart and great people",
    "270115": "Crazy... Could you share little more details about your architecture?",
    "270206": "Sure, it was getting long so I made a thread on it https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47715",
    "270211": "Did you try different mel spectrogram sizes? how much of a difference did that make?  I was wondering how far to go to see if any extra detail could be pulled out of the spectrogram or I would just be slowing down the model.",
    "270260": "Smaller sizes worked less well. This was the biggest we tried."
  },
  "source": "meta"
}