{
  "id": 47574,
  "title": "Following Heng's example",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47574",
  "author_name": "",
  "post_date": "2018-01-16T14:51:28.568893200Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>This is a post exploring some recurrent models and some strange results regarding their convergence. Following Heng's lead by example I am posting here the simple models and the results from the training process.</p>\n\n<p>1st Model:</p>\n\n<pre><code>class RecNet(nn.Module):\ndef __init__(self, in_shape, num_classes=12 ):\n    super(RecNet, self).__init__()\n\n    self.layer1 = nn.GRU(in_shape[-1], 256)\n    self.layer2 = nn.GRU(256, 512)\n    self.fc1 = nn.Linear(512, 256)\n    self.fc2 = nn.Linear(40 * 256, num_classes)\n\ndef forward(self, x):\n    x, h_out = self.layer1(x)\n    x = F.dropout(x, p=0.5)\n    x, h_ou2 = self.layer2(x)\n    x = F.dropout(x, p=0.3)\n    x = self.fc1(x)\n    x = self.fc2(x.view(-1, 40 * 256))\n\n    return x  #logits\n</code></pre>\n\n<p>Results:</p>\n\n<pre><code>0.0100  255.5 k  2560.60  130.8 | 0.7574  0.7880 | 0.5708  0.8218 | 0.6756  0.8047 | 38 hr 58 min  \n255499,255499, torch.Size([512, 40, 101])\n0.0100  256.0 k  2565.61  131.1 | 0.7515  0.7824 | 0.5721  0.8260 | 0.5635  0.8125 | 39 hr 01 min  \n255999,255999, torch.Size([512, 40, 101])\n0.0100  256.5 k  2570.62  131.3 | 0.7604  0.7749 | 0.5928  0.8178 | 0.5597  0.8359 | 39 hr 04 min  \n256499,256499, torch.Size([512, 40, 101])\n0.0100  257.0 k  2575.63  131.6 | 0.7831  0.7777 | 0.5634  0.8272 | 0.5638  0.8184 | 39 hr 07 min  \n256999,256999, torch.Size([512, 40, 101])\n0.0100  257.5 k  2580.65  131.8 | 0.7481  0.7814 | 0.5664  0.8249 | 0.5269  0.8418 | 39 hr 10 min  \n257499,257499, torch.Size([512, 40, 101])\n0.0100  258.0 k  2585.66  132.1 | 0.7748  0.7752 | 0.5889  0.8206 | 0.5863  0.8203 | 39 hr 13 min      \n257999,257999, torch.Size([512, 40, 101])                         |\n0.0100  258.5 k  2590.67  132.4 | 0.7731  0.7727 | 0.5707  0.8066 | 0.5110  0.8379 | 39 hr 16 min  \n258499,258499, torch.Size([512, 40, 101])\n0.0100  259.0 k  2595.68  132.6 | 0.7703  0.7733 | 0.5785  0.8177 | 0.5610  0.8340 | 39 hr 20 min  \n258999,258999, torch.Size([512, 40, 101])\n0.0100  259.5 k  2600.69  132.9 | 0.7675  0.7739 | 0.5747  0.8214 | 0.5562  0.8340 | 39 hr 23 min  \n259499,259499, torch.Size([512, 40, 101])\n0.0100  260.0 k  2605.70  133.1 | 0.7627  0.7786 | 0.5250  0.8320 | 0.5074  0.8398 | 39 hr 26 min  \n259999,259999, torch.Size([512, 40, 101])\n0.0100  260.5 k  2610.71  133.4 | 0.7732  0.7733 | 0.5552  0.8251 | 0.5708  0.8301 | 39 hr 29 min  \n260499,260499, torch.Size([512, 40, 101])\n0.0100  261.0 k  2615.72  133.6 | 0.7728  0.7736 | 0.5742  0.8249 | 0.6212  0.8125 | 39 hr 32 min  260999,260999, torch.Size([512, 40, 101])\n0.0100  261.5 k  2620.73  133.9 | 0.7643  0.7839 | 0.5699  0.8255 | 0.5456  0.8184 | 39 hr 35 min  261499,261499, torch.Size([512, 40, 101])\n0.0100  262.0 k  2625.74  134.1 | 0.7577  0.7792 | 0.5579  0.8296 | 0.4826  0.8574 | 39 hr 38 min  261999,261999, torch.Size([512, 40, 101])\n</code></pre>\n\n<p>As you can see the results keep oscillating back and forth. The loss increases and then decreases again. Which indicates that recurrent models are having some difficulty to converge.</p>\n\n<p>2nd Model:</p>\n\n<pre><code>class RecNet2(nn.Module):\ndef __init__(self, in_shape, num_classes=12 ):\n    super(RecNet2, self).__init__()\n\n    self.layer1 = nn.GRU(in_shape[-1], 256)\n    self.layer2 = nn.LSTM(256, 512)\n    self.fc1 = nn.Linear(512, 256)\n    self.fc2 = nn.Linear(40 * 256, num_classes)\n\ndef forward(self, x):\n    x, h_out = self.layer1(x)\n    x = F.dropout(x, p=0.5)\n    x, h_ou2 = self.layer2(x)\n    x = F.dropout(x, p=0.3)\n    x = self.fc1(x)\n    x = self.fc2(x.view(-1, 40 * 256))\n\n    return x  #logits\n</code></pre>\n\n<p>Results:</p>\n\n<pre><code>0.0100  238.0 k  2385.22  121.9 | 0.7247  0.7855 | 0.5021  0.8462 | 0.5366  0.8418 | 38 hr 11 min  \n237999,237999, torch.Size([512, 40, 101])\n0.0100  238.5 k  2390.23  122.1 | 0.7148  0.7820 | 0.5204  0.8376 | 0.5722  0.8359 | 38 hr 15 min  \n238499,238499, torch.Size([512, 40, 101])\n0.0100  239.0 k  2395.24  122.4 | 0.7133  0.7914 | 0.5043  0.8417 | 0.5279  0.8203 | 38 hr 17 min  \n238999,238999, torch.Size([512, 40, 101])  \n239499,239499, torch.Size([512, 40, 101])\n0.0100  240.0 k  2405.26  122.9 | 0.7407  0.7802 | 0.5226  0.8391 | 0.4966  0.8340 | 38 hr 23 min  \n239999,239999, torch.Size([512, 40, 101])\n0.0100  240.5 k  2410.27  123.1 | 0.7311  0.7861 | 0.5144  0.8438 | 0.4636  0.8633 | 38 hr 26 min  \n240499,240499, torch.Size([512, 40, 101])\n0.0100  241.0 k  2415.28  123.4 | 0.7198  0.7933 | 0.4964  0.8525 | 0.5036  0.8281 | 38 hr 29 min  \n240999,240999, torch.Size([512, 40, 101])\n0.0100  241.5 k  2420.29  123.6 | 0.7377  0.7802 | 0.5068  0.8414 | 0.5212  0.8398 | 38 hr 32 min  \n241499,241499, torch.Size([512, 40, 101])\n0.0100  242.0 k  2425.31  123.9 | 0.7220  0.7892 | 0.5025  0.8483 | 0.4110  0.8711 | 38 hr 35 min  \n241999,241999, torch.Size([512, 40, 101])\n0.0100  242.5 k  2430.32  124.2 | 0.7111  0.8001 | 0.5217  0.8391 | 0.5346  0.8379 | 38 hr 38 min  \n242499,242499, torch.Size([512, 40, 101])\n0.0100  243.0 k  2435.33  124.4 | 0.6963  0.7951 | 0.5064  0.8432 | 0.5590  0.8457 | 38 hr 41 min  \n242999,242999, torch.Size([512, 40, 101])\n0.0100  243.5 k  2440.34  124.7 | 0.7349  0.7898 | 0.5091  0.8421 | 0.5491  0.8184 | 38 hr 44 min  \n243499,243499, torch.Size([512, 40, 101])\n0.0100  244.0 k  2445.35  124.9 | 0.7272  0.7827 | 0.5141  0.8426 | 0.5331  0.8379 | 38 hr 47 min  \n243999,243999, torch.Size([512, 40, 101])\n0.0100  244.5 k  2450.36  125.2 | 0.7296  0.7901 | 0.5211  0.8365 | 0.5010  0.8477 | 38 hr 50 min  \n244499,244499, torch.Size([512, 40, 101])\n0.0100  245.0 k  2455.37  125.4 | 0.7009  0.7951 | 0.5083  0.8164 | 0.5616  0.8438 | 38 hr 53 min  \n244999,244999, torch.Size([512, 40, 101])\n0.0100  245.5 k  2460.38  125.7 | 0.7282  0.7867 | 0.5723  0.8239 | 0.5448  0.8281 | 38 hr 56 min  \n245499,245499, torch.Size([512, 40, 101])\n0.0100  246.0 k  2465.39  126.0 | 0.7250  0.7901 | 0.5079  0.8429 | 0.4877  0.8516 | 38 hr 59 min  \n245999,245999, torch.Size([512, 40, 101])\n0.0100  246.5 k  2470.40  126.2 | 0.7355  0.7817 | 0.5221  0.8385 | 0.4724  0.8691 | 39 hr 01 min  \n246499,246499, torch.Size([512, 40, 101])\n0.0100  247.0 k  2475.41  126.5 | 0.7305  0.7830 | 0.5073  0.8448 | 0.5690  0.8047 | 39 hr 04 min  \n246999,246999, torch.Size([512, 40, 101])\n0.0100  247.5 k  2480.43  126.7 | 0.7105  0.7914 | 0.5217  0.8378 | 0.4530  0.8613 | 39 hr 07 min  \n247499,247499, torch.Size([512, 40, 101])\n0.0100  248.0 k  2485.44  127.0 | 0.7384  0.7795 | 0.5113  0.8452 | 0.5366  0.8262 | 39 hr 12 min  \n247999,247999, torch.Size([512, 40, 101])\n0.0100  248.5 k  2490.45  127.2 | 0.7084  0.7901 | 0.5257  0.8362 | 0.5554  0.8359 | 39 hr 16 min  \n248499,248499, torch.Size([512, 40, 101])\n0.0100  249.0 k  2495.46  127.5 | 0.7141  0.7917 | 0.4956  0.8470 | 0.4636  0.8301 | 39 hr 20 min  \n248999,248999, torch.Size([512, 40, 101])\n0.0100  249.5 k  2500.47  127.7 | 0.7142  0.7973 | 0.5123  0.8398 | 0.4595  0.8613 | 39 hr 26 min  \n249499,249499, torch.Size([512, 40, 101])\n</code></pre>\n\n<p>3rd Model:</p>\n\n<pre><code>class RecNet3(nn.Module):\ndef __init__(self, in_shape, num_classes=12 ):\n    super(RecNet3, self).__init__()\n\n    self.layer1 = nn.GRU(in_shape[-1], 256)\n    self.layer2 = nn.LSTM(256, 512)\n    self.layer3 = nn.GRU(512, 1024)\n    self.layer4 = nn.LSTM(1024, 2048)\n    self.fc1 = nn.Linear(2048, 512)\n    self.fc2 = nn.Linear(40 * 512, num_classes)\n\ndef forward(self, x):\n    x, h_out = self.layer1(x)\n    x = F.dropout(x, p=0.5)\n    x, h_out2 = self.layer2(x)\n    x = F.dropout(x, p=0.3)\n    x, h_out3 = self.layer3(x)\n    x, h_out4 = self.layer4(x)\n    x = F.dropout(x, p=0.3)\n    x = self.fc1(x)\n    x = self.fc2(x.view(-1, 40 * 512))\n\n    return x  #logits\n</code></pre>\n\n<p>Results:</p>\n\n<pre><code>0.0100   90.0 k  901.97  46.1 | 0.7829  0.7643 | 0.7526  0.7660 | 0.7332  0.7539 | 35 hr 45 min  \n89999,89999, torch.Size([512, 40, 101])\n0.0100   90.5 k  906.98  46.3 | 0.7767  0.7664 | 0.7493  0.7664 | 0.6869  0.7773 | 35 hr 53 min  \n90499,90499, torch.Size([512, 40, 101])\n0.0100   91.0 k  911.99  46.6 | 0.8006  0.7596 | 0.7184  0.7765 | 0.7083  0.7617 | 36 hr 05 min  \n90999,90999, torch.Size([512, 40, 101])\n0.0100   91.5 k  917.01  46.8 | 0.7871  0.7655 | 0.7387  0.7714 | 0.7490  0.7695 | 36 hr 19 min  \n91499,91499, torch.Size([512, 40, 101])\n0.0100   92.0 k  922.02  47.1 | 0.7738  0.7630 | 0.7699  0.7578 | 0.7372  0.7812 | 36 hr 27 min  \n91999,91999, torch.Size([512, 40, 101])\n0.0100   92.5 k  927.03  47.4 | 0.7582  0.7749 | 0.7257  0.7653 | 0.6510  0.8125 | 36 hr 39 min  \n92499,92499, torch.Size([512, 40, 101])\n0.0100   93.0 k  932.04  47.6 | 0.7869  0.7661 | 0.7169  0.7782 | 0.6819  0.7754 | 37 hr 34 min  \n92999,92999, torch.Size([512, 40, 101])\n0.0100   93.5 k  937.05  47.9 | 0.7699  0.7646 | 0.7328  0.7727 | 0.7638  0.7578 | 37 hr 43 min  \n93499,93499, torch.Size([512, 40, 101])\n0.0100   94.0 k  942.06  48.1 | 0.7432  0.7736 | 0.7326  0.7723 | 0.7336  0.7656 | 37 hr 50 min  \n93999,93999, torch.Size([512, 40, 101])\n0.0100   94.5 k  947.07  48.4 | 0.7618  0.7714 | 0.7094  0.7786 | 0.8350  0.7480 | 37 hr 58 min  \n94499,94499, torch.Size([512, 40, 101])\n0.0100   95.0 k  952.08  48.6 | 0.7597  0.7696 | 0.7074  0.7778 | 0.7055  0.7793 | 38 hr 06 min  \n94999,94999, torch.Size([512, 40, 101])a\n0.0100   95.5 k  957.09  48.9 | 0.7571  0.7733 | 0.7339  0.7701 | 0.7449  0.7754 | 38 hr 13 min  \n95499,95499, torch.Size([512, 40, 101])\n0.0100   96.0 k  962.10  49.2 | 0.7770  0.7608 | 0.6961  0.7856 | 0.6177  0.8027 | 38 hr 21 min  \n95999,95999, torch.Size([512, 40, 101])\n0.0100   96.5 k  967.12  49.4 | 0.7517  0.7789 | 0.7187  0.7910 | 0.7956  0.7637 | 38 hr 29 min  \n96499,96499, torch.Size([512, 40, 101])\n0.0100   97.0 k  972.13  49.7 | 0.7609  0.7733 | 0.7123  0.7822 | 0.6791  0.7559 | 38 hr 36 min  \n96999,96999, torch.Size([512, 40, 101])\n0.0100   97.5 k  977.14  49.9 | 0.7771  0.7636 | 0.7181  0.7743 | 0.6693  0.7832 | 38 hr 44 min  \n97499,97499, torch.Size([512, 40, 101])\n0.0100   97.6 k  977.87  50.0 | 0.7771  0.7636 | 0.6780  0.7793 | 0.7790  0.7656 | 38 hr 45 min  \n97573,97573, torch.Size([512, 40, 101])\n</code></pre>\n\n<p>4th Model: Modified Heng's resnet with LSTM layer at on top.</p>\n\n<pre><code>class SeResNet3(nn.Module):\ndef __init__(self, in_shape=(1,40,101), num_classes=12 ):\n    super(SeResNet3, self).__init__()\n    in_channels = in_shape[0]\n\n    self.layer1a = ConvBn2d(in_channels, 16, kernel_size=(3, 3), stride=(1, 1))\n    self.layer1b = ResBlock( 16, 16)\n\n    self.layer2a = ConvBn2d(16, 32, kernel_size=(3, 3), stride=(1, 1))\n    self.layer2b = ResBlock(32, 32)\n    self.layer2c = ResBlock(32, 32)\n\n    self.layer3a = ConvBn2d(32, 64, kernel_size=(3, 3), stride=(1, 1))\n    self.layer3b = ResBlock(64, 64)\n    self.layer3c = ResBlock(64, 64)\n\n    self.layer4a = ConvBn2d( 64,128, kernel_size=(3, 3), stride=(1, 1))\n    self.layer4b = ResBlock(128,128)\n    self.layer4c = ResBlock(128,128)\n\n    self.layer5a = ConvBn2d(128, 256, kernel_size=(3, 3), stride=(1, 1))\n    self.layer5ab = nn.LSTMCell(256, 256, 10)\n    self.layer5b = nn.Linear(256,256)\n\n    self.fc = nn.Linear(256,num_classes)\n\n    def forward(self, x):\n\n    x = F.relu(self.layer1a(x),inplace=True)\n    x = self.layer1b(x)\n    x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n    x = F.dropout(x,p=0.1,training=self.training)\n    x = F.relu(self.layer2a(x),inplace=True)\n    x = self.layer2b(x)\n    x = self.layer2c(x)\n    x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.layer3a(x),inplace=True)\n    x = self.layer3b(x)\n    x = self.layer3c(x)\n    x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.layer4a(x),inplace=True)\n    x = self.layer4b(x)\n    x = self.layer4c(x)\n\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.layer5a(x),inplace=True)\n    x = F.adaptive_avg_pool2d(x,1)\n    x, h = F.dropout(F.relu(self.layer5ab(x.view(-1, 40, 256))), p=0.2, self.training)\n    x = x.view(x.size(0), -1)\n    x = F.relu(self.layer5b(x))\n\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = self.fc(x)\n\n    return x  #logits\n</code></pre>\n\n<p>Results:</p>\n\n<pre><code>0.0100  106.5 k  1067.33  54.5 | 0.8049  0.7540 | 0.6601  0.7930 | 0.6240  0.8066 | 19 hr 07 min  \n106499,106499, torch.Size([512, 40, 101])\n0.0100  107.0 k  1072.35  54.8 | 0.7762  0.7671 | 0.6527  0.7935 | 0.7233  0.7871 | 19 hr 09 min  \n106999,106999, torch.Size([512, 40, 101])\n0.0100  107.5 k  1077.36  55.0 | 0.7861  0.7568 | 0.6754  0.7943 | 0.5772  0.8184 | 19 hr 12 min  \n107499,107499, torch.Size([512, 40, 101])\n0.0100  108.0 k  1082.37  55.3 | 0.7906  0.7537 | 0.6592  0.7966 | 0.7249  0.7578 | 19 hr 15 min  \n107999,107999, torch.Size([512, 40, 101])\n0.0100  108.5 k  1087.38  55.6 | 0.8031  0.7627 | 0.6714  0.7917 | 0.7119  0.7734 | 19 hr 17 min  \n108499,108499, torch.Size([512, 40, 101])\n0.0100  109.0 k  1092.39  55.8 | 0.7832  0.7652 | 0.6528  0.7905 | 0.6212  0.8027 | 19 hr 19 min  \n108999,108999, torch.Size([512, 40, 101])\n0.0100  109.5 k  1097.40  56.1 | 0.7807  0.7677 | 0.6528  0.7965 | 0.6304  0.8105 | 19 hr 22 min  \n109499,109499, torch.Size([512, 40, 101])\n0.0100  110.0 k  1102.41  56.3 | 0.7806  0.7674 | 0.6680  0.7891 | 0.6813  0.7871 | 19 hr 24 min  \n109999,109999, torch.Size([512, 40, 101])\n0.0100  110.5 k  1107.42  56.6 | 0.7988  0.7671 | 0.6318  0.8027 | 0.6820  0.7949 | 19 hr 26 min  \n110499,110499, torch.Size([512, 40, 101])\n0.0100  111.0 k  1112.43  56.8 | 0.7735  0.7680 | 0.6892  0.7825 | 0.6944  0.7852 | 19 hr 29 min  \n110999,110999, torch.Size([512, 40, 101])\n0.0100  111.5 k  1117.44  57.1 | 0.7805  0.7661 | 0.6609  0.7952 | 0.6439  0.8066 | 19 hr 31 min  \n111499,111499, torch.Size([512, 40, 101])\n0.0100  112.0 k  1122.46  57.3 | 0.7779  0.7633 | 0.6517  0.7975 | 0.8025  0.7656 | 19 hr 34 min  \n111999,111999, torch.Size([512, 40, 101])\n0.0100  112.5 k  1127.47  57.6 | 0.7480  0.7652 | 0.6583  0.7934 | 0.6775  0.7871 | 19 hr 38 min  \n112499,112499, torch.Size([512, 40, 101])\n0.0100  113.0 k  1132.48  57.9 | 0.7796  0.7574 | 0.6388  0.8022 | 0.5998  0.8105 | 19 hr 42 min  \n112999,112999, torch.Size([512, 40, 101])\n</code></pre>\n\n<p>It would be nice to know if anyone had any success with recurrent networks and reached LB &gt; 0.83?</p>",
  "messages": [
    {
      "id": "269262",
      "postDate": "01/16/2018 14:51:28",
      "content": "<p>This is a post exploring some recurrent models and some strange results regarding their convergence. Following Heng's lead by example I am posting here the simple models and the results from the training process.</p>\n\n<p>1st Model:</p>\n\n<pre><code>class RecNet(nn.Module):\ndef __init__(self, in_shape, num_classes=12 ):\n    super(RecNet, self).__init__()\n\n    self.layer1 = nn.GRU(in_shape[-1], 256)\n    self.layer2 = nn.GRU(256, 512)\n    self.fc1 = nn.Linear(512, 256)\n    self.fc2 = nn.Linear(40 * 256, num_classes)\n\ndef forward(self, x):\n    x, h_out = self.layer1(x)\n    x = F.dropout(x, p=0.5)\n    x, h_ou2 = self.layer2(x)\n    x = F.dropout(x, p=0.3)\n    x = self.fc1(x)\n    x = self.fc2(x.view(-1, 40 * 256))\n\n    return x  #logits\n</code></pre>\n\n<p>Results:</p>\n\n<pre><code>0.0100  255.5 k  2560.60  130.8 | 0.7574  0.7880 | 0.5708  0.8218 | 0.6756  0.8047 | 38 hr 58 min  \n255499,255499, torch.Size([512, 40, 101])\n0.0100  256.0 k  2565.61  131.1 | 0.7515  0.7824 | 0.5721  0.8260 | 0.5635  0.8125 | 39 hr 01 min  \n255999,255999, torch.Size([512, 40, 101])\n0.0100  256.5 k  2570.62  131.3 | 0.7604  0.7749 | 0.5928  0.8178 | 0.5597  0.8359 | 39 hr 04 min  \n256499,256499, torch.Size([512, 40, 101])\n0.0100  257.0 k  2575.63  131.6 | 0.7831  0.7777 | 0.5634  0.8272 | 0.5638  0.8184 | 39 hr 07 min  \n256999,256999, torch.Size([512, 40, 101])\n0.0100  257.5 k  2580.65  131.8 | 0.7481  0.7814 | 0.5664  0.8249 | 0.5269  0.8418 | 39 hr 10 min  \n257499,257499, torch.Size([512, 40, 101])\n0.0100  258.0 k  2585.66  132.1 | 0.7748  0.7752 | 0.5889  0.8206 | 0.5863  0.8203 | 39 hr 13 min      \n257999,257999, torch.Size([512, 40, 101])                         |\n0.0100  258.5 k  2590.67  132.4 | 0.7731  0.7727 | 0.5707  0.8066 | 0.5110  0.8379 | 39 hr 16 min  \n258499,258499, torch.Size([512, 40, 101])\n0.0100  259.0 k  2595.68  132.6 | 0.7703  0.7733 | 0.5785  0.8177 | 0.5610  0.8340 | 39 hr 20 min  \n258999,258999, torch.Size([512, 40, 101])\n0.0100  259.5 k  2600.69  132.9 | 0.7675  0.7739 | 0.5747  0.8214 | 0.5562  0.8340 | 39 hr 23 min  \n259499,259499, torch.Size([512, 40, 101])\n0.0100  260.0 k  2605.70  133.1 | 0.7627  0.7786 | 0.5250  0.8320 | 0.5074  0.8398 | 39 hr 26 min  \n259999,259999, torch.Size([512, 40, 101])\n0.0100  260.5 k  2610.71  133.4 | 0.7732  0.7733 | 0.5552  0.8251 | 0.5708  0.8301 | 39 hr 29 min  \n260499,260499, torch.Size([512, 40, 101])\n0.0100  261.0 k  2615.72  133.6 | 0.7728  0.7736 | 0.5742  0.8249 | 0.6212  0.8125 | 39 hr 32 min  260999,260999, torch.Size([512, 40, 101])\n0.0100  261.5 k  2620.73  133.9 | 0.7643  0.7839 | 0.5699  0.8255 | 0.5456  0.8184 | 39 hr 35 min  261499,261499, torch.Size([512, 40, 101])\n0.0100  262.0 k  2625.74  134.1 | 0.7577  0.7792 | 0.5579  0.8296 | 0.4826  0.8574 | 39 hr 38 min  261999,261999, torch.Size([512, 40, 101])\n</code></pre>\n\n<p>As you can see the results keep oscillating back and forth. The loss increases and then decreases again. Which indicates that recurrent models are having some difficulty to converge.</p>\n\n<p>2nd Model:</p>\n\n<pre><code>class RecNet2(nn.Module):\ndef __init__(self, in_shape, num_classes=12 ):\n    super(RecNet2, self).__init__()\n\n    self.layer1 = nn.GRU(in_shape[-1], 256)\n    self.layer2 = nn.LSTM(256, 512)\n    self.fc1 = nn.Linear(512, 256)\n    self.fc2 = nn.Linear(40 * 256, num_classes)\n\ndef forward(self, x):\n    x, h_out = self.layer1(x)\n    x = F.dropout(x, p=0.5)\n    x, h_ou2 = self.layer2(x)\n    x = F.dropout(x, p=0.3)\n    x = self.fc1(x)\n    x = self.fc2(x.view(-1, 40 * 256))\n\n    return x  #logits\n</code></pre>\n\n<p>Results:</p>\n\n<pre><code>0.0100  238.0 k  2385.22  121.9 | 0.7247  0.7855 | 0.5021  0.8462 | 0.5366  0.8418 | 38 hr 11 min  \n237999,237999, torch.Size([512, 40, 101])\n0.0100  238.5 k  2390.23  122.1 | 0.7148  0.7820 | 0.5204  0.8376 | 0.5722  0.8359 | 38 hr 15 min  \n238499,238499, torch.Size([512, 40, 101])\n0.0100  239.0 k  2395.24  122.4 | 0.7133  0.7914 | 0.5043  0.8417 | 0.5279  0.8203 | 38 hr 17 min  \n238999,238999, torch.Size([512, 40, 101])  \n239499,239499, torch.Size([512, 40, 101])\n0.0100  240.0 k  2405.26  122.9 | 0.7407  0.7802 | 0.5226  0.8391 | 0.4966  0.8340 | 38 hr 23 min  \n239999,239999, torch.Size([512, 40, 101])\n0.0100  240.5 k  2410.27  123.1 | 0.7311  0.7861 | 0.5144  0.8438 | 0.4636  0.8633 | 38 hr 26 min  \n240499,240499, torch.Size([512, 40, 101])\n0.0100  241.0 k  2415.28  123.4 | 0.7198  0.7933 | 0.4964  0.8525 | 0.5036  0.8281 | 38 hr 29 min  \n240999,240999, torch.Size([512, 40, 101])\n0.0100  241.5 k  2420.29  123.6 | 0.7377  0.7802 | 0.5068  0.8414 | 0.5212  0.8398 | 38 hr 32 min  \n241499,241499, torch.Size([512, 40, 101])\n0.0100  242.0 k  2425.31  123.9 | 0.7220  0.7892 | 0.5025  0.8483 | 0.4110  0.8711 | 38 hr 35 min  \n241999,241999, torch.Size([512, 40, 101])\n0.0100  242.5 k  2430.32  124.2 | 0.7111  0.8001 | 0.5217  0.8391 | 0.5346  0.8379 | 38 hr 38 min  \n242499,242499, torch.Size([512, 40, 101])\n0.0100  243.0 k  2435.33  124.4 | 0.6963  0.7951 | 0.5064  0.8432 | 0.5590  0.8457 | 38 hr 41 min  \n242999,242999, torch.Size([512, 40, 101])\n0.0100  243.5 k  2440.34  124.7 | 0.7349  0.7898 | 0.5091  0.8421 | 0.5491  0.8184 | 38 hr 44 min  \n243499,243499, torch.Size([512, 40, 101])\n0.0100  244.0 k  2445.35  124.9 | 0.7272  0.7827 | 0.5141  0.8426 | 0.5331  0.8379 | 38 hr 47 min  \n243999,243999, torch.Size([512, 40, 101])\n0.0100  244.5 k  2450.36  125.2 | 0.7296  0.7901 | 0.5211  0.8365 | 0.5010  0.8477 | 38 hr 50 min  \n244499,244499, torch.Size([512, 40, 101])\n0.0100  245.0 k  2455.37  125.4 | 0.7009  0.7951 | 0.5083  0.8164 | 0.5616  0.8438 | 38 hr 53 min  \n244999,244999, torch.Size([512, 40, 101])\n0.0100  245.5 k  2460.38  125.7 | 0.7282  0.7867 | 0.5723  0.8239 | 0.5448  0.8281 | 38 hr 56 min  \n245499,245499, torch.Size([512, 40, 101])\n0.0100  246.0 k  2465.39  126.0 | 0.7250  0.7901 | 0.5079  0.8429 | 0.4877  0.8516 | 38 hr 59 min  \n245999,245999, torch.Size([512, 40, 101])\n0.0100  246.5 k  2470.40  126.2 | 0.7355  0.7817 | 0.5221  0.8385 | 0.4724  0.8691 | 39 hr 01 min  \n246499,246499, torch.Size([512, 40, 101])\n0.0100  247.0 k  2475.41  126.5 | 0.7305  0.7830 | 0.5073  0.8448 | 0.5690  0.8047 | 39 hr 04 min  \n246999,246999, torch.Size([512, 40, 101])\n0.0100  247.5 k  2480.43  126.7 | 0.7105  0.7914 | 0.5217  0.8378 | 0.4530  0.8613 | 39 hr 07 min  \n247499,247499, torch.Size([512, 40, 101])\n0.0100  248.0 k  2485.44  127.0 | 0.7384  0.7795 | 0.5113  0.8452 | 0.5366  0.8262 | 39 hr 12 min  \n247999,247999, torch.Size([512, 40, 101])\n0.0100  248.5 k  2490.45  127.2 | 0.7084  0.7901 | 0.5257  0.8362 | 0.5554  0.8359 | 39 hr 16 min  \n248499,248499, torch.Size([512, 40, 101])\n0.0100  249.0 k  2495.46  127.5 | 0.7141  0.7917 | 0.4956  0.8470 | 0.4636  0.8301 | 39 hr 20 min  \n248999,248999, torch.Size([512, 40, 101])\n0.0100  249.5 k  2500.47  127.7 | 0.7142  0.7973 | 0.5123  0.8398 | 0.4595  0.8613 | 39 hr 26 min  \n249499,249499, torch.Size([512, 40, 101])\n</code></pre>\n\n<p>3rd Model:</p>\n\n<pre><code>class RecNet3(nn.Module):\ndef __init__(self, in_shape, num_classes=12 ):\n    super(RecNet3, self).__init__()\n\n    self.layer1 = nn.GRU(in_shape[-1], 256)\n    self.layer2 = nn.LSTM(256, 512)\n    self.layer3 = nn.GRU(512, 1024)\n    self.layer4 = nn.LSTM(1024, 2048)\n    self.fc1 = nn.Linear(2048, 512)\n    self.fc2 = nn.Linear(40 * 512, num_classes)\n\ndef forward(self, x):\n    x, h_out = self.layer1(x)\n    x = F.dropout(x, p=0.5)\n    x, h_out2 = self.layer2(x)\n    x = F.dropout(x, p=0.3)\n    x, h_out3 = self.layer3(x)\n    x, h_out4 = self.layer4(x)\n    x = F.dropout(x, p=0.3)\n    x = self.fc1(x)\n    x = self.fc2(x.view(-1, 40 * 512))\n\n    return x  #logits\n</code></pre>\n\n<p>Results:</p>\n\n<pre><code>0.0100   90.0 k  901.97  46.1 | 0.7829  0.7643 | 0.7526  0.7660 | 0.7332  0.7539 | 35 hr 45 min  \n89999,89999, torch.Size([512, 40, 101])\n0.0100   90.5 k  906.98  46.3 | 0.7767  0.7664 | 0.7493  0.7664 | 0.6869  0.7773 | 35 hr 53 min  \n90499,90499, torch.Size([512, 40, 101])\n0.0100   91.0 k  911.99  46.6 | 0.8006  0.7596 | 0.7184  0.7765 | 0.7083  0.7617 | 36 hr 05 min  \n90999,90999, torch.Size([512, 40, 101])\n0.0100   91.5 k  917.01  46.8 | 0.7871  0.7655 | 0.7387  0.7714 | 0.7490  0.7695 | 36 hr 19 min  \n91499,91499, torch.Size([512, 40, 101])\n0.0100   92.0 k  922.02  47.1 | 0.7738  0.7630 | 0.7699  0.7578 | 0.7372  0.7812 | 36 hr 27 min  \n91999,91999, torch.Size([512, 40, 101])\n0.0100   92.5 k  927.03  47.4 | 0.7582  0.7749 | 0.7257  0.7653 | 0.6510  0.8125 | 36 hr 39 min  \n92499,92499, torch.Size([512, 40, 101])\n0.0100   93.0 k  932.04  47.6 | 0.7869  0.7661 | 0.7169  0.7782 | 0.6819  0.7754 | 37 hr 34 min  \n92999,92999, torch.Size([512, 40, 101])\n0.0100   93.5 k  937.05  47.9 | 0.7699  0.7646 | 0.7328  0.7727 | 0.7638  0.7578 | 37 hr 43 min  \n93499,93499, torch.Size([512, 40, 101])\n0.0100   94.0 k  942.06  48.1 | 0.7432  0.7736 | 0.7326  0.7723 | 0.7336  0.7656 | 37 hr 50 min  \n93999,93999, torch.Size([512, 40, 101])\n0.0100   94.5 k  947.07  48.4 | 0.7618  0.7714 | 0.7094  0.7786 | 0.8350  0.7480 | 37 hr 58 min  \n94499,94499, torch.Size([512, 40, 101])\n0.0100   95.0 k  952.08  48.6 | 0.7597  0.7696 | 0.7074  0.7778 | 0.7055  0.7793 | 38 hr 06 min  \n94999,94999, torch.Size([512, 40, 101])a\n0.0100   95.5 k  957.09  48.9 | 0.7571  0.7733 | 0.7339  0.7701 | 0.7449  0.7754 | 38 hr 13 min  \n95499,95499, torch.Size([512, 40, 101])\n0.0100   96.0 k  962.10  49.2 | 0.7770  0.7608 | 0.6961  0.7856 | 0.6177  0.8027 | 38 hr 21 min  \n95999,95999, torch.Size([512, 40, 101])\n0.0100   96.5 k  967.12  49.4 | 0.7517  0.7789 | 0.7187  0.7910 | 0.7956  0.7637 | 38 hr 29 min  \n96499,96499, torch.Size([512, 40, 101])\n0.0100   97.0 k  972.13  49.7 | 0.7609  0.7733 | 0.7123  0.7822 | 0.6791  0.7559 | 38 hr 36 min  \n96999,96999, torch.Size([512, 40, 101])\n0.0100   97.5 k  977.14  49.9 | 0.7771  0.7636 | 0.7181  0.7743 | 0.6693  0.7832 | 38 hr 44 min  \n97499,97499, torch.Size([512, 40, 101])\n0.0100   97.6 k  977.87  50.0 | 0.7771  0.7636 | 0.6780  0.7793 | 0.7790  0.7656 | 38 hr 45 min  \n97573,97573, torch.Size([512, 40, 101])\n</code></pre>\n\n<p>4th Model: Modified Heng's resnet with LSTM layer at on top.</p>\n\n<pre><code>class SeResNet3(nn.Module):\ndef __init__(self, in_shape=(1,40,101), num_classes=12 ):\n    super(SeResNet3, self).__init__()\n    in_channels = in_shape[0]\n\n    self.layer1a = ConvBn2d(in_channels, 16, kernel_size=(3, 3), stride=(1, 1))\n    self.layer1b = ResBlock( 16, 16)\n\n    self.layer2a = ConvBn2d(16, 32, kernel_size=(3, 3), stride=(1, 1))\n    self.layer2b = ResBlock(32, 32)\n    self.layer2c = ResBlock(32, 32)\n\n    self.layer3a = ConvBn2d(32, 64, kernel_size=(3, 3), stride=(1, 1))\n    self.layer3b = ResBlock(64, 64)\n    self.layer3c = ResBlock(64, 64)\n\n    self.layer4a = ConvBn2d( 64,128, kernel_size=(3, 3), stride=(1, 1))\n    self.layer4b = ResBlock(128,128)\n    self.layer4c = ResBlock(128,128)\n\n    self.layer5a = ConvBn2d(128, 256, kernel_size=(3, 3), stride=(1, 1))\n    self.layer5ab = nn.LSTMCell(256, 256, 10)\n    self.layer5b = nn.Linear(256,256)\n\n    self.fc = nn.Linear(256,num_classes)\n\n    def forward(self, x):\n\n    x = F.relu(self.layer1a(x),inplace=True)\n    x = self.layer1b(x)\n    x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n    x = F.dropout(x,p=0.1,training=self.training)\n    x = F.relu(self.layer2a(x),inplace=True)\n    x = self.layer2b(x)\n    x = self.layer2c(x)\n    x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.layer3a(x),inplace=True)\n    x = self.layer3b(x)\n    x = self.layer3c(x)\n    x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.layer4a(x),inplace=True)\n    x = self.layer4b(x)\n    x = self.layer4c(x)\n\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = F.relu(self.layer5a(x),inplace=True)\n    x = F.adaptive_avg_pool2d(x,1)\n    x, h = F.dropout(F.relu(self.layer5ab(x.view(-1, 40, 256))), p=0.2, self.training)\n    x = x.view(x.size(0), -1)\n    x = F.relu(self.layer5b(x))\n\n    x = F.dropout(x,p=0.2,training=self.training)\n    x = self.fc(x)\n\n    return x  #logits\n</code></pre>\n\n<p>Results:</p>\n\n<pre><code>0.0100  106.5 k  1067.33  54.5 | 0.8049  0.7540 | 0.6601  0.7930 | 0.6240  0.8066 | 19 hr 07 min  \n106499,106499, torch.Size([512, 40, 101])\n0.0100  107.0 k  1072.35  54.8 | 0.7762  0.7671 | 0.6527  0.7935 | 0.7233  0.7871 | 19 hr 09 min  \n106999,106999, torch.Size([512, 40, 101])\n0.0100  107.5 k  1077.36  55.0 | 0.7861  0.7568 | 0.6754  0.7943 | 0.5772  0.8184 | 19 hr 12 min  \n107499,107499, torch.Size([512, 40, 101])\n0.0100  108.0 k  1082.37  55.3 | 0.7906  0.7537 | 0.6592  0.7966 | 0.7249  0.7578 | 19 hr 15 min  \n107999,107999, torch.Size([512, 40, 101])\n0.0100  108.5 k  1087.38  55.6 | 0.8031  0.7627 | 0.6714  0.7917 | 0.7119  0.7734 | 19 hr 17 min  \n108499,108499, torch.Size([512, 40, 101])\n0.0100  109.0 k  1092.39  55.8 | 0.7832  0.7652 | 0.6528  0.7905 | 0.6212  0.8027 | 19 hr 19 min  \n108999,108999, torch.Size([512, 40, 101])\n0.0100  109.5 k  1097.40  56.1 | 0.7807  0.7677 | 0.6528  0.7965 | 0.6304  0.8105 | 19 hr 22 min  \n109499,109499, torch.Size([512, 40, 101])\n0.0100  110.0 k  1102.41  56.3 | 0.7806  0.7674 | 0.6680  0.7891 | 0.6813  0.7871 | 19 hr 24 min  \n109999,109999, torch.Size([512, 40, 101])\n0.0100  110.5 k  1107.42  56.6 | 0.7988  0.7671 | 0.6318  0.8027 | 0.6820  0.7949 | 19 hr 26 min  \n110499,110499, torch.Size([512, 40, 101])\n0.0100  111.0 k  1112.43  56.8 | 0.7735  0.7680 | 0.6892  0.7825 | 0.6944  0.7852 | 19 hr 29 min  \n110999,110999, torch.Size([512, 40, 101])\n0.0100  111.5 k  1117.44  57.1 | 0.7805  0.7661 | 0.6609  0.7952 | 0.6439  0.8066 | 19 hr 31 min  \n111499,111499, torch.Size([512, 40, 101])\n0.0100  112.0 k  1122.46  57.3 | 0.7779  0.7633 | 0.6517  0.7975 | 0.8025  0.7656 | 19 hr 34 min  \n111999,111999, torch.Size([512, 40, 101])\n0.0100  112.5 k  1127.47  57.6 | 0.7480  0.7652 | 0.6583  0.7934 | 0.6775  0.7871 | 19 hr 38 min  \n112499,112499, torch.Size([512, 40, 101])\n0.0100  113.0 k  1132.48  57.9 | 0.7796  0.7574 | 0.6388  0.8022 | 0.5998  0.8105 | 19 hr 42 min  \n112999,112999, torch.Size([512, 40, 101])\n</code></pre>\n\n<p>It would be nice to know if anyone had any success with recurrent networks and reached LB &gt; 0.83?</p>",
      "rawMarkdown": "This is a post exploring some recurrent models and some strange results regarding their convergence. Following Heng's lead by example I am posting here the simple models and the results from the training process.\n\n1st Model:\n\n    class RecNet(nn.Module):\n    def __init__(self, in_shape, num_classes=12 ):\n        super(RecNet, self).__init__()\n\n        self.layer1 = nn.GRU(in_shape[-1], 256)\n        self.layer2 = nn.GRU(256, 512)\n        self.fc1 = nn.Linear(512, 256)\n        self.fc2 = nn.Linear(40 * 256, num_classes)\n\n    def forward(self, x):\n        x, h_out = self.layer1(x)\n        x = F.dropout(x, p=0.5)\n        x, h_ou2 = self.layer2(x)\n        x = F.dropout(x, p=0.3)\n        x = self.fc1(x)\n        x = self.fc2(x.view(-1, 40 * 256))\n\n        return x  #logits\n\nResults:\n\n    0.0100  255.5 k  2560.60  130.8 | 0.7574  0.7880 | 0.5708  0.8218 | 0.6756  0.8047 | 38 hr 58 min  \n    255499,255499, torch.Size([512, 40, 101])\n    0.0100  256.0 k  2565.61  131.1 | 0.7515  0.7824 | 0.5721  0.8260 | 0.5635  0.8125 | 39 hr 01 min  \n    255999,255999, torch.Size([512, 40, 101])\n    0.0100  256.5 k  2570.62  131.3 | 0.7604  0.7749 | 0.5928  0.8178 | 0.5597  0.8359 | 39 hr 04 min  \n    256499,256499, torch.Size([512, 40, 101])\n    0.0100  257.0 k  2575.63  131.6 | 0.7831  0.7777 | 0.5634  0.8272 | 0.5638  0.8184 | 39 hr 07 min  \n    256999,256999, torch.Size([512, 40, 101])\n    0.0100  257.5 k  2580.65  131.8 | 0.7481  0.7814 | 0.5664  0.8249 | 0.5269  0.8418 | 39 hr 10 min  \n    257499,257499, torch.Size([512, 40, 101])\n    0.0100  258.0 k  2585.66  132.1 | 0.7748  0.7752 | 0.5889  0.8206 | 0.5863  0.8203 | 39 hr 13 min      \n    257999,257999, torch.Size([512, 40, 101])                         |\n    0.0100  258.5 k  2590.67  132.4 | 0.7731  0.7727 | 0.5707  0.8066 | 0.5110  0.8379 | 39 hr 16 min  \n    258499,258499, torch.Size([512, 40, 101])\n    0.0100  259.0 k  2595.68  132.6 | 0.7703  0.7733 | 0.5785  0.8177 | 0.5610  0.8340 | 39 hr 20 min  \n    258999,258999, torch.Size([512, 40, 101])\n    0.0100  259.5 k  2600.69  132.9 | 0.7675  0.7739 | 0.5747  0.8214 | 0.5562  0.8340 | 39 hr 23 min  \n    259499,259499, torch.Size([512, 40, 101])\n    0.0100  260.0 k  2605.70  133.1 | 0.7627  0.7786 | 0.5250  0.8320 | 0.5074  0.8398 | 39 hr 26 min  \n    259999,259999, torch.Size([512, 40, 101])\n    0.0100  260.5 k  2610.71  133.4 | 0.7732  0.7733 | 0.5552  0.8251 | 0.5708  0.8301 | 39 hr 29 min  \n    260499,260499, torch.Size([512, 40, 101])\n    0.0100  261.0 k  2615.72  133.6 | 0.7728  0.7736 | 0.5742  0.8249 | 0.6212  0.8125 | 39 hr 32 min  260999,260999, torch.Size([512, 40, 101])\n    0.0100  261.5 k  2620.73  133.9 | 0.7643  0.7839 | 0.5699  0.8255 | 0.5456  0.8184 | 39 hr 35 min  261499,261499, torch.Size([512, 40, 101])\n    0.0100  262.0 k  2625.74  134.1 | 0.7577  0.7792 | 0.5579  0.8296 | 0.4826  0.8574 | 39 hr 38 min  261999,261999, torch.Size([512, 40, 101])\n\nAs you can see the results keep oscillating back and forth. The loss increases and then decreases again. Which indicates that recurrent models are having some difficulty to converge.\n\n 2nd Model:\n\n    class RecNet2(nn.Module):\n    def __init__(self, in_shape, num_classes=12 ):\n        super(RecNet2, self).__init__()\n\n        self.layer1 = nn.GRU(in_shape[-1], 256)\n        self.layer2 = nn.LSTM(256, 512)\n        self.fc1 = nn.Linear(512, 256)\n        self.fc2 = nn.Linear(40 * 256, num_classes)\n\n    def forward(self, x):\n        x, h_out = self.layer1(x)\n        x = F.dropout(x, p=0.5)\n        x, h_ou2 = self.layer2(x)\n        x = F.dropout(x, p=0.3)\n        x = self.fc1(x)\n        x = self.fc2(x.view(-1, 40 * 256))\n\n        return x  #logits\n\nResults:\n\n    0.0100  238.0 k  2385.22  121.9 | 0.7247  0.7855 | 0.5021  0.8462 | 0.5366  0.8418 | 38 hr 11 min  \n    237999,237999, torch.Size([512, 40, 101])\n    0.0100  238.5 k  2390.23  122.1 | 0.7148  0.7820 | 0.5204  0.8376 | 0.5722  0.8359 | 38 hr 15 min  \n    238499,238499, torch.Size([512, 40, 101])\n    0.0100  239.0 k  2395.24  122.4 | 0.7133  0.7914 | 0.5043  0.8417 | 0.5279  0.8203 | 38 hr 17 min  \n    238999,238999, torch.Size([512, 40, 101])  \n    239499,239499, torch.Size([512, 40, 101])\n    0.0100  240.0 k  2405.26  122.9 | 0.7407  0.7802 | 0.5226  0.8391 | 0.4966  0.8340 | 38 hr 23 min  \n    239999,239999, torch.Size([512, 40, 101])\n    0.0100  240.5 k  2410.27  123.1 | 0.7311  0.7861 | 0.5144  0.8438 | 0.4636  0.8633 | 38 hr 26 min  \n    240499,240499, torch.Size([512, 40, 101])\n    0.0100  241.0 k  2415.28  123.4 | 0.7198  0.7933 | 0.4964  0.8525 | 0.5036  0.8281 | 38 hr 29 min  \n    240999,240999, torch.Size([512, 40, 101])\n    0.0100  241.5 k  2420.29  123.6 | 0.7377  0.7802 | 0.5068  0.8414 | 0.5212  0.8398 | 38 hr 32 min  \n    241499,241499, torch.Size([512, 40, 101])\n    0.0100  242.0 k  2425.31  123.9 | 0.7220  0.7892 | 0.5025  0.8483 | 0.4110  0.8711 | 38 hr 35 min  \n    241999,241999, torch.Size([512, 40, 101])\n    0.0100  242.5 k  2430.32  124.2 | 0.7111  0.8001 | 0.5217  0.8391 | 0.5346  0.8379 | 38 hr 38 min  \n    242499,242499, torch.Size([512, 40, 101])\n    0.0100  243.0 k  2435.33  124.4 | 0.6963  0.7951 | 0.5064  0.8432 | 0.5590  0.8457 | 38 hr 41 min  \n    242999,242999, torch.Size([512, 40, 101])\n    0.0100  243.5 k  2440.34  124.7 | 0.7349  0.7898 | 0.5091  0.8421 | 0.5491  0.8184 | 38 hr 44 min  \n    243499,243499, torch.Size([512, 40, 101])\n    0.0100  244.0 k  2445.35  124.9 | 0.7272  0.7827 | 0.5141  0.8426 | 0.5331  0.8379 | 38 hr 47 min  \n    243999,243999, torch.Size([512, 40, 101])\n    0.0100  244.5 k  2450.36  125.2 | 0.7296  0.7901 | 0.5211  0.8365 | 0.5010  0.8477 | 38 hr 50 min  \n    244499,244499, torch.Size([512, 40, 101])\n    0.0100  245.0 k  2455.37  125.4 | 0.7009  0.7951 | 0.5083  0.8164 | 0.5616  0.8438 | 38 hr 53 min  \n    244999,244999, torch.Size([512, 40, 101])\n    0.0100  245.5 k  2460.38  125.7 | 0.7282  0.7867 | 0.5723  0.8239 | 0.5448  0.8281 | 38 hr 56 min  \n    245499,245499, torch.Size([512, 40, 101])\n    0.0100  246.0 k  2465.39  126.0 | 0.7250  0.7901 | 0.5079  0.8429 | 0.4877  0.8516 | 38 hr 59 min  \n    245999,245999, torch.Size([512, 40, 101])\n    0.0100  246.5 k  2470.40  126.2 | 0.7355  0.7817 | 0.5221  0.8385 | 0.4724  0.8691 | 39 hr 01 min  \n    246499,246499, torch.Size([512, 40, 101])\n    0.0100  247.0 k  2475.41  126.5 | 0.7305  0.7830 | 0.5073  0.8448 | 0.5690  0.8047 | 39 hr 04 min  \n    246999,246999, torch.Size([512, 40, 101])\n    0.0100  247.5 k  2480.43  126.7 | 0.7105  0.7914 | 0.5217  0.8378 | 0.4530  0.8613 | 39 hr 07 min  \n    247499,247499, torch.Size([512, 40, 101])\n    0.0100  248.0 k  2485.44  127.0 | 0.7384  0.7795 | 0.5113  0.8452 | 0.5366  0.8262 | 39 hr 12 min  \n    247999,247999, torch.Size([512, 40, 101])\n    0.0100  248.5 k  2490.45  127.2 | 0.7084  0.7901 | 0.5257  0.8362 | 0.5554  0.8359 | 39 hr 16 min  \n    248499,248499, torch.Size([512, 40, 101])\n    0.0100  249.0 k  2495.46  127.5 | 0.7141  0.7917 | 0.4956  0.8470 | 0.4636  0.8301 | 39 hr 20 min  \n    248999,248999, torch.Size([512, 40, 101])\n    0.0100  249.5 k  2500.47  127.7 | 0.7142  0.7973 | 0.5123  0.8398 | 0.4595  0.8613 | 39 hr 26 min  \n    249499,249499, torch.Size([512, 40, 101])\n\n3rd Model:\n\n    class RecNet3(nn.Module):\n    def __init__(self, in_shape, num_classes=12 ):\n        super(RecNet3, self).__init__()\n\n        self.layer1 = nn.GRU(in_shape[-1], 256)\n        self.layer2 = nn.LSTM(256, 512)\n        self.layer3 = nn.GRU(512, 1024)\n        self.layer4 = nn.LSTM(1024, 2048)\n        self.fc1 = nn.Linear(2048, 512)\n        self.fc2 = nn.Linear(40 * 512, num_classes)\n\n    def forward(self, x):\n        x, h_out = self.layer1(x)\n        x = F.dropout(x, p=0.5)\n        x, h_out2 = self.layer2(x)\n        x = F.dropout(x, p=0.3)\n        x, h_out3 = self.layer3(x)\n        x, h_out4 = self.layer4(x)\n        x = F.dropout(x, p=0.3)\n        x = self.fc1(x)\n        x = self.fc2(x.view(-1, 40 * 512))\n\n        return x  #logits\n\nResults:\n\n    0.0100   90.0 k  901.97  46.1 | 0.7829  0.7643 | 0.7526  0.7660 | 0.7332  0.7539 | 35 hr 45 min  \n    89999,89999, torch.Size([512, 40, 101])\n    0.0100   90.5 k  906.98  46.3 | 0.7767  0.7664 | 0.7493  0.7664 | 0.6869  0.7773 | 35 hr 53 min  \n    90499,90499, torch.Size([512, 40, 101])\n    0.0100   91.0 k  911.99  46.6 | 0.8006  0.7596 | 0.7184  0.7765 | 0.7083  0.7617 | 36 hr 05 min  \n    90999,90999, torch.Size([512, 40, 101])\n    0.0100   91.5 k  917.01  46.8 | 0.7871  0.7655 | 0.7387  0.7714 | 0.7490  0.7695 | 36 hr 19 min  \n    91499,91499, torch.Size([512, 40, 101])\n    0.0100   92.0 k  922.02  47.1 | 0.7738  0.7630 | 0.7699  0.7578 | 0.7372  0.7812 | 36 hr 27 min  \n    91999,91999, torch.Size([512, 40, 101])\n    0.0100   92.5 k  927.03  47.4 | 0.7582  0.7749 | 0.7257  0.7653 | 0.6510  0.8125 | 36 hr 39 min  \n    92499,92499, torch.Size([512, 40, 101])\n    0.0100   93.0 k  932.04  47.6 | 0.7869  0.7661 | 0.7169  0.7782 | 0.6819  0.7754 | 37 hr 34 min  \n    92999,92999, torch.Size([512, 40, 101])\n    0.0100   93.5 k  937.05  47.9 | 0.7699  0.7646 | 0.7328  0.7727 | 0.7638  0.7578 | 37 hr 43 min  \n    93499,93499, torch.Size([512, 40, 101])\n    0.0100   94.0 k  942.06  48.1 | 0.7432  0.7736 | 0.7326  0.7723 | 0.7336  0.7656 | 37 hr 50 min  \n    93999,93999, torch.Size([512, 40, 101])\n    0.0100   94.5 k  947.07  48.4 | 0.7618  0.7714 | 0.7094  0.7786 | 0.8350  0.7480 | 37 hr 58 min  \n    94499,94499, torch.Size([512, 40, 101])\n    0.0100   95.0 k  952.08  48.6 | 0.7597  0.7696 | 0.7074  0.7778 | 0.7055  0.7793 | 38 hr 06 min  \n    94999,94999, torch.Size([512, 40, 101])a\n    0.0100   95.5 k  957.09  48.9 | 0.7571  0.7733 | 0.7339  0.7701 | 0.7449  0.7754 | 38 hr 13 min  \n    95499,95499, torch.Size([512, 40, 101])\n    0.0100   96.0 k  962.10  49.2 | 0.7770  0.7608 | 0.6961  0.7856 | 0.6177  0.8027 | 38 hr 21 min  \n    95999,95999, torch.Size([512, 40, 101])\n    0.0100   96.5 k  967.12  49.4 | 0.7517  0.7789 | 0.7187  0.7910 | 0.7956  0.7637 | 38 hr 29 min  \n    96499,96499, torch.Size([512, 40, 101])\n    0.0100   97.0 k  972.13  49.7 | 0.7609  0.7733 | 0.7123  0.7822 | 0.6791  0.7559 | 38 hr 36 min  \n    96999,96999, torch.Size([512, 40, 101])\n    0.0100   97.5 k  977.14  49.9 | 0.7771  0.7636 | 0.7181  0.7743 | 0.6693  0.7832 | 38 hr 44 min  \n    97499,97499, torch.Size([512, 40, 101])\n    0.0100   97.6 k  977.87  50.0 | 0.7771  0.7636 | 0.6780  0.7793 | 0.7790  0.7656 | 38 hr 45 min  \n    97573,97573, torch.Size([512, 40, 101])\n\n4th Model: Modified Heng's resnet with LSTM layer at on top.\n\n    class SeResNet3(nn.Module):\n    def __init__(self, in_shape=(1,40,101), num_classes=12 ):\n        super(SeResNet3, self).__init__()\n        in_channels = in_shape[0]\n\n        self.layer1a = ConvBn2d(in_channels, 16, kernel_size=(3, 3), stride=(1, 1))\n        self.layer1b = ResBlock( 16, 16)\n\n        self.layer2a = ConvBn2d(16, 32, kernel_size=(3, 3), stride=(1, 1))\n        self.layer2b = ResBlock(32, 32)\n        self.layer2c = ResBlock(32, 32)\n\n        self.layer3a = ConvBn2d(32, 64, kernel_size=(3, 3), stride=(1, 1))\n        self.layer3b = ResBlock(64, 64)\n        self.layer3c = ResBlock(64, 64)\n\n        self.layer4a = ConvBn2d( 64,128, kernel_size=(3, 3), stride=(1, 1))\n        self.layer4b = ResBlock(128,128)\n        self.layer4c = ResBlock(128,128)\n\n        self.layer5a = ConvBn2d(128, 256, kernel_size=(3, 3), stride=(1, 1))\n        self.layer5ab = nn.LSTMCell(256, 256, 10)\n        self.layer5b = nn.Linear(256,256)\n\n        self.fc = nn.Linear(256,num_classes)\n\n        def forward(self, x):\n\n        x = F.relu(self.layer1a(x),inplace=True)\n        x = self.layer1b(x)\n        x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n        x = F.dropout(x,p=0.1,training=self.training)\n        x = F.relu(self.layer2a(x),inplace=True)\n        x = self.layer2b(x)\n        x = self.layer2c(x)\n        x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.layer3a(x),inplace=True)\n        x = self.layer3b(x)\n        x = self.layer3c(x)\n        x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.layer4a(x),inplace=True)\n        x = self.layer4b(x)\n        x = self.layer4c(x)\n\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.layer5a(x),inplace=True)\n        x = F.adaptive_avg_pool2d(x,1)\n        x, h = F.dropout(F.relu(self.layer5ab(x.view(-1, 40, 256))), p=0.2, self.training)\n        x = x.view(x.size(0), -1)\n        x = F.relu(self.layer5b(x))\n\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = self.fc(x)\n\n        return x  #logits\n\nResults:\n\n    0.0100  106.5 k  1067.33  54.5 | 0.8049  0.7540 | 0.6601  0.7930 | 0.6240  0.8066 | 19 hr 07 min  \n    106499,106499, torch.Size([512, 40, 101])\n    0.0100  107.0 k  1072.35  54.8 | 0.7762  0.7671 | 0.6527  0.7935 | 0.7233  0.7871 | 19 hr 09 min  \n    106999,106999, torch.Size([512, 40, 101])\n    0.0100  107.5 k  1077.36  55.0 | 0.7861  0.7568 | 0.6754  0.7943 | 0.5772  0.8184 | 19 hr 12 min  \n    107499,107499, torch.Size([512, 40, 101])\n    0.0100  108.0 k  1082.37  55.3 | 0.7906  0.7537 | 0.6592  0.7966 | 0.7249  0.7578 | 19 hr 15 min  \n    107999,107999, torch.Size([512, 40, 101])\n    0.0100  108.5 k  1087.38  55.6 | 0.8031  0.7627 | 0.6714  0.7917 | 0.7119  0.7734 | 19 hr 17 min  \n    108499,108499, torch.Size([512, 40, 101])\n    0.0100  109.0 k  1092.39  55.8 | 0.7832  0.7652 | 0.6528  0.7905 | 0.6212  0.8027 | 19 hr 19 min  \n    108999,108999, torch.Size([512, 40, 101])\n    0.0100  109.5 k  1097.40  56.1 | 0.7807  0.7677 | 0.6528  0.7965 | 0.6304  0.8105 | 19 hr 22 min  \n    109499,109499, torch.Size([512, 40, 101])\n    0.0100  110.0 k  1102.41  56.3 | 0.7806  0.7674 | 0.6680  0.7891 | 0.6813  0.7871 | 19 hr 24 min  \n    109999,109999, torch.Size([512, 40, 101])\n    0.0100  110.5 k  1107.42  56.6 | 0.7988  0.7671 | 0.6318  0.8027 | 0.6820  0.7949 | 19 hr 26 min  \n    110499,110499, torch.Size([512, 40, 101])\n    0.0100  111.0 k  1112.43  56.8 | 0.7735  0.7680 | 0.6892  0.7825 | 0.6944  0.7852 | 19 hr 29 min  \n    110999,110999, torch.Size([512, 40, 101])\n    0.0100  111.5 k  1117.44  57.1 | 0.7805  0.7661 | 0.6609  0.7952 | 0.6439  0.8066 | 19 hr 31 min  \n    111499,111499, torch.Size([512, 40, 101])\n    0.0100  112.0 k  1122.46  57.3 | 0.7779  0.7633 | 0.6517  0.7975 | 0.8025  0.7656 | 19 hr 34 min  \n    111999,111999, torch.Size([512, 40, 101])\n    0.0100  112.5 k  1127.47  57.6 | 0.7480  0.7652 | 0.6583  0.7934 | 0.6775  0.7871 | 19 hr 38 min  \n    112499,112499, torch.Size([512, 40, 101])\n    0.0100  113.0 k  1132.48  57.9 | 0.7796  0.7574 | 0.6388  0.8022 | 0.5998  0.8105 | 19 hr 42 min  \n    112999,112999, torch.Size([512, 40, 101])\n\nIt would be nice to know if anyone had any success with recurrent networks and reached LB &gt; 0.83?",
      "votes": null
    },
    {
      "id": "269328",
      "postDate": "01/16/2018 16:32:59",
      "content": "<p>I had tried Recurrent Networks and experienced problems exactly like what you have encountered. My first submission, which was an RNN, did exactly 83%. Subsequent RNN networks haven't been able to do better than that. Maybe an 84 would be in the cards. I wouldn't know because I stopped measuring the performance of every single network of mine on the LB at least 50 submissions ago :D :D</p>",
      "rawMarkdown": "I had tried Recurrent Networks and experienced problems exactly like what you have encountered. My first submission, which was an RNN, did exactly 83%. Subsequent RNN networks haven't been able to do better than that. Maybe an 84 would be in the cards. I wouldn't know because I stopped measuring the performance of every single network of mine on the LB at least 50 submissions ago :D :D",
      "votes": null
    },
    {
      "id": "269405",
      "postDate": "01/16/2018 19:08:40",
      "content": "<p>Thanks for the input Sarthak. There's must be some limitations regarding convergence of RNN in this setting. At least it seems that way from our experiences.</p>",
      "rawMarkdown": "Thanks for the input Sarthak. There's must be some limitations regarding convergence of RNN in this setting. At least it seems that way from our experiences.",
      "votes": null
    },
    {
      "id": "269407",
      "postDate": "01/16/2018 19:21:50",
      "content": "<p>I had a conv-RNN that got 87%, but it didn't end up being helpful even in an ensemble (although I'm pretty new to ensembling, more advanced techniques probably could have gained from it)</p>\n\n<p>The RNN bit was a two-layer GRU a la the Hello Edge paper, tried a one-layer LSTM too that did great in local validation but got 86% LB</p>",
      "rawMarkdown": "I had a conv-RNN that got 87%, but it didn't end up being helpful even in an ensemble (although I'm pretty new to ensembling, more advanced techniques probably could have gained from it)\n\nThe RNN bit was a two-layer GRU a la the Hello Edge paper, tried a one-layer LSTM too that did great in local validation but got 86% LB",
      "votes": null
    },
    {
      "id": "269422",
      "postDate": "01/16/2018 20:12:15",
      "content": "<p>Hey Thomas thanks for sharing your experience. I find it odd because I've had as you can see from the examples, one of the models was exactly that. A 2 layer GRU but couldn't get pas the 83% LB mark sometimes even worse 82%. If I may ask are you splitting  the data the same way as @Heng? or are you using some other strategy? What models have given you better resutls? pure conv ones or mixed ones?</p>",
      "rawMarkdown": "Hey Thomas thanks for sharing your experience. I find it odd because I've had as you can see from the examples, one of the models was exactly that. A 2 layer GRU but couldn't get pas the 83% LB mark sometimes even worse 82%. If I may ask are you splitting  the data the same way as @Heng? or are you using some other strategy? What models have given you better resutls? pure conv ones or mixed ones?",
      "votes": null
    },
    {
      "id": "269423",
      "postDate": "01/16/2018 20:17:34",
      "content": "<p>I built a LSTM-HMM that got 89% in LB... but I was not able get any improvement with ensembles.</p>",
      "rawMarkdown": "I built a LSTM-HMM that got 89% in LB... but I was not able get any improvement with ensembles.",
      "votes": null
    },
    {
      "id": "269426",
      "postDate": "01/16/2018 20:32:54",
      "content": "<p>Can you explain how you used the LSTM-HMM?link to an article explaining your method? \nThanks</p>",
      "rawMarkdown": "Can you explain how you used the LSTM-HMM?link to an article explaining your method? \nThanks",
      "votes": null
    },
    {
      "id": "269455",
      "postDate": "01/16/2018 21:16:51",
      "content": "<p>Well, this is a kind of usual architecture in current automatic speech recognition systems. The idea is that the probabilities generated by \"something\" feed a Hidden Markov Model. And normally Viterbi is used to find out a path with the \"best probability\". This \"something\" may be a GMM (as it used to be before 2012) or a DNN. In my experiments, I tried a many combinations of LSTM, CNN and TDNN. In the end, I got 0.89 with each one of these nets, but I was not able to combine them in a better system. I tried the strategy shown <a href=\"https://pdfs.semanticscholar.org/8201/55ecb57325503183253b8796de5f4535eb16.pdf\">here</a>, but without success. </p>\n\n<p>And some other references about <a href=\"https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/HintonDengYuEtAl-SPM2012.pdf\">acoustic models using DNN</a> and, more specifically, <a href=\"https://arxiv.org/pdf/1303.5778.pdf\">LSTM in ASR</a></p>",
      "rawMarkdown": "Well, this is a kind of usual architecture in current automatic speech recognition systems. The idea is that the probabilities generated by \"something\" feed a Hidden Markov Model. And normally Viterbi is used to find out a path with the \"best probability\". This \"something\" may be a GMM (as it used to be before 2012) or a DNN. In my experiments, I tried a many combinations of LSTM, CNN and TDNN. In the end, I got 0.89 with each one of these nets, but I was not able to combine them in a better system. I tried the strategy shown [here][1], but without success. \n\nAnd some other references about [acoustic models using DNN][2] and, more specifically, [LSTM in ASR][3]\n\n\n  [1]: https://pdfs.semanticscholar.org/8201/55ecb57325503183253b8796de5f4535eb16.pdf\n  [2]: https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/HintonDengYuEtAl-SPM2012.pdf\n  [3]: https://arxiv.org/pdf/1303.5778.pdf",
      "votes": null
    },
    {
      "id": "269580",
      "postDate": "01/17/2018 02:27:18",
      "content": "<p>Hey I'm using similar method! What do you use for Conv-part? my best single model has a Inception-like head and 2 GRUs attached and got .87 in public .88 in privage LB. no-augmentation or test-prob used.</p>",
      "rawMarkdown": "Hey I'm using similar method! What do you use for Conv-part? my best single model has a Inception-like head and 2 GRUs attached and got .87 in public .88 in privage LB. no-augmentation or test-prob used.",
      "votes": null
    },
    {
      "id": "269959",
      "postDate": "01/17/2018 15:23:47",
      "content": "<p>Ensembling Heng's resnet, private LB: .85 and the gru model above: LB .67 gives a private LB score of 0.86. Didn't had the time to submit before competition end.</p>",
      "rawMarkdown": "Ensembling Heng's resnet, private LB: .85 and the gru model above: LB .67 gives a private LB score of 0.86. Didn't had the time to submit before competition end.",
      "votes": null
    },
    {
      "id": "270238",
      "postDate": "01/18/2018 00:05:31",
      "content": "<p>thanks for your info! .67 is a little bit weird, I found recent paper prefer the CNN-GRU architecture and in practice it do perform well. I'm running out of time to explore more possibility.</p>",
      "rawMarkdown": "thanks for your info! .67 is a little bit weird, I found recent paper prefer the CNN-GRU architecture and in practice it do perform well. I'm running out of time to explore more possibility.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 269328,
      "author_name": "yadavsarthak",
      "author_url": "",
      "post_date": "01/16/2018 16:32:59",
      "content": "<p>I had tried Recurrent Networks and experienced problems exactly like what you have encountered. My first submission, which was an RNN, did exactly 83%. Subsequent RNN networks haven't been able to do better than that. Maybe an 84 would be in the cards. I wouldn't know because I stopped measuring the performance of every single network of mine on the LB at least 50 submissions ago :D :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 269405,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/16/2018 19:08:40",
          "content": "<p>Thanks for the input Sarthak. There's must be some limitations regarding convergence of RNN in this setting. At least it seems that way from our experiences.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269407,
      "author_name": "omalleyt",
      "author_url": "",
      "post_date": "01/16/2018 19:21:50",
      "content": "<p>I had a conv-RNN that got 87%, but it didn't end up being helpful even in an ensemble (although I'm pretty new to ensembling, more advanced techniques probably could have gained from it)</p>\n\n<p>The RNN bit was a two-layer GRU a la the Hello Edge paper, tried a one-layer LSTM too that did great in local validation but got 86% LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 269422,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/16/2018 20:12:15",
          "content": "<p>Hey Thomas thanks for sharing your experience. I find it odd because I've had as you can see from the examples, one of the models was exactly that. A 2 layer GRU but couldn't get pas the 83% LB mark sometimes even worse 82%. If I may ask are you splitting  the data the same way as @Heng? or are you using some other strategy? What models have given you better resutls? pure conv ones or mixed ones?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269580,
          "author_name": "philipxue",
          "author_url": "",
          "post_date": "01/17/2018 02:27:18",
          "content": "<p>Hey I'm using similar method! What do you use for Conv-part? my best single model has a Inception-like head and 2 GRUs attached and got .87 in public .88 in privage LB. no-augmentation or test-prob used.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269423,
      "author_name": "jcsilva",
      "author_url": "",
      "post_date": "01/16/2018 20:17:34",
      "content": "<p>I built a LSTM-HMM that got 89% in LB... but I was not able get any improvement with ensembles.</p>",
      "votes": null,
      "replies": [
        {
          "id": 269426,
          "author_name": "ori226",
          "author_url": "",
          "post_date": "01/16/2018 20:32:54",
          "content": "<p>Can you explain how you used the LSTM-HMM?link to an article explaining your method? \nThanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269455,
          "author_name": "jcsilva",
          "author_url": "",
          "post_date": "01/16/2018 21:16:51",
          "content": "<p>Well, this is a kind of usual architecture in current automatic speech recognition systems. The idea is that the probabilities generated by \"something\" feed a Hidden Markov Model. And normally Viterbi is used to find out a path with the \"best probability\". This \"something\" may be a GMM (as it used to be before 2012) or a DNN. In my experiments, I tried a many combinations of LSTM, CNN and TDNN. In the end, I got 0.89 with each one of these nets, but I was not able to combine them in a better system. I tried the strategy shown <a href=\"https://pdfs.semanticscholar.org/8201/55ecb57325503183253b8796de5f4535eb16.pdf\">here</a>, but without success. </p>\n\n<p>And some other references about <a href=\"https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/HintonDengYuEtAl-SPM2012.pdf\">acoustic models using DNN</a> and, more specifically, <a href=\"https://arxiv.org/pdf/1303.5778.pdf\">LSTM in ASR</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269959,
      "author_name": "kirk86",
      "author_url": "",
      "post_date": "01/17/2018 15:23:47",
      "content": "<p>Ensembling Heng's resnet, private LB: .85 and the gru model above: LB .67 gives a private LB score of 0.86. Didn't had the time to submit before competition end.</p>",
      "votes": null,
      "replies": [
        {
          "id": 270238,
          "author_name": "philipxue",
          "author_url": "",
          "post_date": "01/18/2018 00:05:31",
          "content": "<p>thanks for your info! .67 is a little bit weird, I found recent paper prefer the CNN-GRU architecture and in practice it do perform well. I'm running out of time to explore more possibility.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "269262": "This is a post exploring some recurrent models and some strange results regarding their convergence. Following Heng's lead by example I am posting here the simple models and the results from the training process.\n\n1st Model:\n\n    class RecNet(nn.Module):\n    def __init__(self, in_shape, num_classes=12 ):\n        super(RecNet, self).__init__()\n\n        self.layer1 = nn.GRU(in_shape[-1], 256)\n        self.layer2 = nn.GRU(256, 512)\n        self.fc1 = nn.Linear(512, 256)\n        self.fc2 = nn.Linear(40 * 256, num_classes)\n\n    def forward(self, x):\n        x, h_out = self.layer1(x)\n        x = F.dropout(x, p=0.5)\n        x, h_ou2 = self.layer2(x)\n        x = F.dropout(x, p=0.3)\n        x = self.fc1(x)\n        x = self.fc2(x.view(-1, 40 * 256))\n\n        return x  #logits\n\nResults:\n\n    0.0100  255.5 k  2560.60  130.8 | 0.7574  0.7880 | 0.5708  0.8218 | 0.6756  0.8047 | 38 hr 58 min  \n    255499,255499, torch.Size([512, 40, 101])\n    0.0100  256.0 k  2565.61  131.1 | 0.7515  0.7824 | 0.5721  0.8260 | 0.5635  0.8125 | 39 hr 01 min  \n    255999,255999, torch.Size([512, 40, 101])\n    0.0100  256.5 k  2570.62  131.3 | 0.7604  0.7749 | 0.5928  0.8178 | 0.5597  0.8359 | 39 hr 04 min  \n    256499,256499, torch.Size([512, 40, 101])\n    0.0100  257.0 k  2575.63  131.6 | 0.7831  0.7777 | 0.5634  0.8272 | 0.5638  0.8184 | 39 hr 07 min  \n    256999,256999, torch.Size([512, 40, 101])\n    0.0100  257.5 k  2580.65  131.8 | 0.7481  0.7814 | 0.5664  0.8249 | 0.5269  0.8418 | 39 hr 10 min  \n    257499,257499, torch.Size([512, 40, 101])\n    0.0100  258.0 k  2585.66  132.1 | 0.7748  0.7752 | 0.5889  0.8206 | 0.5863  0.8203 | 39 hr 13 min      \n    257999,257999, torch.Size([512, 40, 101])                         |\n    0.0100  258.5 k  2590.67  132.4 | 0.7731  0.7727 | 0.5707  0.8066 | 0.5110  0.8379 | 39 hr 16 min  \n    258499,258499, torch.Size([512, 40, 101])\n    0.0100  259.0 k  2595.68  132.6 | 0.7703  0.7733 | 0.5785  0.8177 | 0.5610  0.8340 | 39 hr 20 min  \n    258999,258999, torch.Size([512, 40, 101])\n    0.0100  259.5 k  2600.69  132.9 | 0.7675  0.7739 | 0.5747  0.8214 | 0.5562  0.8340 | 39 hr 23 min  \n    259499,259499, torch.Size([512, 40, 101])\n    0.0100  260.0 k  2605.70  133.1 | 0.7627  0.7786 | 0.5250  0.8320 | 0.5074  0.8398 | 39 hr 26 min  \n    259999,259999, torch.Size([512, 40, 101])\n    0.0100  260.5 k  2610.71  133.4 | 0.7732  0.7733 | 0.5552  0.8251 | 0.5708  0.8301 | 39 hr 29 min  \n    260499,260499, torch.Size([512, 40, 101])\n    0.0100  261.0 k  2615.72  133.6 | 0.7728  0.7736 | 0.5742  0.8249 | 0.6212  0.8125 | 39 hr 32 min  260999,260999, torch.Size([512, 40, 101])\n    0.0100  261.5 k  2620.73  133.9 | 0.7643  0.7839 | 0.5699  0.8255 | 0.5456  0.8184 | 39 hr 35 min  261499,261499, torch.Size([512, 40, 101])\n    0.0100  262.0 k  2625.74  134.1 | 0.7577  0.7792 | 0.5579  0.8296 | 0.4826  0.8574 | 39 hr 38 min  261999,261999, torch.Size([512, 40, 101])\n\nAs you can see the results keep oscillating back and forth. The loss increases and then decreases again. Which indicates that recurrent models are having some difficulty to converge.\n\n 2nd Model:\n\n    class RecNet2(nn.Module):\n    def __init__(self, in_shape, num_classes=12 ):\n        super(RecNet2, self).__init__()\n\n        self.layer1 = nn.GRU(in_shape[-1], 256)\n        self.layer2 = nn.LSTM(256, 512)\n        self.fc1 = nn.Linear(512, 256)\n        self.fc2 = nn.Linear(40 * 256, num_classes)\n\n    def forward(self, x):\n        x, h_out = self.layer1(x)\n        x = F.dropout(x, p=0.5)\n        x, h_ou2 = self.layer2(x)\n        x = F.dropout(x, p=0.3)\n        x = self.fc1(x)\n        x = self.fc2(x.view(-1, 40 * 256))\n\n        return x  #logits\n\nResults:\n\n    0.0100  238.0 k  2385.22  121.9 | 0.7247  0.7855 | 0.5021  0.8462 | 0.5366  0.8418 | 38 hr 11 min  \n    237999,237999, torch.Size([512, 40, 101])\n    0.0100  238.5 k  2390.23  122.1 | 0.7148  0.7820 | 0.5204  0.8376 | 0.5722  0.8359 | 38 hr 15 min  \n    238499,238499, torch.Size([512, 40, 101])\n    0.0100  239.0 k  2395.24  122.4 | 0.7133  0.7914 | 0.5043  0.8417 | 0.5279  0.8203 | 38 hr 17 min  \n    238999,238999, torch.Size([512, 40, 101])  \n    239499,239499, torch.Size([512, 40, 101])\n    0.0100  240.0 k  2405.26  122.9 | 0.7407  0.7802 | 0.5226  0.8391 | 0.4966  0.8340 | 38 hr 23 min  \n    239999,239999, torch.Size([512, 40, 101])\n    0.0100  240.5 k  2410.27  123.1 | 0.7311  0.7861 | 0.5144  0.8438 | 0.4636  0.8633 | 38 hr 26 min  \n    240499,240499, torch.Size([512, 40, 101])\n    0.0100  241.0 k  2415.28  123.4 | 0.7198  0.7933 | 0.4964  0.8525 | 0.5036  0.8281 | 38 hr 29 min  \n    240999,240999, torch.Size([512, 40, 101])\n    0.0100  241.5 k  2420.29  123.6 | 0.7377  0.7802 | 0.5068  0.8414 | 0.5212  0.8398 | 38 hr 32 min  \n    241499,241499, torch.Size([512, 40, 101])\n    0.0100  242.0 k  2425.31  123.9 | 0.7220  0.7892 | 0.5025  0.8483 | 0.4110  0.8711 | 38 hr 35 min  \n    241999,241999, torch.Size([512, 40, 101])\n    0.0100  242.5 k  2430.32  124.2 | 0.7111  0.8001 | 0.5217  0.8391 | 0.5346  0.8379 | 38 hr 38 min  \n    242499,242499, torch.Size([512, 40, 101])\n    0.0100  243.0 k  2435.33  124.4 | 0.6963  0.7951 | 0.5064  0.8432 | 0.5590  0.8457 | 38 hr 41 min  \n    242999,242999, torch.Size([512, 40, 101])\n    0.0100  243.5 k  2440.34  124.7 | 0.7349  0.7898 | 0.5091  0.8421 | 0.5491  0.8184 | 38 hr 44 min  \n    243499,243499, torch.Size([512, 40, 101])\n    0.0100  244.0 k  2445.35  124.9 | 0.7272  0.7827 | 0.5141  0.8426 | 0.5331  0.8379 | 38 hr 47 min  \n    243999,243999, torch.Size([512, 40, 101])\n    0.0100  244.5 k  2450.36  125.2 | 0.7296  0.7901 | 0.5211  0.8365 | 0.5010  0.8477 | 38 hr 50 min  \n    244499,244499, torch.Size([512, 40, 101])\n    0.0100  245.0 k  2455.37  125.4 | 0.7009  0.7951 | 0.5083  0.8164 | 0.5616  0.8438 | 38 hr 53 min  \n    244999,244999, torch.Size([512, 40, 101])\n    0.0100  245.5 k  2460.38  125.7 | 0.7282  0.7867 | 0.5723  0.8239 | 0.5448  0.8281 | 38 hr 56 min  \n    245499,245499, torch.Size([512, 40, 101])\n    0.0100  246.0 k  2465.39  126.0 | 0.7250  0.7901 | 0.5079  0.8429 | 0.4877  0.8516 | 38 hr 59 min  \n    245999,245999, torch.Size([512, 40, 101])\n    0.0100  246.5 k  2470.40  126.2 | 0.7355  0.7817 | 0.5221  0.8385 | 0.4724  0.8691 | 39 hr 01 min  \n    246499,246499, torch.Size([512, 40, 101])\n    0.0100  247.0 k  2475.41  126.5 | 0.7305  0.7830 | 0.5073  0.8448 | 0.5690  0.8047 | 39 hr 04 min  \n    246999,246999, torch.Size([512, 40, 101])\n    0.0100  247.5 k  2480.43  126.7 | 0.7105  0.7914 | 0.5217  0.8378 | 0.4530  0.8613 | 39 hr 07 min  \n    247499,247499, torch.Size([512, 40, 101])\n    0.0100  248.0 k  2485.44  127.0 | 0.7384  0.7795 | 0.5113  0.8452 | 0.5366  0.8262 | 39 hr 12 min  \n    247999,247999, torch.Size([512, 40, 101])\n    0.0100  248.5 k  2490.45  127.2 | 0.7084  0.7901 | 0.5257  0.8362 | 0.5554  0.8359 | 39 hr 16 min  \n    248499,248499, torch.Size([512, 40, 101])\n    0.0100  249.0 k  2495.46  127.5 | 0.7141  0.7917 | 0.4956  0.8470 | 0.4636  0.8301 | 39 hr 20 min  \n    248999,248999, torch.Size([512, 40, 101])\n    0.0100  249.5 k  2500.47  127.7 | 0.7142  0.7973 | 0.5123  0.8398 | 0.4595  0.8613 | 39 hr 26 min  \n    249499,249499, torch.Size([512, 40, 101])\n\n3rd Model:\n\n    class RecNet3(nn.Module):\n    def __init__(self, in_shape, num_classes=12 ):\n        super(RecNet3, self).__init__()\n\n        self.layer1 = nn.GRU(in_shape[-1], 256)\n        self.layer2 = nn.LSTM(256, 512)\n        self.layer3 = nn.GRU(512, 1024)\n        self.layer4 = nn.LSTM(1024, 2048)\n        self.fc1 = nn.Linear(2048, 512)\n        self.fc2 = nn.Linear(40 * 512, num_classes)\n\n    def forward(self, x):\n        x, h_out = self.layer1(x)\n        x = F.dropout(x, p=0.5)\n        x, h_out2 = self.layer2(x)\n        x = F.dropout(x, p=0.3)\n        x, h_out3 = self.layer3(x)\n        x, h_out4 = self.layer4(x)\n        x = F.dropout(x, p=0.3)\n        x = self.fc1(x)\n        x = self.fc2(x.view(-1, 40 * 512))\n\n        return x  #logits\n\nResults:\n\n    0.0100   90.0 k  901.97  46.1 | 0.7829  0.7643 | 0.7526  0.7660 | 0.7332  0.7539 | 35 hr 45 min  \n    89999,89999, torch.Size([512, 40, 101])\n    0.0100   90.5 k  906.98  46.3 | 0.7767  0.7664 | 0.7493  0.7664 | 0.6869  0.7773 | 35 hr 53 min  \n    90499,90499, torch.Size([512, 40, 101])\n    0.0100   91.0 k  911.99  46.6 | 0.8006  0.7596 | 0.7184  0.7765 | 0.7083  0.7617 | 36 hr 05 min  \n    90999,90999, torch.Size([512, 40, 101])\n    0.0100   91.5 k  917.01  46.8 | 0.7871  0.7655 | 0.7387  0.7714 | 0.7490  0.7695 | 36 hr 19 min  \n    91499,91499, torch.Size([512, 40, 101])\n    0.0100   92.0 k  922.02  47.1 | 0.7738  0.7630 | 0.7699  0.7578 | 0.7372  0.7812 | 36 hr 27 min  \n    91999,91999, torch.Size([512, 40, 101])\n    0.0100   92.5 k  927.03  47.4 | 0.7582  0.7749 | 0.7257  0.7653 | 0.6510  0.8125 | 36 hr 39 min  \n    92499,92499, torch.Size([512, 40, 101])\n    0.0100   93.0 k  932.04  47.6 | 0.7869  0.7661 | 0.7169  0.7782 | 0.6819  0.7754 | 37 hr 34 min  \n    92999,92999, torch.Size([512, 40, 101])\n    0.0100   93.5 k  937.05  47.9 | 0.7699  0.7646 | 0.7328  0.7727 | 0.7638  0.7578 | 37 hr 43 min  \n    93499,93499, torch.Size([512, 40, 101])\n    0.0100   94.0 k  942.06  48.1 | 0.7432  0.7736 | 0.7326  0.7723 | 0.7336  0.7656 | 37 hr 50 min  \n    93999,93999, torch.Size([512, 40, 101])\n    0.0100   94.5 k  947.07  48.4 | 0.7618  0.7714 | 0.7094  0.7786 | 0.8350  0.7480 | 37 hr 58 min  \n    94499,94499, torch.Size([512, 40, 101])\n    0.0100   95.0 k  952.08  48.6 | 0.7597  0.7696 | 0.7074  0.7778 | 0.7055  0.7793 | 38 hr 06 min  \n    94999,94999, torch.Size([512, 40, 101])a\n    0.0100   95.5 k  957.09  48.9 | 0.7571  0.7733 | 0.7339  0.7701 | 0.7449  0.7754 | 38 hr 13 min  \n    95499,95499, torch.Size([512, 40, 101])\n    0.0100   96.0 k  962.10  49.2 | 0.7770  0.7608 | 0.6961  0.7856 | 0.6177  0.8027 | 38 hr 21 min  \n    95999,95999, torch.Size([512, 40, 101])\n    0.0100   96.5 k  967.12  49.4 | 0.7517  0.7789 | 0.7187  0.7910 | 0.7956  0.7637 | 38 hr 29 min  \n    96499,96499, torch.Size([512, 40, 101])\n    0.0100   97.0 k  972.13  49.7 | 0.7609  0.7733 | 0.7123  0.7822 | 0.6791  0.7559 | 38 hr 36 min  \n    96999,96999, torch.Size([512, 40, 101])\n    0.0100   97.5 k  977.14  49.9 | 0.7771  0.7636 | 0.7181  0.7743 | 0.6693  0.7832 | 38 hr 44 min  \n    97499,97499, torch.Size([512, 40, 101])\n    0.0100   97.6 k  977.87  50.0 | 0.7771  0.7636 | 0.6780  0.7793 | 0.7790  0.7656 | 38 hr 45 min  \n    97573,97573, torch.Size([512, 40, 101])\n\n4th Model: Modified Heng's resnet with LSTM layer at on top.\n\n    class SeResNet3(nn.Module):\n    def __init__(self, in_shape=(1,40,101), num_classes=12 ):\n        super(SeResNet3, self).__init__()\n        in_channels = in_shape[0]\n\n        self.layer1a = ConvBn2d(in_channels, 16, kernel_size=(3, 3), stride=(1, 1))\n        self.layer1b = ResBlock( 16, 16)\n\n        self.layer2a = ConvBn2d(16, 32, kernel_size=(3, 3), stride=(1, 1))\n        self.layer2b = ResBlock(32, 32)\n        self.layer2c = ResBlock(32, 32)\n\n        self.layer3a = ConvBn2d(32, 64, kernel_size=(3, 3), stride=(1, 1))\n        self.layer3b = ResBlock(64, 64)\n        self.layer3c = ResBlock(64, 64)\n\n        self.layer4a = ConvBn2d( 64,128, kernel_size=(3, 3), stride=(1, 1))\n        self.layer4b = ResBlock(128,128)\n        self.layer4c = ResBlock(128,128)\n\n        self.layer5a = ConvBn2d(128, 256, kernel_size=(3, 3), stride=(1, 1))\n        self.layer5ab = nn.LSTMCell(256, 256, 10)\n        self.layer5b = nn.Linear(256,256)\n\n        self.fc = nn.Linear(256,num_classes)\n\n        def forward(self, x):\n\n        x = F.relu(self.layer1a(x),inplace=True)\n        x = self.layer1b(x)\n        x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n        x = F.dropout(x,p=0.1,training=self.training)\n        x = F.relu(self.layer2a(x),inplace=True)\n        x = self.layer2b(x)\n        x = self.layer2c(x)\n        x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.layer3a(x),inplace=True)\n        x = self.layer3b(x)\n        x = self.layer3c(x)\n        x = F.max_pool2d(x,kernel_size=(2,2),stride=(2,2))\n\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.layer4a(x),inplace=True)\n        x = self.layer4b(x)\n        x = self.layer4c(x)\n\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = F.relu(self.layer5a(x),inplace=True)\n        x = F.adaptive_avg_pool2d(x,1)\n        x, h = F.dropout(F.relu(self.layer5ab(x.view(-1, 40, 256))), p=0.2, self.training)\n        x = x.view(x.size(0), -1)\n        x = F.relu(self.layer5b(x))\n\n        x = F.dropout(x,p=0.2,training=self.training)\n        x = self.fc(x)\n\n        return x  #logits\n\nResults:\n\n    0.0100  106.5 k  1067.33  54.5 | 0.8049  0.7540 | 0.6601  0.7930 | 0.6240  0.8066 | 19 hr 07 min  \n    106499,106499, torch.Size([512, 40, 101])\n    0.0100  107.0 k  1072.35  54.8 | 0.7762  0.7671 | 0.6527  0.7935 | 0.7233  0.7871 | 19 hr 09 min  \n    106999,106999, torch.Size([512, 40, 101])\n    0.0100  107.5 k  1077.36  55.0 | 0.7861  0.7568 | 0.6754  0.7943 | 0.5772  0.8184 | 19 hr 12 min  \n    107499,107499, torch.Size([512, 40, 101])\n    0.0100  108.0 k  1082.37  55.3 | 0.7906  0.7537 | 0.6592  0.7966 | 0.7249  0.7578 | 19 hr 15 min  \n    107999,107999, torch.Size([512, 40, 101])\n    0.0100  108.5 k  1087.38  55.6 | 0.8031  0.7627 | 0.6714  0.7917 | 0.7119  0.7734 | 19 hr 17 min  \n    108499,108499, torch.Size([512, 40, 101])\n    0.0100  109.0 k  1092.39  55.8 | 0.7832  0.7652 | 0.6528  0.7905 | 0.6212  0.8027 | 19 hr 19 min  \n    108999,108999, torch.Size([512, 40, 101])\n    0.0100  109.5 k  1097.40  56.1 | 0.7807  0.7677 | 0.6528  0.7965 | 0.6304  0.8105 | 19 hr 22 min  \n    109499,109499, torch.Size([512, 40, 101])\n    0.0100  110.0 k  1102.41  56.3 | 0.7806  0.7674 | 0.6680  0.7891 | 0.6813  0.7871 | 19 hr 24 min  \n    109999,109999, torch.Size([512, 40, 101])\n    0.0100  110.5 k  1107.42  56.6 | 0.7988  0.7671 | 0.6318  0.8027 | 0.6820  0.7949 | 19 hr 26 min  \n    110499,110499, torch.Size([512, 40, 101])\n    0.0100  111.0 k  1112.43  56.8 | 0.7735  0.7680 | 0.6892  0.7825 | 0.6944  0.7852 | 19 hr 29 min  \n    110999,110999, torch.Size([512, 40, 101])\n    0.0100  111.5 k  1117.44  57.1 | 0.7805  0.7661 | 0.6609  0.7952 | 0.6439  0.8066 | 19 hr 31 min  \n    111499,111499, torch.Size([512, 40, 101])\n    0.0100  112.0 k  1122.46  57.3 | 0.7779  0.7633 | 0.6517  0.7975 | 0.8025  0.7656 | 19 hr 34 min  \n    111999,111999, torch.Size([512, 40, 101])\n    0.0100  112.5 k  1127.47  57.6 | 0.7480  0.7652 | 0.6583  0.7934 | 0.6775  0.7871 | 19 hr 38 min  \n    112499,112499, torch.Size([512, 40, 101])\n    0.0100  113.0 k  1132.48  57.9 | 0.7796  0.7574 | 0.6388  0.8022 | 0.5998  0.8105 | 19 hr 42 min  \n    112999,112999, torch.Size([512, 40, 101])\n\nIt would be nice to know if anyone had any success with recurrent networks and reached LB &gt; 0.83?",
    "269328": "I had tried Recurrent Networks and experienced problems exactly like what you have encountered. My first submission, which was an RNN, did exactly 83%. Subsequent RNN networks haven't been able to do better than that. Maybe an 84 would be in the cards. I wouldn't know because I stopped measuring the performance of every single network of mine on the LB at least 50 submissions ago :D :D",
    "269405": "Thanks for the input Sarthak. There's must be some limitations regarding convergence of RNN in this setting. At least it seems that way from our experiences.",
    "269407": "I had a conv-RNN that got 87%, but it didn't end up being helpful even in an ensemble (although I'm pretty new to ensembling, more advanced techniques probably could have gained from it)\n\nThe RNN bit was a two-layer GRU a la the Hello Edge paper, tried a one-layer LSTM too that did great in local validation but got 86% LB",
    "269422": "Hey Thomas thanks for sharing your experience. I find it odd because I've had as you can see from the examples, one of the models was exactly that. A 2 layer GRU but couldn't get pas the 83% LB mark sometimes even worse 82%. If I may ask are you splitting  the data the same way as @Heng? or are you using some other strategy? What models have given you better resutls? pure conv ones or mixed ones?",
    "269423": "I built a LSTM-HMM that got 89% in LB... but I was not able get any improvement with ensembles.",
    "269426": "Can you explain how you used the LSTM-HMM?link to an article explaining your method? \nThanks",
    "269455": "Well, this is a kind of usual architecture in current automatic speech recognition systems. The idea is that the probabilities generated by \"something\" feed a Hidden Markov Model. And normally Viterbi is used to find out a path with the \"best probability\". This \"something\" may be a GMM (as it used to be before 2012) or a DNN. In my experiments, I tried a many combinations of LSTM, CNN and TDNN. In the end, I got 0.89 with each one of these nets, but I was not able to combine them in a better system. I tried the strategy shown [here][1], but without success. \n\nAnd some other references about [acoustic models using DNN][2] and, more specifically, [LSTM in ASR][3]\n\n\n  [1]: https://pdfs.semanticscholar.org/8201/55ecb57325503183253b8796de5f4535eb16.pdf\n  [2]: https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/HintonDengYuEtAl-SPM2012.pdf\n  [3]: https://arxiv.org/pdf/1303.5778.pdf",
    "269580": "Hey I'm using similar method! What do you use for Conv-part? my best single model has a Inception-like head and 2 GRUs attached and got .87 in public .88 in privage LB. no-augmentation or test-prob used.",
    "269959": "Ensembling Heng's resnet, private LB: .85 and the gru model above: LB .67 gives a private LB score of 0.86. Didn't had the time to submit before competition end.",
    "270238": "thanks for your info! .67 is a little bit weird, I found recent paper prefer the CNN-GRU architecture and in practice it do perform well. I'm running out of time to explore more possibility."
  },
  "source": "meta"
}