{
  "id": 47746,
  "title": "ConvLSTM",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47746",
  "author_name": "",
  "post_date": "2018-01-18T10:34:03.699849300Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello and congratulations to everyone for very demanding competition. But.... :D</p>\n\n<p>Does anyone try to train a ConvLSTM approach?</p>\n\n<p>In ConvLSTM we substitute all FC parts of LSTM with Convolution 2D. It's a useful model for processing video data or similar.</p>\n\n<p>Paper \"Very Deep Convolutional Networks for End-to-End Speech Recognition\":\n<a href=\"https://arxiv.org/abs/1610.03022v1\">https://arxiv.org/abs/1610.03022v1</a></p>\n\n<p>Here is my implementation:  <a href=\"https://gist.github.com/Raalsky/270beace32d7f2d37a95334a657a89b7\">https://gist.github.com/Raalsky/270beace32d7f2d37a95334a657a89b7</a>\n(1DConv Model -&gt; Time splitting into \"animation\" -&gt; 3 x ConvLSTM + BN -&gt; TimeDistributed Global Average (:D) -&gt; 3 * LSTM + BN)</p>\n\n<p>Raw training without transfer learning took a huge amount of time. After 100 epochs (~10h on Tesla P100)  I killed training with results aprox. 85% both on training and validation without submitting.</p>\n\n<p>Any advise or results in training such different approach?</p>",
  "messages": [
    {
      "id": "270489",
      "postDate": "01/18/2018 10:34:03",
      "content": "<p>Hello and congratulations to everyone for very demanding competition. But.... :D</p>\n\n<p>Does anyone try to train a ConvLSTM approach?</p>\n\n<p>In ConvLSTM we substitute all FC parts of LSTM with Convolution 2D. It's a useful model for processing video data or similar.</p>\n\n<p>Paper \"Very Deep Convolutional Networks for End-to-End Speech Recognition\":\n<a href=\"https://arxiv.org/abs/1610.03022v1\">https://arxiv.org/abs/1610.03022v1</a></p>\n\n<p>Here is my implementation:  <a href=\"https://gist.github.com/Raalsky/270beace32d7f2d37a95334a657a89b7\">https://gist.github.com/Raalsky/270beace32d7f2d37a95334a657a89b7</a>\n(1DConv Model -&gt; Time splitting into \"animation\" -&gt; 3 x ConvLSTM + BN -&gt; TimeDistributed Global Average (:D) -&gt; 3 * LSTM + BN)</p>\n\n<p>Raw training without transfer learning took a huge amount of time. After 100 epochs (~10h on Tesla P100)  I killed training with results aprox. 85% both on training and validation without submitting.</p>\n\n<p>Any advise or results in training such different approach?</p>",
      "rawMarkdown": "Hello and congratulations to everyone for very demanding competition. But.... :D\n\nDoes anyone try to train a ConvLSTM approach?\n\nIn ConvLSTM we substitute all FC parts of LSTM with Convolution 2D. It's a useful model for processing video data or similar.\n\nPaper \"Very Deep Convolutional Networks for End-to-End Speech Recognition\":\nhttps://arxiv.org/abs/1610.03022v1\n\nHere is my implementation:  https://gist.github.com/Raalsky/270beace32d7f2d37a95334a657a89b7\n(1DConv Model -&gt; Time splitting into \"animation\" -&gt; 3 x ConvLSTM + BN -&gt; TimeDistributed Global Average (:D) -&gt; 3 * LSTM + BN)\n\nRaw training without transfer learning took a huge amount of time. After 100 epochs (~10h on Tesla P100)  I killed training with results aprox. 85% both on training and validation without submitting.\n\nAny advise or results in training such different approach?",
      "votes": null
    },
    {
      "id": "270536",
      "postDate": "01/18/2018 13:02:48",
      "content": "<p>I tried the same paper, validation score was 97% however leaderboard scored only 82%. The implementation in available on  <a href=\"https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution\">https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution</a></p>",
      "rawMarkdown": "I tried the same paper, validation score was 97% however leaderboard scored only 82%. The implementation in available on  https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution",
      "votes": null
    },
    {
      "id": "985436",
      "postDate": "08/25/2020 18:11:56",
      "content": "<p>To train a ConvLSTM, you can try a similar approach to what was found in this paper by Xingjian Shi in 2015: <a href=\"url\" target=\"_blank\">https://papers.nips.cc/paper/5955-convolutional-lstm-network-a-machine-learning-approach-for-precipitation-nowcasting.pdf</a>. Set the label for a training sample as the image at the next timestep and then shift over each sample by one frame to create multiple training samples with labels. Then, the model will learn that the label for an image at a current timestep is the next timestep's image, thus preserving the temporal-correlation while also utilizing spatial-correlation from the convolution in the ConvLSTM.</p>",
      "rawMarkdown": "To train a ConvLSTM, you can try a similar approach to what was found in this paper by Xingjian Shi in 2015: [https://papers.nips.cc/paper/5955-convolutional-lstm-network-a-machine-learning-approach-for-precipitation-nowcasting.pdf](url). Set the label for a training sample as the image at the next timestep and then shift over each sample by one frame to create multiple training samples with labels. Then, the model will learn that the label for an image at a current timestep is the next timestep's image, thus preserving the temporal-correlation while also utilizing spatial-correlation from the convolution in the ConvLSTM.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 985436,
      "author_name": "pannumuthu",
      "author_url": "",
      "post_date": "08/25/2020 18:11:56",
      "content": "<p>To train a ConvLSTM, you can try a similar approach to what was found in this paper by Xingjian Shi in 2015: <a href=\"url\" target=\"_blank\">https://papers.nips.cc/paper/5955-convolutional-lstm-network-a-machine-learning-approach-for-precipitation-nowcasting.pdf</a>. Set the label for a training sample as the image at the next timestep and then shift over each sample by one frame to create multiple training samples with labels. Then, the model will learn that the label for an image at a current timestep is the next timestep's image, thus preserving the temporal-correlation while also utilizing spatial-correlation from the convolution in the ConvLSTM.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270536,
      "author_name": "subho406",
      "author_url": "",
      "post_date": "01/18/2018 13:02:48",
      "content": "<p>I tried the same paper, validation score was 97% however leaderboard scored only 82%. The implementation in available on  <a href=\"https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution\">https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "270489": "Hello and congratulations to everyone for very demanding competition. But.... :D\n\nDoes anyone try to train a ConvLSTM approach?\n\nIn ConvLSTM we substitute all FC parts of LSTM with Convolution 2D. It's a useful model for processing video data or similar.\n\nPaper \"Very Deep Convolutional Networks for End-to-End Speech Recognition\":\nhttps://arxiv.org/abs/1610.03022v1\n\nHere is my implementation:  https://gist.github.com/Raalsky/270beace32d7f2d37a95334a657a89b7\n(1DConv Model -&gt; Time splitting into \"animation\" -&gt; 3 x ConvLSTM + BN -&gt; TimeDistributed Global Average (:D) -&gt; 3 * LSTM + BN)\n\nRaw training without transfer learning took a huge amount of time. After 100 epochs (~10h on Tesla P100)  I killed training with results aprox. 85% both on training and validation without submitting.\n\nAny advise or results in training such different approach?",
    "270536": "I tried the same paper, validation score was 97% however leaderboard scored only 82%. The implementation in available on  https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution",
    "985436": "To train a ConvLSTM, you can try a similar approach to what was found in this paper by Xingjian Shi in 2015: [https://papers.nips.cc/paper/5955-convolutional-lstm-network-a-machine-learning-approach-for-precipitation-nowcasting.pdf](url). Set the label for a training sample as the image at the next timestep and then shift over each sample by one frame to create multiple training samples with labels. Then, the model will learn that the label for an image at a current timestep is the next timestep's image, thus preserving the temporal-correlation while also utilizing spatial-correlation from the convolution in the ConvLSTM."
  },
  "source": "meta"
}