{
  "id": 47085,
  "title": "Attention Models",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47085",
  "author_name": "",
  "post_date": "2018-01-08T11:10:54.434235500Z",
  "votes": 3,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Has anyone tried any attention models with an RNN for example?\nI do not have enough time to implement a model and would be interesting if one could share some insights regarding that.</p>",
  "messages": [
    {
      "id": "266283",
      "postDate": "01/08/2018 11:10:54",
      "content": "<p>Has anyone tried any attention models with an RNN for example?\nI do not have enough time to implement a model and would be interesting if one could share some insights regarding that.</p>",
      "rawMarkdown": "Has anyone tried any attention models with an RNN for example?\nI do not have enough time to implement a model and would be interesting if one could share some insights regarding that.",
      "votes": null
    },
    {
      "id": "266285",
      "postDate": "01/08/2018 11:24:05",
      "content": "<p>you have any link to paper that use RNN attention models for keyword spotting?</p>",
      "rawMarkdown": "you have any link to paper that use RNN attention models for keyword spotting?",
      "votes": null
    },
    {
      "id": "266291",
      "postDate": "01/08/2018 11:41:02",
      "content": "<p>These are the two papers that I am currently reading\n<a href=\"https://arxiv.org/pdf/1601.06823.pdf\">https://arxiv.org/pdf/1601.06823.pdf</a>\n<a href=\"https://arxiv.org/pdf/1506.07503.pdf\">https://arxiv.org/pdf/1506.07503.pdf</a></p>",
      "rawMarkdown": "These are the two papers that I am currently reading\nhttps://arxiv.org/pdf/1601.06823.pdf\nhttps://arxiv.org/pdf/1506.07503.pdf",
      "votes": null
    },
    {
      "id": "266294",
      "postDate": "01/08/2018 11:46:06",
      "content": "<p>There is also a relatively simple implementation using Keras\n<a href=\"https://github.com/philipperemy/keras-attention-mechanism\">https://github.com/philipperemy/keras-attention-mechanism</a></p>\n\n<p>I am just curious if they can reach a 0.90LB assuming you have a model greater than 0.86LB</p>",
      "rawMarkdown": "There is also a relatively simple implementation using Keras\nhttps://github.com/philipperemy/keras-attention-mechanism\n\nI am just curious if they can reach a 0.90LB assuming you have a model greater than 0.86LB",
      "votes": null
    },
    {
      "id": "266297",
      "postDate": "01/08/2018 11:52:21",
      "content": "<p>Are you familiar with attention model? I can put a CNN model on spectrogram with LB=0.86 here for you to try. But it is a pytorch model. (You probably can import it to tensorflow using onnx).</p>\n\n<p>Let me know if you are interested to try.</p>",
      "rawMarkdown": "Are you familiar with attention model? I can put a CNN model on spectrogram with LB=0.86 here for you to try. But it is a pytorch model. (You probably can import it to tensorflow using onnx).\n\nLet me know if you are interested to try.",
      "votes": null
    },
    {
      "id": "266300",
      "postDate": "01/08/2018 12:00:54",
      "content": "<p>I am familiar with the attention model in general and I would like to try. However, it would not be fair for you if I run your 0.86LB code. I am still at 0.84.</p>",
      "rawMarkdown": "I am familiar with the attention model in general and I would like to try. However, it would not be fair for you if I run your 0.86LB code. I am still at 0.84.",
      "votes": null
    },
    {
      "id": "266302",
      "postDate": "01/08/2018 12:16:45",
      "content": "<p>No problem. I am also interested to know the results of RNN vs attentive model too.\nI am now using a \"similar\" approach in between, i called it \"selective gating\" </p>\n\n<p>I intent to upload the trained CNN model and inference/train code. You can use it as feature extraction for your attention model, i.e. with the attention blocks added on top the CNN model. Then you can share some initial results here. (You can keep your final results  with fine adjustment, hacks, etc)</p>\n\n<p>The model and results will be shared, so it is fair to everyone. It is also a faster way to try different architecture for everyone.</p>\n\n<p>Please wait for a while when i prepare the code and model</p>",
      "rawMarkdown": "No problem. I am also interested to know the results of RNN vs attentive model too.\nI am now using a \"similar\" approach in between, i called it \"selective gating\" \n\nI intent to upload the trained CNN model and inference/train code. You can use it as feature extraction for your attention model, i.e. with the attention blocks added on top the CNN model. Then you can share some initial results here. (You can keep your final results  with fine adjustment, hacks, etc)\n\nThe model and results will be shared, so it is fair to everyone. It is also a faster way to try different architecture for everyone.\n\nPlease wait for a while when i prepare the code and model",
      "votes": null
    },
    {
      "id": "266303",
      "postDate": "01/08/2018 12:18:03",
      "content": "<p>Thanks for your help. I appreciate it!</p>",
      "rawMarkdown": "Thanks for your help. I appreciate it!",
      "votes": null
    },
    {
      "id": "266364",
      "postDate": "01/08/2018 16:03:22",
      "content": "<p>@Tasos Vafeiadis</p>\n\n<p>Please refer to this post: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46988\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46988</a></p>\n\n<p>look for the below message:</p>\n\n<p>\"please refer to ppt for details.</p>\n\n<p>here, the files contain:</p>\n\n<ul>\n<li><p>full pycharm project, include train, evaluate, submit code</p></li>\n<li><p>trained model at  LB=0.86</p></li>\n<li><p>data split</p>\n\n<p>build.tar.gz (82.37 KB)</p>\n\n<p>data.tar.gz (2.27 MB)</p>\n\n<p>results.tar.gz (4.45 MB)</p>\n\n<p>readme.pptx (524.8 KB)\"</p></li>\n</ul>",
      "rawMarkdown": "Tasos Vafeiadis\n\nPlease refer to this post: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46988\n\nlook for the below message:\n\n\"please refer to ppt for details.\n\nhere, the files contain:\n\n - full pycharm project, include train, evaluate, submit code\n\n - trained model at  LB=0.86\n\n - data split\n\n build.tar.gz (82.37 KB)\n\n data.tar.gz (2.27 MB)\n\n results.tar.gz (4.45 MB)\n\n readme.pptx (524.8 KB)\"",
      "votes": null
    },
    {
      "id": "266629",
      "postDate": "01/09/2018 08:07:54",
      "content": "<p>I will let you know as soon as I get the first result. Thank you again for the help!</p>",
      "rawMarkdown": "I will let you know as soon as I get the first result. Thank you again for the help!",
      "votes": null
    },
    {
      "id": "267280",
      "postDate": "01/11/2018 00:09:09",
      "content": "<p>Hi, how is your result testing attention models?\nmy experience adding attention layer either before or after recurrent layers not only make converges harder but gives no improvement (most of the time degradation in LB), replace recurrent layers with attention also don't help.</p>",
      "rawMarkdown": "Hi, how is your result testing attention models?\nmy experience adding attention layer either before or after recurrent layers not only make converges harder but gives no improvement (most of the time degradation in LB), replace recurrent layers with attention also don't help.",
      "votes": null
    },
    {
      "id": "267458",
      "postDate": "01/11/2018 12:21:49",
      "content": "<p>I am at 0.83-0.84 with attention. Without it at 0.86LB.</p>",
      "rawMarkdown": "I am at 0.83-0.84 with attention. Without it at 0.86LB.",
      "votes": null
    },
    {
      "id": "267490",
      "postDate": "01/11/2018 14:32:49",
      "content": "<p>thanks for the results!</p>",
      "rawMarkdown": "thanks for the results!",
      "votes": null
    },
    {
      "id": "268662",
      "postDate": "01/15/2018 06:56:16",
      "content": "<p>My degradation is not that huge, most of the times drop from .87 to .86 for a single model. So still I use the model with attention for ensembling.</p>",
      "rawMarkdown": "My degradation is not that huge, most of the times drop from .87 to .86 for a single model. So still I use the model with attention for ensembling.",
      "votes": null
    },
    {
      "id": "268684",
      "postDate": "01/15/2018 07:47:39",
      "content": "<p>Can anyone share the attention model code? I'm not familiar with the idea and it sound interesting.</p>",
      "rawMarkdown": "Can anyone share the attention model code? I'm not familiar with the idea and it sound interesting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 266285,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/08/2018 11:24:05",
      "content": "<p>you have any link to paper that use RNN attention models for keyword spotting?</p>",
      "votes": null,
      "replies": [
        {
          "id": 266291,
          "author_name": "anasvaf",
          "author_url": "",
          "post_date": "01/08/2018 11:41:02",
          "content": "<p>These are the two papers that I am currently reading\n<a href=\"https://arxiv.org/pdf/1601.06823.pdf\">https://arxiv.org/pdf/1601.06823.pdf</a>\n<a href=\"https://arxiv.org/pdf/1506.07503.pdf\">https://arxiv.org/pdf/1506.07503.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266294,
          "author_name": "anasvaf",
          "author_url": "",
          "post_date": "01/08/2018 11:46:06",
          "content": "<p>There is also a relatively simple implementation using Keras\n<a href=\"https://github.com/philipperemy/keras-attention-mechanism\">https://github.com/philipperemy/keras-attention-mechanism</a></p>\n\n<p>I am just curious if they can reach a 0.90LB assuming you have a model greater than 0.86LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266297,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/08/2018 11:52:21",
          "content": "<p>Are you familiar with attention model? I can put a CNN model on spectrogram with LB=0.86 here for you to try. But it is a pytorch model. (You probably can import it to tensorflow using onnx).</p>\n\n<p>Let me know if you are interested to try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266300,
          "author_name": "anasvaf",
          "author_url": "",
          "post_date": "01/08/2018 12:00:54",
          "content": "<p>I am familiar with the attention model in general and I would like to try. However, it would not be fair for you if I run your 0.86LB code. I am still at 0.84.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266302,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/08/2018 12:16:45",
          "content": "<p>No problem. I am also interested to know the results of RNN vs attentive model too.\nI am now using a \"similar\" approach in between, i called it \"selective gating\" </p>\n\n<p>I intent to upload the trained CNN model and inference/train code. You can use it as feature extraction for your attention model, i.e. with the attention blocks added on top the CNN model. Then you can share some initial results here. (You can keep your final results  with fine adjustment, hacks, etc)</p>\n\n<p>The model and results will be shared, so it is fair to everyone. It is also a faster way to try different architecture for everyone.</p>\n\n<p>Please wait for a while when i prepare the code and model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266303,
          "author_name": "anasvaf",
          "author_url": "",
          "post_date": "01/08/2018 12:18:03",
          "content": "<p>Thanks for your help. I appreciate it!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266364,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/08/2018 16:03:22",
          "content": "<p>@Tasos Vafeiadis</p>\n\n<p>Please refer to this post: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46988\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46988</a></p>\n\n<p>look for the below message:</p>\n\n<p>\"please refer to ppt for details.</p>\n\n<p>here, the files contain:</p>\n\n<ul>\n<li><p>full pycharm project, include train, evaluate, submit code</p></li>\n<li><p>trained model at  LB=0.86</p></li>\n<li><p>data split</p>\n\n<p>build.tar.gz (82.37 KB)</p>\n\n<p>data.tar.gz (2.27 MB)</p>\n\n<p>results.tar.gz (4.45 MB)</p>\n\n<p>readme.pptx (524.8 KB)\"</p></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266629,
          "author_name": "anasvaf",
          "author_url": "",
          "post_date": "01/09/2018 08:07:54",
          "content": "<p>I will let you know as soon as I get the first result. Thank you again for the help!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267280,
      "author_name": "philipxue",
      "author_url": "",
      "post_date": "01/11/2018 00:09:09",
      "content": "<p>Hi, how is your result testing attention models?\nmy experience adding attention layer either before or after recurrent layers not only make converges harder but gives no improvement (most of the time degradation in LB), replace recurrent layers with attention also don't help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 267458,
          "author_name": "anasvaf",
          "author_url": "",
          "post_date": "01/11/2018 12:21:49",
          "content": "<p>I am at 0.83-0.84 with attention. Without it at 0.86LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267490,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/11/2018 14:32:49",
          "content": "<p>thanks for the results!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268662,
          "author_name": "philipxue",
          "author_url": "",
          "post_date": "01/15/2018 06:56:16",
          "content": "<p>My degradation is not that huge, most of the times drop from .87 to .86 for a single model. So still I use the model with attention for ensembling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268684,
          "author_name": "ori226",
          "author_url": "",
          "post_date": "01/15/2018 07:47:39",
          "content": "<p>Can anyone share the attention model code? I'm not familiar with the idea and it sound interesting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "266283": "Has anyone tried any attention models with an RNN for example?\nI do not have enough time to implement a model and would be interesting if one could share some insights regarding that.",
    "266285": "you have any link to paper that use RNN attention models for keyword spotting?",
    "266291": "These are the two papers that I am currently reading\nhttps://arxiv.org/pdf/1601.06823.pdf\nhttps://arxiv.org/pdf/1506.07503.pdf",
    "266294": "There is also a relatively simple implementation using Keras\nhttps://github.com/philipperemy/keras-attention-mechanism\n\nI am just curious if they can reach a 0.90LB assuming you have a model greater than 0.86LB",
    "266297": "Are you familiar with attention model? I can put a CNN model on spectrogram with LB=0.86 here for you to try. But it is a pytorch model. (You probably can import it to tensorflow using onnx).\n\nLet me know if you are interested to try.",
    "266300": "I am familiar with the attention model in general and I would like to try. However, it would not be fair for you if I run your 0.86LB code. I am still at 0.84.",
    "266302": "No problem. I am also interested to know the results of RNN vs attentive model too.\nI am now using a \"similar\" approach in between, i called it \"selective gating\" \n\nI intent to upload the trained CNN model and inference/train code. You can use it as feature extraction for your attention model, i.e. with the attention blocks added on top the CNN model. Then you can share some initial results here. (You can keep your final results  with fine adjustment, hacks, etc)\n\nThe model and results will be shared, so it is fair to everyone. It is also a faster way to try different architecture for everyone.\n\nPlease wait for a while when i prepare the code and model",
    "266303": "Thanks for your help. I appreciate it!",
    "266364": "Tasos Vafeiadis\n\nPlease refer to this post: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46988\n\nlook for the below message:\n\n\"please refer to ppt for details.\n\nhere, the files contain:\n\n - full pycharm project, include train, evaluate, submit code\n\n - trained model at  LB=0.86\n\n - data split\n\n build.tar.gz (82.37 KB)\n\n data.tar.gz (2.27 MB)\n\n results.tar.gz (4.45 MB)\n\n readme.pptx (524.8 KB)\"",
    "266629": "I will let you know as soon as I get the first result. Thank you again for the help!",
    "267280": "Hi, how is your result testing attention models?\nmy experience adding attention layer either before or after recurrent layers not only make converges harder but gives no improvement (most of the time degradation in LB), replace recurrent layers with attention also don't help.",
    "267458": "I am at 0.83-0.84 with attention. Without it at 0.86LB.",
    "267490": "thanks for the results!",
    "268662": "My degradation is not that huge, most of the times drop from .87 to .86 for a single model. So still I use the model with attention for ensembling.",
    "268684": "Can anyone share the attention model code? I'm not familiar with the idea and it sound interesting."
  },
  "source": "meta"
}