{
  "id": 229992,
  "title": "Help With TF2 Image Captioning Pipeline",
  "url": "/competitions/bms-molecular-translation/discussion/229992",
  "author_name": "Darien Schettler",
  "post_date": "2021-04-01T15:02:09.450000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi there, I am struggling with my implementation (TF 2+) of an image captioning pipeline. I wanted to reach out for help from the Kaggle community.</p>\n<p>I have gotten it to 'work', but the scores are less than stellar.</p>\n<p><br></p>\n<p>Here is a link to the notebook: <a href=\"https://www.kaggle.com/dschettler8845/bms-image-captioning-w-attention-train/edit/run/58303170\" target=\"_blank\"><strong>NOTEBOOK</strong></a></p>\n<p>Here is a link to the tutorials that inspired the notebook: <a href=\"https://www.tensorflow.org/tutorials/text/image_captioning\" target=\"_blank\"><strong>TF TUTORIAL 1</strong></a>, <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\"><strong>TF TUTORIAL 2</strong></a></p>\n<p>Here is a link to the paper the decoder architecture is based off of: <a href=\"https://arxiv.org/pdf/1502.03044.pdf\" target=\"_blank\"><strong>Show, Attend and Tell</strong></a></p>\n<p><br></p>\n<p>Any feedback would be appreciated… I want to make a simple, easy-to-understand notebook to teach people while learning as I go (similar to <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\"><strong>this notebook</strong></a> by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\"><strong>YNakuma</strong></a> in Pytorch). I believe I should be able to achieve a distance of ~20 or less with the method I am attempting. I was on the fence about posting this but I figured I should reach out for help while in parallel also troubleshooting it on my own. </p>\n<p><br></p>\n<p>I've compiled my old updates below.</p>\n<p><br><br></p>\n<hr>\n<p><strong>UPDATE APRIL 1, 2021</strong> </p>\n<hr>\n<p>I have attempted a revamp using some of my learnings from other notebooks to no avail. The loss decreases and everything looks good, but there is still way too much error present. I will fork <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> wonderful notebook and try and replicate it (form it) to be even more similar to what I am doing. Ideally, this will allow me to identify what I am doing wrong.</p>\n<p>I'll post more updates here as I go</p>\n<hr>\n<p><br></p>\n<p><strong>UPDATE MARCH 29, 2021</strong> </p>\n<hr>\n<p>Seeing the success of other similar image captioning architectures in this competition I'm fairly certain I'm doing something wrong. I will be revamping and re-releasing this notebook. I will comment here when that is complete.</p>\n<hr>\n<p><br></p>\n<p>Thanks in advance!! If there's anything I can do to provide more information or context please let me know. Also, if this is uncouth or violates any of the written (or unwritten) rules of Kaggle, please let me know and I will delete it.</p>",
  "messages": [
    {
      "id": 1259647,
      "postDate": "2021-04-01T15:02:09.450Z",
      "content": "<p>Hi there, I am struggling with my implementation (TF 2+) of an image captioning pipeline. I wanted to reach out for help from the Kaggle community.</p>\n<p>I have gotten it to 'work', but the scores are less than stellar.</p>\n<p><br></p>\n<p>Here is a link to the notebook: <a href=\"https://www.kaggle.com/dschettler8845/bms-image-captioning-w-attention-train/edit/run/58303170\" target=\"_blank\"><strong>NOTEBOOK</strong></a></p>\n<p>Here is a link to the tutorials that inspired the notebook: <a href=\"https://www.tensorflow.org/tutorials/text/image_captioning\" target=\"_blank\"><strong>TF TUTORIAL 1</strong></a>, <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\"><strong>TF TUTORIAL 2</strong></a></p>\n<p>Here is a link to the paper the decoder architecture is based off of: <a href=\"https://arxiv.org/pdf/1502.03044.pdf\" target=\"_blank\"><strong>Show, Attend and Tell</strong></a></p>\n<p><br></p>\n<p>Any feedback would be appreciated… I want to make a simple, easy-to-understand notebook to teach people while learning as I go (similar to <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\"><strong>this notebook</strong></a> by <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\"><strong>YNakuma</strong></a> in Pytorch). I believe I should be able to achieve a distance of ~20 or less with the method I am attempting. I was on the fence about posting this but I figured I should reach out for help while in parallel also troubleshooting it on my own. </p>\n<p><br></p>\n<p>I've compiled my old updates below.</p>\n<p><br><br></p>\n<hr>\n<p><strong>UPDATE APRIL 1, 2021</strong> </p>\n<hr>\n<p>I have attempted a revamp using some of my learnings from other notebooks to no avail. The loss decreases and everything looks good, but there is still way too much error present. I will fork <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> wonderful notebook and try and replicate it (form it) to be even more similar to what I am doing. Ideally, this will allow me to identify what I am doing wrong.</p>\n<p>I'll post more updates here as I go</p>\n<hr>\n<p><br></p>\n<p><strong>UPDATE MARCH 29, 2021</strong> </p>\n<hr>\n<p>Seeing the success of other similar image captioning architectures in this competition I'm fairly certain I'm doing something wrong. I will be revamping and re-releasing this notebook. I will comment here when that is complete.</p>\n<hr>\n<p><br></p>\n<p>Thanks in advance!! If there's anything I can do to provide more information or context please let me know. Also, if this is uncouth or violates any of the written (or unwritten) rules of Kaggle, please let me know and I will delete it.</p>",
      "rawMarkdown": "Hi there, I am struggling with my implementation (TF 2+) of an image captioning pipeline. I wanted to reach out for help from the Kaggle community.\n\nI have gotten it to 'work', but the scores are less than stellar.\n\n<br>\n\nHere is a link to the notebook: [**NOTEBOOK**](https://www.kaggle.com/dschettler8845/bms-image-captioning-w-attention-train/edit/run/58303170)\n\nHere is a link to the tutorials that inspired the notebook: [**TF TUTORIAL 1**](https://www.tensorflow.org/tutorials/text/image_captioning), [**TF TUTORIAL 2**](https://www.tensorflow.org/tutorials/text/transformer)\n\nHere is a link to the paper the decoder architecture is based off of: [**Show, Attend and Tell**](https://arxiv.org/pdf/1502.03044.pdf)\n\n<br>\n\nAny feedback would be appreciated... I want to make a simple, easy-to-understand notebook to teach people while learning as I go (similar to [**this notebook**](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter) by [**YNakuma**](https://www.kaggle.com/yasufuminakama) in Pytorch). I believe I should be able to achieve a distance of ~20 or less with the method I am attempting. I was on the fence about posting this but I figured I should reach out for help while in parallel also troubleshooting it on my own. \n\n<br>\n\nI've compiled my old updates below.\n\n<br><br>\n\n---\n\n**UPDATE APRIL 1, 2021** \n\n---\n\nI have attempted a revamp using some of my learnings from other notebooks to no avail. The loss decreases and everything looks good, but there is still way too much error present. I will fork @markwijkhuizen wonderful notebook and try and replicate it (form it) to be even more similar to what I am doing. Ideally, this will allow me to identify what I am doing wrong.\n\nI'll post more updates here as I go\n\n---\n\n<br>\n\n**UPDATE MARCH 29, 2021** \n\n---\n\nSeeing the success of other similar image captioning architectures in this competition I'm fairly certain I'm doing something wrong. I will be revamping and re-releasing this notebook. I will comment here when that is complete.\n\n---\n\n<br>\n\nThanks in advance!! If there's anything I can do to provide more information or context please let me know. Also, if this is uncouth or violates any of the written (or unwritten) rules of Kaggle, please let me know and I will delete it.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1259647": "Hi there, I am struggling with my implementation (TF 2+) of an image captioning pipeline. I wanted to reach out for help from the Kaggle community.\n\nI have gotten it to 'work', but the scores are less than stellar.\n\n<br>\n\nHere is a link to the notebook: [**NOTEBOOK**](https://www.kaggle.com/dschettler8845/bms-image-captioning-w-attention-train/edit/run/58303170)\n\nHere is a link to the tutorials that inspired the notebook: [**TF TUTORIAL 1**](https://www.tensorflow.org/tutorials/text/image_captioning), [**TF TUTORIAL 2**](https://www.tensorflow.org/tutorials/text/transformer)\n\nHere is a link to the paper the decoder architecture is based off of: [**Show, Attend and Tell**](https://arxiv.org/pdf/1502.03044.pdf)\n\n<br>\n\nAny feedback would be appreciated... I want to make a simple, easy-to-understand notebook to teach people while learning as I go (similar to [**this notebook**](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter) by [**YNakuma**](https://www.kaggle.com/yasufuminakama) in Pytorch). I believe I should be able to achieve a distance of ~20 or less with the method I am attempting. I was on the fence about posting this but I figured I should reach out for help while in parallel also troubleshooting it on my own. \n\n<br>\n\nI've compiled my old updates below.\n\n<br><br>\n\n---\n\n**UPDATE APRIL 1, 2021** \n\n---\n\nI have attempted a revamp using some of my learnings from other notebooks to no avail. The loss decreases and everything looks good, but there is still way too much error present. I will fork @markwijkhuizen wonderful notebook and try and replicate it (form it) to be even more similar to what I am doing. Ideally, this will allow me to identify what I am doing wrong.\n\nI'll post more updates here as I go\n\n---\n\n<br>\n\n**UPDATE MARCH 29, 2021** \n\n---\n\nSeeing the success of other similar image captioning architectures in this competition I'm fairly certain I'm doing something wrong. I will be revamping and re-releasing this notebook. I will comment here when that is complete.\n\n---\n\n<br>\n\nThanks in advance!! If there's anything I can do to provide more information or context please let me know. Also, if this is uncouth or violates any of the written (or unwritten) rules of Kaggle, please let me know and I will delete it."
  }
}