{
  "id": 227244,
  "title": "Suitable (pretrained) CNNs for image captioning of molecules?",
  "url": "/competitions/bms-molecular-translation/discussion/227244",
  "author_name": "",
  "post_date": "2021-03-19T13:44:17.686769400Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi there!</p>\n<p>This is my first kaggle competition :)</p>\n<p>I'm working on a kernel following the <a href=\"https://www.tensorflow.org/tutorials/text/image_captioning\" target=\"_blank\">Tensorflow Tutorial on Image Captioning</a> and wondered what CNN would be suitable for recognizing the molecule structures. </p>\n<p>I've started with the pretrained InceptionV3 CNN from the tutorial predicting just the chemical formula. In a first try the results were not too bad.</p>\n<p>However in above tutorial and I guess in other typical applications of InceptionV3 one deals with photographic images, not the seemingly much simpler black and white structural images we have here.</p>\n<p>I guess there are pretrained CNNs more suitable for this task which might even be simpler and faster to train later on.  Do you have any hints / experiences you would like to share?</p>\n<p>Any insights are very much appreciated :)</p>",
  "messages": [
    {
      "id": "1245140",
      "postDate": "03/19/2021 13:44:17",
      "content": "<p>Hi there!</p>\n<p>This is my first kaggle competition :)</p>\n<p>I'm working on a kernel following the <a href=\"https://www.tensorflow.org/tutorials/text/image_captioning\" target=\"_blank\">Tensorflow Tutorial on Image Captioning</a> and wondered what CNN would be suitable for recognizing the molecule structures. </p>\n<p>I've started with the pretrained InceptionV3 CNN from the tutorial predicting just the chemical formula. In a first try the results were not too bad.</p>\n<p>However in above tutorial and I guess in other typical applications of InceptionV3 one deals with photographic images, not the seemingly much simpler black and white structural images we have here.</p>\n<p>I guess there are pretrained CNNs more suitable for this task which might even be simpler and faster to train later on.  Do you have any hints / experiences you would like to share?</p>\n<p>Any insights are very much appreciated :)</p>",
      "rawMarkdown": "Hi there!\n\nThis is my first kaggle competition :)\n\nI'm working on a kernel following the [Tensorflow Tutorial on Image Captioning](https://www.tensorflow.org/tutorials/text/image_captioning) and wondered what CNN would be suitable for recognizing the molecule structures. \n\nI've started with the pretrained InceptionV3 CNN from the tutorial predicting just the chemical formula. In a first try the results were not too bad.\n\nHowever in above tutorial and I guess in other typical applications of InceptionV3 one deals with photographic images, not the seemingly much simpler black and white structural images we have here.\n\nI guess there are pretrained CNNs more suitable for this task which might even be simpler and faster to train later on.  Do you have any hints / experiences you would like to share?\n\nAny insights are very much appreciated :)",
      "votes": null
    },
    {
      "id": "1245305",
      "postDate": "03/19/2021 16:56:03",
      "content": "<p>Good question fundamentally. I think the initial layers of pre-trained Imagenet weights should be useful, not too sure about the later ones. Folks here have probably tried CNN weights of SMILES-based algorithms that already exist, so you can check them out. But what I feel is there might be no better way than fine-tuning the models ourselves 😅</p>",
      "rawMarkdown": "Good question fundamentally. I think the initial layers of pre-trained Imagenet weights should be useful, not too sure about the later ones. Folks here have probably tried CNN weights of SMILES-based algorithms that already exist, so you can check them out. But what I feel is there might be no better way than fine-tuning the models ourselves 😅",
      "votes": null
    },
    {
      "id": "1245867",
      "postDate": "03/20/2021 09:35:19",
      "content": "<p>I have experimented with the EfficientNetB0 model and found training from scratch gave better results than using ImageNet weights using the same training method. This is probably due to the fact that ImageNet does not contain any images related to the molecule structure images in the train dataset. It should also be noted ImageNet is trained on RGB images and the molecule images are grayscale.<br>\nIn short, it looks like the ImageNet images have too little in common with the molecule images.</p>\n<table>\n<thead>\n<tr>\n<th>Weights</th>\n<th>Training Loss</th>\n<th>Validation Loss</th>\n<th>Validation Levenshtein distance</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Random Initialized</td>\n<td>0.133</td>\n<td>4.304</td>\n<td>41.4</td>\n</tr>\n<tr>\n<td>Noisy Student</td>\n<td>0.197</td>\n<td>5.387</td>\n<td>58.2</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I have experimented with the EfficientNetB0 model and found training from scratch gave better results than using ImageNet weights using the same training method. This is probably due to the fact that ImageNet does not contain any images related to the molecule structure images in the train dataset. It should also be noted ImageNet is trained on RGB images and the molecule images are grayscale.\nIn short, it looks like the ImageNet images have too little in common with the molecule images.\n\n\n| Weights | Training Loss | Validation Loss | Validation Levenshtein distance |\n| -- | -- | -- | -- |\n| Random Initialized | 0.133 | 4.304 | 41.4 |\n| Noisy Student | 0.197 | 5.387 | 58.2 |",
      "votes": null
    },
    {
      "id": "1245902",
      "postDate": "03/20/2021 10:19:01",
      "content": "<p>Thanks for your help. I think I'll move away from using a pretrained model in the next try.</p>",
      "rawMarkdown": "Thanks for your help. I think I'll move away from using a pretrained model in the next try.",
      "votes": null
    },
    {
      "id": "1245905",
      "postDate": "03/20/2021 10:20:55",
      "content": "<p>Thanks for the information! I'll give EfficientNetB0 a try.</p>",
      "rawMarkdown": "Thanks for the information! I'll give EfficientNetB0 a try.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1245305,
      "author_name": "arka47",
      "author_url": "",
      "post_date": "03/19/2021 16:56:03",
      "content": "<p>Good question fundamentally. I think the initial layers of pre-trained Imagenet weights should be useful, not too sure about the later ones. Folks here have probably tried CNN weights of SMILES-based algorithms that already exist, so you can check them out. But what I feel is there might be no better way than fine-tuning the models ourselves 😅</p>",
      "votes": null,
      "replies": [
        {
          "id": 1245902,
          "author_name": "michaelwolff",
          "author_url": "",
          "post_date": "03/20/2021 10:19:01",
          "content": "<p>Thanks for your help. I think I'll move away from using a pretrained model in the next try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1245867,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "03/20/2021 09:35:19",
      "content": "<p>I have experimented with the EfficientNetB0 model and found training from scratch gave better results than using ImageNet weights using the same training method. This is probably due to the fact that ImageNet does not contain any images related to the molecule structure images in the train dataset. It should also be noted ImageNet is trained on RGB images and the molecule images are grayscale.<br>\nIn short, it looks like the ImageNet images have too little in common with the molecule images.</p>\n<table>\n<thead>\n<tr>\n<th>Weights</th>\n<th>Training Loss</th>\n<th>Validation Loss</th>\n<th>Validation Levenshtein distance</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Random Initialized</td>\n<td>0.133</td>\n<td>4.304</td>\n<td>41.4</td>\n</tr>\n<tr>\n<td>Noisy Student</td>\n<td>0.197</td>\n<td>5.387</td>\n<td>58.2</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 1245905,
          "author_name": "michaelwolff",
          "author_url": "",
          "post_date": "03/20/2021 10:20:55",
          "content": "<p>Thanks for the information! I'll give EfficientNetB0 a try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1245140": "Hi there!\n\nThis is my first kaggle competition :)\n\nI'm working on a kernel following the [Tensorflow Tutorial on Image Captioning](https://www.tensorflow.org/tutorials/text/image_captioning) and wondered what CNN would be suitable for recognizing the molecule structures. \n\nI've started with the pretrained InceptionV3 CNN from the tutorial predicting just the chemical formula. In a first try the results were not too bad.\n\nHowever in above tutorial and I guess in other typical applications of InceptionV3 one deals with photographic images, not the seemingly much simpler black and white structural images we have here.\n\nI guess there are pretrained CNNs more suitable for this task which might even be simpler and faster to train later on.  Do you have any hints / experiences you would like to share?\n\nAny insights are very much appreciated :)",
    "1245305": "Good question fundamentally. I think the initial layers of pre-trained Imagenet weights should be useful, not too sure about the later ones. Folks here have probably tried CNN weights of SMILES-based algorithms that already exist, so you can check them out. But what I feel is there might be no better way than fine-tuning the models ourselves 😅",
    "1245867": "I have experimented with the EfficientNetB0 model and found training from scratch gave better results than using ImageNet weights using the same training method. This is probably due to the fact that ImageNet does not contain any images related to the molecule structure images in the train dataset. It should also be noted ImageNet is trained on RGB images and the molecule images are grayscale.\nIn short, it looks like the ImageNet images have too little in common with the molecule images.\n\n\n| Weights | Training Loss | Validation Loss | Validation Levenshtein distance |\n| -- | -- | -- | -- |\n| Random Initialized | 0.133 | 4.304 | 41.4 |\n| Noisy Student | 0.197 | 5.387 | 58.2 |",
    "1245902": "Thanks for your help. I think I'll move away from using a pretrained model in the next try.",
    "1245905": "Thanks for the information! I'll give EfficientNetB0 a try."
  },
  "source": "meta"
}