{
  "id": 233387,
  "title": "there is something fishy about image size",
  "url": "/competitions/bms-molecular-translation/discussion/233387",
  "author_name": "",
  "post_date": "2021-04-19T01:04:55.135257200Z",
  "votes": 18,
  "comment_count": 25,
  "views": 0,
  "content": "<p>i was trying to analyze the relationship for  prediction verus sequence length and image size.</p>\n<p>i discover something fishy ….</p>\n<h1>🐟</h1>\n<p><img src=\"https://i.ibb.co/KmD952w/Selection-041.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/qdg981H/Selection-038.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1277591",
      "postDate": "04/19/2021 01:04:55",
      "content": "<p>i was trying to analyze the relationship for  prediction verus sequence length and image size.</p>\n<p>i discover something fishy ….</p>\n<h1>🐟</h1>\n<p><img src=\"https://i.ibb.co/KmD952w/Selection-041.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/qdg981H/Selection-038.png\" alt=\"\"></p>",
      "rawMarkdown": "i was trying to analyze the relationship for  prediction verus sequence length and image size.\n\ni discover something fishy ....\n\n# 🐟\n\n![](https://i.ibb.co/KmD952w/Selection-041.png)\n![](https://i.ibb.co/qdg981H/Selection-038.png)",
      "votes": null
    },
    {
      "id": "1277611",
      "postDate": "04/19/2021 01:50:14",
      "content": "<p>This is very interesting…I am thinking we can use this in a couple ways:</p>\n<p><strong>First idea</strong><br>\n Use original image size as a meta feature - might help model determine the desired sequence length from the image size alone.</p>\n<p><strong>Second idea</strong><br>\nTrain different models on different sequence length ranges and then use test image sizes to determine how to infer with these models. For example:</p>\n<ul>\n<li>Train 3 models. 1st with sequence length min length - 100, 2nd with 101 - 150, 3rd with 151 - max length</li>\n<li>Estimate test image sequence length based on previous submission coupled with information about their image sizes (as you've shown above)</li>\n<li>Infer with 1st model on test images with sequence length min length - 100, infer with 2nd on sequence lengths 101 - 150, etc.</li>\n</ul>\n<p><strong>Third idea</strong><br>\nSame as second idea, but infer on all images with all three \"specialized\" models but use sequence length and test image sizes to determine best blend. </p>\n<hr>\n<p>You can apply the second and third ideas without this sequence length / image size correlation, but perhaps we will get better results using both the sequence length of a previous submission and information about the image size together.</p>\n<p>We could also train these groups with different resolutions: smaller sequences use lower resolution so you can save time, but the larger sequence models can be trained at fairly high resolution. </p>",
      "rawMarkdown": "This is very interesting...I am thinking we can use this in a couple ways:\n\n**First idea**\n Use original image size as a meta feature - might help model determine the desired sequence length from the image size alone.\n\n**Second idea**\nTrain different models on different sequence length ranges and then use test image sizes to determine how to infer with these models. For example:\n* Train 3 models. 1st with sequence length min length - 100, 2nd with 101 - 150, 3rd with 151 - max length\n* Estimate test image sequence length based on previous submission coupled with information about their image sizes (as you've shown above)\n* Infer with 1st model on test images with sequence length min length - 100, infer with 2nd on sequence lengths 101 - 150, etc.\n\n**Third idea**\nSame as second idea, but infer on all images with all three \"specialized\" models but use sequence length and test image sizes to determine best blend. \n\n---------------------------------------------------------------------------------------------------------------------------\nYou can apply the second and third ideas without this sequence length / image size correlation, but perhaps we will get better results using both the sequence length of a previous submission and information about the image size together.\n\nWe could also train these groups with different resolutions: smaller sequences use lower resolution so you can save time, but the larger sequence models can be trained at fairly high resolution.",
      "votes": null
    },
    {
      "id": "1277628",
      "postDate": "04/19/2021 02:25:22",
      "content": "<p>i was training and validating only big images.</p>\n<p>then i train another model using big + some small images. I want to see if small images can help to improve the prediction of big images. if there is no improvement, then one might as well separate big and small images</p>",
      "rawMarkdown": "i was training and validating only big images.\n\nthen i train another model using big + some small images. I want to see if small images can help to improve the prediction of big images. if there is no improvement, then one might as well separate big and small images",
      "votes": null
    },
    {
      "id": "1277633",
      "postDate": "04/19/2021 02:35:26",
      "content": "<p>add the coordinate layer as input may help.</p>\n<p>coordinate layer  can be normalised to local image size and/or to common size.<br>\nbecause of the \"flaw\"  the data is prepared, pixel location has information.</p>\n<p>mean image of similar size images show this</p>\n<p><img src=\"https://1fykyq3mdn5r21tpna3wkdyi-wpengine.netdna-ssl.com/wp-content/uploads/2018/07/Featured-Image-1-768x329.png\" alt=\"\"></p>",
      "rawMarkdown": "add the coordinate layer as input may help.\n\ncoordinate layer  can be normalised to local image size and/or to common size.\nbecause of the \"flaw\"  the data is prepared, pixel location has information.\n\nmean image of similar size images show this\n\n![](https://1fykyq3mdn5r21tpna3wkdyi-wpengine.netdna-ssl.com/wp-content/uploads/2018/07/Featured-Image-1-768x329.png)",
      "votes": null
    },
    {
      "id": "1277639",
      "postDate": "04/19/2021 02:53:09",
      "content": "<p>I had not heard of a coordconv layer before…thanks for sharing. I found a PyTorch implementation of it and will run some tests </p>",
      "rawMarkdown": "I had not heard of a coordconv layer before...thanks for sharing. I found a PyTorch implementation of it and will run some tests",
      "votes": null
    },
    {
      "id": "1277644",
      "postDate": "04/19/2021 03:10:28",
      "content": "<p><img src=\"https://i.ibb.co/tKTFvNP/Selection-044.png\" alt=\"\"></p>\n<p>a better idea is to get rid of the image copletely<br>\n</p>\n<p>make a mistake. should be 2**(16x16)</p>\n<p>much of the image are empty, so we can encode it sparsely. i think we can can run very fast</p>",
      "rawMarkdown": "![](https://i.ibb.co/tKTFvNP/Selection-044.png)\n\n\n\na better idea is to get rid of the image copletely\n~~even without auto encoder, because our image is binary, we only have 2**16 = 65536 discrete values. just by clustering, we can retrieve top 4k word~~\n\nmake a mistake. should be 2**(16x16)\n\nmuch of the image are empty, so we can encode it sparsely. i think we can can run very fast",
      "votes": null
    },
    {
      "id": "1277824",
      "postDate": "04/19/2021 09:11:28",
      "content": "<p>very neat! 🙂 could I please ask how you are detecting orientation?</p>\n<p>EDIT: from taking a look at the data it doesn't seem that they are ever rotated counter-clockwise or upside down… in this case that would be straightforward by just looking at the dimensions…</p>",
      "rawMarkdown": "very neat! 🙂 could I please ask how you are detecting orientation?\n\nEDIT: from taking a look at the data it doesn't seem that they are ever rotated counter-clockwise or upside down... in this case that would be straightforward by just looking at the dimensions...",
      "votes": null
    },
    {
      "id": "1277840",
      "postDate": "04/19/2021 09:35:52",
      "content": "<p>h&gt;w -&gt; rotate; but there are some images that don't follow this rule</p>",
      "rawMarkdown": "h>w -> rotate; but there are some images that don't follow this rule",
      "votes": null
    },
    {
      "id": "1277848",
      "postDate": "04/19/2021 09:42:57",
      "content": "<p>I train an orientation detection and scale detector using resnet34.</p>",
      "rawMarkdown": "I train an orientation detection and scale detector using resnet34.",
      "votes": null
    },
    {
      "id": "1278007",
      "postDate": "04/19/2021 13:30:22",
      "content": "<p>You have to keep the positions of those image patches in mind. And as the encoded patches still would need some kind of embedding vectors I don't see a big difference to given setups. I mean you would do Autoencoder -&gt; image patch encoding -&gt; patch embedding sequence. Now we have: image model -&gt; image features -&gt; shape-transformed feature sequence</p>",
      "rawMarkdown": "You have to keep the positions of those image patches in mind. And as the encoded patches still would need some kind of embedding vectors I don't see a big difference to given setups. I mean you would do Autoencoder -> image patch encoding -> patch embedding sequence. Now we have: image model -> image features -> shape-transformed feature sequence",
      "votes": null
    },
    {
      "id": "1278012",
      "postDate": "04/19/2021 13:33:50",
      "content": "<p>there is a difference if you have big and small images.<br>\n(in this case, you have to set image size to the largest image size. or you can train in batches of similar sizes, from smallest to largest)</p>\n<p>i am skipping all empty patches</p>",
      "rawMarkdown": "there is a difference if you have big and small images.\n(in this case, you have to set image size to the largest image size. or you can train in batches of similar sizes, from smallest to largest)\n\ni am skipping all empty patches",
      "votes": null
    },
    {
      "id": "1278019",
      "postDate": "04/19/2021 13:39:54",
      "content": "<p><code>coordinate layer</code><br>\nThanks! I knew that this existed and searched for it, but couldn't find it anymore.<br>\n<a href=\"https://paperswithcode.com/paper/an-intriguing-failing-of-convolutional-neural\" target=\"_blank\">paperswithcode - Here are some implementations and the paper</a></p>\n<p>Edit: I have my doubts, that coordinates are worth the effort. The images have different scales. And while you can correct rotation, the sizes are still different. And the molecular structures of the same kind are all over the place. So I predict that you get only a very small performance gain.</p>",
      "rawMarkdown": "`coordinate layer`\nThanks! I knew that this existed and searched for it, but couldn't find it anymore.\n[paperswithcode - Here are some implementations and the paper](https://paperswithcode.com/paper/an-intriguing-failing-of-convolutional-neural)\n\nEdit: I have my doubts, that coordinates are worth the effort. The images have different scales. And while you can correct rotation, the sizes are still different. And the molecular structures of the same kind are all over the place. So I predict that you get only a very small performance gain.",
      "votes": null
    },
    {
      "id": "1278044",
      "postDate": "04/19/2021 14:11:13",
      "content": "<p>Watch out that being too greedy with \"smart batching\" can lead to unstable training. I've lost a hole week because of it (smartbatched by inchi length).</p>",
      "rawMarkdown": "Watch out that being too greedy with \"smart batching\" can lead to unstable training. I've lost a hole week because of it (smartbatched by inchi length).",
      "votes": null
    },
    {
      "id": "1278114",
      "postDate": "04/19/2021 15:07:09",
      "content": "<blockquote>\n  <p>(smartbatched by inchi length).</p>\n</blockquote>\n<p>How so? I have been doing this for years for text data, and it works great. What you have to keep in mind though is: Don't sort after length over your whole training data, do it for windows where a window equals x batch sizes. Then you split the sorted windows again into batches and randomize them within the window. The latter, because so you don't get windows with batches of increasing lengths.<br>\nIf you do it over your whole training data, then yes, it won't generalize well. It might not even learn well in the first place and be unstable.</p>\n<blockquote>\n  <p>there is a difference if you have big and small images.</p>\n</blockquote>\n<p>Resizing it to the largest size would still make no real difference to given approach. The scale is nevertheless different for the images. And empty patches should be ignored by image model anyway.<br>\nImportant question: Does it actually work better than an image model? In this case you would be right.</p>",
      "rawMarkdown": "> (smartbatched by inchi length).\n\nHow so? I have been doing this for years for text data, and it works great. What you have to keep in mind though is: Don't sort after length over your whole training data, do it for windows where a window equals x batch sizes. Then you split the sorted windows again into batches and randomize them within the window. The latter, because so you don't get windows with batches of increasing lengths.\nIf you do it over your whole training data, then yes, it won't generalize well. It might not even learn well in the first place and be unstable.\n\n> there is a difference if you have big and small images.\n\nResizing it to the largest size would still make no real difference to given approach. The scale is nevertheless different for the images. And empty patches should be ignored by image model anyway.\nImportant question: Does it actually work better than an image model? In this case you would be right.",
      "votes": null
    },
    {
      "id": "1278130",
      "postDate": "04/19/2021 15:28:40",
      "content": "<blockquote>\n  <p>How so?</p>\n</blockquote>\n<p>As you just described it. :)</p>\n<p>Another way of the randomization: sorted(lambda x: x[1], key=len(x)*(1+random.random()*0.1-0.05))</p>",
      "rawMarkdown": "> How so?\n\nAs you just described it. :)\n\nAnother way of the randomization: sorted(lambda x: x[1], key=len(x)\\*(1+random.random()\\*0.1-0.05))",
      "votes": null
    },
    {
      "id": "1278132",
      "postDate": "04/19/2021 15:29:14",
      "content": "<p>\"The scale is nevertheless different for the images.\"</p>\n<p>no. there is only 2 scales in the test and train images.<br>\nthe bigger molecule image has a border of 58 pixel and the smaller one has a border of 36 pixel</p>\n<p>\"Does it actually work better than an image model? \"<br>\nI don't think it will be more accurate. But i hope it would be much faster</p>",
      "rawMarkdown": "\"The scale is nevertheless different for the images.\"\n\nno. there is only 2 scales in the test and train images.\nthe bigger molecule image has a border of 58 pixel and the smaller one has a border of 36 pixel\n\n\"Does it actually work better than an image model? \"\nI don't think it will be more accurate. But i hope it would be much faster",
      "votes": null
    },
    {
      "id": "1278217",
      "postDate": "04/19/2021 16:53:33",
      "content": "<p><code>As you just described it. :)</code><br>\nI meant, how does it come to be unstable. Unless I described that also with my last part.</p>\n<p><code>there is only 2 scales in the test and train images.</code><br>\nI see more than 2 different sizes of molecules. You can see it through the size of the letters or pentagons very well. Maybe the amount they where scaled afterwards has only two different sizes, but that doesn't change the absolute size.</p>\n<p><code>But i hope it would be much faster</code><br>\nI see. But I doubt that a lot. You still need the AutoEncoder + Sequence-Encoder VS. (CNN) Image Model. <br>\nEdit: I think you could go with a Transformer for images and just exclude the white patches and use a smart positional embedding which takes that into account. That should work as well as a classical approach and be faster (than at least other transformers).</p>",
      "rawMarkdown": "`As you just described it. :)`\nI meant, how does it come to be unstable. Unless I described that also with my last part.\n\n`there is only 2 scales in the test and train images.`\nI see more than 2 different sizes of molecules. You can see it through the size of the letters or pentagons very well. Maybe the amount they where scaled afterwards has only two different sizes, but that doesn't change the absolute size.\n\n`But i hope it would be much faster`\nI see. But I doubt that a lot. You still need the AutoEncoder + Sequence-Encoder VS. (CNN) Image Model. \nEdit: I think you could go with a Transformer for images and just exclude the white patches and use a smart positional embedding which takes that into account. That should work as well as a classical approach and be faster (than at least other transformers).",
      "votes": null
    },
    {
      "id": "1278222",
      "postDate": "04/19/2021 16:57:52",
      "content": "<blockquote>\n  <p>Unless I described that also with my last part.</p>\n  <blockquote>\n    <p>If you do it over your whole training data, then yes, it won't generalize well. It might not even learn well in the first place and be unstable.</p>\n  </blockquote>\n</blockquote>\n<p>I guess it comes to that in a batch it sees that all sequences are of length L and it probably drives the model a lot in that direction.</p>",
      "rawMarkdown": "> Unless I described that also with my last part.\n> > If you do it over your whole training data, then yes, it won't generalize well. It might not even learn well in the first place and be unstable.\n\nI guess it comes to that in a batch it sees that all sequences are of length L and it probably drives the model a lot in that direction.",
      "votes": null
    },
    {
      "id": "1278261",
      "postDate": "04/19/2021 17:40:14",
      "content": "<p><code>that in a batch it sees that all sequences are of length L...</code><br>\nYes, the gradient goes to much in a specific, <em>specialized</em> direction. And if you don't permute if afterwards, the length increases/decreases the whole epoch. This means the model forgets earlier patters for different lengths. <br>\nThat's why I use sorting within a window of x*batch_size. And then permute the order of those locally sorted batches again. That can accelerate the training speed usually by 2x (for the sequence model) and also improve the loss slightly. Although the optimal window sizes for both objectives might be different. <br>\nx is usually in [2, 30]. So it sorts elements of up to 30 batches. For this competition I use currently x = 4, but haven't tuned it yet.</p>\n<p>Why can local length sorting improve the loss slightly? Well, first the sequences are more similar, but still random enough so it gives a slightly better training signal. Second, you have less padding per sequence on average which means less noise (You would have to pad and probably reset to zero otherwise). The second part is even valid for a auto-regressive model because the weights that process the hidden embeddings are shared over all time-steps. Of course you just can mask the padding always (with zeros) - which I don't do.</p>\n<p>That worked well also for such benchmarks as Wikitext-103 where you</p>",
      "rawMarkdown": "`that in a batch it sees that all sequences are of length L...`\nYes, the gradient goes to much in a specific, *specialized* direction. And if you don't permute if afterwards, the length increases/decreases the whole epoch. This means the model forgets earlier patters for different lengths. \nThat's why I use sorting within a window of x*batch_size. And then permute the order of those locally sorted batches again. That can accelerate the training speed usually by 2x (for the sequence model) and also improve the loss slightly. Although the optimal window sizes for both objectives might be different. \nx is usually in [2, 30]. So it sorts elements of up to 30 batches. For this competition I use currently x = 4, but haven't tuned it yet.\n\nWhy can local length sorting improve the loss slightly? Well, first the sequences are more similar, but still random enough so it gives a slightly better training signal. Second, you have less padding per sequence on average which means less noise (You would have to pad and probably reset to zero otherwise). The second part is even valid for a auto-regressive model because the weights that process the hidden embeddings are shared over all time-steps. Of course you just can mask the padding always (with zeros) - which I don't do.\n\nThat worked well also for such benchmarks as Wikitext-103 where you",
      "votes": null
    },
    {
      "id": "1278269",
      "postDate": "04/19/2021 17:51:39",
      "content": "<blockquote>\n  <p>accelerate the training speed usually by 2x</p>\n</blockquote>\n<p>I doubt that. The speedup is around 2x just for the decoder part. If you use fixed sized images, then the time spent on the encoder is the same regardless of smart batching.</p>\n<p>Also, can you please compress your thoughts?</p>",
      "rawMarkdown": "> accelerate the training speed usually by 2x\n\nI doubt that. The speedup is around 2x just for the decoder part. If you use fixed sized images, then the time spent on the encoder is the same regardless of smart batching.\n\nAlso, can you please compress your thoughts?",
      "votes": null
    },
    {
      "id": "1278276",
      "postDate": "04/19/2021 17:55:39",
      "content": "<p>Your are right, in this case it isn't 2x. But I meant just the sequence part, forgot to add that.<br>\nI will add a TLDR in the future.</p>",
      "rawMarkdown": "Your are right, in this case it isn't 2x. But I meant just the sequence part, forgot to add that.\nI will add a TLDR in the future.",
      "votes": null
    },
    {
      "id": "1278291",
      "postDate": "04/19/2021 18:02:05",
      "content": "<p>I hope you don't mind the critique!</p>",
      "rawMarkdown": "I hope you don't mind the critique!",
      "votes": null
    },
    {
      "id": "1278298",
      "postDate": "04/19/2021 18:05:00",
      "content": "<p><code>I hope you don't mind the critique!</code><br>\nNo. Yours is based on arguments (apart from the compress part), so it's welcome.<br>\nI also hope I could help with my length sorting description.</p>",
      "rawMarkdown": "`I hope you don't mind the critique!`\nNo. Yours is based on arguments (apart from the compress part), so it's welcome.\nI also hope I could help with my length sorting description.",
      "votes": null
    },
    {
      "id": "1278300",
      "postDate": "04/19/2021 18:09:39",
      "content": "<blockquote>\n  <p>(apart from the compress part)</p>\n</blockquote>\n<p>I was referring to that under <code>critique</code> :D<br>\nSo I hope you don't mind that!</p>",
      "rawMarkdown": "> (apart from the compress part)\n\nI was referring to that under `critique` :D\nSo I hope you don't mind that!",
      "votes": null
    },
    {
      "id": "1278335",
      "postDate": "04/19/2021 18:55:12",
      "content": "<p>It's fine. Some people don't understand it with a long explanation, others don't want/need those. You get a TLDR in case of long posts.</p>",
      "rawMarkdown": "It's fine. Some people don't understand it with a long explanation, others don't want/need those. You get a TLDR in case of long posts.",
      "votes": null
    },
    {
      "id": "1280476",
      "postDate": "04/22/2021 02:52:19",
      "content": "<p>another way to do text-to-text:<br>\n(use ascii art)</p>\n<pre><code>          O\n         //\n    Cl--{/\n         \\\n          \\_____\n          / --- \\\n         /       \\\n         \\\\     //\n          \\_____/\n</code></pre>\n<pre><code>        OH\n       /\n  C --C       O\n //    \\\\    //\nC       C---C\n \\  __ /     \\\n  C --C             OH\n</code></pre>",
      "rawMarkdown": "another way to do text-to-text:\n(use ascii art)\n\n```\n\n          O\n         //\n    Cl--{/\n         \\\n          \\_____\n          / --- \\\n         /       \\\n         \\\\     //\n          \\_____/\n\n\n```\n\n```\n        OH\n       /\n  C --C       O\n //    \\\\    //\nC       C---C\n \\  __ /     \\\n  C --C             OH\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1277611,
      "author_name": "tuckerarrants",
      "author_url": "",
      "post_date": "04/19/2021 01:50:14",
      "content": "<p>This is very interesting…I am thinking we can use this in a couple ways:</p>\n<p><strong>First idea</strong><br>\n Use original image size as a meta feature - might help model determine the desired sequence length from the image size alone.</p>\n<p><strong>Second idea</strong><br>\nTrain different models on different sequence length ranges and then use test image sizes to determine how to infer with these models. For example:</p>\n<ul>\n<li>Train 3 models. 1st with sequence length min length - 100, 2nd with 101 - 150, 3rd with 151 - max length</li>\n<li>Estimate test image sequence length based on previous submission coupled with information about their image sizes (as you've shown above)</li>\n<li>Infer with 1st model on test images with sequence length min length - 100, infer with 2nd on sequence lengths 101 - 150, etc.</li>\n</ul>\n<p><strong>Third idea</strong><br>\nSame as second idea, but infer on all images with all three \"specialized\" models but use sequence length and test image sizes to determine best blend. </p>\n<hr>\n<p>You can apply the second and third ideas without this sequence length / image size correlation, but perhaps we will get better results using both the sequence length of a previous submission and information about the image size together.</p>\n<p>We could also train these groups with different resolutions: smaller sequences use lower resolution so you can save time, but the larger sequence models can be trained at fairly high resolution. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1277628,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/19/2021 02:25:22",
          "content": "<p>i was training and validating only big images.</p>\n<p>then i train another model using big + some small images. I want to see if small images can help to improve the prediction of big images. if there is no improvement, then one might as well separate big and small images</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1277633,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/19/2021 02:35:26",
          "content": "<p>add the coordinate layer as input may help.</p>\n<p>coordinate layer  can be normalised to local image size and/or to common size.<br>\nbecause of the \"flaw\"  the data is prepared, pixel location has information.</p>\n<p>mean image of similar size images show this</p>\n<p><img src=\"https://1fykyq3mdn5r21tpna3wkdyi-wpengine.netdna-ssl.com/wp-content/uploads/2018/07/Featured-Image-1-768x329.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1277639,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "04/19/2021 02:53:09",
          "content": "<p>I had not heard of a coordconv layer before…thanks for sharing. I found a PyTorch implementation of it and will run some tests </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278019,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/19/2021 13:39:54",
          "content": "<p><code>coordinate layer</code><br>\nThanks! I knew that this existed and searched for it, but couldn't find it anymore.<br>\n<a href=\"https://paperswithcode.com/paper/an-intriguing-failing-of-convolutional-neural\" target=\"_blank\">paperswithcode - Here are some implementations and the paper</a></p>\n<p>Edit: I have my doubts, that coordinates are worth the effort. The images have different scales. And while you can correct rotation, the sizes are still different. And the molecular structures of the same kind are all over the place. So I predict that you get only a very small performance gain.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1277644,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/19/2021 03:10:28",
      "content": "<p><img src=\"https://i.ibb.co/tKTFvNP/Selection-044.png\" alt=\"\"></p>\n<p>a better idea is to get rid of the image copletely<br>\n</p>\n<p>make a mistake. should be 2**(16x16)</p>\n<p>much of the image are empty, so we can encode it sparsely. i think we can can run very fast</p>",
      "votes": null,
      "replies": [
        {
          "id": 1278007,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/19/2021 13:30:22",
          "content": "<p>You have to keep the positions of those image patches in mind. And as the encoded patches still would need some kind of embedding vectors I don't see a big difference to given setups. I mean you would do Autoencoder -&gt; image patch encoding -&gt; patch embedding sequence. Now we have: image model -&gt; image features -&gt; shape-transformed feature sequence</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278012,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/19/2021 13:33:50",
          "content": "<p>there is a difference if you have big and small images.<br>\n(in this case, you have to set image size to the largest image size. or you can train in batches of similar sizes, from smallest to largest)</p>\n<p>i am skipping all empty patches</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278044,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/19/2021 14:11:13",
          "content": "<p>Watch out that being too greedy with \"smart batching\" can lead to unstable training. I've lost a hole week because of it (smartbatched by inchi length).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278114,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/19/2021 15:07:09",
          "content": "<blockquote>\n  <p>(smartbatched by inchi length).</p>\n</blockquote>\n<p>How so? I have been doing this for years for text data, and it works great. What you have to keep in mind though is: Don't sort after length over your whole training data, do it for windows where a window equals x batch sizes. Then you split the sorted windows again into batches and randomize them within the window. The latter, because so you don't get windows with batches of increasing lengths.<br>\nIf you do it over your whole training data, then yes, it won't generalize well. It might not even learn well in the first place and be unstable.</p>\n<blockquote>\n  <p>there is a difference if you have big and small images.</p>\n</blockquote>\n<p>Resizing it to the largest size would still make no real difference to given approach. The scale is nevertheless different for the images. And empty patches should be ignored by image model anyway.<br>\nImportant question: Does it actually work better than an image model? In this case you would be right.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278130,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/19/2021 15:28:40",
          "content": "<blockquote>\n  <p>How so?</p>\n</blockquote>\n<p>As you just described it. :)</p>\n<p>Another way of the randomization: sorted(lambda x: x[1], key=len(x)*(1+random.random()*0.1-0.05))</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278132,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/19/2021 15:29:14",
          "content": "<p>\"The scale is nevertheless different for the images.\"</p>\n<p>no. there is only 2 scales in the test and train images.<br>\nthe bigger molecule image has a border of 58 pixel and the smaller one has a border of 36 pixel</p>\n<p>\"Does it actually work better than an image model? \"<br>\nI don't think it will be more accurate. But i hope it would be much faster</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278217,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/19/2021 16:53:33",
          "content": "<p><code>As you just described it. :)</code><br>\nI meant, how does it come to be unstable. Unless I described that also with my last part.</p>\n<p><code>there is only 2 scales in the test and train images.</code><br>\nI see more than 2 different sizes of molecules. You can see it through the size of the letters or pentagons very well. Maybe the amount they where scaled afterwards has only two different sizes, but that doesn't change the absolute size.</p>\n<p><code>But i hope it would be much faster</code><br>\nI see. But I doubt that a lot. You still need the AutoEncoder + Sequence-Encoder VS. (CNN) Image Model. <br>\nEdit: I think you could go with a Transformer for images and just exclude the white patches and use a smart positional embedding which takes that into account. That should work as well as a classical approach and be faster (than at least other transformers).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278222,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/19/2021 16:57:52",
          "content": "<blockquote>\n  <p>Unless I described that also with my last part.</p>\n  <blockquote>\n    <p>If you do it over your whole training data, then yes, it won't generalize well. It might not even learn well in the first place and be unstable.</p>\n  </blockquote>\n</blockquote>\n<p>I guess it comes to that in a batch it sees that all sequences are of length L and it probably drives the model a lot in that direction.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278261,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/19/2021 17:40:14",
          "content": "<p><code>that in a batch it sees that all sequences are of length L...</code><br>\nYes, the gradient goes to much in a specific, <em>specialized</em> direction. And if you don't permute if afterwards, the length increases/decreases the whole epoch. This means the model forgets earlier patters for different lengths. <br>\nThat's why I use sorting within a window of x*batch_size. And then permute the order of those locally sorted batches again. That can accelerate the training speed usually by 2x (for the sequence model) and also improve the loss slightly. Although the optimal window sizes for both objectives might be different. <br>\nx is usually in [2, 30]. So it sorts elements of up to 30 batches. For this competition I use currently x = 4, but haven't tuned it yet.</p>\n<p>Why can local length sorting improve the loss slightly? Well, first the sequences are more similar, but still random enough so it gives a slightly better training signal. Second, you have less padding per sequence on average which means less noise (You would have to pad and probably reset to zero otherwise). The second part is even valid for a auto-regressive model because the weights that process the hidden embeddings are shared over all time-steps. Of course you just can mask the padding always (with zeros) - which I don't do.</p>\n<p>That worked well also for such benchmarks as Wikitext-103 where you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278269,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/19/2021 17:51:39",
          "content": "<blockquote>\n  <p>accelerate the training speed usually by 2x</p>\n</blockquote>\n<p>I doubt that. The speedup is around 2x just for the decoder part. If you use fixed sized images, then the time spent on the encoder is the same regardless of smart batching.</p>\n<p>Also, can you please compress your thoughts?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278276,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/19/2021 17:55:39",
          "content": "<p>Your are right, in this case it isn't 2x. But I meant just the sequence part, forgot to add that.<br>\nI will add a TLDR in the future.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278291,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/19/2021 18:02:05",
          "content": "<p>I hope you don't mind the critique!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278298,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/19/2021 18:05:00",
          "content": "<p><code>I hope you don't mind the critique!</code><br>\nNo. Yours is based on arguments (apart from the compress part), so it's welcome.<br>\nI also hope I could help with my length sorting description.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278300,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/19/2021 18:09:39",
          "content": "<blockquote>\n  <p>(apart from the compress part)</p>\n</blockquote>\n<p>I was referring to that under <code>critique</code> :D<br>\nSo I hope you don't mind that!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278335,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "04/19/2021 18:55:12",
          "content": "<p>It's fine. Some people don't understand it with a long explanation, others don't want/need those. You get a TLDR in case of long posts.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1277824,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "04/19/2021 09:11:28",
      "content": "<p>very neat! 🙂 could I please ask how you are detecting orientation?</p>\n<p>EDIT: from taking a look at the data it doesn't seem that they are ever rotated counter-clockwise or upside down… in this case that would be straightforward by just looking at the dimensions…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1277840,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/19/2021 09:35:52",
          "content": "<p>h&gt;w -&gt; rotate; but there are some images that don't follow this rule</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1277848,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/19/2021 09:42:57",
          "content": "<p>I train an orientation detection and scale detector using resnet34.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1280476,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/22/2021 02:52:19",
      "content": "<p>another way to do text-to-text:<br>\n(use ascii art)</p>\n<pre><code>          O\n         //\n    Cl--{/\n         \\\n          \\_____\n          / --- \\\n         /       \\\n         \\\\     //\n          \\_____/\n</code></pre>\n<pre><code>        OH\n       /\n  C --C       O\n //    \\\\    //\nC       C---C\n \\  __ /     \\\n  C --C             OH\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1277591": "i was trying to analyze the relationship for  prediction verus sequence length and image size.\n\ni discover something fishy ....\n\n# 🐟\n\n![](https://i.ibb.co/KmD952w/Selection-041.png)\n![](https://i.ibb.co/qdg981H/Selection-038.png)",
    "1277611": "This is very interesting...I am thinking we can use this in a couple ways:\n\n**First idea**\n Use original image size as a meta feature - might help model determine the desired sequence length from the image size alone.\n\n**Second idea**\nTrain different models on different sequence length ranges and then use test image sizes to determine how to infer with these models. For example:\n* Train 3 models. 1st with sequence length min length - 100, 2nd with 101 - 150, 3rd with 151 - max length\n* Estimate test image sequence length based on previous submission coupled with information about their image sizes (as you've shown above)\n* Infer with 1st model on test images with sequence length min length - 100, infer with 2nd on sequence lengths 101 - 150, etc.\n\n**Third idea**\nSame as second idea, but infer on all images with all three \"specialized\" models but use sequence length and test image sizes to determine best blend. \n\n---------------------------------------------------------------------------------------------------------------------------\nYou can apply the second and third ideas without this sequence length / image size correlation, but perhaps we will get better results using both the sequence length of a previous submission and information about the image size together.\n\nWe could also train these groups with different resolutions: smaller sequences use lower resolution so you can save time, but the larger sequence models can be trained at fairly high resolution.",
    "1277628": "i was training and validating only big images.\n\nthen i train another model using big + some small images. I want to see if small images can help to improve the prediction of big images. if there is no improvement, then one might as well separate big and small images",
    "1277633": "add the coordinate layer as input may help.\n\ncoordinate layer  can be normalised to local image size and/or to common size.\nbecause of the \"flaw\"  the data is prepared, pixel location has information.\n\nmean image of similar size images show this\n\n![](https://1fykyq3mdn5r21tpna3wkdyi-wpengine.netdna-ssl.com/wp-content/uploads/2018/07/Featured-Image-1-768x329.png)",
    "1277639": "I had not heard of a coordconv layer before...thanks for sharing. I found a PyTorch implementation of it and will run some tests",
    "1277644": "![](https://i.ibb.co/tKTFvNP/Selection-044.png)\n\n\n\na better idea is to get rid of the image copletely\n~~even without auto encoder, because our image is binary, we only have 2**16 = 65536 discrete values. just by clustering, we can retrieve top 4k word~~\n\nmake a mistake. should be 2**(16x16)\n\nmuch of the image are empty, so we can encode it sparsely. i think we can can run very fast",
    "1277824": "very neat! 🙂 could I please ask how you are detecting orientation?\n\nEDIT: from taking a look at the data it doesn't seem that they are ever rotated counter-clockwise or upside down... in this case that would be straightforward by just looking at the dimensions...",
    "1277840": "h>w -> rotate; but there are some images that don't follow this rule",
    "1277848": "I train an orientation detection and scale detector using resnet34.",
    "1278007": "You have to keep the positions of those image patches in mind. And as the encoded patches still would need some kind of embedding vectors I don't see a big difference to given setups. I mean you would do Autoencoder -> image patch encoding -> patch embedding sequence. Now we have: image model -> image features -> shape-transformed feature sequence",
    "1278012": "there is a difference if you have big and small images.\n(in this case, you have to set image size to the largest image size. or you can train in batches of similar sizes, from smallest to largest)\n\ni am skipping all empty patches",
    "1278019": "`coordinate layer`\nThanks! I knew that this existed and searched for it, but couldn't find it anymore.\n[paperswithcode - Here are some implementations and the paper](https://paperswithcode.com/paper/an-intriguing-failing-of-convolutional-neural)\n\nEdit: I have my doubts, that coordinates are worth the effort. The images have different scales. And while you can correct rotation, the sizes are still different. And the molecular structures of the same kind are all over the place. So I predict that you get only a very small performance gain.",
    "1278044": "Watch out that being too greedy with \"smart batching\" can lead to unstable training. I've lost a hole week because of it (smartbatched by inchi length).",
    "1278114": "> (smartbatched by inchi length).\n\nHow so? I have been doing this for years for text data, and it works great. What you have to keep in mind though is: Don't sort after length over your whole training data, do it for windows where a window equals x batch sizes. Then you split the sorted windows again into batches and randomize them within the window. The latter, because so you don't get windows with batches of increasing lengths.\nIf you do it over your whole training data, then yes, it won't generalize well. It might not even learn well in the first place and be unstable.\n\n> there is a difference if you have big and small images.\n\nResizing it to the largest size would still make no real difference to given approach. The scale is nevertheless different for the images. And empty patches should be ignored by image model anyway.\nImportant question: Does it actually work better than an image model? In this case you would be right.",
    "1278130": "> How so?\n\nAs you just described it. :)\n\nAnother way of the randomization: sorted(lambda x: x[1], key=len(x)\\*(1+random.random()\\*0.1-0.05))",
    "1278132": "\"The scale is nevertheless different for the images.\"\n\nno. there is only 2 scales in the test and train images.\nthe bigger molecule image has a border of 58 pixel and the smaller one has a border of 36 pixel\n\n\"Does it actually work better than an image model? \"\nI don't think it will be more accurate. But i hope it would be much faster",
    "1278217": "`As you just described it. :)`\nI meant, how does it come to be unstable. Unless I described that also with my last part.\n\n`there is only 2 scales in the test and train images.`\nI see more than 2 different sizes of molecules. You can see it through the size of the letters or pentagons very well. Maybe the amount they where scaled afterwards has only two different sizes, but that doesn't change the absolute size.\n\n`But i hope it would be much faster`\nI see. But I doubt that a lot. You still need the AutoEncoder + Sequence-Encoder VS. (CNN) Image Model. \nEdit: I think you could go with a Transformer for images and just exclude the white patches and use a smart positional embedding which takes that into account. That should work as well as a classical approach and be faster (than at least other transformers).",
    "1278222": "> Unless I described that also with my last part.\n> > If you do it over your whole training data, then yes, it won't generalize well. It might not even learn well in the first place and be unstable.\n\nI guess it comes to that in a batch it sees that all sequences are of length L and it probably drives the model a lot in that direction.",
    "1278261": "`that in a batch it sees that all sequences are of length L...`\nYes, the gradient goes to much in a specific, *specialized* direction. And if you don't permute if afterwards, the length increases/decreases the whole epoch. This means the model forgets earlier patters for different lengths. \nThat's why I use sorting within a window of x*batch_size. And then permute the order of those locally sorted batches again. That can accelerate the training speed usually by 2x (for the sequence model) and also improve the loss slightly. Although the optimal window sizes for both objectives might be different. \nx is usually in [2, 30]. So it sorts elements of up to 30 batches. For this competition I use currently x = 4, but haven't tuned it yet.\n\nWhy can local length sorting improve the loss slightly? Well, first the sequences are more similar, but still random enough so it gives a slightly better training signal. Second, you have less padding per sequence on average which means less noise (You would have to pad and probably reset to zero otherwise). The second part is even valid for a auto-regressive model because the weights that process the hidden embeddings are shared over all time-steps. Of course you just can mask the padding always (with zeros) - which I don't do.\n\nThat worked well also for such benchmarks as Wikitext-103 where you",
    "1278269": "> accelerate the training speed usually by 2x\n\nI doubt that. The speedup is around 2x just for the decoder part. If you use fixed sized images, then the time spent on the encoder is the same regardless of smart batching.\n\nAlso, can you please compress your thoughts?",
    "1278276": "Your are right, in this case it isn't 2x. But I meant just the sequence part, forgot to add that.\nI will add a TLDR in the future.",
    "1278291": "I hope you don't mind the critique!",
    "1278298": "`I hope you don't mind the critique!`\nNo. Yours is based on arguments (apart from the compress part), so it's welcome.\nI also hope I could help with my length sorting description.",
    "1278300": "> (apart from the compress part)\n\nI was referring to that under `critique` :D\nSo I hope you don't mind that!",
    "1278335": "It's fine. Some people don't understand it with a long explanation, others don't want/need those. You get a TLDR in case of long posts.",
    "1280476": "another way to do text-to-text:\n(use ascii art)\n\n```\n\n          O\n         //\n    Cl--{/\n         \\\n          \\_____\n          / --- \\\n         /       \\\n         \\\\     //\n          \\_____/\n\n\n```\n\n```\n        OH\n       /\n  C --C       O\n //    \\\\    //\nC       C---C\n \\  __ /     \\\n  C --C             OH\n```"
  },
  "source": "meta"
}