{
  "id": 226245,
  "title": "Difficulties training a CNN",
  "url": "/competitions/bms-molecular-translation/discussion/226245",
  "author_name": "",
  "post_date": "2021-03-15T18:52:12.163872900Z",
  "votes": 4,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi! It's my first Kaggle contest, and I'm having some difficulty that I suspect is due to some basic error. </p>\n<p>I'm trying to build a CNN to give me a count of how many atoms of a given element are in each molecule - starting with oxygen. I'm treating it as a regression problem rather than classification.</p>\n<ul>\n<li>Load in all the images starting \"000\", which is about 592. (I know I'll need to do more eventually but while I'm still fiddling around so much I want to use a more manageable set.)</li>\n<li>Trimmed whitespace around the molecules, and then padded the images back out to get to uniform size.</li>\n<li>Parsed out the InChI labels to get the # of oxygen atoms in each molecule.</li>\n<li>Built a CNN with the following layers:</li>\n</ul>\n<ol>\n<li>Convolution, 10x10, 1 input channel =&gt; 16 output channels, ReLu activation</li>\n<li>Max Pooling, 2x2, stride 2</li>\n<li>Convolution, 10x10, 16 input channels =&gt; 32 output channels, ReLu activation</li>\n<li>Max Pooling, 2x2, stride 2</li>\n<li>sum</li>\n</ol>\n<ul>\n<li>Defined loss function as mean squared error</li>\n<li>Trained on all 592 images against their oxygen counts (I know eventually I will need to do train/test splitting to avoid overfitting, but for now I'm just trying to get something working)</li>\n<li>Compare predictions on the training data with actuals:<br>\n![<a href=\"https://photos.app.goo.gl/gusw7GmvmJSztdo28](url\" target=\"_blank\">https://photos.app.goo.gl/gusw7GmvmJSztdo28](url</a> to embed)<br>\nWell, that's not good. That's pretty much just noise. My fear with a small set was overfitting, and this is not that. Correlation coefficient between the predictions and actuals is 0.328. </li>\n</ul>\n<p>I'm not sure where I'm going wrong. Thoughts are:</p>\n<ol>\n<li>I generally haven't seen CNNs used for counting by regression - is this a plausible approach or should I classify instead? Seems like I could end up with lots and lots of classes if I do that.</li>\n<li>If I am regressing, is the sum() at the end the right way to express that? Or do I need a dense layer instead? </li>\n<li>Do I have the right # of layers in my CNN? How do I figure that out?</li>\n<li>Do I need to configure them differently? Convolution sizes and channels and strides and stuff?</li>\n<li>Do I just need to use more of the training data to get a meaningful result?</li>\n<li>Is padding the images to be the same size the right approach, or should I scale them down instead? Or is there some way of managing them being different sizes?</li>\n<li>Is there a bug somewhere in my code, such that it's not doing what I think it's doing? (I've been working in Julia instead of Python - happy to share the code if this is the most likely issue.)</li>\n<li>Are CNNs just the wrong tool for this portion of the problem?</li>\n</ol>\n<p>I have approaches that I could take towards exploring each of these more, but they are each time consuming, and I'm concerned about putting a lot of effort into fixing the wrong issue. (e.g. I could do cross-validation over a bunch of different convolution parameters, but if the fundamental problem is that my training set is too small, that could be hours thrown away for no benefit.) I'd be very grateful if anyone can spot a fundamental error in what I'm doing here.</p>",
  "messages": [
    {
      "id": "1239482",
      "postDate": "03/15/2021 18:52:12",
      "content": "<p>Hi! It's my first Kaggle contest, and I'm having some difficulty that I suspect is due to some basic error. </p>\n<p>I'm trying to build a CNN to give me a count of how many atoms of a given element are in each molecule - starting with oxygen. I'm treating it as a regression problem rather than classification.</p>\n<ul>\n<li>Load in all the images starting \"000\", which is about 592. (I know I'll need to do more eventually but while I'm still fiddling around so much I want to use a more manageable set.)</li>\n<li>Trimmed whitespace around the molecules, and then padded the images back out to get to uniform size.</li>\n<li>Parsed out the InChI labels to get the # of oxygen atoms in each molecule.</li>\n<li>Built a CNN with the following layers:</li>\n</ul>\n<ol>\n<li>Convolution, 10x10, 1 input channel =&gt; 16 output channels, ReLu activation</li>\n<li>Max Pooling, 2x2, stride 2</li>\n<li>Convolution, 10x10, 16 input channels =&gt; 32 output channels, ReLu activation</li>\n<li>Max Pooling, 2x2, stride 2</li>\n<li>sum</li>\n</ol>\n<ul>\n<li>Defined loss function as mean squared error</li>\n<li>Trained on all 592 images against their oxygen counts (I know eventually I will need to do train/test splitting to avoid overfitting, but for now I'm just trying to get something working)</li>\n<li>Compare predictions on the training data with actuals:<br>\n![<a href=\"https://photos.app.goo.gl/gusw7GmvmJSztdo28](url\" target=\"_blank\">https://photos.app.goo.gl/gusw7GmvmJSztdo28](url</a> to embed)<br>\nWell, that's not good. That's pretty much just noise. My fear with a small set was overfitting, and this is not that. Correlation coefficient between the predictions and actuals is 0.328. </li>\n</ul>\n<p>I'm not sure where I'm going wrong. Thoughts are:</p>\n<ol>\n<li>I generally haven't seen CNNs used for counting by regression - is this a plausible approach or should I classify instead? Seems like I could end up with lots and lots of classes if I do that.</li>\n<li>If I am regressing, is the sum() at the end the right way to express that? Or do I need a dense layer instead? </li>\n<li>Do I have the right # of layers in my CNN? How do I figure that out?</li>\n<li>Do I need to configure them differently? Convolution sizes and channels and strides and stuff?</li>\n<li>Do I just need to use more of the training data to get a meaningful result?</li>\n<li>Is padding the images to be the same size the right approach, or should I scale them down instead? Or is there some way of managing them being different sizes?</li>\n<li>Is there a bug somewhere in my code, such that it's not doing what I think it's doing? (I've been working in Julia instead of Python - happy to share the code if this is the most likely issue.)</li>\n<li>Are CNNs just the wrong tool for this portion of the problem?</li>\n</ol>\n<p>I have approaches that I could take towards exploring each of these more, but they are each time consuming, and I'm concerned about putting a lot of effort into fixing the wrong issue. (e.g. I could do cross-validation over a bunch of different convolution parameters, but if the fundamental problem is that my training set is too small, that could be hours thrown away for no benefit.) I'd be very grateful if anyone can spot a fundamental error in what I'm doing here.</p>",
      "rawMarkdown": "Hi! It's my first Kaggle contest, and I'm having some difficulty that I suspect is due to some basic error. \n\nI'm trying to build a CNN to give me a count of how many atoms of a given element are in each molecule - starting with oxygen. I'm treating it as a regression problem rather than classification.\n\n- Load in all the images starting \"000\", which is about 592. (I know I'll need to do more eventually but while I'm still fiddling around so much I want to use a more manageable set.)\n- Trimmed whitespace around the molecules, and then padded the images back out to get to uniform size.\n- Parsed out the InChI labels to get the # of oxygen atoms in each molecule.\n- Built a CNN with the following layers:\n1. Convolution, 10x10, 1 input channel => 16 output channels, ReLu activation\n2. Max Pooling, 2x2, stride 2\n3. Convolution, 10x10, 16 input channels => 32 output channels, ReLu activation\n4. Max Pooling, 2x2, stride 2\n5. sum\n- Defined loss function as mean squared error\n- Trained on all 592 images against their oxygen counts (I know eventually I will need to do train/test splitting to avoid overfitting, but for now I'm just trying to get something working)\n- Compare predictions on the training data with actuals:\n![https://photos.app.goo.gl/gusw7GmvmJSztdo28](url to embed)\nWell, that's not good. That's pretty much just noise. My fear with a small set was overfitting, and this is not that. Correlation coefficient between the predictions and actuals is 0.328. \n\nI'm not sure where I'm going wrong. Thoughts are:\n1. I generally haven't seen CNNs used for counting by regression - is this a plausible approach or should I classify instead? Seems like I could end up with lots and lots of classes if I do that.\n2. If I am regressing, is the sum() at the end the right way to express that? Or do I need a dense layer instead? \n3. Do I have the right # of layers in my CNN? How do I figure that out?\n4. Do I need to configure them differently? Convolution sizes and channels and strides and stuff?\n5. Do I just need to use more of the training data to get a meaningful result?\n6. Is padding the images to be the same size the right approach, or should I scale them down instead? Or is there some way of managing them being different sizes?\n7. Is there a bug somewhere in my code, such that it's not doing what I think it's doing? (I've been working in Julia instead of Python - happy to share the code if this is the most likely issue.)\n8. Are CNNs just the wrong tool for this portion of the problem?\n\nI have approaches that I could take towards exploring each of these more, but they are each time consuming, and I'm concerned about putting a lot of effort into fixing the wrong issue. (e.g. I could do cross-validation over a bunch of different convolution parameters, but if the fundamental problem is that my training set is too small, that could be hours thrown away for no benefit.) I'd be very grateful if anyone can spot a fundamental error in what I'm doing here.",
      "votes": null
    },
    {
      "id": "1241937",
      "postDate": "03/17/2021 09:39:00",
      "content": "<p>Hi Jeremy,<br>\nI'll try to summarize what I've understood from your experiments. The approach to do a regression for each atom type is ok as long as you have a technique in mind to combine all of them together eventually.</p>\n<ol>\n<li><p>You are padding them all to be the same size, so assuming you get them all to 512x512 size. </p></li>\n<li><p>I think the issue is that the <em>receptive region</em>* of the conv net that you've defined is too low at the moment. I believe doing more convolutions with smaller kernel sizes would be better than doing 2 convolutions with 10x10 filter sizes. Currently after all the conv+pool layers your dimensions are somewhere around (121, 121). Try something like either 3 or 4 blocks of  (Conv (3x3), Conv (3x3), Max_pool). </p></li>\n<li><p>I would suggest that you sum along the height and width dimensions after the CNNs and then 1-2 dense layers to get the output. You'll need to flatten the output of this summing layer before using the dense layers.</p></li>\n<li><p>Since you are doing an aggregation after the CNNs layers (sum here), you can also try doing it on original image sizes provided that you can code you model to be dynamic enough to deal with sized inputs. </p></li>\n<li><p>Check if your model is underfitting or not, you need more data in any case if you want to debug the issue at hand. Try to cover more training data in your experiments. Like checking error after training on 5,000 images on same model architecture. If the error is reduced significantly, then the issue is with the training size. If the error remains in similar range, then the model architecture may be an issue.</p></li>\n</ol>\n<p>*<em>Here I mean the receptive region of a neuron in the last CNN layer.</em></p>",
      "rawMarkdown": "Hi Jeremy,\nI'll try to summarize what I've understood from your experiments. The approach to do a regression for each atom type is ok as long as you have a technique in mind to combine all of them together eventually.\n\n\n1. You are padding them all to be the same size, so assuming you get them all to 512x512 size. \n\n1. I think the issue is that the *receptive region** of the conv net that you've defined is too low at the moment. I believe doing more convolutions with smaller kernel sizes would be better than doing 2 convolutions with 10x10 filter sizes. Currently after all the conv+pool layers your dimensions are somewhere around (121, 121). Try something like either 3 or 4 blocks of  (Conv (3x3), Conv (3x3), Max_pool). \n\n1. I would suggest that you sum along the height and width dimensions after the CNNs and then 1-2 dense layers to get the output. You'll need to flatten the output of this summing layer before using the dense layers.\n\n1. Since you are doing an aggregation after the CNNs layers (sum here), you can also try doing it on original image sizes provided that you can code you model to be dynamic enough to deal with sized inputs. \n\n1. Check if your model is underfitting or not, you need more data in any case if you want to debug the issue at hand. Try to cover more training data in your experiments. Like checking error after training on 5,000 images on same model architecture. If the error is reduced significantly, then the issue is with the training size. If the error remains in similar range, then the model architecture may be an issue.\n\n\n**Here I mean the receptive region of a neuron in the last CNN layer.*",
      "votes": null
    },
    {
      "id": "1242561",
      "postDate": "03/17/2021 17:23:15",
      "content": "<p>Thank you, this is very helpful! </p>\n<p>To make sure I understand what you're saying conceptually, is this correct?: </p>\n<ul>\n<li>the images going in each have HeightxWidthx1 Channel</li>\n<li>as I convolve and pool, the height and width decrease, and the number of channels increases to N</li>\n<li>the summing layer collapses the height and width, leaving each image as height 1, width 1, N channels</li>\n<li>the dense layer (or layers) takes the N channels and reduces them to 1</li>\n</ul>\n<p>Trying to read in all 9550 images that start with \"00\" was giving me OutOfMemory errors, so I have been fiddling with ways of working around that. I've increased the learning rate (using Adam with eta 0.1 instead of 0.001) so that it doesn't take forever to evaluate model changes. (It's still taking a pretty long time to run.)</p>",
      "rawMarkdown": "Thank you, this is very helpful! \n\nTo make sure I understand what you're saying conceptually, is this correct?: \n- the images going in each have HeightxWidthx1 Channel\n- as I convolve and pool, the height and width decrease, and the number of channels increases to N\n- the summing layer collapses the height and width, leaving each image as height 1, width 1, N channels\n- the dense layer (or layers) takes the N channels and reduces them to 1\n\nTrying to read in all 9550 images that start with \"00\" was giving me OutOfMemory errors, so I have been fiddling with ways of working around that. I've increased the learning rate (using Adam with eta 0.1 instead of 0.001) so that it doesn't take forever to evaluate model changes. (It's still taking a pretty long time to run.)",
      "votes": null
    },
    {
      "id": "1242778",
      "postDate": "03/17/2021 19:41:37",
      "content": "<p>Yeah, that's correct.</p>\n<p>Are you using Kaggle notebooks? I'm quite sure reading 10k images here shouldn't be causing OOM Errors.</p>",
      "rawMarkdown": "Yeah, that's correct.\n\nAre you using Kaggle notebooks? I'm quite sure reading 10k images here shouldn't be causing OOM Errors.",
      "votes": null
    },
    {
      "id": "1243555",
      "postDate": "03/18/2021 10:07:00",
      "content": "<p>Ah, I've been using my own machine. I'll see if the notebooks will run Julia, or if not, I guess I may have to go back to Python.</p>",
      "rawMarkdown": "Ah, I've been using my own machine. I'll see if the notebooks will run Julia, or if not, I guess I may have to go back to Python.",
      "votes": null
    },
    {
      "id": "1243596",
      "postDate": "03/18/2021 10:50:30",
      "content": "<p>Oh, best of luck in either case! Do let me know if you have any further queries.</p>",
      "rawMarkdown": "Oh, best of luck in either case! Do let me know if you have any further queries.",
      "votes": null
    },
    {
      "id": "1243967",
      "postDate": "03/18/2021 16:05:36",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>, I made a notebook that shows you how to run Julia on Kaggle which may be helpful for you.</p>\n<p><a href=\"https://www.kaggle.com/marketneutral/julia-live-on-kaggle\" target=\"_blank\">https://www.kaggle.com/marketneutral/julia-live-on-kaggle</a></p>",
      "rawMarkdown": "abdurrafae, I made a notebook that shows you how to run Julia on Kaggle which may be helpful for you.\n\nhttps://www.kaggle.com/marketneutral/julia-live-on-kaggle",
      "votes": null
    },
    {
      "id": "1248066",
      "postDate": "03/22/2021 10:10:43",
      "content": "<p>There seems to be a large number of papers out there on counting (e.g. number of people in a crowd, number of cells in a microscopic image etc.). I guess one interesting thing here is that we don't have bounding boxes/segmentation for individual atoms available as an annotation (but perhaps that not that hard to auto-generate?). So, using CNNs for this type of regression task seems natural enough, but figuring out what out of all this literature is truly useful seems like a key step.</p>\n<p>Did you use a pre-trained neural network (it sounds like a custom architecture, so I assume not?)? I'd imagine pre-training on something like ImageNet should help with recognizing letters (given <a href=\"https://www.theverge.com/2021/3/8/22319173/openai-machine-vision-adversarial-typographic-attacka-clip-multimodal-neuron\" target=\"_blank\">recent \"funny\" adversarial examples</a>), especially if you work with such a small dataset for your initial experiments. Why not experiment with a small pre-trained model like a ResNet-18 or so?</p>\n<p>The comments by others seem to highlight some other obvious questions, but I also wondered about using the sum function. Would that not treat 10 weakish activations in different regions of the image as the same as a single strong activation in a single location (as long as they sum to the same thing)? That seems non-ideal to me. You may want to instead go straight to a single output (that can form any linear combination of what comes out of the pooling layer) or have some linear (+perhaps dropout, ReLU, BatchNorm etc. layers) before the output. Also, is mean squared error the obvious choice, especially when you let your predictions even be negative (perhaps exponentiating the output easily deals with that)?</p>",
      "rawMarkdown": "There seems to be a large number of papers out there on counting (e.g. number of people in a crowd, number of cells in a microscopic image etc.). I guess one interesting thing here is that we don't have bounding boxes/segmentation for individual atoms available as an annotation (but perhaps that not that hard to auto-generate?). So, using CNNs for this type of regression task seems natural enough, but figuring out what out of all this literature is truly useful seems like a key step.\n\nDid you use a pre-trained neural network (it sounds like a custom architecture, so I assume not?)? I'd imagine pre-training on something like ImageNet should help with recognizing letters (given [recent \"funny\" adversarial examples](https://www.theverge.com/2021/3/8/22319173/openai-machine-vision-adversarial-typographic-attacka-clip-multimodal-neuron)), especially if you work with such a small dataset for your initial experiments. Why not experiment with a small pre-trained model like a ResNet-18 or so?\n\nThe comments by others seem to highlight some other obvious questions, but I also wondered about using the sum function. Would that not treat 10 weakish activations in different regions of the image as the same as a single strong activation in a single location (as long as they sum to the same thing)? That seems non-ideal to me. You may want to instead go straight to a single output (that can form any linear combination of what comes out of the pooling layer) or have some linear (+perhaps dropout, ReLU, BatchNorm etc. layers) before the output. Also, is mean squared error the obvious choice, especially when you let your predictions even be negative (perhaps exponentiating the output easily deals with that)?",
      "votes": null
    },
    {
      "id": "1248068",
      "postDate": "03/22/2021 10:13:39",
      "content": "<p>Regarding out of memory: Are you loading all training data into memory at once? Esp. once you use more data, you really want to process small batches at a time. I don't know much about Julia, but I'd assume any sensible deep learning set-up would have the ability for training in batches.</p>",
      "rawMarkdown": "Regarding out of memory: Are you loading all training data into memory at once? Esp. once you use more data, you really want to process small batches at a time. I don't know much about Julia, but I'd assume any sensible deep learning set-up would have the ability for training in batches.",
      "votes": null
    },
    {
      "id": "1255151",
      "postDate": "03/28/2021 13:56:53",
      "content": "<p>Sorry for delay - I started a new job this week which took my attention. Yes, <a href=\"https://www.kaggle.com/marketneutral\" target=\"_blank\">@marketneutral</a> 's notebook was super helpful! </p>\n<p>I was doing batches of around 64 - when I reduced the batch size to 16 the memory errors stopped. </p>",
      "rawMarkdown": "Sorry for delay - I started a new job this week which took my attention. Yes, @marketneutral 's notebook was super helpful! \n\nI was doing batches of around 64 - when I reduced the batch size to 16 the memory errors stopped.",
      "votes": null
    },
    {
      "id": "1255224",
      "postDate": "03/28/2021 15:28:03",
      "content": "<p>A lot of the counting problems use Poisson loss, but that didn't seem appropriate here because this isn't a random assortment of stuff that might include oxygen atoms, it's molecules with structure. I haven't been pre-training but I can look into that. </p>\n<p>I did make some progress last week. I read a suggestion somewhere to get it working on a very small training set, and then expand and adjust as needed to generalize. So I cut down to training on 10 images in 2 batches of 5. </p>\n<p>I added a Dense layer after the sum. </p>\n<p>The next thing I realized was that I needed to train for more than a single epoch. Since my training set was now much smaller, I was able to experiment with that and found that by about epoch 40, the training error mostly stabilized, and by epoch 75 it was pretty much converged. </p>\n<p>The problem was, it converged to always predicting the mean.</p>\n<p>I tried combinations of optimizers and learning rates - found that AdaGrad converged marginally faster than Adam, but SGD with and without momentum gave me NaNs. I decided not to pursue that further and just stuck with Adam. </p>\n<p>I found something online that talked about dying ReLUs, and experimented with that a bit. Changing the ReLUs for the last block into leaky ReLUs worked - it was predicting different values for different images! I then tweaked some more - put a normalization layer on the front (although I'd think that <em>shouldn't</em> be necessary?) and a batch normalization layer after the first pooling layer, and that got me pretty good results on my 10 training images. </p>\n<p>My next plan is to generalize it and get it working on larger sets, do some cross-validation on hyperparameters, proper train/test split, etc. And then do the other letters. </p>",
      "rawMarkdown": "A lot of the counting problems use Poisson loss, but that didn't seem appropriate here because this isn't a random assortment of stuff that might include oxygen atoms, it's molecules with structure. I haven't been pre-training but I can look into that. \n\nI did make some progress last week. I read a suggestion somewhere to get it working on a very small training set, and then expand and adjust as needed to generalize. So I cut down to training on 10 images in 2 batches of 5. \n\nI added a Dense layer after the sum. \n\nThe next thing I realized was that I needed to train for more than a single epoch. Since my training set was now much smaller, I was able to experiment with that and found that by about epoch 40, the training error mostly stabilized, and by epoch 75 it was pretty much converged. \n\nThe problem was, it converged to always predicting the mean.\n\nI tried combinations of optimizers and learning rates - found that AdaGrad converged marginally faster than Adam, but SGD with and without momentum gave me NaNs. I decided not to pursue that further and just stuck with Adam. \n\nI found something online that talked about dying ReLUs, and experimented with that a bit. Changing the ReLUs for the last block into leaky ReLUs worked - it was predicting different values for different images! I then tweaked some more - put a normalization layer on the front (although I'd think that *shouldn't* be necessary?) and a batch normalization layer after the first pooling layer, and that got me pretty good results on my 10 training images. \n\nMy next plan is to generalize it and get it working on larger sets, do some cross-validation on hyperparameters, proper train/test split, etc. And then do the other letters.",
      "votes": null
    },
    {
      "id": "2772397",
      "postDate": "04/24/2024 17:46:10",
      "content": "<p>Yeah, that's correct.</p>\n<p>Are you using Kaggle notebooks?</p>",
      "rawMarkdown": "Yeah, that's correct.\n\nAre you using Kaggle notebooks?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1241937,
      "author_name": "abdurrafae",
      "author_url": "",
      "post_date": "03/17/2021 09:39:00",
      "content": "<p>Hi Jeremy,<br>\nI'll try to summarize what I've understood from your experiments. The approach to do a regression for each atom type is ok as long as you have a technique in mind to combine all of them together eventually.</p>\n<ol>\n<li><p>You are padding them all to be the same size, so assuming you get them all to 512x512 size. </p></li>\n<li><p>I think the issue is that the <em>receptive region</em>* of the conv net that you've defined is too low at the moment. I believe doing more convolutions with smaller kernel sizes would be better than doing 2 convolutions with 10x10 filter sizes. Currently after all the conv+pool layers your dimensions are somewhere around (121, 121). Try something like either 3 or 4 blocks of  (Conv (3x3), Conv (3x3), Max_pool). </p></li>\n<li><p>I would suggest that you sum along the height and width dimensions after the CNNs and then 1-2 dense layers to get the output. You'll need to flatten the output of this summing layer before using the dense layers.</p></li>\n<li><p>Since you are doing an aggregation after the CNNs layers (sum here), you can also try doing it on original image sizes provided that you can code you model to be dynamic enough to deal with sized inputs. </p></li>\n<li><p>Check if your model is underfitting or not, you need more data in any case if you want to debug the issue at hand. Try to cover more training data in your experiments. Like checking error after training on 5,000 images on same model architecture. If the error is reduced significantly, then the issue is with the training size. If the error remains in similar range, then the model architecture may be an issue.</p></li>\n</ol>\n<p>*<em>Here I mean the receptive region of a neuron in the last CNN layer.</em></p>",
      "votes": null,
      "replies": [
        {
          "id": 1242561,
          "author_name": "jezsadler",
          "author_url": "",
          "post_date": "03/17/2021 17:23:15",
          "content": "<p>Thank you, this is very helpful! </p>\n<p>To make sure I understand what you're saying conceptually, is this correct?: </p>\n<ul>\n<li>the images going in each have HeightxWidthx1 Channel</li>\n<li>as I convolve and pool, the height and width decrease, and the number of channels increases to N</li>\n<li>the summing layer collapses the height and width, leaving each image as height 1, width 1, N channels</li>\n<li>the dense layer (or layers) takes the N channels and reduces them to 1</li>\n</ul>\n<p>Trying to read in all 9550 images that start with \"00\" was giving me OutOfMemory errors, so I have been fiddling with ways of working around that. I've increased the learning rate (using Adam with eta 0.1 instead of 0.001) so that it doesn't take forever to evaluate model changes. (It's still taking a pretty long time to run.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242778,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "03/17/2021 19:41:37",
          "content": "<p>Yeah, that's correct.</p>\n<p>Are you using Kaggle notebooks? I'm quite sure reading 10k images here shouldn't be causing OOM Errors.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243555,
          "author_name": "jezsadler",
          "author_url": "",
          "post_date": "03/18/2021 10:07:00",
          "content": "<p>Ah, I've been using my own machine. I'll see if the notebooks will run Julia, or if not, I guess I may have to go back to Python.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243596,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "03/18/2021 10:50:30",
          "content": "<p>Oh, best of luck in either case! Do let me know if you have any further queries.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243967,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "03/18/2021 16:05:36",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>, I made a notebook that shows you how to run Julia on Kaggle which may be helpful for you.</p>\n<p><a href=\"https://www.kaggle.com/marketneutral/julia-live-on-kaggle\" target=\"_blank\">https://www.kaggle.com/marketneutral/julia-live-on-kaggle</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1248068,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/22/2021 10:13:39",
          "content": "<p>Regarding out of memory: Are you loading all training data into memory at once? Esp. once you use more data, you really want to process small batches at a time. I don't know much about Julia, but I'd assume any sensible deep learning set-up would have the ability for training in batches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255151,
          "author_name": "jezsadler",
          "author_url": "",
          "post_date": "03/28/2021 13:56:53",
          "content": "<p>Sorry for delay - I started a new job this week which took my attention. Yes, <a href=\"https://www.kaggle.com/marketneutral\" target=\"_blank\">@marketneutral</a> 's notebook was super helpful! </p>\n<p>I was doing batches of around 64 - when I reduced the batch size to 16 the memory errors stopped. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1248066,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/22/2021 10:10:43",
      "content": "<p>There seems to be a large number of papers out there on counting (e.g. number of people in a crowd, number of cells in a microscopic image etc.). I guess one interesting thing here is that we don't have bounding boxes/segmentation for individual atoms available as an annotation (but perhaps that not that hard to auto-generate?). So, using CNNs for this type of regression task seems natural enough, but figuring out what out of all this literature is truly useful seems like a key step.</p>\n<p>Did you use a pre-trained neural network (it sounds like a custom architecture, so I assume not?)? I'd imagine pre-training on something like ImageNet should help with recognizing letters (given <a href=\"https://www.theverge.com/2021/3/8/22319173/openai-machine-vision-adversarial-typographic-attacka-clip-multimodal-neuron\" target=\"_blank\">recent \"funny\" adversarial examples</a>), especially if you work with such a small dataset for your initial experiments. Why not experiment with a small pre-trained model like a ResNet-18 or so?</p>\n<p>The comments by others seem to highlight some other obvious questions, but I also wondered about using the sum function. Would that not treat 10 weakish activations in different regions of the image as the same as a single strong activation in a single location (as long as they sum to the same thing)? That seems non-ideal to me. You may want to instead go straight to a single output (that can form any linear combination of what comes out of the pooling layer) or have some linear (+perhaps dropout, ReLU, BatchNorm etc. layers) before the output. Also, is mean squared error the obvious choice, especially when you let your predictions even be negative (perhaps exponentiating the output easily deals with that)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1255224,
          "author_name": "jezsadler",
          "author_url": "",
          "post_date": "03/28/2021 15:28:03",
          "content": "<p>A lot of the counting problems use Poisson loss, but that didn't seem appropriate here because this isn't a random assortment of stuff that might include oxygen atoms, it's molecules with structure. I haven't been pre-training but I can look into that. </p>\n<p>I did make some progress last week. I read a suggestion somewhere to get it working on a very small training set, and then expand and adjust as needed to generalize. So I cut down to training on 10 images in 2 batches of 5. </p>\n<p>I added a Dense layer after the sum. </p>\n<p>The next thing I realized was that I needed to train for more than a single epoch. Since my training set was now much smaller, I was able to experiment with that and found that by about epoch 40, the training error mostly stabilized, and by epoch 75 it was pretty much converged. </p>\n<p>The problem was, it converged to always predicting the mean.</p>\n<p>I tried combinations of optimizers and learning rates - found that AdaGrad converged marginally faster than Adam, but SGD with and without momentum gave me NaNs. I decided not to pursue that further and just stuck with Adam. </p>\n<p>I found something online that talked about dying ReLUs, and experimented with that a bit. Changing the ReLUs for the last block into leaky ReLUs worked - it was predicting different values for different images! I then tweaked some more - put a normalization layer on the front (although I'd think that <em>shouldn't</em> be necessary?) and a batch normalization layer after the first pooling layer, and that got me pretty good results on my 10 training images. </p>\n<p>My next plan is to generalize it and get it working on larger sets, do some cross-validation on hyperparameters, proper train/test split, etc. And then do the other letters. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2772397,
      "author_name": "sonialikhan",
      "author_url": "",
      "post_date": "04/24/2024 17:46:10",
      "content": "<p>Yeah, that's correct.</p>\n<p>Are you using Kaggle notebooks?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1239482": "Hi! It's my first Kaggle contest, and I'm having some difficulty that I suspect is due to some basic error. \n\nI'm trying to build a CNN to give me a count of how many atoms of a given element are in each molecule - starting with oxygen. I'm treating it as a regression problem rather than classification.\n\n- Load in all the images starting \"000\", which is about 592. (I know I'll need to do more eventually but while I'm still fiddling around so much I want to use a more manageable set.)\n- Trimmed whitespace around the molecules, and then padded the images back out to get to uniform size.\n- Parsed out the InChI labels to get the # of oxygen atoms in each molecule.\n- Built a CNN with the following layers:\n1. Convolution, 10x10, 1 input channel => 16 output channels, ReLu activation\n2. Max Pooling, 2x2, stride 2\n3. Convolution, 10x10, 16 input channels => 32 output channels, ReLu activation\n4. Max Pooling, 2x2, stride 2\n5. sum\n- Defined loss function as mean squared error\n- Trained on all 592 images against their oxygen counts (I know eventually I will need to do train/test splitting to avoid overfitting, but for now I'm just trying to get something working)\n- Compare predictions on the training data with actuals:\n![https://photos.app.goo.gl/gusw7GmvmJSztdo28](url to embed)\nWell, that's not good. That's pretty much just noise. My fear with a small set was overfitting, and this is not that. Correlation coefficient between the predictions and actuals is 0.328. \n\nI'm not sure where I'm going wrong. Thoughts are:\n1. I generally haven't seen CNNs used for counting by regression - is this a plausible approach or should I classify instead? Seems like I could end up with lots and lots of classes if I do that.\n2. If I am regressing, is the sum() at the end the right way to express that? Or do I need a dense layer instead? \n3. Do I have the right # of layers in my CNN? How do I figure that out?\n4. Do I need to configure them differently? Convolution sizes and channels and strides and stuff?\n5. Do I just need to use more of the training data to get a meaningful result?\n6. Is padding the images to be the same size the right approach, or should I scale them down instead? Or is there some way of managing them being different sizes?\n7. Is there a bug somewhere in my code, such that it's not doing what I think it's doing? (I've been working in Julia instead of Python - happy to share the code if this is the most likely issue.)\n8. Are CNNs just the wrong tool for this portion of the problem?\n\nI have approaches that I could take towards exploring each of these more, but they are each time consuming, and I'm concerned about putting a lot of effort into fixing the wrong issue. (e.g. I could do cross-validation over a bunch of different convolution parameters, but if the fundamental problem is that my training set is too small, that could be hours thrown away for no benefit.) I'd be very grateful if anyone can spot a fundamental error in what I'm doing here.",
    "1241937": "Hi Jeremy,\nI'll try to summarize what I've understood from your experiments. The approach to do a regression for each atom type is ok as long as you have a technique in mind to combine all of them together eventually.\n\n\n1. You are padding them all to be the same size, so assuming you get them all to 512x512 size. \n\n1. I think the issue is that the *receptive region** of the conv net that you've defined is too low at the moment. I believe doing more convolutions with smaller kernel sizes would be better than doing 2 convolutions with 10x10 filter sizes. Currently after all the conv+pool layers your dimensions are somewhere around (121, 121). Try something like either 3 or 4 blocks of  (Conv (3x3), Conv (3x3), Max_pool). \n\n1. I would suggest that you sum along the height and width dimensions after the CNNs and then 1-2 dense layers to get the output. You'll need to flatten the output of this summing layer before using the dense layers.\n\n1. Since you are doing an aggregation after the CNNs layers (sum here), you can also try doing it on original image sizes provided that you can code you model to be dynamic enough to deal with sized inputs. \n\n1. Check if your model is underfitting or not, you need more data in any case if you want to debug the issue at hand. Try to cover more training data in your experiments. Like checking error after training on 5,000 images on same model architecture. If the error is reduced significantly, then the issue is with the training size. If the error remains in similar range, then the model architecture may be an issue.\n\n\n**Here I mean the receptive region of a neuron in the last CNN layer.*",
    "1242561": "Thank you, this is very helpful! \n\nTo make sure I understand what you're saying conceptually, is this correct?: \n- the images going in each have HeightxWidthx1 Channel\n- as I convolve and pool, the height and width decrease, and the number of channels increases to N\n- the summing layer collapses the height and width, leaving each image as height 1, width 1, N channels\n- the dense layer (or layers) takes the N channels and reduces them to 1\n\nTrying to read in all 9550 images that start with \"00\" was giving me OutOfMemory errors, so I have been fiddling with ways of working around that. I've increased the learning rate (using Adam with eta 0.1 instead of 0.001) so that it doesn't take forever to evaluate model changes. (It's still taking a pretty long time to run.)",
    "1242778": "Yeah, that's correct.\n\nAre you using Kaggle notebooks? I'm quite sure reading 10k images here shouldn't be causing OOM Errors.",
    "1243555": "Ah, I've been using my own machine. I'll see if the notebooks will run Julia, or if not, I guess I may have to go back to Python.",
    "1243596": "Oh, best of luck in either case! Do let me know if you have any further queries.",
    "1243967": "abdurrafae, I made a notebook that shows you how to run Julia on Kaggle which may be helpful for you.\n\nhttps://www.kaggle.com/marketneutral/julia-live-on-kaggle",
    "1248066": "There seems to be a large number of papers out there on counting (e.g. number of people in a crowd, number of cells in a microscopic image etc.). I guess one interesting thing here is that we don't have bounding boxes/segmentation for individual atoms available as an annotation (but perhaps that not that hard to auto-generate?). So, using CNNs for this type of regression task seems natural enough, but figuring out what out of all this literature is truly useful seems like a key step.\n\nDid you use a pre-trained neural network (it sounds like a custom architecture, so I assume not?)? I'd imagine pre-training on something like ImageNet should help with recognizing letters (given [recent \"funny\" adversarial examples](https://www.theverge.com/2021/3/8/22319173/openai-machine-vision-adversarial-typographic-attacka-clip-multimodal-neuron)), especially if you work with such a small dataset for your initial experiments. Why not experiment with a small pre-trained model like a ResNet-18 or so?\n\nThe comments by others seem to highlight some other obvious questions, but I also wondered about using the sum function. Would that not treat 10 weakish activations in different regions of the image as the same as a single strong activation in a single location (as long as they sum to the same thing)? That seems non-ideal to me. You may want to instead go straight to a single output (that can form any linear combination of what comes out of the pooling layer) or have some linear (+perhaps dropout, ReLU, BatchNorm etc. layers) before the output. Also, is mean squared error the obvious choice, especially when you let your predictions even be negative (perhaps exponentiating the output easily deals with that)?",
    "1248068": "Regarding out of memory: Are you loading all training data into memory at once? Esp. once you use more data, you really want to process small batches at a time. I don't know much about Julia, but I'd assume any sensible deep learning set-up would have the ability for training in batches.",
    "1255151": "Sorry for delay - I started a new job this week which took my attention. Yes, @marketneutral 's notebook was super helpful! \n\nI was doing batches of around 64 - when I reduced the batch size to 16 the memory errors stopped.",
    "1255224": "A lot of the counting problems use Poisson loss, but that didn't seem appropriate here because this isn't a random assortment of stuff that might include oxygen atoms, it's molecules with structure. I haven't been pre-training but I can look into that. \n\nI did make some progress last week. I read a suggestion somewhere to get it working on a very small training set, and then expand and adjust as needed to generalize. So I cut down to training on 10 images in 2 batches of 5. \n\nI added a Dense layer after the sum. \n\nThe next thing I realized was that I needed to train for more than a single epoch. Since my training set was now much smaller, I was able to experiment with that and found that by about epoch 40, the training error mostly stabilized, and by epoch 75 it was pretty much converged. \n\nThe problem was, it converged to always predicting the mean.\n\nI tried combinations of optimizers and learning rates - found that AdaGrad converged marginally faster than Adam, but SGD with and without momentum gave me NaNs. I decided not to pursue that further and just stuck with Adam. \n\nI found something online that talked about dying ReLUs, and experimented with that a bit. Changing the ReLUs for the last block into leaky ReLUs worked - it was predicting different values for different images! I then tweaked some more - put a normalization layer on the front (although I'd think that *shouldn't* be necessary?) and a batch normalization layer after the first pooling layer, and that got me pretty good results on my 10 training images. \n\nMy next plan is to generalize it and get it working on larger sets, do some cross-validation on hyperparameters, proper train/test split, etc. And then do the other letters.",
    "2772397": "Yeah, that's correct.\n\nAre you using Kaggle notebooks?"
  },
  "source": "meta"
}