{
  "id": 122398,
  "title": "New to Machine Learning or Kaggle?",
  "url": "/competitions/bengaliai-cv19/discussion/122398",
  "author_name": "Addison Howard",
  "post_date": "2019-12-19T23:57:22.344000",
  "votes": 8,
  "comment_count": 29,
  "views": 0,
  "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n\n<p>New to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\">how to enter a competition using Kaggle Notebooks</a>.</p>",
  "messages": [
    {
      "id": 698949,
      "postDate": "2019-12-19T23:57:22.343Z",
      "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n\n<p>New to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\">how to enter a competition using Kaggle Notebooks</a>.</p>",
      "rawMarkdown": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ).",
      "votes": 7
    },
    {
      "id": 721157,
      "postDate": "2020-01-17T05:33:41.120Z",
      "content": "<p>Hi,</p>\n\n<p>In competitions like this, when people say they are using Resnet, Densenet, Inception, etc. are they using the pretrained model and training dense layers on top or training the architecture from scratch?</p>\n\n<p>It seems infeasible to train from scratch on kaggle kernels?</p>",
      "rawMarkdown": "Hi,\n\nIn competitions like this, when people say they are using Resnet, Densenet, Inception, etc. are they using the pretrained model and training dense layers on top or training the architecture from scratch?\n\nIt seems infeasible to train from scratch on kaggle kernels?",
      "votes": 1,
      "replies": [
        {
          "id": 727100,
          "postDate": "2020-01-23T13:09:45.710Z",
          "content": "<p>Hey <a href=\"/nathanpw\">@nathanpw</a> ,</p>\n\n<p>They are indeed using the pretrained convolutional layers. The choice of the dataset on which it has been pretrained is a matter of choice (usually <strong>ImageNet</strong>). </p>\n\n<p>They just get rid of the fully connected layers (also called top layers) and setup their own that they train for scratch as indeed, training from scratch a whole model is way too hard for this task. </p>\n\n<p>If you use an already existing architecture like <strong>ResNet</strong>, you could design your own CNN architecture, not too deep as a starter submission to give you a baseline score. Hence, this CNN could be trained from scratch. </p>\n\n<p>I hope I helped</p>",
          "rawMarkdown": "Hey @nathanpw ,\n\nThey are indeed using the pretrained convolutional layers. The choice of the dataset on which it has been pretrained is a matter of choice (usually **ImageNet**). \n\nThey just get rid of the fully connected layers (also called top layers) and setup their own that they train for scratch as indeed, training from scratch a whole model is way too hard for this task. \n\nIf you use an already existing architecture like **ResNet**, you could design your own CNN architecture, not too deep as a starter submission to give you a baseline score. Hence, this CNN could be trained from scratch. \n\nI hope I helped",
          "votes": 3
        },
        {
          "id": 734557,
          "postDate": "2020-02-01T16:49:06.750Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 737069,
          "postDate": "2020-02-04T21:53:31.930Z",
          "content": "<p>Hi Esteban,\nI've done a bit of experimenting since i asked this question and think i have the answers. People generally use models with weights trained on imagenet, don't freeze any layers (i.e. all layers are trainable) and retrain the entire model on the new data.</p>\n\n<p>It's really easy with Keras, you just need to import from keras.applications (see: <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>). If you google something like \"pretrained keras models transfer learning\" you'll find a tonne of examples. </p>\n\n<p>I'm using Keras for my models so you can look through my notebooks if you like. In saying that, my code is junk and there are a lot of better examples online. </p>\n\n<p>good luck :) </p>",
          "rawMarkdown": "Hi Esteban,\nI've done a bit of experimenting since i asked this question and think i have the answers. People generally use models with weights trained on imagenet, don't freeze any layers (i.e. all layers are trainable) and retrain the entire model on the new data.\n\nIt's really easy with Keras, you just need to import from keras.applications (see: https://keras.io/applications/). If you google something like \"pretrained keras models transfer learning\" you'll find a tonne of examples. \n\nI'm using Keras for my models so you can look through my notebooks if you like. In saying that, my code is junk and there are a lot of better examples online. \n\ngood luck :) ",
          "votes": 1
        },
        {
          "id": 737326,
          "postDate": "2020-02-05T07:26:31.120Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 708693,
      "postDate": "2020-01-02T14:58:49.933Z",
      "content": "<p>Hello,</p>\n\n<p>This might seem like a really noobish question but I'm new at machine learning and since this topic is specifically for this maybe someone can explain some stuff to me :).\nWhen training in the beginning the accuracy starts growing slowly with each batch (like it goes 0.100 -&gt; 0.110 -&gt; 0.1120 and so on) also the loss goes down slowly. This is ok, it's what I expect to happen. However at the end of the epoch there is a big jump for both the accuracy as well as the loss function (the accuracy goes up and the loss goes down). I'm not complaining as this seems good, but I don't understand exactly why this is happening.\nFrom what I know training is basically for each batch do forwardprop and then backprop. When the epoch ends check the validation set and then start again. Is there some training step at the end of the epoch or something that's causing the accuracy to improve a lot at the end of the epoch?</p>\n\n<p>[Later Edit]\nHere is an example of how the keras output looks at the end of the epoch:</p>\n\n<p>```\n383/390 [============================&gt;.] - ETA: 11s - loss: 0.7742 - accuracy: 0.7450\n384/390 [============================&gt;.] - ETA: 9s - loss: 0.7729 - accuracy: 0.7455 \n385/390 [============================&gt;.] - ETA: 8s - loss: 0.7716 - accuracy: 0.7460\n386/390 [============================&gt;.] - ETA: 6s - loss: 0.7705 - accuracy: 0.7464\n387/390 [============================&gt;.] - ETA: 4s - loss: 0.7693 - accuracy: 0.7468\n388/390 [============================&gt;.] - ETA: 3s - loss: 0.7682 - accuracy: 0.7472\n389/390 [============================&gt;.] - ETA: 1s - loss: 0.7671 - accuracy: 0.7476</p>\n\n<p>Epoch 00001: val_accuracy improved from -inf to 0.87509, saving model to data/consonantd_model.h5</p>\n\n<p>390/390 [==============================] - 638s 2s/step - loss: 0.7659 - accuracy: 0.7481 - val_loss: 0.0520 - val_accuracy: 0.8751\nEpoch 2/15</p>\n\n<p>1/390 [..............................] - ETA: 4:41 - loss: 0.2990 - accuracy: 0.9219\n  2/390 [..............................] - ETA: 4:42 - loss: 0.2972 - accuracy: 0.9180\n  3/390 [..............................] - ETA: 4:42 - loss: 0.2933 - accuracy: 0.9206\n  4/390 [..............................] - ETA: 4:41 - loss: 0.3102 - accuracy: 0.9160\n```</p>",
      "rawMarkdown": "Hello,\n\nThis might seem like a really noobish question but I'm new at machine learning and since this topic is specifically for this maybe someone can explain some stuff to me :).\nWhen training in the beginning the accuracy starts growing slowly with each batch (like it goes 0.100 -&gt; 0.110 -&gt; 0.1120 and so on) also the loss goes down slowly. This is ok, it's what I expect to happen. However at the end of the epoch there is a big jump for both the accuracy as well as the loss function (the accuracy goes up and the loss goes down). I'm not complaining as this seems good, but I don't understand exactly why this is happening.\nFrom what I know training is basically for each batch do forwardprop and then backprop. When the epoch ends check the validation set and then start again. Is there some training step at the end of the epoch or something that's causing the accuracy to improve a lot at the end of the epoch?\n\n[Later Edit]\nHere is an example of how the keras output looks at the end of the epoch:\n\n```\n383/390 [============================&gt;.] - ETA: 11s - loss: 0.7742 - accuracy: 0.7450\n384/390 [============================&gt;.] - ETA: 9s - loss: 0.7729 - accuracy: 0.7455 \n385/390 [============================&gt;.] - ETA: 8s - loss: 0.7716 - accuracy: 0.7460\n386/390 [============================&gt;.] - ETA: 6s - loss: 0.7705 - accuracy: 0.7464\n387/390 [============================&gt;.] - ETA: 4s - loss: 0.7693 - accuracy: 0.7468\n388/390 [============================&gt;.] - ETA: 3s - loss: 0.7682 - accuracy: 0.7472\n389/390 [============================&gt;.] - ETA: 1s - loss: 0.7671 - accuracy: 0.7476\n\nEpoch 00001: val_accuracy improved from -inf to 0.87509, saving model to data/consonantd_model.h5\n\n390/390 [==============================] - 638s 2s/step - loss: 0.7659 - accuracy: 0.7481 - val_loss: 0.0520 - val_accuracy: 0.8751\nEpoch 2/15\n\n  1/390 [..............................] - ETA: 4:41 - loss: 0.2990 - accuracy: 0.9219\n  2/390 [..............................] - ETA: 4:42 - loss: 0.2972 - accuracy: 0.9180\n  3/390 [..............................] - ETA: 4:42 - loss: 0.2933 - accuracy: 0.9206\n  4/390 [..............................] - ETA: 4:41 - loss: 0.3102 - accuracy: 0.9160\n```",
      "votes": 1,
      "replies": [
        {
          "id": 709345,
          "postDate": "2020-01-03T11:50:20.113Z",
          "content": "<p>I may be wrong and it may be possible to correct me but from my understanding here is what happens:</p>\n\n<p>To evaluate loss during training, Keras calculate an average loss (same for accuracy measure) over all the already seen batches (explained <a href=\"https://github.com/keras-team/keras/issues/10426\">here</a>). So it can explain what you observe in this way:</p>\n\n<p>Epoch 1/15</p>\n\n<p>1/390 [..............................] - ETA: 4:41 - loss: X1 - accuracy: Y1</p>\n\n<p>Here, we have the first batch of data, our <strong>network will perform awfully on it</strong>, resulting in the <strong>highest loss and accuracy from all the possible batches</strong> of all the epochs of the training session. But the <strong>network will start learning and it will start learning fast</strong>, even more at the start of the training. The batches passes by and we arrive at the last batches of the first epoch:</p>\n\n<p>387/390 [============================&gt;.] - ETA: 4s - loss: X_387 - accuracy: Y_387\n388/390 [============================&gt;.] - ETA: 3s - loss: X_388 - accuracy: Y_388\n389/390 [============================&gt;.] - ETA: 1s - loss: X_389 - accuracy: Y_389</p>\n\n<p>At this point, the network should perform way better and on batches 387, 388 and 389, the loss calculated for each of these batch, independently, will be much lower than for the loss of <em>X1</em>.</p>\n\n<p>However, the value of X_387 is calculated as follows:</p>\n\n<p>$$ \\dfrac{1}{n} * \\sum_{i=1}^{n} \\mathbb{B(i)} $$</p>\n\n<p>where B(i) is the loss value of the batch number <em>i</em> and n equals 387.\nHence, when you arrive at the X_387 value, to evaluate it, you do the mean of the batches loss from the 1st batch to the last. However, the first is sooo bad compared to your last batches (hence so high), that it biased the overall mean of the epoch. \nThis is a weird choice of implementation to show this accumulative mean and not the loss of the batch 387.\nHowever, you can justify it by the fact that you don't care how your network behave on a subsample of data, you want its performance to be the best in a general situation =&gt; <strong>use of a mean</strong>.</p>\n\n<p>So in your situation, you could guess that, on your last batches of your 1st epoch, your network had loss similar to the first batches of the second epoch. \nYou have to note that this behaviour will always exist throughout the training but should be less and less noticeable as batches pass by !</p>\n\n<p>I hope that I solved your question, if you have any more interrogation, don't mind asking !</p>\n\n<p>I wish you good luck for this competition and for your future in Machine Learning ;) !</p>",
          "rawMarkdown": "I may be wrong and it may be possible to correct me but from my understanding here is what happens:\n\nTo evaluate loss during training, Keras calculate an average loss (same for accuracy measure) over all the already seen batches (explained [here](https://github.com/keras-team/keras/issues/10426)). So it can explain what you observe in this way:\n\nEpoch 1/15\n\n1/390 [..............................] - ETA: 4:41 - loss: X1 - accuracy: Y1\n\nHere, we have the first batch of data, our **network will perform awfully on it**, resulting in the **highest loss and accuracy from all the possible batches** of all the epochs of the training session. But the **network will start learning and it will start learning fast**, even more at the start of the training. The batches passes by and we arrive at the last batches of the first epoch:\n\n387/390 [============================&gt;.] - ETA: 4s - loss: X_387 - accuracy: Y_387\n388/390 [============================&gt;.] - ETA: 3s - loss: X_388 - accuracy: Y_388\n389/390 [============================&gt;.] - ETA: 1s - loss: X_389 - accuracy: Y_389\n\nAt this point, the network should perform way better and on batches 387, 388 and 389, the loss calculated for each of these batch, independently, will be much lower than for the loss of *X1*.\n\nHowever, the value of X_387 is calculated as follows:\n\n$$ \\dfrac{1}{n} * \\sum_{i=1}^{n} \\mathbb{B(i)} $$\n\nwhere B(i) is the loss value of the batch number *i* and n equals 387.\nHence, when you arrive at the X_387 value, to evaluate it, you do the mean of the batches loss from the 1st batch to the last. However, the first is sooo bad compared to your last batches (hence so high), that it biased the overall mean of the epoch. \nThis is a weird choice of implementation to show this accumulative mean and not the loss of the batch 387.\nHowever, you can justify it by the fact that you don't care how your network behave on a subsample of data, you want its performance to be the best in a general situation =&gt; **use of a mean**.\n\nSo in your situation, you could guess that, on your last batches of your 1st epoch, your network had loss similar to the first batches of the second epoch. \nYou have to note that this behaviour will always exist throughout the training but should be less and less noticeable as batches pass by !\n\nI hope that I solved your question, if you have any more interrogation, don't mind asking !\n\nI wish you good luck for this competition and for your future in Machine Learning ;) !",
          "votes": 2
        },
        {
          "id": 709727,
          "postDate": "2020-01-03T21:52:51.010Z",
          "content": "<p>Oh I see. Thank you very much. Indeed I was under the impression that the loss and accuracy were the from that specific mini_batch not cumulative across the entire epoch. \nThis does indeed explain the behaviour I'm seeing and it makes sense. Indeed this behaviour becomes less noticeable for higher epochs.\nAgain thanks for the explanation kind sir.</p>",
          "rawMarkdown": "Oh I see. Thank you very much. Indeed I was under the impression that the loss and accuracy were the from that specific mini_batch not cumulative across the entire epoch. \nThis does indeed explain the behaviour I'm seeing and it makes sense. Indeed this behaviour becomes less noticeable for higher epochs.\nAgain thanks for the explanation kind sir.",
          "votes": 1
        },
        {
          "id": 762579,
          "postDate": "2020-03-03T16:10:56.413Z",
          "content": "<p>Thanks for you answer~</p>",
          "rawMarkdown": "Thanks for you answer~"
        }
      ]
    },
    {
      "id": 765506,
      "postDate": "2020-03-06T17:58:32.773Z",
      "content": "<p>Hello, <a href=\"/addisonhoward\">@addisonhoward</a> ! Not completely new, but definitely a beginner.</p>\n\n<p>Can you explain how does one solve this \"Notebook Exceeded Allowed Compute\"? What does this really mean?\nThank you!</p>",
      "rawMarkdown": "Hello, @addisonhoward ! Not completely new, but definitely a beginner.\n\nCan you explain how does one solve this \"Notebook Exceeded Allowed Compute\"? What does this really mean?\nThank you!",
      "replies": [
        {
          "id": 766401,
          "postDate": "2020-03-08T05:06:00.610Z",
          "content": "<p>Hi João,</p>\n\n<p>Each Kaggle notebook has a present limit of computing power. Look <a href=\"https://www.kaggle.com/c/bengaliai-cv19/overview/notebooks-requirements\">here</a> for details. It sounds like your notebook may have run over the limits allowed.</p>",
          "rawMarkdown": "Hi João,\n\nEach Kaggle notebook has a present limit of computing power. Look [here](https://www.kaggle.com/c/bengaliai-cv19/overview/notebooks-requirements) for details. It sounds like your notebook may have run over the limits allowed.\n\n",
          "votes": 1
        },
        {
          "id": 766960,
          "postDate": "2020-03-09T01:30:07.807Z",
          "content": "<p>Thank you for the kind reply.</p>\n\n<p>I don't understand. My inference notebook runs in ~100 seconds with a GPU in my kaggles private area. However, the submission fails.</p>",
          "rawMarkdown": "Thank you for the kind reply.\n\nI don't understand. My inference notebook runs in ~100 seconds with a GPU in my kaggles private area. However, the submission fails."
        },
        {
          "id": 767402,
          "postDate": "2020-03-09T15:31:04.053Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4186383%2F7f9c8b664249961e5e3ede088207bc20%2Fsubmission.png?generation=1583767668140589&amp;alt=media\" alt=\"\"></p>\n\n<p>For example, this notebook gives me the error \"Notebook Exceeded Allowed Compute\". I don't understand why.</p>\n\n<p>Once again, thank you for your kind attention, <a href=\"/addisonhoward\">@addisonhoward</a> </p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4186383%2F7f9c8b664249961e5e3ede088207bc20%2Fsubmission.png?generation=1583767668140589&amp;alt=media)\n\nFor example, this notebook gives me the error \"Notebook Exceeded Allowed Compute\". I don't understand why.\n\nOnce again, thank you for your kind attention, @addisonhoward "
        }
      ]
    },
    {
      "id": 761781,
      "postDate": "2020-03-02T23:35:33.367Z",
      "content": "<p>Hi,how to load weights trained on multi gpu model to the single gpu model? </p>",
      "rawMarkdown": "Hi,how to load weights trained on multi gpu model to the single gpu model? "
    },
    {
      "id": 753675,
      "postDate": "2020-02-22T14:36:15.967Z",
      "content": "<p>Hi! I feel overwhelmed with the kind of progress other Kagglers have already made. How long do you think I will be a Master Kaggler?</p>",
      "rawMarkdown": "Hi! I feel overwhelmed with the kind of progress other Kagglers have already made. How long do you think I will be a Master Kaggler?"
    },
    {
      "id": 747197,
      "postDate": "2020-02-16T05:14:57.260Z",
      "content": "<p>As a beginner trying to self-teach himself in Natural Language processing and computer vision,what would you suggest as some good reference sites to explore these topics in addition to Kaggle? </p>",
      "rawMarkdown": "As a beginner trying to self-teach himself in Natural Language processing and computer vision,what would you suggest as some good reference sites to explore these topics in addition to Kaggle? "
    },
    {
      "id": 699620,
      "postDate": "2019-12-20T16:37:58.797Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4232764%2Fa5f7fe2224a0d90b502344708d194a8a%2FScreen%20Shot%202019-12-20%20at%2012.55.46.png?generation=1576859875129967&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4232764%2Fa5f7fe2224a0d90b502344708d194a8a%2FScreen%20Shot%202019-12-20%20at%2012.55.46.png?generation=1576859875129967&amp;alt=media)\n"
    },
    {
      "id": 760473,
      "postDate": "2020-03-01T10:47:24.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 747126,
      "postDate": "2020-02-16T01:46:39.003Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 735288,
      "postDate": "2020-02-02T20:26:00.173Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 733507,
      "postDate": "2020-01-31T08:27:32.923Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 735287,
          "postDate": "2020-02-02T20:24:03.473Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 742316,
          "postDate": "2020-02-11T08:23:43.693Z",
          "content": "<p><a href=\"/e60315\">@e60315</a> your reply is really helpful and also solved my question. But I wonder if you know how to upload the offline pretrained model or how to use it as an input to my kaggle kernel? Thanks a lot!</p>",
          "rawMarkdown": "@e60315 your reply is really helpful and also solved my question. But I wonder if you know how to upload the offline pretrained model or how to use it as an input to my kaggle kernel? Thanks a lot!"
        },
        {
          "id": 743969,
          "postDate": "2020-02-12T12:53:47.813Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 745015,
          "postDate": "2020-02-13T11:50:05.063Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 699593,
      "postDate": "2019-12-20T16:10:09.810Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 699219,
      "postDate": "2019-12-20T07:23:01.657Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": 1
    },
    {
      "id": 709944,
      "postDate": "2020-01-04T05:19:50.340Z",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!"
    },
    {
      "id": 700054,
      "postDate": "2019-12-21T11:16:45.757Z",
      "content": "<p>Thank you very much.</p>",
      "rawMarkdown": "Thank you very much."
    }
  ],
  "comments": [
    {
      "id": 721157,
      "author_name": "nathanpw",
      "author_url": "",
      "post_date": "2020-01-17T05:33:41.120000",
      "content": "<p>Hi,</p>\n\n<p>In competitions like this, when people say they are using Resnet, Densenet, Inception, etc. are they using the pretrained model and training dense layers on top or training the architecture from scratch?</p>\n\n<p>It seems infeasible to train from scratch on kaggle kernels?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 727100,
          "author_name": "Thomas Di Martino",
          "author_url": "",
          "post_date": "2020-01-23T13:09:45.710000",
          "content": "<p>Hey <a href=\"/nathanpw\">@nathanpw</a> ,</p>\n\n<p>They are indeed using the pretrained convolutional layers. The choice of the dataset on which it has been pretrained is a matter of choice (usually <strong>ImageNet</strong>). </p>\n\n<p>They just get rid of the fully connected layers (also called top layers) and setup their own that they train for scratch as indeed, training from scratch a whole model is way too hard for this task. </p>\n\n<p>If you use an already existing architecture like <strong>ResNet</strong>, you could design your own CNN architecture, not too deep as a starter submission to give you a baseline score. Hence, this CNN could be trained from scratch. </p>\n\n<p>I hope I helped</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 734557,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-01T16:49:06.750000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737069,
          "author_name": "nathanpw",
          "author_url": "",
          "post_date": "2020-02-04T21:53:31.930000",
          "content": "<p>Hi Esteban,\nI've done a bit of experimenting since i asked this question and think i have the answers. People generally use models with weights trained on imagenet, don't freeze any layers (i.e. all layers are trainable) and retrain the entire model on the new data.</p>\n\n<p>It's really easy with Keras, you just need to import from keras.applications (see: <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>). If you google something like \"pretrained keras models transfer learning\" you'll find a tonne of examples. </p>\n\n<p>I'm using Keras for my models so you can look through my notebooks if you like. In saying that, my code is junk and there are a lot of better examples online. </p>\n\n<p>good luck :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 737326,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T07:26:31.120000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 708693,
      "author_name": "Tiberiu Savin",
      "author_url": "",
      "post_date": "2020-01-02T14:58:49.933000",
      "content": "<p>Hello,</p>\n\n<p>This might seem like a really noobish question but I'm new at machine learning and since this topic is specifically for this maybe someone can explain some stuff to me :).\nWhen training in the beginning the accuracy starts growing slowly with each batch (like it goes 0.100 -&gt; 0.110 -&gt; 0.1120 and so on) also the loss goes down slowly. This is ok, it's what I expect to happen. However at the end of the epoch there is a big jump for both the accuracy as well as the loss function (the accuracy goes up and the loss goes down). I'm not complaining as this seems good, but I don't understand exactly why this is happening.\nFrom what I know training is basically for each batch do forwardprop and then backprop. When the epoch ends check the validation set and then start again. Is there some training step at the end of the epoch or something that's causing the accuracy to improve a lot at the end of the epoch?</p>\n\n<p>[Later Edit]\nHere is an example of how the keras output looks at the end of the epoch:</p>\n\n<p>```\n383/390 [============================&gt;.] - ETA: 11s - loss: 0.7742 - accuracy: 0.7450\n384/390 [============================&gt;.] - ETA: 9s - loss: 0.7729 - accuracy: 0.7455 \n385/390 [============================&gt;.] - ETA: 8s - loss: 0.7716 - accuracy: 0.7460\n386/390 [============================&gt;.] - ETA: 6s - loss: 0.7705 - accuracy: 0.7464\n387/390 [============================&gt;.] - ETA: 4s - loss: 0.7693 - accuracy: 0.7468\n388/390 [============================&gt;.] - ETA: 3s - loss: 0.7682 - accuracy: 0.7472\n389/390 [============================&gt;.] - ETA: 1s - loss: 0.7671 - accuracy: 0.7476</p>\n\n<p>Epoch 00001: val_accuracy improved from -inf to 0.87509, saving model to data/consonantd_model.h5</p>\n\n<p>390/390 [==============================] - 638s 2s/step - loss: 0.7659 - accuracy: 0.7481 - val_loss: 0.0520 - val_accuracy: 0.8751\nEpoch 2/15</p>\n\n<p>1/390 [..............................] - ETA: 4:41 - loss: 0.2990 - accuracy: 0.9219\n  2/390 [..............................] - ETA: 4:42 - loss: 0.2972 - accuracy: 0.9180\n  3/390 [..............................] - ETA: 4:42 - loss: 0.2933 - accuracy: 0.9206\n  4/390 [..............................] - ETA: 4:41 - loss: 0.3102 - accuracy: 0.9160\n```</p>",
      "votes": 1,
      "replies": [
        {
          "id": 709345,
          "author_name": "Thomas Di Martino",
          "author_url": "",
          "post_date": "2020-01-03T11:50:20.113000",
          "content": "<p>I may be wrong and it may be possible to correct me but from my understanding here is what happens:</p>\n\n<p>To evaluate loss during training, Keras calculate an average loss (same for accuracy measure) over all the already seen batches (explained <a href=\"https://github.com/keras-team/keras/issues/10426\">here</a>). So it can explain what you observe in this way:</p>\n\n<p>Epoch 1/15</p>\n\n<p>1/390 [..............................] - ETA: 4:41 - loss: X1 - accuracy: Y1</p>\n\n<p>Here, we have the first batch of data, our <strong>network will perform awfully on it</strong>, resulting in the <strong>highest loss and accuracy from all the possible batches</strong> of all the epochs of the training session. But the <strong>network will start learning and it will start learning fast</strong>, even more at the start of the training. The batches passes by and we arrive at the last batches of the first epoch:</p>\n\n<p>387/390 [============================&gt;.] - ETA: 4s - loss: X_387 - accuracy: Y_387\n388/390 [============================&gt;.] - ETA: 3s - loss: X_388 - accuracy: Y_388\n389/390 [============================&gt;.] - ETA: 1s - loss: X_389 - accuracy: Y_389</p>\n\n<p>At this point, the network should perform way better and on batches 387, 388 and 389, the loss calculated for each of these batch, independently, will be much lower than for the loss of <em>X1</em>.</p>\n\n<p>However, the value of X_387 is calculated as follows:</p>\n\n<p>$$ \\dfrac{1}{n} * \\sum_{i=1}^{n} \\mathbb{B(i)} $$</p>\n\n<p>where B(i) is the loss value of the batch number <em>i</em> and n equals 387.\nHence, when you arrive at the X_387 value, to evaluate it, you do the mean of the batches loss from the 1st batch to the last. However, the first is sooo bad compared to your last batches (hence so high), that it biased the overall mean of the epoch. \nThis is a weird choice of implementation to show this accumulative mean and not the loss of the batch 387.\nHowever, you can justify it by the fact that you don't care how your network behave on a subsample of data, you want its performance to be the best in a general situation =&gt; <strong>use of a mean</strong>.</p>\n\n<p>So in your situation, you could guess that, on your last batches of your 1st epoch, your network had loss similar to the first batches of the second epoch. \nYou have to note that this behaviour will always exist throughout the training but should be less and less noticeable as batches pass by !</p>\n\n<p>I hope that I solved your question, if you have any more interrogation, don't mind asking !</p>\n\n<p>I wish you good luck for this competition and for your future in Machine Learning ;) !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 709727,
          "author_name": "Tiberiu Savin",
          "author_url": "",
          "post_date": "2020-01-03T21:52:51.010000",
          "content": "<p>Oh I see. Thank you very much. Indeed I was under the impression that the loss and accuracy were the from that specific mini_batch not cumulative across the entire epoch. \nThis does indeed explain the behaviour I'm seeing and it makes sense. Indeed this behaviour becomes less noticeable for higher epochs.\nAgain thanks for the explanation kind sir.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762579,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-03-03T16:10:56.413000",
          "content": "<p>Thanks for you answer~</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 765506,
      "author_name": "Vasco Mano",
      "author_url": "",
      "post_date": "2020-03-06T17:58:32.773000",
      "content": "<p>Hello, <a href=\"/addisonhoward\">@addisonhoward</a> ! Not completely new, but definitely a beginner.</p>\n\n<p>Can you explain how does one solve this \"Notebook Exceeded Allowed Compute\"? What does this really mean?\nThank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 766401,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2020-03-08T05:06:00.610000",
          "content": "<p>Hi João,</p>\n\n<p>Each Kaggle notebook has a present limit of computing power. Look <a href=\"https://www.kaggle.com/c/bengaliai-cv19/overview/notebooks-requirements\">here</a> for details. It sounds like your notebook may have run over the limits allowed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 766960,
          "author_name": "Vasco Mano",
          "author_url": "",
          "post_date": "2020-03-09T01:30:07.807000",
          "content": "<p>Thank you for the kind reply.</p>\n\n<p>I don't understand. My inference notebook runs in ~100 seconds with a GPU in my kaggles private area. However, the submission fails.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767402,
          "author_name": "Vasco Mano",
          "author_url": "",
          "post_date": "2020-03-09T15:31:04.053000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4186383%2F7f9c8b664249961e5e3ede088207bc20%2Fsubmission.png?generation=1583767668140589&amp;alt=media\" alt=\"\"></p>\n\n<p>For example, this notebook gives me the error \"Notebook Exceeded Allowed Compute\". I don't understand why.</p>\n\n<p>Once again, thank you for your kind attention, <a href=\"/addisonhoward\">@addisonhoward</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 761781,
      "author_name": "ermacov",
      "author_url": "",
      "post_date": "2020-03-02T23:35:33.367000",
      "content": "<p>Hi,how to load weights trained on multi gpu model to the single gpu model? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 753675,
      "author_name": "Devarshi Aggarwal",
      "author_url": "",
      "post_date": "2020-02-22T14:36:15.967000",
      "content": "<p>Hi! I feel overwhelmed with the kind of progress other Kagglers have already made. How long do you think I will be a Master Kaggler?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 747197,
      "author_name": "Akhil Sharma",
      "author_url": "",
      "post_date": "2020-02-16T05:14:57.260000",
      "content": "<p>As a beginner trying to self-teach himself in Natural Language processing and computer vision,what would you suggest as some good reference sites to explore these topics in addition to Kaggle? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 699620,
      "author_name": "christian schaedel",
      "author_url": "",
      "post_date": "2019-12-20T16:37:58.797000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4232764%2Fa5f7fe2224a0d90b502344708d194a8a%2FScreen%20Shot%202019-12-20%20at%2012.55.46.png?generation=1576859875129967&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 760473,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-01T10:47:24.433000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 747126,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-16T01:46:39.003000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 735288,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-02T20:26:00.173000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 733507,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-31T08:27:32.923000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 735287,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-02T20:24:03.473000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742316,
          "author_name": "Sylvia Chan",
          "author_url": "",
          "post_date": "2020-02-11T08:23:43.693000",
          "content": "<p><a href=\"/e60315\">@e60315</a> your reply is really helpful and also solved my question. But I wonder if you know how to upload the offline pretrained model or how to use it as an input to my kaggle kernel? Thanks a lot!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743969,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-12T12:53:47.813000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 745015,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-13T11:50:05.063000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 699593,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-20T16:10:09.810000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 699219,
      "author_name": "Tianyi Wang",
      "author_url": "",
      "post_date": "2019-12-20T07:23:01.657000",
      "content": "<p>Thanks a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 709944,
      "author_name": "William Li",
      "author_url": "",
      "post_date": "2020-01-04T05:19:50.340000",
      "content": "<p>thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 700054,
      "author_name": "LUKMAN AFOLABI",
      "author_url": "",
      "post_date": "2019-12-21T11:16:45.757000",
      "content": "<p>Thank you very much.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "698949": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ).",
    "721157": "Hi,\n\nIn competitions like this, when people say they are using Resnet, Densenet, Inception, etc. are they using the pretrained model and training dense layers on top or training the architecture from scratch?\n\nIt seems infeasible to train from scratch on kaggle kernels?",
    "708693": "Hello,\n\nThis might seem like a really noobish question but I'm new at machine learning and since this topic is specifically for this maybe someone can explain some stuff to me :).\nWhen training in the beginning the accuracy starts growing slowly with each batch (like it goes 0.100 -&gt; 0.110 -&gt; 0.1120 and so on) also the loss goes down slowly. This is ok, it's what I expect to happen. However at the end of the epoch there is a big jump for both the accuracy as well as the loss function (the accuracy goes up and the loss goes down). I'm not complaining as this seems good, but I don't understand exactly why this is happening.\nFrom what I know training is basically for each batch do forwardprop and then backprop. When the epoch ends check the validation set and then start again. Is there some training step at the end of the epoch or something that's causing the accuracy to improve a lot at the end of the epoch?\n\n[Later Edit]\nHere is an example of how the keras output looks at the end of the epoch:\n\n```\n383/390 [============================&gt;.] - ETA: 11s - loss: 0.7742 - accuracy: 0.7450\n384/390 [============================&gt;.] - ETA: 9s - loss: 0.7729 - accuracy: 0.7455 \n385/390 [============================&gt;.] - ETA: 8s - loss: 0.7716 - accuracy: 0.7460\n386/390 [============================&gt;.] - ETA: 6s - loss: 0.7705 - accuracy: 0.7464\n387/390 [============================&gt;.] - ETA: 4s - loss: 0.7693 - accuracy: 0.7468\n388/390 [============================&gt;.] - ETA: 3s - loss: 0.7682 - accuracy: 0.7472\n389/390 [============================&gt;.] - ETA: 1s - loss: 0.7671 - accuracy: 0.7476\n\nEpoch 00001: val_accuracy improved from -inf to 0.87509, saving model to data/consonantd_model.h5\n\n390/390 [==============================] - 638s 2s/step - loss: 0.7659 - accuracy: 0.7481 - val_loss: 0.0520 - val_accuracy: 0.8751\nEpoch 2/15\n\n  1/390 [..............................] - ETA: 4:41 - loss: 0.2990 - accuracy: 0.9219\n  2/390 [..............................] - ETA: 4:42 - loss: 0.2972 - accuracy: 0.9180\n  3/390 [..............................] - ETA: 4:42 - loss: 0.2933 - accuracy: 0.9206\n  4/390 [..............................] - ETA: 4:41 - loss: 0.3102 - accuracy: 0.9160\n```",
    "765506": "Hello, @addisonhoward ! Not completely new, but definitely a beginner.\n\nCan you explain how does one solve this \"Notebook Exceeded Allowed Compute\"? What does this really mean?\nThank you!",
    "761781": "Hi,how to load weights trained on multi gpu model to the single gpu model? ",
    "753675": "Hi! I feel overwhelmed with the kind of progress other Kagglers have already made. How long do you think I will be a Master Kaggler?",
    "747197": "As a beginner trying to self-teach himself in Natural Language processing and computer vision,what would you suggest as some good reference sites to explore these topics in addition to Kaggle? ",
    "699620": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4232764%2Fa5f7fe2224a0d90b502344708d194a8a%2FScreen%20Shot%202019-12-20%20at%2012.55.46.png?generation=1576859875129967&amp;alt=media)\n",
    "760473": "",
    "747126": "",
    "735288": "",
    "733507": "",
    "699593": "",
    "699219": "Thanks a lot!",
    "709944": "thanks!",
    "700054": "Thank you very much."
  }
}