{
  "id": 123658,
  "title": "Submission Score worst than expected",
  "url": "/competitions/bengaliai-cv19/discussion/123658",
  "author_name": "",
  "post_date": "2019-12-29T12:13:41.740533600Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I had train 3 separated models (one for each category) using only first image train file with ~50k images. My best 'val categorical accuracy' are 0.95, 0.89 and 0.96. Those models give me a submission score of 0.26 in the private test, it is a lot worst that I expected. At the moment I am not using any kind of data augmentation. The validation split is 20% of total images in first train file, and I am not getting a lot of overfit.\nMy questions on this:\n1. Using only first train file is not representative of the test dataset therefore the submission score is very low?\n2. I saw people training using the 4 files in a for loop, fitting the model for each file at a time (using all 4 merged in same fit is not possible because memory usage). Doesn't the second loop 'override' the information learned on the first loop images? \n3. If 1. is not true than maybe I have something wrong in the submission code, I already looked closely but didn't find anything wrong.</p>\n\n<p>Thank you in advance,\nKeep it up!</p>",
  "messages": [
    {
      "id": "705749",
      "postDate": "12/29/2019 12:13:41",
      "content": "<p>I had train 3 separated models (one for each category) using only first image train file with ~50k images. My best 'val categorical accuracy' are 0.95, 0.89 and 0.96. Those models give me a submission score of 0.26 in the private test, it is a lot worst that I expected. At the moment I am not using any kind of data augmentation. The validation split is 20% of total images in first train file, and I am not getting a lot of overfit.\nMy questions on this:\n1. Using only first train file is not representative of the test dataset therefore the submission score is very low?\n2. I saw people training using the 4 files in a for loop, fitting the model for each file at a time (using all 4 merged in same fit is not possible because memory usage). Doesn't the second loop 'override' the information learned on the first loop images? \n3. If 1. is not true than maybe I have something wrong in the submission code, I already looked closely but didn't find anything wrong.</p>\n\n<p>Thank you in advance,\nKeep it up!</p>",
      "rawMarkdown": "I had train 3 separated models (one for each category) using only first image train file with ~50k images. My best 'val categorical accuracy' are 0.95, 0.89 and 0.96. Those models give me a submission score of 0.26 in the private test, it is a lot worst that I expected. At the moment I am not using any kind of data augmentation. The validation split is 20% of total images in first train file, and I am not getting a lot of overfit.\nMy questions on this:\n1. Using only first train file is not representative of the test dataset therefore the submission score is very low?\n2. I saw people training using the 4 files in a for loop, fitting the model for each file at a time (using all 4 merged in same fit is not possible because memory usage). Doesn't the second loop 'override' the information learned on the first loop images? \n3. If 1. is not true than maybe I have something wrong in the submission code, I already looked closely but didn't find anything wrong.\n\nThank you in advance,\nKeep it up!",
      "votes": null
    },
    {
      "id": "705982",
      "postDate": "12/29/2019 19:14:18",
      "content": "<p>I am a newbie to CV too so I tried my best to share my 2 cents.</p>\n\n<ol>\n<li>Why don't use all four and create a train/valid split for validation? Ideally you should use as much data for training as possible. But at the same time you may want to leave some data for validation (and strictly not for training).</li>\n<li>What do you mean by loop? Do you mean by each epoch?</li>\n<li>I think you can answer this question yourself.</li>\n</ol>\n\n<p>I hope it helps.</p>",
      "rawMarkdown": "I am a newbie to CV too so I tried my best to share my 2 cents.\n\n1. Why don't use all four and create a train/valid split for validation? Ideally you should use as much data for training as possible. But at the same time you may want to leave some data for validation (and strictly not for training).\n2. What do you mean by loop? Do you mean by each epoch?\n3. I think you can answer this question yourself.\n\nI hope it helps.",
      "votes": null
    },
    {
      "id": "706051",
      "postDate": "12/29/2019 21:17:55",
      "content": "<ol>\n<li>I am not using all 4 because it uses a lot of memory, I am trying to get a decent result using only this, and train later in a good model architecture.</li>\n<li>Something like:\nfor file in range(4):\n    model.fit(file)</li>\n</ol>",
      "rawMarkdown": "1. I am not using all 4 because it uses a lot of memory, I am trying to get a decent result using only this, and train later in a good model architecture.\n2. Something like:\n    for file in range(4):\n        model.fit(file)",
      "votes": null
    },
    {
      "id": "706053",
      "postDate": "12/29/2019 21:22:12",
      "content": "<p>I don't have a very high score for the moment with only 0.74. However, it is first interesting to note that score used by the competition is a weighted <a href=\"https://en.wikipedia.org/wiki/Precision_and_recall\">recall</a>, hence it will be different to the accuracy you are using for the validation process (I achieve roughly from 0.95 to 0.99 val_acc for each output).</p>\n\n<p>What I would advise, if you take the <em>NeuralNet</em> solution, is to save your images into .png and use a generator to load them dynamically without having to worry about a for loop. Do most of your computation outside of a Kaggle notebook, then limit the use of a Kaggle notebook to inference. This will make you worry less about MemoryConsumption but will indeed take more time to compute. However, code is more maintainable and easy to reuse for you !</p>\n\n<p>To answer your second point, no, it does not \"override\" per se the first loop as you will be training a network that would already perform pretty well on your, for now, unseen second parquet file. But it would fit your model, previously trained on the 1st parquet file, to the data specific to the 2 parquet file meaning that your epochs would run only over this file in particular and not the 3 others, so you would ultimately loose some regularisation performance. </p>\n\n<p>So ultimately, my advice would be summarized as:</p>\n\n<p><strong>Mix all 4 parquets file into one folder</strong> with your images inside (they should be <strong>preprocessed by cropping and resizing</strong> as it has been presented in a few different kernels). From this folder, do a <strong>CV process so that you have images from all 4 parquets file</strong> used in every fold. \nAnd concerning the score calculation, remember that they <a href=\"https://www.kaggle.com/c/bengaliai-cv19/overview/evaluation\"><strong>use weighted recall</strong></a>.  </p>\n\n<p>I wish you good luck as this contest seems way harder that it could look like at first: it is no MNIST ahah.</p>",
      "rawMarkdown": "I don't have a very high score for the moment with only 0.74. However, it is first interesting to note that score used by the competition is a weighted [recall](https://en.wikipedia.org/wiki/Precision_and_recall), hence it will be different to the accuracy you are using for the validation process (I achieve roughly from 0.95 to 0.99 val_acc for each output).\n\nWhat I would advise, if you take the *NeuralNet* solution, is to save your images into .png and use a generator to load them dynamically without having to worry about a for loop. Do most of your computation outside of a Kaggle notebook, then limit the use of a Kaggle notebook to inference. This will make you worry less about MemoryConsumption but will indeed take more time to compute. However, code is more maintainable and easy to reuse for you !\n\nTo answer your second point, no, it does not \"override\" per se the first loop as you will be training a network that would already perform pretty well on your, for now, unseen second parquet file. But it would fit your model, previously trained on the 1st parquet file, to the data specific to the 2 parquet file meaning that your epochs would run only over this file in particular and not the 3 others, so you would ultimately loose some regularisation performance. \n\n\n\nSo ultimately, my advice would be summarized as:\n\n**Mix all 4 parquets file into one folder** with your images inside (they should be **preprocessed by cropping and resizing** as it has been presented in a few different kernels). From this folder, do a **CV process so that you have images from all 4 parquets file** used in every fold. \nAnd concerning the score calculation, remember that they [**use weighted recall**](https://www.kaggle.com/c/bengaliai-cv19/overview/evaluation).  \n\nI wish you good luck as this contest seems way harder that it could look like at first: it is no MNIST ahah.",
      "votes": null
    },
    {
      "id": "706055",
      "postDate": "12/29/2019 21:35:10",
      "content": "<p>Thank you very much for the detailed answer, I was indeed forgetting the recall was used, not the validation, but thought it will be better anyway. I will try to take a better approach overall </p>",
      "rawMarkdown": "Thank you very much for the detailed answer, I was indeed forgetting the recall was used, not the validation, but thought it will be better anyway. I will try to take a better approach overall",
      "votes": null
    },
    {
      "id": "706506",
      "postDate": "12/30/2019 13:34:36",
      "content": "<p>Solved - I was using pd.get_dummies().astype(str) and this make the column order 0,1,10,11 instead of 0,1,2,3. When converting the model prediction to class number for the submission file I was using argmax</p>",
      "rawMarkdown": "Solved - I was using pd.get_dummies().astype(str) and this make the column order 0,1,10,11 instead of 0,1,2,3. When converting the model prediction to class number for the submission file I was using argmax",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 705982,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "12/29/2019 19:14:18",
      "content": "<p>I am a newbie to CV too so I tried my best to share my 2 cents.</p>\n\n<ol>\n<li>Why don't use all four and create a train/valid split for validation? Ideally you should use as much data for training as possible. But at the same time you may want to leave some data for validation (and strictly not for training).</li>\n<li>What do you mean by loop? Do you mean by each epoch?</li>\n<li>I think you can answer this question yourself.</li>\n</ol>\n\n<p>I hope it helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 706051,
          "author_name": "macarrony00",
          "author_url": "",
          "post_date": "12/29/2019 21:17:55",
          "content": "<ol>\n<li>I am not using all 4 because it uses a lot of memory, I am trying to get a decent result using only this, and train later in a good model architecture.</li>\n<li>Something like:\nfor file in range(4):\n    model.fit(file)</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 706053,
      "author_name": "dimartinot",
      "author_url": "",
      "post_date": "12/29/2019 21:22:12",
      "content": "<p>I don't have a very high score for the moment with only 0.74. However, it is first interesting to note that score used by the competition is a weighted <a href=\"https://en.wikipedia.org/wiki/Precision_and_recall\">recall</a>, hence it will be different to the accuracy you are using for the validation process (I achieve roughly from 0.95 to 0.99 val_acc for each output).</p>\n\n<p>What I would advise, if you take the <em>NeuralNet</em> solution, is to save your images into .png and use a generator to load them dynamically without having to worry about a for loop. Do most of your computation outside of a Kaggle notebook, then limit the use of a Kaggle notebook to inference. This will make you worry less about MemoryConsumption but will indeed take more time to compute. However, code is more maintainable and easy to reuse for you !</p>\n\n<p>To answer your second point, no, it does not \"override\" per se the first loop as you will be training a network that would already perform pretty well on your, for now, unseen second parquet file. But it would fit your model, previously trained on the 1st parquet file, to the data specific to the 2 parquet file meaning that your epochs would run only over this file in particular and not the 3 others, so you would ultimately loose some regularisation performance. </p>\n\n<p>So ultimately, my advice would be summarized as:</p>\n\n<p><strong>Mix all 4 parquets file into one folder</strong> with your images inside (they should be <strong>preprocessed by cropping and resizing</strong> as it has been presented in a few different kernels). From this folder, do a <strong>CV process so that you have images from all 4 parquets file</strong> used in every fold. \nAnd concerning the score calculation, remember that they <a href=\"https://www.kaggle.com/c/bengaliai-cv19/overview/evaluation\"><strong>use weighted recall</strong></a>.  </p>\n\n<p>I wish you good luck as this contest seems way harder that it could look like at first: it is no MNIST ahah.</p>",
      "votes": null,
      "replies": [
        {
          "id": 706055,
          "author_name": "macarrony00",
          "author_url": "",
          "post_date": "12/29/2019 21:35:10",
          "content": "<p>Thank you very much for the detailed answer, I was indeed forgetting the recall was used, not the validation, but thought it will be better anyway. I will try to take a better approach overall </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 706506,
      "author_name": "macarrony00",
      "author_url": "",
      "post_date": "12/30/2019 13:34:36",
      "content": "<p>Solved - I was using pd.get_dummies().astype(str) and this make the column order 0,1,10,11 instead of 0,1,2,3. When converting the model prediction to class number for the submission file I was using argmax</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "705749": "I had train 3 separated models (one for each category) using only first image train file with ~50k images. My best 'val categorical accuracy' are 0.95, 0.89 and 0.96. Those models give me a submission score of 0.26 in the private test, it is a lot worst that I expected. At the moment I am not using any kind of data augmentation. The validation split is 20% of total images in first train file, and I am not getting a lot of overfit.\nMy questions on this:\n1. Using only first train file is not representative of the test dataset therefore the submission score is very low?\n2. I saw people training using the 4 files in a for loop, fitting the model for each file at a time (using all 4 merged in same fit is not possible because memory usage). Doesn't the second loop 'override' the information learned on the first loop images? \n3. If 1. is not true than maybe I have something wrong in the submission code, I already looked closely but didn't find anything wrong.\n\nThank you in advance,\nKeep it up!",
    "705982": "I am a newbie to CV too so I tried my best to share my 2 cents.\n\n1. Why don't use all four and create a train/valid split for validation? Ideally you should use as much data for training as possible. But at the same time you may want to leave some data for validation (and strictly not for training).\n2. What do you mean by loop? Do you mean by each epoch?\n3. I think you can answer this question yourself.\n\nI hope it helps.",
    "706051": "1. I am not using all 4 because it uses a lot of memory, I am trying to get a decent result using only this, and train later in a good model architecture.\n2. Something like:\n    for file in range(4):\n        model.fit(file)",
    "706053": "I don't have a very high score for the moment with only 0.74. However, it is first interesting to note that score used by the competition is a weighted [recall](https://en.wikipedia.org/wiki/Precision_and_recall), hence it will be different to the accuracy you are using for the validation process (I achieve roughly from 0.95 to 0.99 val_acc for each output).\n\nWhat I would advise, if you take the *NeuralNet* solution, is to save your images into .png and use a generator to load them dynamically without having to worry about a for loop. Do most of your computation outside of a Kaggle notebook, then limit the use of a Kaggle notebook to inference. This will make you worry less about MemoryConsumption but will indeed take more time to compute. However, code is more maintainable and easy to reuse for you !\n\nTo answer your second point, no, it does not \"override\" per se the first loop as you will be training a network that would already perform pretty well on your, for now, unseen second parquet file. But it would fit your model, previously trained on the 1st parquet file, to the data specific to the 2 parquet file meaning that your epochs would run only over this file in particular and not the 3 others, so you would ultimately loose some regularisation performance. \n\n\n\nSo ultimately, my advice would be summarized as:\n\n**Mix all 4 parquets file into one folder** with your images inside (they should be **preprocessed by cropping and resizing** as it has been presented in a few different kernels). From this folder, do a **CV process so that you have images from all 4 parquets file** used in every fold. \nAnd concerning the score calculation, remember that they [**use weighted recall**](https://www.kaggle.com/c/bengaliai-cv19/overview/evaluation).  \n\nI wish you good luck as this contest seems way harder that it could look like at first: it is no MNIST ahah.",
    "706055": "Thank you very much for the detailed answer, I was indeed forgetting the recall was used, not the validation, but thought it will be better anyway. I will try to take a better approach overall",
    "706506": "Solved - I was using pd.get_dummies().astype(str) and this make the column order 0,1,10,11 instead of 0,1,2,3. When converting the model prediction to class number for the submission file I was using argmax"
  },
  "source": "meta"
}