{
  "id": 228610,
  "title": "Are you using all the training data at this time?",
  "url": "/competitions/bms-molecular-translation/discussion/228610",
  "author_name": "",
  "post_date": "2021-03-25T13:57:35.102906500Z",
  "votes": 28,
  "comment_count": 12,
  "views": 0,
  "content": "<p>This competition provides us very large dataset(training: 2.42 million, test: 1.61 million). It takes so long long time to train models by all the training data.</p>\n<p>Should we use all data at this time？ I don't think so because we still have more than two months until the competition deadline.<br>\nIt may be better to do trial and error with some of data, and train models by all the data sometimes.<br>\n<br><br>\nFor your reference, here are the results of training <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/228139\" target=\"_blank\">chemical formula prediction</a> in my local environment. </p>\n<table>\n<thead>\n<tr>\n<th>% of data</th>\n<th>training time (1 fold, 20 epoch)</th>\n<th>OOF Levenshtein Distance (5 Fold)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>10%</td>\n<td>1.2 hours</td>\n<td>0.499</td>\n</tr>\n<tr>\n<td>20%</td>\n<td>2.4 hours</td>\n<td>0.343</td>\n</tr>\n<tr>\n<td>50%</td>\n<td>6 hours</td>\n<td>0.210</td>\n</tr>\n<tr>\n<td>100%</td>\n<td>12 hours</td>\n<td>0.154</td>\n</tr>\n</tbody>\n</table>\n<p>I think at most 50% of data is sufficient for trial and error.</p>",
  "messages": [
    {
      "id": "1252210",
      "postDate": "03/25/2021 13:57:35",
      "content": "<p>This competition provides us very large dataset(training: 2.42 million, test: 1.61 million). It takes so long long time to train models by all the training data.</p>\n<p>Should we use all data at this time？ I don't think so because we still have more than two months until the competition deadline.<br>\nIt may be better to do trial and error with some of data, and train models by all the data sometimes.<br>\n<br><br>\nFor your reference, here are the results of training <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/228139\" target=\"_blank\">chemical formula prediction</a> in my local environment. </p>\n<table>\n<thead>\n<tr>\n<th>% of data</th>\n<th>training time (1 fold, 20 epoch)</th>\n<th>OOF Levenshtein Distance (5 Fold)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>10%</td>\n<td>1.2 hours</td>\n<td>0.499</td>\n</tr>\n<tr>\n<td>20%</td>\n<td>2.4 hours</td>\n<td>0.343</td>\n</tr>\n<tr>\n<td>50%</td>\n<td>6 hours</td>\n<td>0.210</td>\n</tr>\n<tr>\n<td>100%</td>\n<td>12 hours</td>\n<td>0.154</td>\n</tr>\n</tbody>\n</table>\n<p>I think at most 50% of data is sufficient for trial and error.</p>",
      "rawMarkdown": "This competition provides us very large dataset(training: 2.42 million, test: 1.61 million). It takes so long long time to train models by all the training data.\n\nShould we use all data at this time？ I don't think so because we still have more than two months until the competition deadline.\nIt may be better to do trial and error with some of data, and train models by all the data sometimes.\n<br>\nFor your reference, here are the results of training [chemical formula prediction](https://www.kaggle.com/c/bms-molecular-translation/discussion/228139) in my local environment. \n\n| % of data | training time (1 fold, 20 epoch) | OOF Levenshtein Distance (5 Fold) |\n|:---------:|:---------:|:---------:|\n| 10% | 1.2 hours | 0.499 |\n| 20% | 2.4 hours | 0.343 |\n| 50% | 6 hours | 0.210 |\n| 100% | 12 hours | 0.154 |\n\n  I think at most 50% of data is sufficient for trial and error.",
      "votes": null
    },
    {
      "id": "1252635",
      "postDate": "03/25/2021 21:32:42",
      "content": "<p>Can confirm.. I run experiments with 1 million images … it usually translate well when upscaling =)</p>",
      "rawMarkdown": "Can confirm.. I run experiments with 1 million images ... it usually translate well when upscaling =)",
      "votes": null
    },
    {
      "id": "1252697",
      "postDate": "03/25/2021 23:52:47",
      "content": "<p>Thank you for sharing.</p>\n<p>Do you choose 1 million data in a specfic manner? I choose data by random sampling now but think this is not a good way.</p>",
      "rawMarkdown": "Thank you for sharing.\n\nDo you choose 1 million data in a specfic manner? I choose data by random sampling now but think this is not a good way.",
      "votes": null
    },
    {
      "id": "1252714",
      "postDate": "03/26/2021 00:41:05",
      "content": "<p>Thanks for sharing, I'm curious, what hardware do you use? TPUs?</p>",
      "rawMarkdown": "Thanks for sharing, I'm curious, what hardware do you use? TPUs?",
      "votes": null
    },
    {
      "id": "1252721",
      "postDate": "03/26/2021 01:00:22",
      "content": "<p>It was random… But I agree its not a optimal … </p>",
      "rawMarkdown": "It was random... But I agree its not a optimal ...",
      "votes": null
    },
    {
      "id": "1252740",
      "postDate": "03/26/2021 01:45:24",
      "content": "<p>OK, thanks.</p>\n<p>I think considering number of atoms or InChI length may be an easy and good way.</p>",
      "rawMarkdown": "OK, thanks.\n\nI think considering number of atoms or InChI length may be an easy and good way.",
      "votes": null
    },
    {
      "id": "1252746",
      "postDate": "03/26/2021 01:50:17",
      "content": "<p>I use a single TitanRTX(24GB). But only 3~4GB was used for the above experiments because image size and batch size are not large.</p>",
      "rawMarkdown": "I use a single TitanRTX(24GB). But only 3~4GB was used for the above experiments because image size and batch size are not large.",
      "votes": null
    },
    {
      "id": "1252967",
      "postDate": "03/26/2021 08:33:34",
      "content": "<p>I use all the train data.(2.42 million images val 1fold) So it takes about 2~5 hours to train 1epoch. <br>\nAlso, I have experienced poor accuracy when training with batch size larger than 256 in my case.<br>\nI feel that training with a large batch size is difficult.</p>",
      "rawMarkdown": "I use all the train data.(2.42 million images val 1fold) So it takes about 2~5 hours to train 1epoch. \nAlso, I have experienced poor accuracy when training with batch size larger than 256 in my case.\nI feel that training with a large batch size is difficult.",
      "votes": null
    },
    {
      "id": "1252998",
      "postDate": "03/26/2021 09:25:56",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for the insight<br>\nI believe we can take a percentage of the sample (30, 40, or 50%) in random using numpy random at this stage. This get us close to 80% to 90&amp; of our target goal from where we can fine tune for greater performance. </p>",
      "rawMarkdown": "Thank you @ttahara for the insight\nI believe we can take a percentage of the sample (30, 40, or 50%) in random using numpy random at this stage. This get us close to 80% to 90& of our target goal from where we can fine tune for greater performance.",
      "votes": null
    },
    {
      "id": "1253180",
      "postDate": "03/26/2021 12:48:29",
      "content": "<p>Thanks for sharing.</p>\n<p>I train models with batch size of 64.</p>",
      "rawMarkdown": "Thanks for sharing.\n\nI train models with batch size of 64.",
      "votes": null
    },
    {
      "id": "1254102",
      "postDate": "03/27/2021 10:55:18",
      "content": "<p>How are you spitting the data for train and validation? randomly ?</p>",
      "rawMarkdown": "How are you spitting the data for train and validation? randomly ?",
      "votes": null
    },
    {
      "id": "1254265",
      "postDate": "03/27/2021 13:33:36",
      "content": "<p>I split data into train and valid in Multi-Label Stratified KFold(K=5) manner, using whether or not each atom is included in molecular as multi-label.</p>\n<p>I'm sharing splitting process in <a href=\"https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-training\" target=\"_blank\">my notebook</a>.</p>",
      "rawMarkdown": "I split data into train and valid in Multi-Label Stratified KFold(K=5) manner, using whether or not each atom is included in molecular as multi-label.\n\nI'm sharing splitting process in [my notebook](https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-training).",
      "votes": null
    },
    {
      "id": "1296173",
      "postDate": "05/07/2021 03:43:55",
      "content": "<p>I have done some experiments with different Encoders (CNN models) and different percentage of the training data and I have concluded that experimenting with 10% training samples is ok and reliable when scaling. Maybe less than 10% is ok as well.</p>\n<p>But for the final score and submission it's crucial to train with all data available, it is not worthy to reduce the training set.</p>",
      "rawMarkdown": "I have done some experiments with different Encoders (CNN models) and different percentage of the training data and I have concluded that experimenting with 10% training samples is ok and reliable when scaling. Maybe less than 10% is ok as well.\n\nBut for the final score and submission it's crucial to train with all data available, it is not worthy to reduce the training set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1252635,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "03/25/2021 21:32:42",
      "content": "<p>Can confirm.. I run experiments with 1 million images … it usually translate well when upscaling =)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1252697,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "03/25/2021 23:52:47",
          "content": "<p>Thank you for sharing.</p>\n<p>Do you choose 1 million data in a specfic manner? I choose data by random sampling now but think this is not a good way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252721,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/26/2021 01:00:22",
          "content": "<p>It was random… But I agree its not a optimal … </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252740,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "03/26/2021 01:45:24",
          "content": "<p>OK, thanks.</p>\n<p>I think considering number of atoms or InChI length may be an easy and good way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1252714,
      "author_name": "michaelwolff",
      "author_url": "",
      "post_date": "03/26/2021 00:41:05",
      "content": "<p>Thanks for sharing, I'm curious, what hardware do you use? TPUs?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1252746,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "03/26/2021 01:50:17",
          "content": "<p>I use a single TitanRTX(24GB). But only 3~4GB was used for the above experiments because image size and batch size are not large.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1252967,
      "author_name": "kanbehmw",
      "author_url": "",
      "post_date": "03/26/2021 08:33:34",
      "content": "<p>I use all the train data.(2.42 million images val 1fold) So it takes about 2~5 hours to train 1epoch. <br>\nAlso, I have experienced poor accuracy when training with batch size larger than 256 in my case.<br>\nI feel that training with a large batch size is difficult.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1253180,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "03/26/2021 12:48:29",
          "content": "<p>Thanks for sharing.</p>\n<p>I train models with batch size of 64.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1252998,
      "author_name": "olusesiadebisi",
      "author_url": "",
      "post_date": "03/26/2021 09:25:56",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for the insight<br>\nI believe we can take a percentage of the sample (30, 40, or 50%) in random using numpy random at this stage. This get us close to 80% to 90&amp; of our target goal from where we can fine tune for greater performance. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1254102,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "03/27/2021 10:55:18",
      "content": "<p>How are you spitting the data for train and validation? randomly ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1254265,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "03/27/2021 13:33:36",
          "content": "<p>I split data into train and valid in Multi-Label Stratified KFold(K=5) manner, using whether or not each atom is included in molecular as multi-label.</p>\n<p>I'm sharing splitting process in <a href=\"https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-training\" target=\"_blank\">my notebook</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1296173,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "05/07/2021 03:43:55",
      "content": "<p>I have done some experiments with different Encoders (CNN models) and different percentage of the training data and I have concluded that experimenting with 10% training samples is ok and reliable when scaling. Maybe less than 10% is ok as well.</p>\n<p>But for the final score and submission it's crucial to train with all data available, it is not worthy to reduce the training set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1252210": "This competition provides us very large dataset(training: 2.42 million, test: 1.61 million). It takes so long long time to train models by all the training data.\n\nShould we use all data at this time？ I don't think so because we still have more than two months until the competition deadline.\nIt may be better to do trial and error with some of data, and train models by all the data sometimes.\n<br>\nFor your reference, here are the results of training [chemical formula prediction](https://www.kaggle.com/c/bms-molecular-translation/discussion/228139) in my local environment. \n\n| % of data | training time (1 fold, 20 epoch) | OOF Levenshtein Distance (5 Fold) |\n|:---------:|:---------:|:---------:|\n| 10% | 1.2 hours | 0.499 |\n| 20% | 2.4 hours | 0.343 |\n| 50% | 6 hours | 0.210 |\n| 100% | 12 hours | 0.154 |\n\n  I think at most 50% of data is sufficient for trial and error.",
    "1252635": "Can confirm.. I run experiments with 1 million images ... it usually translate well when upscaling =)",
    "1252697": "Thank you for sharing.\n\nDo you choose 1 million data in a specfic manner? I choose data by random sampling now but think this is not a good way.",
    "1252714": "Thanks for sharing, I'm curious, what hardware do you use? TPUs?",
    "1252721": "It was random... But I agree its not a optimal ...",
    "1252740": "OK, thanks.\n\nI think considering number of atoms or InChI length may be an easy and good way.",
    "1252746": "I use a single TitanRTX(24GB). But only 3~4GB was used for the above experiments because image size and batch size are not large.",
    "1252967": "I use all the train data.(2.42 million images val 1fold) So it takes about 2~5 hours to train 1epoch. \nAlso, I have experienced poor accuracy when training with batch size larger than 256 in my case.\nI feel that training with a large batch size is difficult.",
    "1252998": "Thank you @ttahara for the insight\nI believe we can take a percentage of the sample (30, 40, or 50%) in random using numpy random at this stage. This get us close to 80% to 90& of our target goal from where we can fine tune for greater performance.",
    "1253180": "Thanks for sharing.\n\nI train models with batch size of 64.",
    "1254102": "How are you spitting the data for train and validation? randomly ?",
    "1254265": "I split data into train and valid in Multi-Label Stratified KFold(K=5) manner, using whether or not each atom is included in molecular as multi-label.\n\nI'm sharing splitting process in [my notebook](https://www.kaggle.com/ttahara/bms-mt-chemical-formula-regression-training).",
    "1296173": "I have done some experiments with different Encoders (CNN models) and different percentage of the training data and I have concluded that experimenting with 10% training samples is ok and reliable when scaling. Maybe less than 10% is ok as well.\n\nBut for the final score and submission it's crucial to train with all data available, it is not worthy to reduce the training set."
  },
  "source": "meta"
}