{
  "id": 191191,
  "title": "Are you using all train data?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191191",
  "author_name": "",
  "post_date": "2020-10-15T05:54:50.840071700Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have tried to use all data to fit Keras model. Data were divided into 1M row chunks.</p>\n<p>Validation score graph for single pass was:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1508055%2F033427dd92251286006b93eab379fab0%2F1.png?generation=1602740880760447&amp;alt=media\" alt=\"Single pass\"></p>\n<p>When I did 3 epochs (feed the same data to model 3 times) I got:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1508055%2F775bd35aea309ec911e69577fda215db%2F2.png?generation=1602741201256798&amp;alt=media\" alt=\"3 epochs\"></p>\n<p>It seems there is no need to use all data, adding more epochs increases accuracy.</p>\n<p>Do you observe anything like this?</p>",
  "messages": [
    {
      "id": "1050150",
      "postDate": "10/15/2020 05:54:50",
      "content": "<p>I have tried to use all data to fit Keras model. Data were divided into 1M row chunks.</p>\n<p>Validation score graph for single pass was:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1508055%2F033427dd92251286006b93eab379fab0%2F1.png?generation=1602740880760447&amp;alt=media\" alt=\"Single pass\"></p>\n<p>When I did 3 epochs (feed the same data to model 3 times) I got:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1508055%2F775bd35aea309ec911e69577fda215db%2F2.png?generation=1602741201256798&amp;alt=media\" alt=\"3 epochs\"></p>\n<p>It seems there is no need to use all data, adding more epochs increases accuracy.</p>\n<p>Do you observe anything like this?</p>",
      "rawMarkdown": "I have tried to use all data to fit Keras model. Data were divided into 1M row chunks.\n\nValidation score graph for single pass was:\n![Single pass](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1508055%2F033427dd92251286006b93eab379fab0%2F1.png?generation=1602740880760447&alt=media)\n\nWhen I did 3 epochs (feed the same data to model 3 times) I got:\n![3 epochs](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1508055%2F775bd35aea309ec911e69577fda215db%2F2.png?generation=1602741201256798&alt=media)\n\nIt seems there is no need to use all data, adding more epochs increases accuracy.\n\nDo you observe anything like this?",
      "votes": null
    },
    {
      "id": "1050157",
      "postDate": "10/15/2020 06:00:40",
      "content": "<p>If the selected chunk can correctly represent the whole dataset, then you don't need to use all data to train the model. Thats a lucky situation.</p>",
      "rawMarkdown": "If the selected chunk can correctly represent the whole dataset, then you don't need to use all data to train the model. Thats a lucky situation.",
      "votes": null
    },
    {
      "id": "1050739",
      "postDate": "10/15/2020 17:17:26",
      "content": "<p>The idea is to use 80/20 for training and testing. Often we use validation data as well from that 80 percent.</p>",
      "rawMarkdown": "The idea is to use 80/20 for training and testing. Often we use validation data as well from that 80 percent.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1050157,
      "author_name": "mineshjethva",
      "author_url": "",
      "post_date": "10/15/2020 06:00:40",
      "content": "<p>If the selected chunk can correctly represent the whole dataset, then you don't need to use all data to train the model. Thats a lucky situation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1050739,
      "author_name": "mobasshir",
      "author_url": "",
      "post_date": "10/15/2020 17:17:26",
      "content": "<p>The idea is to use 80/20 for training and testing. Often we use validation data as well from that 80 percent.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1050150": "I have tried to use all data to fit Keras model. Data were divided into 1M row chunks.\n\nValidation score graph for single pass was:\n![Single pass](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1508055%2F033427dd92251286006b93eab379fab0%2F1.png?generation=1602740880760447&alt=media)\n\nWhen I did 3 epochs (feed the same data to model 3 times) I got:\n![3 epochs](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1508055%2F775bd35aea309ec911e69577fda215db%2F2.png?generation=1602741201256798&alt=media)\n\nIt seems there is no need to use all data, adding more epochs increases accuracy.\n\nDo you observe anything like this?",
    "1050157": "If the selected chunk can correctly represent the whole dataset, then you don't need to use all data to train the model. Thats a lucky situation.",
    "1050739": "The idea is to use 80/20 for training and testing. Often we use validation data as well from that 80 percent."
  },
  "source": "meta"
}