{
  "id": 203411,
  "title": "saving data from test",
  "url": "/competitions/riiid-test-answer-prediction/discussion/203411",
  "author_name": "",
  "post_date": "2020-12-15T06:47:00.056103600Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>is it possible to save some data during the actual test iteration, so I can later see, for example, whether new questions and/or new users appear in the test (that have not appeared in the train), etc? if so, can you give advice on how to do that or link to a notebook that does that?</p>",
  "messages": [
    {
      "id": "1113067",
      "postDate": "12/15/2020 06:47:00",
      "content": "<p>is it possible to save some data during the actual test iteration, so I can later see, for example, whether new questions and/or new users appear in the test (that have not appeared in the train), etc? if so, can you give advice on how to do that or link to a notebook that does that?</p>",
      "rawMarkdown": "is it possible to save some data during the actual test iteration, so I can later see, for example, whether new questions and/or new users appear in the test (that have not appeared in the train), etc? if so, can you give advice on how to do that or link to a notebook that does that?",
      "votes": null
    },
    {
      "id": "1113666",
      "postDate": "12/15/2020 16:09:46",
      "content": "<p>Well I don't get the need of saving some of the data for rest of the testing process <br>\nSince it will be done in the training data when you use <br>\nXtrain,ytrain,Xtest,ytest<br>\n but still if you want some fresh data for testing and something not seen by the model in the training data <br>\nWhat you may do is<br>\nShorten some test data <br>\nSo you do <br>\nTrain=pd.read_csv('../kaggle/input/Train.csv')<br>\n now you use the code to save last 1000 rows seperately </p>\n<p>Train= Train.iloc[:-1000, :]<br>\nThen declare Xtrain, ytrain,Xtest,ytest<br>\nHence,<br>\nyou save the data not repeated in training and you use as a fresh data </p>\n<p>I hope this helps </p>",
      "rawMarkdown": "Well I don't get the need of saving some of the data for rest of the testing process \nSince it will be done in the training data when you use \nXtrain,ytrain,Xtest,ytest\n but still if you want some fresh data for testing and something not seen by the model in the training data \nWhat you may do is\nShorten some test data \nSo you do \nTrain=pd.read_csv('../kaggle/input/Train.csv')\n now you use the code to save last 1000 rows seperately \n\nTrain= Train.iloc[:-1000, :]\nThen declare Xtrain, ytrain,Xtest,ytest\nHence,\nyou save the data not repeated in training and you use as a fresh data \n\nI hope this helps",
      "votes": null
    },
    {
      "id": "1113918",
      "postDate": "12/15/2020 20:01:13",
      "content": "<p>I guess why you are asking so may be to know how an untrained observation will perform? <br>\nYou can use train, validation and test data. So, divide around 20% of the data as test data. Then, with the 80% of the remaining data divide it into train and validation data using train_test_split. I think this can help.</p>",
      "rawMarkdown": "I guess why you are asking so may be to know how an untrained observation will perform? \nYou can use train, validation and test data. So, divide around 20% of the data as test data. Then, with the 80% of the remaining data divide it into train and validation data using train_test_split. I think this can help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1113666,
      "author_name": "shyamgupta196",
      "author_url": "",
      "post_date": "12/15/2020 16:09:46",
      "content": "<p>Well I don't get the need of saving some of the data for rest of the testing process <br>\nSince it will be done in the training data when you use <br>\nXtrain,ytrain,Xtest,ytest<br>\n but still if you want some fresh data for testing and something not seen by the model in the training data <br>\nWhat you may do is<br>\nShorten some test data <br>\nSo you do <br>\nTrain=pd.read_csv('../kaggle/input/Train.csv')<br>\n now you use the code to save last 1000 rows seperately </p>\n<p>Train= Train.iloc[:-1000, :]<br>\nThen declare Xtrain, ytrain,Xtest,ytest<br>\nHence,<br>\nyou save the data not repeated in training and you use as a fresh data </p>\n<p>I hope this helps </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1113918,
      "author_name": "cassin4996",
      "author_url": "",
      "post_date": "12/15/2020 20:01:13",
      "content": "<p>I guess why you are asking so may be to know how an untrained observation will perform? <br>\nYou can use train, validation and test data. So, divide around 20% of the data as test data. Then, with the 80% of the remaining data divide it into train and validation data using train_test_split. I think this can help.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1113067": "is it possible to save some data during the actual test iteration, so I can later see, for example, whether new questions and/or new users appear in the test (that have not appeared in the train), etc? if so, can you give advice on how to do that or link to a notebook that does that?",
    "1113666": "Well I don't get the need of saving some of the data for rest of the testing process \nSince it will be done in the training data when you use \nXtrain,ytrain,Xtest,ytest\n but still if you want some fresh data for testing and something not seen by the model in the training data \nWhat you may do is\nShorten some test data \nSo you do \nTrain=pd.read_csv('../kaggle/input/Train.csv')\n now you use the code to save last 1000 rows seperately \n\nTrain= Train.iloc[:-1000, :]\nThen declare Xtrain, ytrain,Xtest,ytest\nHence,\nyou save the data not repeated in training and you use as a fresh data \n\nI hope this helps",
    "1113918": "I guess why you are asking so may be to know how an untrained observation will perform? \nYou can use train, validation and test data. So, divide around 20% of the data as test data. Then, with the 80% of the remaining data divide it into train and validation data using train_test_split. I think this can help."
  },
  "source": "meta"
}