{
  "id": 55196,
  "title": "Noob remark about training data",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55196",
  "author_name": "Mihai Cvasnievschi",
  "post_date": "2018-04-23T11:43:55.939000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Probably most of you already thought about this, but I'm throwing it out as it was a huge improvement of my training.</p>\n\n<p>Just so you know, I'm training an NN and use batches - not sure if what I'm saying would make sens if you use other methods.</p>\n\n<p>I used to train my model using batches from training data. </p>\n\n<p>The main problem with training data is that is_attributed = 0 is obviously more frequent than is_attributed = 1. </p>\n\n<p>Being a noob, it took me some time to realize that I need to balance the data used during training. Without that, the NN model would basically learn to say that is_attributed is around 0. Add some noise, and the validation set would think you're scoring OK, when in reality you're not. </p>\n\n<p>Now I've created an addition training set (lets call it downloads set) that only contains is_attributed = 1 data. For each batch from the main training data I also sample an equal number of rows from downloads set. This way in the end, each batch would have an equal number of rows with is_attributed = 0 and 1. </p>\n\n<p>So, lesson learnt from my end. The question is, do you think it would improve your model if you'd have an equal number of rows with is_attributed=1 in your training set? Has somebody tried training by doing it?</p>\n\n<p>PS 1: I know, everybody is complaining about memory ... if your not using batching it would obviously almost double the memory usage. </p>\n\n<p>PS 2: Simple code for balancing the data, if you don't use any fancy data loader</p>\n\n<p>download_data = train_df[df.is_attributed == 1]</p>\n\n<p>...</p>\n\n<p>in batch loop</p>\n\n<p>df = train_df[start:end]</p>\n\n<p>df1 = download_data.sample(len(df[df.is_attributed == 0]) - len(df[df.is_attributed == 1]))</p>\n\n<p>df = pd.concat([df, df1]).sample(frac=1) #I always shuffle my data</p>",
  "messages": [
    {
      "id": 318203,
      "postDate": "2018-04-23T11:43:55.940Z",
      "content": "<p>Probably most of you already thought about this, but I'm throwing it out as it was a huge improvement of my training.</p>\n\n<p>Just so you know, I'm training an NN and use batches - not sure if what I'm saying would make sens if you use other methods.</p>\n\n<p>I used to train my model using batches from training data. </p>\n\n<p>The main problem with training data is that is_attributed = 0 is obviously more frequent than is_attributed = 1. </p>\n\n<p>Being a noob, it took me some time to realize that I need to balance the data used during training. Without that, the NN model would basically learn to say that is_attributed is around 0. Add some noise, and the validation set would think you're scoring OK, when in reality you're not. </p>\n\n<p>Now I've created an addition training set (lets call it downloads set) that only contains is_attributed = 1 data. For each batch from the main training data I also sample an equal number of rows from downloads set. This way in the end, each batch would have an equal number of rows with is_attributed = 0 and 1. </p>\n\n<p>So, lesson learnt from my end. The question is, do you think it would improve your model if you'd have an equal number of rows with is_attributed=1 in your training set? Has somebody tried training by doing it?</p>\n\n<p>PS 1: I know, everybody is complaining about memory ... if your not using batching it would obviously almost double the memory usage. </p>\n\n<p>PS 2: Simple code for balancing the data, if you don't use any fancy data loader</p>\n\n<p>download_data = train_df[df.is_attributed == 1]</p>\n\n<p>...</p>\n\n<p>in batch loop</p>\n\n<p>df = train_df[start:end]</p>\n\n<p>df1 = download_data.sample(len(df[df.is_attributed == 0]) - len(df[df.is_attributed == 1]))</p>\n\n<p>df = pd.concat([df, df1]).sample(frac=1) #I always shuffle my data</p>",
      "rawMarkdown": "Probably most of you already thought about this, but I'm throwing it out as it was a huge improvement of my training.\n\nJust so you know, I'm training an NN and use batches - not sure if what I'm saying would make sens if you use other methods.\n\nI used to train my model using batches from training data. \n\nThe main problem with training data is that is_attributed = 0 is obviously more frequent than is_attributed = 1. \n\nBeing a noob, it took me some time to realize that I need to balance the data used during training. Without that, the NN model would basically learn to say that is_attributed is around 0. Add some noise, and the validation set would think you're scoring OK, when in reality you're not. \n\nNow I've created an addition training set (lets call it downloads set) that only contains is_attributed = 1 data. For each batch from the main training data I also sample an equal number of rows from downloads set. This way in the end, each batch would have an equal number of rows with is_attributed = 0 and 1. \n\nSo, lesson learnt from my end. The question is, do you think it would improve your model if you'd have an equal number of rows with is_attributed=1 in your training set? Has somebody tried training by doing it?\n\nPS 1: I know, everybody is complaining about memory ... if your not using batching it would obviously almost double the memory usage. \n\nPS 2: Simple code for balancing the data, if you don't use any fancy data loader\n\n\ndownload_data = train_df[df.is_attributed == 1]\n\n...\n\nin batch loop\n\ndf = train_df[start:end]\n\ndf1 = download_data.sample(len(df[df.is_attributed == 0]) - len(df[df.is_attributed == 1]))\n\ndf = pd.concat([df, df1]).sample(frac=1) #I always shuffle my data\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "318203": "Probably most of you already thought about this, but I'm throwing it out as it was a huge improvement of my training.\n\nJust so you know, I'm training an NN and use batches - not sure if what I'm saying would make sens if you use other methods.\n\nI used to train my model using batches from training data. \n\nThe main problem with training data is that is_attributed = 0 is obviously more frequent than is_attributed = 1. \n\nBeing a noob, it took me some time to realize that I need to balance the data used during training. Without that, the NN model would basically learn to say that is_attributed is around 0. Add some noise, and the validation set would think you're scoring OK, when in reality you're not. \n\nNow I've created an addition training set (lets call it downloads set) that only contains is_attributed = 1 data. For each batch from the main training data I also sample an equal number of rows from downloads set. This way in the end, each batch would have an equal number of rows with is_attributed = 0 and 1. \n\nSo, lesson learnt from my end. The question is, do you think it would improve your model if you'd have an equal number of rows with is_attributed=1 in your training set? Has somebody tried training by doing it?\n\nPS 1: I know, everybody is complaining about memory ... if your not using batching it would obviously almost double the memory usage. \n\nPS 2: Simple code for balancing the data, if you don't use any fancy data loader\n\n\ndownload_data = train_df[df.is_attributed == 1]\n\n...\n\nin batch loop\n\ndf = train_df[start:end]\n\ndf1 = download_data.sample(len(df[df.is_attributed == 0]) - len(df[df.is_attributed == 1]))\n\ndf = pd.concat([df, df1]).sample(frac=1) #I always shuffle my data\n"
  }
}