{
  "id": 18726,
  "title": "nolearn BatchIterator question",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18726",
  "author_name": "",
  "post_date": "2016-02-03T18:38:23.410Z",
  "votes": null,
  "comment_count": 1,
  "views": 886,
  "content": "<p>Hi all, I just started this competition and already learned a lot the past week.</p>\n\n<p>For those using nolearn in this competition, did you run into out of memory issues when trying to load the CSV file produced by the mxnet Python script?</p>\n\n<p>I ended up running out of memory on my AWS spot instance so I thought to try out subclassing the BatchIterator and load the single CSV in batches. This seemed to work as far as the iterator implementation is concerned but doesn't seem to fit well with how nolearn is generating the Train/Validation split. Mxnet uses two CSV iterators but I'm not sure if nolearn supports this type of input. I'm sure it's possible to change the internals of nolearn a bit, but I don't think I would generate a submission in time :)</p>\n\n<p>I wanted to use nolearn since I could do data augmentation on the fly. Mxnet supports this with the RecordIO iterator but it seems like the im2rec tool only works on images with 3 color channels (correct me if wrong here), when we could have 30 here.</p>\n\n<p>So I'm stuck in a bind where both Mxnet and nolearn seem to support some but not all of what I want to do.</p>\n\n<p>I think as a last resort I can write out the augmented data to files if using Mxnet, and split up the single input CSV into multiple ones if using nolearn, but was wondering if there was any off-the-shelf functionality I'm missing.</p>",
  "messages": [
    {
      "id": "106761",
      "postDate": "02/03/2016 18:38:23",
      "content": "<p>Hi all, I just started this competition and already learned a lot the past week.</p>\n\n<p>For those using nolearn in this competition, did you run into out of memory issues when trying to load the CSV file produced by the mxnet Python script?</p>\n\n<p>I ended up running out of memory on my AWS spot instance so I thought to try out subclassing the BatchIterator and load the single CSV in batches. This seemed to work as far as the iterator implementation is concerned but doesn't seem to fit well with how nolearn is generating the Train/Validation split. Mxnet uses two CSV iterators but I'm not sure if nolearn supports this type of input. I'm sure it's possible to change the internals of nolearn a bit, but I don't think I would generate a submission in time :)</p>\n\n<p>I wanted to use nolearn since I could do data augmentation on the fly. Mxnet supports this with the RecordIO iterator but it seems like the im2rec tool only works on images with 3 color channels (correct me if wrong here), when we could have 30 here.</p>\n\n<p>So I'm stuck in a bind where both Mxnet and nolearn seem to support some but not all of what I want to do.</p>\n\n<p>I think as a last resort I can write out the augmented data to files if using Mxnet, and split up the single input CSV into multiple ones if using nolearn, but was wondering if there was any off-the-shelf functionality I'm missing.</p>",
      "rawMarkdown": "Hi all, I just started this competition and already learned a lot the past week.\r\n\r\nFor those using nolearn in this competition, did you run into out of memory issues when trying to load the CSV file produced by the mxnet Python script?\r\n\r\nI ended up running out of memory on my AWS spot instance so I thought to try out subclassing the BatchIterator and load the single CSV in batches. This seemed to work as far as the iterator implementation is concerned but doesn't seem to fit well with how nolearn is generating the Train/Validation split. Mxnet uses two CSV iterators but I'm not sure if nolearn supports this type of input. I'm sure it's possible to change the internals of nolearn a bit, but I don't think I would generate a submission in time :)\r\n\r\nI wanted to use nolearn since I could do data augmentation on the fly. Mxnet supports this with the RecordIO iterator but it seems like the im2rec tool only works on images with 3 color channels (correct me if wrong here), when we could have 30 here.\r\n\r\nSo I'm stuck in a bind where both Mxnet and nolearn seem to support some but not all of what I want to do.\r\n\r\nI think as a last resort I can write out the augmented data to files if using Mxnet, and split up the single input CSV into multiple ones if using nolearn, but was wondering if there was any off-the-shelf functionality I'm missing.",
      "votes": null
    },
    {
      "id": "106923",
      "postDate": "02/04/2016 22:38:55",
      "content": "<p>I think I got both the train and test batch iterators working in nolearn now by splitting up the input CSV into files with 256 lines each. However, the training speed seems much slower than mxnet. I'm not sure if it's because mxnet is using the GPU better or it's because its doing concurrent I/O. I have a relatively simple network but its taking about 5 min per epoch with nolearn/lasagne/theano vs. 15 sec for mxnet. I believe it's a config issue on my side, but outside of trying to implement my own prefetch iterator, I'm not sure where to look.</p>",
      "rawMarkdown": "I think I got both the train and test batch iterators working in nolearn now by splitting up the input CSV into files with 256 lines each. However, the training speed seems much slower than mxnet. I'm not sure if it's because mxnet is using the GPU better or it's because its doing concurrent I/O. I have a relatively simple network but its taking about 5 min per epoch with nolearn/lasagne/theano vs. 15 sec for mxnet. I believe it's a config issue on my side, but outside of trying to implement my own prefetch iterator, I'm not sure where to look.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 106923,
      "author_name": "bshang",
      "author_url": "",
      "post_date": "02/04/2016 22:38:55",
      "content": "<p>I think I got both the train and test batch iterators working in nolearn now by splitting up the input CSV into files with 256 lines each. However, the training speed seems much slower than mxnet. I'm not sure if it's because mxnet is using the GPU better or it's because its doing concurrent I/O. I have a relatively simple network but its taking about 5 min per epoch with nolearn/lasagne/theano vs. 15 sec for mxnet. I believe it's a config issue on my side, but outside of trying to implement my own prefetch iterator, I'm not sure where to look.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "106761": "Hi all, I just started this competition and already learned a lot the past week.\r\n\r\nFor those using nolearn in this competition, did you run into out of memory issues when trying to load the CSV file produced by the mxnet Python script?\r\n\r\nI ended up running out of memory on my AWS spot instance so I thought to try out subclassing the BatchIterator and load the single CSV in batches. This seemed to work as far as the iterator implementation is concerned but doesn't seem to fit well with how nolearn is generating the Train/Validation split. Mxnet uses two CSV iterators but I'm not sure if nolearn supports this type of input. I'm sure it's possible to change the internals of nolearn a bit, but I don't think I would generate a submission in time :)\r\n\r\nI wanted to use nolearn since I could do data augmentation on the fly. Mxnet supports this with the RecordIO iterator but it seems like the im2rec tool only works on images with 3 color channels (correct me if wrong here), when we could have 30 here.\r\n\r\nSo I'm stuck in a bind where both Mxnet and nolearn seem to support some but not all of what I want to do.\r\n\r\nI think as a last resort I can write out the augmented data to files if using Mxnet, and split up the single input CSV into multiple ones if using nolearn, but was wondering if there was any off-the-shelf functionality I'm missing.",
    "106923": "I think I got both the train and test batch iterators working in nolearn now by splitting up the input CSV into files with 256 lines each. However, the training speed seems much slower than mxnet. I'm not sure if it's because mxnet is using the GPU better or it's because its doing concurrent I/O. I have a relatively simple network but its taking about 5 min per epoch with nolearn/lasagne/theano vs. 15 sec for mxnet. I believe it's a config issue on my side, but outside of trying to implement my own prefetch iterator, I'm not sure where to look."
  },
  "source": "meta"
}