{
  "id": 165496,
  "title": "Efficiency",
  "url": "/competitions/birdsong-recognition/discussion/165496",
  "author_name": "",
  "post_date": "2020-07-10T00:14:06.632534900Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In the beginning of the competition, data loading was very slow for me. I later found a module more suited for loading;  now my kernel can load, create features, and train on  1000 samples within 4 1/2 minutes. I am not using a GPU or TPU. Firstly I would like to know how others kernels compare and what methods you're using to boost loading and training efficiency.\nSecondly, a question to the people running the site, how do you usually adjust the site or servers to make loading quicker or more accommodating apart from diverting more resources?</p>",
  "messages": [
    {
      "id": "922239",
      "postDate": "07/10/2020 00:14:06",
      "content": "<p>In the beginning of the competition, data loading was very slow for me. I later found a module more suited for loading;  now my kernel can load, create features, and train on  1000 samples within 4 1/2 minutes. I am not using a GPU or TPU. Firstly I would like to know how others kernels compare and what methods you're using to boost loading and training efficiency.\nSecondly, a question to the people running the site, how do you usually adjust the site or servers to make loading quicker or more accommodating apart from diverting more resources?</p>",
      "rawMarkdown": "In the beginning of the competition, data loading was very slow for me. I later found a module more suited for loading;  now my kernel can load, create features, and train on  1000 samples within 4 1/2 minutes. I am not using a GPU or TPU. Firstly I would like to know how others kernels compare and what methods you're using to boost loading and training efficiency.\nSecondly, a question to the people running the site, how do you usually adjust the site or servers to make loading quicker or more accommodating apart from diverting more resources?",
      "votes": null
    },
    {
      "id": "923068",
      "postDate": "07/10/2020 14:08:30",
      "content": "<p>I did all the preprocessing and saved it as a .npy file which took about 7 hours. Training was faster though. took me about 1 1/2 hours to train the model and get an score of about 47 percent. I think it's better to spend more time preprocessing than to do everything at training time.</p>",
      "rawMarkdown": "I did all the preprocessing and saved it as a .npy file which took about 7 hours. Training was faster though. took me about 1 1/2 hours to train the model and get an score of about 47 percent. I think it's better to spend more time preprocessing than to do everything at training time.",
      "votes": null
    },
    {
      "id": "923247",
      "postDate": "07/10/2020 16:37:34",
      "content": "<p>I keep my preprocessing the way I do, just because I want to make the process as easy as possible for myself to be able to keep track of it, use, and change, it's more of a preference. Also, while my training time would take around as much time to train as your model using all samples, my score on the leaderboard (0.544) is only reflective of training on exactly one thousand samples. You have me beat by a hidden digit though. If my 1000 &amp; 22,000 sample models perform at the same .544 and yours also the same 22,000 are .544, it my assumption is that we are both using relatively less significant features (unless you are using a neural net). I used mfcc in my first go around, how about you?</p>",
      "rawMarkdown": "I keep my preprocessing the way I do, just because I want to make the process as easy as possible for myself to be able to keep track of it, use, and change, it's more of a preference. Also, while my training time would take around as much time to train as your model using all samples, my score on the leaderboard (0.544) is only reflective of training on exactly one thousand samples. You have me beat by a hidden digit though. If my 1000 &amp; 22,000 sample models perform at the same .544 and yours also the same 22,000 are .544, it my assumption is that we are both using relatively less significant features (unless you are using a neural net). I used mfcc in my first go around, how about you?",
      "votes": null
    },
    {
      "id": "923282",
      "postDate": "07/10/2020 17:25:20",
      "content": "<p>I did use a neural network, a simple CNN with a few layers. It seemed to do okay. I had about 52% accuracy on the validation set and got around 0.48 on the LB</p>",
      "rawMarkdown": "I did use a neural network, a simple CNN with a few layers. It seemed to do okay. I had about 52% accuracy on the validation set and got around 0.48 on the LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 923068,
      "author_name": "jonykarki",
      "author_url": "",
      "post_date": "07/10/2020 14:08:30",
      "content": "<p>I did all the preprocessing and saved it as a .npy file which took about 7 hours. Training was faster though. took me about 1 1/2 hours to train the model and get an score of about 47 percent. I think it's better to spend more time preprocessing than to do everything at training time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 923247,
          "author_name": "eladwar",
          "author_url": "",
          "post_date": "07/10/2020 16:37:34",
          "content": "<p>I keep my preprocessing the way I do, just because I want to make the process as easy as possible for myself to be able to keep track of it, use, and change, it's more of a preference. Also, while my training time would take around as much time to train as your model using all samples, my score on the leaderboard (0.544) is only reflective of training on exactly one thousand samples. You have me beat by a hidden digit though. If my 1000 &amp; 22,000 sample models perform at the same .544 and yours also the same 22,000 are .544, it my assumption is that we are both using relatively less significant features (unless you are using a neural net). I used mfcc in my first go around, how about you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 923282,
          "author_name": "jonykarki",
          "author_url": "",
          "post_date": "07/10/2020 17:25:20",
          "content": "<p>I did use a neural network, a simple CNN with a few layers. It seemed to do okay. I had about 52% accuracy on the validation set and got around 0.48 on the LB</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "922239": "In the beginning of the competition, data loading was very slow for me. I later found a module more suited for loading;  now my kernel can load, create features, and train on  1000 samples within 4 1/2 minutes. I am not using a GPU or TPU. Firstly I would like to know how others kernels compare and what methods you're using to boost loading and training efficiency.\nSecondly, a question to the people running the site, how do you usually adjust the site or servers to make loading quicker or more accommodating apart from diverting more resources?",
    "923068": "I did all the preprocessing and saved it as a .npy file which took about 7 hours. Training was faster though. took me about 1 1/2 hours to train the model and get an score of about 47 percent. I think it's better to spend more time preprocessing than to do everything at training time.",
    "923247": "I keep my preprocessing the way I do, just because I want to make the process as easy as possible for myself to be able to keep track of it, use, and change, it's more of a preference. Also, while my training time would take around as much time to train as your model using all samples, my score on the leaderboard (0.544) is only reflective of training on exactly one thousand samples. You have me beat by a hidden digit though. If my 1000 &amp; 22,000 sample models perform at the same .544 and yours also the same 22,000 are .544, it my assumption is that we are both using relatively less significant features (unless you are using a neural net). I used mfcc in my first go around, how about you?",
    "923282": "I did use a neural network, a simple CNN with a few layers. It seemed to do okay. I had about 52% accuracy on the validation set and got around 0.48 on the LB"
  },
  "source": "meta"
}