{
  "id": 441990,
  "title": "Time taken for the test data",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/441990",
  "author_name": "",
  "post_date": "2023-09-20T21:20:57.880076800Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>So I know it has to answer these sorts of questions without much context. I've developed my model which I seem quite happy with in preliminary testing but the slight problem arises with the time taken to allow the model to work on test data. <br>\nFor a million predictions it takes ~7000 seconds (from my testing of different sized subsets of data this was the average time for this size) and obviously there are so many more predictions to be made.</p>\n<p>Does anyone know any fundamental points I can go over to drastically reduce this time. </p>\n<p>For a little context, its trained and saved and tested on GPUs, handles the sequences as mapped arrays and for basic architecture it uses:</p>\n<p>self.conv1 = xxxx(num_features, 128)<br>\nself.conv2 = xxxx(128, 64)<br>\nself.fc = torch.nn.Linear(64, 1)</p>\n<p>p.s. this is a very interesting competition.</p>",
  "messages": [
    {
      "id": "2448921",
      "postDate": "09/20/2023 21:20:57",
      "content": "<p>So I know it has to answer these sorts of questions without much context. I've developed my model which I seem quite happy with in preliminary testing but the slight problem arises with the time taken to allow the model to work on test data. <br>\nFor a million predictions it takes ~7000 seconds (from my testing of different sized subsets of data this was the average time for this size) and obviously there are so many more predictions to be made.</p>\n<p>Does anyone know any fundamental points I can go over to drastically reduce this time. </p>\n<p>For a little context, its trained and saved and tested on GPUs, handles the sequences as mapped arrays and for basic architecture it uses:</p>\n<p>self.conv1 = xxxx(num_features, 128)<br>\nself.conv2 = xxxx(128, 64)<br>\nself.fc = torch.nn.Linear(64, 1)</p>\n<p>p.s. this is a very interesting competition.</p>",
      "rawMarkdown": "So I know it has to answer these sorts of questions without much context. I've developed my model which I seem quite happy with in preliminary testing but the slight problem arises with the time taken to allow the model to work on test data. \nFor a million predictions it takes ~7000 seconds (from my testing of different sized subsets of data this was the average time for this size) and obviously there are so many more predictions to be made.\n\nDoes anyone know any fundamental points I can go over to drastically reduce this time. \n\nFor a little context, its trained and saved and tested on GPUs, handles the sequences as mapped arrays and for basic architecture it uses:\n\nself.conv1 = xxxx(num_features, 128)\nself.conv2 = xxxx(128, 64)\nself.fc = torch.nn.Linear(64, 1)\n \np.s. this is a very interesting competition.",
      "votes": null
    },
    {
      "id": "2449498",
      "postDate": "09/21/2023 08:10:14",
      "content": "<p>Hi, I'v had the same problem, 7000 sec is to long….</p>\n<p>But now avarage time to submit near 40 minutes.</p>\n<p>Let's start.<br>\n1) We have 1 343 823 test sequences, so if we will pass every single sequence throunght our models, and then write result to file, it would take ~7000 seconds (:<br>\n2) Model works required near 85-90% of time, the rest 10-15% - for reading and saving data<br>\n3) I grouped test sequences in batches, and passed batch of sequences throught the models once.<br>\n4) Even average CPU can pass 10 sequences batch without wasting time, on GPU you can use batches up to 2000<br>\n5) So my method is: group sequences in batches=2000, then submit from kaggle notebook  with acceleration GPU T4 x2, it takes near 40 minutes :)</p>\n<p>Good Luck, Have Fun !</p>",
      "rawMarkdown": "Hi, I'v had the same problem, 7000 sec is to long....\n\nBut now avarage time to submit near 40 minutes.\n\nLet's start.\n1) We have 1 343 823 test sequences, so if we will pass every single sequence throunght our models, and then write result to file, it would take ~7000 seconds (:\n2) Model works required near 85-90% of time, the rest 10-15% - for reading and saving data\n3) I grouped test sequences in batches, and passed batch of sequences throught the models once.\n4) Even average CPU can pass 10 sequences batch without wasting time, on GPU you can use batches up to 2000\n5) So my method is: group sequences in batches=2000, then submit from kaggle notebook  with acceleration GPU T4 x2, it takes near 40 minutes :)\n\nGood Luck, Have Fun !",
      "votes": null
    },
    {
      "id": "2453185",
      "postDate": "09/23/2023 21:18:02",
      "content": "<p>Thanks! I've done this using GPU and batch size ~2000 but each time and modification of the code my GPU runs out! My thoughts are it's not properly utilising the GPU?</p>",
      "rawMarkdown": "Thanks! I've done this using GPU and batch size ~2000 but each time and modification of the code my GPU runs out! My thoughts are it's not properly utilising the GPU?",
      "votes": null
    },
    {
      "id": "2453418",
      "postDate": "09/24/2023 03:20:09",
      "content": "<p>Try different batch sizes, I'v ran submissions on kaggle CPU recently, it takes ~180 min with batch size 2000.</p>\n<p>Sometimes GPU runs out (out of memory), it hapens as garbage collector doesn't work properly on GPU's (in my case too).<br>\nYou should collect garbage manually:</p>\n<p>import gc<br>\ngc.collect()   # after every pass for example</p>\n<p>Good luck!</p>",
      "rawMarkdown": "Try different batch sizes, I'v ran submissions on kaggle CPU recently, it takes ~180 min with batch size 2000.\n\nSometimes GPU runs out (out of memory), it hapens as garbage collector doesn't work properly on GPU's (in my case too).\nYou should collect garbage manually:\n\nimport gc\ngc.collect()   # after every pass for example\n\nGood luck!",
      "votes": null
    },
    {
      "id": "2454066",
      "postDate": "09/24/2023 14:50:59",
      "content": "<p>Thank you! Good luck to you too!</p>",
      "rawMarkdown": "Thank you! Good luck to you too!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2449498,
      "author_name": "alexg5",
      "author_url": "",
      "post_date": "09/21/2023 08:10:14",
      "content": "<p>Hi, I'v had the same problem, 7000 sec is to long….</p>\n<p>But now avarage time to submit near 40 minutes.</p>\n<p>Let's start.<br>\n1) We have 1 343 823 test sequences, so if we will pass every single sequence throunght our models, and then write result to file, it would take ~7000 seconds (:<br>\n2) Model works required near 85-90% of time, the rest 10-15% - for reading and saving data<br>\n3) I grouped test sequences in batches, and passed batch of sequences throught the models once.<br>\n4) Even average CPU can pass 10 sequences batch without wasting time, on GPU you can use batches up to 2000<br>\n5) So my method is: group sequences in batches=2000, then submit from kaggle notebook  with acceleration GPU T4 x2, it takes near 40 minutes :)</p>\n<p>Good Luck, Have Fun !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2453185,
          "author_name": "marcusbrady",
          "author_url": "",
          "post_date": "09/23/2023 21:18:02",
          "content": "<p>Thanks! I've done this using GPU and batch size ~2000 but each time and modification of the code my GPU runs out! My thoughts are it's not properly utilising the GPU?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2453418,
              "author_name": "alexg5",
              "author_url": "",
              "post_date": "09/24/2023 03:20:09",
              "content": "<p>Try different batch sizes, I'v ran submissions on kaggle CPU recently, it takes ~180 min with batch size 2000.</p>\n<p>Sometimes GPU runs out (out of memory), it hapens as garbage collector doesn't work properly on GPU's (in my case too).<br>\nYou should collect garbage manually:</p>\n<p>import gc<br>\ngc.collect()   # after every pass for example</p>\n<p>Good luck!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2454066,
                  "author_name": "marcusbrady",
                  "author_url": "",
                  "post_date": "09/24/2023 14:50:59",
                  "content": "<p>Thank you! Good luck to you too!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2448921": "So I know it has to answer these sorts of questions without much context. I've developed my model which I seem quite happy with in preliminary testing but the slight problem arises with the time taken to allow the model to work on test data. \nFor a million predictions it takes ~7000 seconds (from my testing of different sized subsets of data this was the average time for this size) and obviously there are so many more predictions to be made.\n\nDoes anyone know any fundamental points I can go over to drastically reduce this time. \n\nFor a little context, its trained and saved and tested on GPUs, handles the sequences as mapped arrays and for basic architecture it uses:\n\nself.conv1 = xxxx(num_features, 128)\nself.conv2 = xxxx(128, 64)\nself.fc = torch.nn.Linear(64, 1)\n \np.s. this is a very interesting competition.",
    "2449498": "Hi, I'v had the same problem, 7000 sec is to long....\n\nBut now avarage time to submit near 40 minutes.\n\nLet's start.\n1) We have 1 343 823 test sequences, so if we will pass every single sequence throunght our models, and then write result to file, it would take ~7000 seconds (:\n2) Model works required near 85-90% of time, the rest 10-15% - for reading and saving data\n3) I grouped test sequences in batches, and passed batch of sequences throught the models once.\n4) Even average CPU can pass 10 sequences batch without wasting time, on GPU you can use batches up to 2000\n5) So my method is: group sequences in batches=2000, then submit from kaggle notebook  with acceleration GPU T4 x2, it takes near 40 minutes :)\n\nGood Luck, Have Fun !",
    "2453185": "Thanks! I've done this using GPU and batch size ~2000 but each time and modification of the code my GPU runs out! My thoughts are it's not properly utilising the GPU?",
    "2453418": "Try different batch sizes, I'v ran submissions on kaggle CPU recently, it takes ~180 min with batch size 2000.\n\nSometimes GPU runs out (out of memory), it hapens as garbage collector doesn't work properly on GPU's (in my case too).\nYou should collect garbage manually:\n\nimport gc\ngc.collect()   # after every pass for example\n\nGood luck!",
    "2454066": "Thank you! Good luck to you too!"
  },
  "source": "meta"
}