{
  "id": 399030,
  "title": "Optimizing run time of a cell",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/399030",
  "author_name": "",
  "post_date": "2023-04-02T05:52:55.168349300Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>My approach uses multiple linear regression and as a part of the submission to this contest, I considered using the coefficients I got from running on a different notebook used for training. The function used to get the required azimuth and zenith angles takes 0.518 s and forms the crucial part of the submission. I got to know that ~1400k entries make the public scoring dataset, and my cell needs to run in 0.06s to satisfy the given submission criteria. Is there a way to reduce the reading time of Parquet files to make it 10X faster as an alternative to using fastparquet or pandas.read_parquet , or to get the regression fit done 10X faster ? Please suggest on this.</p>",
  "messages": [
    {
      "id": "2205946",
      "postDate": "04/02/2023 05:52:55",
      "content": "<p>My approach uses multiple linear regression and as a part of the submission to this contest, I considered using the coefficients I got from running on a different notebook used for training. The function used to get the required azimuth and zenith angles takes 0.518 s and forms the crucial part of the submission. I got to know that ~1400k entries make the public scoring dataset, and my cell needs to run in 0.06s to satisfy the given submission criteria. Is there a way to reduce the reading time of Parquet files to make it 10X faster as an alternative to using fastparquet or pandas.read_parquet , or to get the regression fit done 10X faster ? Please suggest on this.</p>",
      "rawMarkdown": "My approach uses multiple linear regression and as a part of the submission to this contest, I considered using the coefficients I got from running on a different notebook used for training. The function used to get the required azimuth and zenith angles takes 0.518 s and forms the crucial part of the submission. I got to know that ~1400k entries make the public scoring dataset, and my cell needs to run in 0.06s to satisfy the given submission criteria. Is there a way to reduce the reading time of Parquet files to make it 10X faster as an alternative to using fastparquet or pandas.read_parquet , or to get the regression fit done 10X faster ? Please suggest on this.",
      "votes": null
    },
    {
      "id": "2206275",
      "postDate": "04/02/2023 12:55:27",
      "content": "<p>You should be able to read the full test dataset in under an hour.  I don't know what function you are evaluating that takes 0.518s per event, but that seems like the issue, not reading the dataset.  I would see if you can vectorize that function - that is, do not compute each event within a for loop, but try to use numpy or pandas to do a set of events at once.</p>",
      "rawMarkdown": "You should be able to read the full test dataset in under an hour.  I don't know what function you are evaluating that takes 0.518s per event, but that seems like the issue, not reading the dataset.  I would see if you can vectorize that function - that is, do not compute each event within a for loop, but try to use numpy or pandas to do a set of events at once.",
      "votes": null
    },
    {
      "id": "2206407",
      "postDate": "04/02/2023 14:29:03",
      "content": "<p>Ok, I will try vectorizing. Thanks a lot! I was using a for loop to fetch the events specified against each id in test_meta, though I am fetching events through a batch at once. </p>",
      "rawMarkdown": "Ok, I will try vectorizing. Thanks a lot! I was using a for loop to fetch the events specified against each id in test_meta, though I am fetching events through a batch at once.",
      "votes": null
    },
    {
      "id": "2210760",
      "postDate": "04/05/2023 15:48:17",
      "content": "<p><a href=\"https://www.kaggle.com/sumamallapragada\" target=\"_blank\">@sumamallapragada</a> Take a look at some of the public inference notebooks. They contain various ways the events are processed in a performant way.</p>\n<p>There is likely one of them that you can use to build upon and add your linear regression code into.</p>",
      "rawMarkdown": "sumamallapragada Take a look at some of the public inference notebooks. They contain various ways the events are processed in a performant way.\n\nThere is likely one of them that you can use to build upon and add your linear regression code into.",
      "votes": null
    },
    {
      "id": "2211332",
      "postDate": "04/06/2023 01:02:10",
      "content": "<p>Yeah, checking on them. Thank you!</p>",
      "rawMarkdown": "Yeah, checking on them. Thank you!",
      "votes": null
    },
    {
      "id": "2226829",
      "postDate": "04/19/2023 09:08:33",
      "content": "<p>You could do something like <br>\n`import numpy as np # linear algebra<br>\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)<br>\nimport pyarrow </p>\n<p>trainMeta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet', <br>\n                           engine=\"pyarrow\", use_threads=True, columns=['batch_id', 'event_id', 'azimuth', 'zenith'])<br>\ntestMeta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/test_meta.parquet', <br>\n                           engine=\"pyarrow\", use_threads=True, columns=['batch_id', 'event_id'])</p>\n<p>OUTPATH = '/kaggle/working/train_meta_batches/'<br>\nimport os<br>\nos.makedirs(OUTPATH)<br>\ndef get_metadata_for_batch(batch, write=False):<br>\n    metadata = trainMeta.loc[trainMeta['batch_id'] == batch]<br>\n    if write == True:<br>\n        outfile = OUTPATH + f'batch_{batch}.parquet'<br>\n        metadata.to_parquet(outfile)<br>\n    else:<br>\n        return metadata</p>\n<p>batches = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, ]    <br>\nfor i in batches:<br>\n    get_metadata_for_batch(i, write=True)<code>\n</code>OUTPATH = '/kaggle/working/test_meta_batches/'<br>\nimport os<br>\nos.makedirs(OUTPATH)<br>\nmetadata = testMeta.loc[testMeta['batch_id'] == 661]<br>\noutfile = OUTPATH + f'batch_{661}.parquet'<br>\nmetadata.to_parquet(outfile)`</p>\n<p>To not have to read the train_metadata set each time you run an actual regression fitting notebook but instead take the outputs from a notebooks that fragments the train_metadata praquet file into batch metadata files and then create a dataset out of that. I've only found reading train_meta.parquet to be time and memory intensive. Sorry for horrible formatting btw.</p>",
      "rawMarkdown": "You could do something like \n`import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport pyarrow \n\ntrainMeta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet', \n                           engine=\"pyarrow\", use_threads=True, columns=['batch_id', 'event_id', 'azimuth', 'zenith'])\ntestMeta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/test_meta.parquet', \n                           engine=\"pyarrow\", use_threads=True, columns=['batch_id', 'event_id'])\n\nOUTPATH = '/kaggle/working/train_meta_batches/'\nimport os\nos.makedirs(OUTPATH)\ndef get_metadata_for_batch(batch, write=False):\n    metadata = trainMeta.loc[trainMeta['batch_id'] == batch]\n    if write == True:\n        outfile = OUTPATH + f'batch_{batch}.parquet'\n        metadata.to_parquet(outfile)\n    else:\n        return metadata\n    \nbatches = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, ]    \nfor i in batches:\n    get_metadata_for_batch(i, write=True)`\n`OUTPATH = '/kaggle/working/test_meta_batches/'\nimport os\nos.makedirs(OUTPATH)\nmetadata = testMeta.loc[testMeta['batch_id'] == 661]\noutfile = OUTPATH + f'batch_{661}.parquet'\nmetadata.to_parquet(outfile)`\n\nTo not have to read the train_metadata set each time you run an actual regression fitting notebook but instead take the outputs from a notebooks that fragments the train_metadata praquet file into batch metadata files and then create a dataset out of that. I've only found reading train_meta.parquet to be time and memory intensive. Sorry for horrible formatting btw.",
      "votes": null
    },
    {
      "id": "2227052",
      "postDate": "04/19/2023 13:32:23",
      "content": "<p>Thanks for this approach. Will give it a try. I could optimize the run time to 0.026s per event using polars and other forms of vectorization, but the issue in submission seems to be due to integrity.</p>",
      "rawMarkdown": "Thanks for this approach. Will give it a try. I could optimize the run time to 0.026s per event using polars and other forms of vectorization, but the issue in submission seems to be due to integrity.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2206275,
      "author_name": "solverworld",
      "author_url": "",
      "post_date": "04/02/2023 12:55:27",
      "content": "<p>You should be able to read the full test dataset in under an hour.  I don't know what function you are evaluating that takes 0.518s per event, but that seems like the issue, not reading the dataset.  I would see if you can vectorize that function - that is, do not compute each event within a for loop, but try to use numpy or pandas to do a set of events at once.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2206407,
          "author_name": "sumamallapragada",
          "author_url": "",
          "post_date": "04/02/2023 14:29:03",
          "content": "<p>Ok, I will try vectorizing. Thanks a lot! I was using a for loop to fetch the events specified against each id in test_meta, though I am fetching events through a batch at once. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2210760,
              "author_name": "rsmits",
              "author_url": "",
              "post_date": "04/05/2023 15:48:17",
              "content": "<p><a href=\"https://www.kaggle.com/sumamallapragada\" target=\"_blank\">@sumamallapragada</a> Take a look at some of the public inference notebooks. They contain various ways the events are processed in a performant way.</p>\n<p>There is likely one of them that you can use to build upon and add your linear regression code into.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2211332,
                  "author_name": "sumamallapragada",
                  "author_url": "",
                  "post_date": "04/06/2023 01:02:10",
                  "content": "<p>Yeah, checking on them. Thank you!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2226829,
      "author_name": "monikagetsova",
      "author_url": "",
      "post_date": "04/19/2023 09:08:33",
      "content": "<p>You could do something like <br>\n`import numpy as np # linear algebra<br>\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)<br>\nimport pyarrow </p>\n<p>trainMeta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet', <br>\n                           engine=\"pyarrow\", use_threads=True, columns=['batch_id', 'event_id', 'azimuth', 'zenith'])<br>\ntestMeta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/test_meta.parquet', <br>\n                           engine=\"pyarrow\", use_threads=True, columns=['batch_id', 'event_id'])</p>\n<p>OUTPATH = '/kaggle/working/train_meta_batches/'<br>\nimport os<br>\nos.makedirs(OUTPATH)<br>\ndef get_metadata_for_batch(batch, write=False):<br>\n    metadata = trainMeta.loc[trainMeta['batch_id'] == batch]<br>\n    if write == True:<br>\n        outfile = OUTPATH + f'batch_{batch}.parquet'<br>\n        metadata.to_parquet(outfile)<br>\n    else:<br>\n        return metadata</p>\n<p>batches = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, ]    <br>\nfor i in batches:<br>\n    get_metadata_for_batch(i, write=True)<code>\n</code>OUTPATH = '/kaggle/working/test_meta_batches/'<br>\nimport os<br>\nos.makedirs(OUTPATH)<br>\nmetadata = testMeta.loc[testMeta['batch_id'] == 661]<br>\noutfile = OUTPATH + f'batch_{661}.parquet'<br>\nmetadata.to_parquet(outfile)`</p>\n<p>To not have to read the train_metadata set each time you run an actual regression fitting notebook but instead take the outputs from a notebooks that fragments the train_metadata praquet file into batch metadata files and then create a dataset out of that. I've only found reading train_meta.parquet to be time and memory intensive. Sorry for horrible formatting btw.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2227052,
          "author_name": "sumamallapragada",
          "author_url": "",
          "post_date": "04/19/2023 13:32:23",
          "content": "<p>Thanks for this approach. Will give it a try. I could optimize the run time to 0.026s per event using polars and other forms of vectorization, but the issue in submission seems to be due to integrity.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2205946": "My approach uses multiple linear regression and as a part of the submission to this contest, I considered using the coefficients I got from running on a different notebook used for training. The function used to get the required azimuth and zenith angles takes 0.518 s and forms the crucial part of the submission. I got to know that ~1400k entries make the public scoring dataset, and my cell needs to run in 0.06s to satisfy the given submission criteria. Is there a way to reduce the reading time of Parquet files to make it 10X faster as an alternative to using fastparquet or pandas.read_parquet , or to get the regression fit done 10X faster ? Please suggest on this.",
    "2206275": "You should be able to read the full test dataset in under an hour.  I don't know what function you are evaluating that takes 0.518s per event, but that seems like the issue, not reading the dataset.  I would see if you can vectorize that function - that is, do not compute each event within a for loop, but try to use numpy or pandas to do a set of events at once.",
    "2206407": "Ok, I will try vectorizing. Thanks a lot! I was using a for loop to fetch the events specified against each id in test_meta, though I am fetching events through a batch at once.",
    "2210760": "sumamallapragada Take a look at some of the public inference notebooks. They contain various ways the events are processed in a performant way.\n\nThere is likely one of them that you can use to build upon and add your linear regression code into.",
    "2211332": "Yeah, checking on them. Thank you!",
    "2226829": "You could do something like \n`import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport pyarrow \n\ntrainMeta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/train_meta.parquet', \n                           engine=\"pyarrow\", use_threads=True, columns=['batch_id', 'event_id', 'azimuth', 'zenith'])\ntestMeta = pd.read_parquet('/kaggle/input/icecube-neutrinos-in-deep-ice/test_meta.parquet', \n                           engine=\"pyarrow\", use_threads=True, columns=['batch_id', 'event_id'])\n\nOUTPATH = '/kaggle/working/train_meta_batches/'\nimport os\nos.makedirs(OUTPATH)\ndef get_metadata_for_batch(batch, write=False):\n    metadata = trainMeta.loc[trainMeta['batch_id'] == batch]\n    if write == True:\n        outfile = OUTPATH + f'batch_{batch}.parquet'\n        metadata.to_parquet(outfile)\n    else:\n        return metadata\n    \nbatches = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, ]    \nfor i in batches:\n    get_metadata_for_batch(i, write=True)`\n`OUTPATH = '/kaggle/working/test_meta_batches/'\nimport os\nos.makedirs(OUTPATH)\nmetadata = testMeta.loc[testMeta['batch_id'] == 661]\noutfile = OUTPATH + f'batch_{661}.parquet'\nmetadata.to_parquet(outfile)`\n\nTo not have to read the train_metadata set each time you run an actual regression fitting notebook but instead take the outputs from a notebooks that fragments the train_metadata praquet file into batch metadata files and then create a dataset out of that. I've only found reading train_meta.parquet to be time and memory intensive. Sorry for horrible formatting btw.",
    "2227052": "Thanks for this approach. Will give it a try. I could optimize the run time to 0.026s per event using polars and other forms of vectorization, but the issue in submission seems to be due to integrity."
  },
  "source": "meta"
}