{
  "id": 446223,
  "title": "Quicker workflow for iteration?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/446223",
  "author_name": "",
  "post_date": "2023-10-10T22:53:52.273009300Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>What are teams doing to make they're workflow quicker for training as well as submission (the submission file being huge). </p>\n<p>Currently I'm using Kaggle notebooks, and for a small change, I'm having to re-run the entire notebook that involves reading the training csv files, building the model on the large dataset (which results in CPU utilization maxed out if run interactively) followed by the creation of the submissions.csv which takes another 30-40 minutes. </p>\n<p>I've read notebooks in which people pre-convert this data to parquet and operate on that, but it seems like Kaggle's environment is transient, and the parquet's saved do not persist. </p>\n<p>I'm looking for any solution that makes this process a little less painful for quicker model iteration. I also cannot afford too much money in terms of Cloud compute or Deep Learning Machines.</p>",
  "messages": [
    {
      "id": "2476747",
      "postDate": "10/10/2023 22:53:52",
      "content": "<p>What are teams doing to make they're workflow quicker for training as well as submission (the submission file being huge). </p>\n<p>Currently I'm using Kaggle notebooks, and for a small change, I'm having to re-run the entire notebook that involves reading the training csv files, building the model on the large dataset (which results in CPU utilization maxed out if run interactively) followed by the creation of the submissions.csv which takes another 30-40 minutes. </p>\n<p>I've read notebooks in which people pre-convert this data to parquet and operate on that, but it seems like Kaggle's environment is transient, and the parquet's saved do not persist. </p>\n<p>I'm looking for any solution that makes this process a little less painful for quicker model iteration. I also cannot afford too much money in terms of Cloud compute or Deep Learning Machines.</p>",
      "rawMarkdown": "What are teams doing to make they're workflow quicker for training as well as submission (the submission file being huge). \n\nCurrently I'm using Kaggle notebooks, and for a small change, I'm having to re-run the entire notebook that involves reading the training csv files, building the model on the large dataset (which results in CPU utilization maxed out if run interactively) followed by the creation of the submissions.csv which takes another 30-40 minutes. \n\nI've read notebooks in which people pre-convert this data to parquet and operate on that, but it seems like Kaggle's environment is transient, and the parquet's saved do not persist. \n\nI'm looking for any solution that makes this process a little less painful for quicker model iteration. I also cannot afford too much money in terms of Cloud compute or Deep Learning Machines.",
      "votes": null
    },
    {
      "id": "2476771",
      "postDate": "10/11/2023 00:09:18",
      "content": "<p>Hi! This competition is hardware hungry, but here is what you can do</p>\n<p>1) Monitor CV metric and submit only when it improves, - public LB correlates well with CV<br>\n2) Use polars to save .csv, - this library has built-in multi-threaded .csv write and also 3x faster than pandas when using single core, <a href=\"https://www.kaggle.com/code/martynoveduard/save-csv-faster-using-polars?scriptVersionId=146049633\" target=\"_blank\">example</a></p>",
      "rawMarkdown": "Hi! This competition is hardware hungry, but here is what you can do\n\n1) Monitor CV metric and submit only when it improves, - public LB correlates well with CV\n2) Use polars to save .csv, - this library has built-in multi-threaded .csv write and also 3x faster than pandas when using single core, [example](https://www.kaggle.com/code/martynoveduard/save-csv-faster-using-polars?scriptVersionId=146049633)",
      "votes": null
    },
    {
      "id": "2476885",
      "postDate": "10/11/2023 01:57:40",
      "content": "<p>Thanks for the demo!!! Love how fast polar writes. </p>",
      "rawMarkdown": "Thanks for the demo!!! Love how fast polar writes.",
      "votes": null
    },
    {
      "id": "2483647",
      "postDate": "10/15/2023 21:13:19",
      "content": "<blockquote>\n  <p>I've read notebooks in which people pre-convert this data to parquet and operate on that, but it seems like Kaggle's environment is transient, and the parquet's saved do not persist.</p>\n</blockquote>\n<p>You can preprocess the data in another notebook and save the output. Then you can import that output to another notebook and use it without re-creating the parquet from scratch every time.</p>",
      "rawMarkdown": ">I've read notebooks in which people pre-convert this data to parquet and operate on that, but it seems like Kaggle's environment is transient, and the parquet's saved do not persist.\n\nYou can preprocess the data in another notebook and save the output. Then you can import that output to another notebook and use it without re-creating the parquet from scratch every time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2476771,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "10/11/2023 00:09:18",
      "content": "<p>Hi! This competition is hardware hungry, but here is what you can do</p>\n<p>1) Monitor CV metric and submit only when it improves, - public LB correlates well with CV<br>\n2) Use polars to save .csv, - this library has built-in multi-threaded .csv write and also 3x faster than pandas when using single core, <a href=\"https://www.kaggle.com/code/martynoveduard/save-csv-faster-using-polars?scriptVersionId=146049633\" target=\"_blank\">example</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2476885,
          "author_name": "akashshah59",
          "author_url": "",
          "post_date": "10/11/2023 01:57:40",
          "content": "<p>Thanks for the demo!!! Love how fast polar writes. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2483647,
      "author_name": "vbogach",
      "author_url": "",
      "post_date": "10/15/2023 21:13:19",
      "content": "<blockquote>\n  <p>I've read notebooks in which people pre-convert this data to parquet and operate on that, but it seems like Kaggle's environment is transient, and the parquet's saved do not persist.</p>\n</blockquote>\n<p>You can preprocess the data in another notebook and save the output. Then you can import that output to another notebook and use it without re-creating the parquet from scratch every time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2476747": "What are teams doing to make they're workflow quicker for training as well as submission (the submission file being huge). \n\nCurrently I'm using Kaggle notebooks, and for a small change, I'm having to re-run the entire notebook that involves reading the training csv files, building the model on the large dataset (which results in CPU utilization maxed out if run interactively) followed by the creation of the submissions.csv which takes another 30-40 minutes. \n\nI've read notebooks in which people pre-convert this data to parquet and operate on that, but it seems like Kaggle's environment is transient, and the parquet's saved do not persist. \n\nI'm looking for any solution that makes this process a little less painful for quicker model iteration. I also cannot afford too much money in terms of Cloud compute or Deep Learning Machines.",
    "2476771": "Hi! This competition is hardware hungry, but here is what you can do\n\n1) Monitor CV metric and submit only when it improves, - public LB correlates well with CV\n2) Use polars to save .csv, - this library has built-in multi-threaded .csv write and also 3x faster than pandas when using single core, [example](https://www.kaggle.com/code/martynoveduard/save-csv-faster-using-polars?scriptVersionId=146049633)",
    "2476885": "Thanks for the demo!!! Love how fast polar writes.",
    "2483647": ">I've read notebooks in which people pre-convert this data to parquet and operate on that, but it seems like Kaggle's environment is transient, and the parquet's saved do not persist.\n\nYou can preprocess the data in another notebook and save the output. Then you can import that output to another notebook and use it without re-creating the parquet from scratch every time."
  },
  "source": "meta"
}