{
  "id": 75373,
  "title": "Stable, Performant Code",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/75373",
  "author_name": "",
  "post_date": "2018-12-20T23:13:52.990487100Z",
  "votes": 29,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The dataset isn't larger than others we've hosted for competitions (~20 GB in memory), but the signal matrix is on the larger side (~23 billion data points). This can lead to issues with tools that have overhead per row or column, like pandas. In the spirit of @CPMP's <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398\">thread on efficient code</a> for Plasticc competition, here are a few tricks I found. I'm looking forward to hearing what else you find!</p>\n\n<ul>\n<li>Drop out of pandas and down to numpy for improved stability. This can be as simple as calling <code>np.max(df[col].values)</code> instead of <code>df[col].max()</code>. With pandas I experienced crashes from sequences of simple calculations, like checking the columnwise min and then the max, even though each operation only took a few seconds on my machine.</li>\n<li>There are some cases where numpy's overhead is excessive. When I was converting the raw signal files from csv to parquet, I was able to load the csv files 2-3 orders of magnitude faster with a ~10 line custom file parser than with either <code>pd.read_csv</code> or <code>np.loadtxt</code>.</li>\n<li>Cache your features and checkpoint your models. This can be as simple as littering your code with calls like <code>if os.path.exists('feature_1.csv') then load('feature_1.csv') else calculate_feature_1()</code> or as structured as using a dedicated pipeline tool such as <a href=\"https://luigi.readthedocs.io/en/stable/index.html\">Luigi</a>. I've personally found the Luigi API to be a bit of a hassle and ended up in the middle for my data preparation pipeline, however.</li>\n</ul>",
  "messages": [
    {
      "id": "443042",
      "postDate": "12/20/2018 23:13:52",
      "content": "<p>The dataset isn't larger than others we've hosted for competitions (~20 GB in memory), but the signal matrix is on the larger side (~23 billion data points). This can lead to issues with tools that have overhead per row or column, like pandas. In the spirit of @CPMP's <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398\">thread on efficient code</a> for Plasticc competition, here are a few tricks I found. I'm looking forward to hearing what else you find!</p>\n\n<ul>\n<li>Drop out of pandas and down to numpy for improved stability. This can be as simple as calling <code>np.max(df[col].values)</code> instead of <code>df[col].max()</code>. With pandas I experienced crashes from sequences of simple calculations, like checking the columnwise min and then the max, even though each operation only took a few seconds on my machine.</li>\n<li>There are some cases where numpy's overhead is excessive. When I was converting the raw signal files from csv to parquet, I was able to load the csv files 2-3 orders of magnitude faster with a ~10 line custom file parser than with either <code>pd.read_csv</code> or <code>np.loadtxt</code>.</li>\n<li>Cache your features and checkpoint your models. This can be as simple as littering your code with calls like <code>if os.path.exists('feature_1.csv') then load('feature_1.csv') else calculate_feature_1()</code> or as structured as using a dedicated pipeline tool such as <a href=\"https://luigi.readthedocs.io/en/stable/index.html\">Luigi</a>. I've personally found the Luigi API to be a bit of a hassle and ended up in the middle for my data preparation pipeline, however.</li>\n</ul>",
      "rawMarkdown": "The dataset isn't larger than others we've hosted for competitions (~20 GB in memory), but the signal matrix is on the larger side (~23 billion data points). This can lead to issues with tools that have overhead per row or column, like pandas. In the spirit of @CPMP's [thread on efficient code](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398) for Plasticc competition, here are a few tricks I found. I'm looking forward to hearing what else you find!\n\n- Drop out of pandas and down to numpy for improved stability. This can be as simple as calling `np.max(df[col].values)` instead of `df[col].max()`. With pandas I experienced crashes from sequences of simple calculations, like checking the columnwise min and then the max, even though each operation only took a few seconds on my machine.\n- There are some cases where numpy's overhead is excessive. When I was converting the raw signal files from csv to parquet, I was able to load the csv files 2-3 orders of magnitude faster with a ~10 line custom file parser than with either `pd.read_csv` or `np.loadtxt`.\n- Cache your features and checkpoint your models. This can be as simple as littering your code with calls like `if os.path.exists('feature_1.csv') then load('feature_1.csv') else calculate_feature_1()` or as structured as using a dedicated pipeline tool such as [Luigi](https://luigi.readthedocs.io/en/stable/index.html). I've personally found the Luigi API to be a bit of a hassle and ended up in the middle for my data preparation pipeline, however.",
      "votes": null
    },
    {
      "id": "454549",
      "postDate": "01/11/2019 20:22:09",
      "content": "<p>Thank you. what about dask? It can work with parquet as well</p>",
      "rawMarkdown": "Thank you. what about dask? It can work with parquet as well",
      "votes": null
    },
    {
      "id": "454583",
      "postDate": "01/11/2019 21:11:24",
      "content": "<p>I haven't tried dask with parquet.  The best speedups I've gotten from dask over pandas in the past were from loading the data, but parquet should already be using all available cores by default.</p>",
      "rawMarkdown": "I haven't tried dask with parquet.  The best speedups I've gotten from dask over pandas in the past were from loading the data, but parquet should already be using all available cores by default.",
      "votes": null
    },
    {
      "id": "471959",
      "postDate": "02/15/2019 06:27:21",
      "content": "<p>Thanks. </p>",
      "rawMarkdown": "Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454549,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "01/11/2019 20:22:09",
      "content": "<p>Thank you. what about dask? It can work with parquet as well</p>",
      "votes": null,
      "replies": [
        {
          "id": 454583,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "01/11/2019 21:11:24",
          "content": "<p>I haven't tried dask with parquet.  The best speedups I've gotten from dask over pandas in the past were from loading the data, but parquet should already be using all available cores by default.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471959,
      "author_name": "longyin2",
      "author_url": "",
      "post_date": "02/15/2019 06:27:21",
      "content": "<p>Thanks. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "443042": "The dataset isn't larger than others we've hosted for competitions (~20 GB in memory), but the signal matrix is on the larger side (~23 billion data points). This can lead to issues with tools that have overhead per row or column, like pandas. In the spirit of @CPMP's [thread on efficient code](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398) for Plasticc competition, here are a few tricks I found. I'm looking forward to hearing what else you find!\n\n- Drop out of pandas and down to numpy for improved stability. This can be as simple as calling `np.max(df[col].values)` instead of `df[col].max()`. With pandas I experienced crashes from sequences of simple calculations, like checking the columnwise min and then the max, even though each operation only took a few seconds on my machine.\n- There are some cases where numpy's overhead is excessive. When I was converting the raw signal files from csv to parquet, I was able to load the csv files 2-3 orders of magnitude faster with a ~10 line custom file parser than with either `pd.read_csv` or `np.loadtxt`.\n- Cache your features and checkpoint your models. This can be as simple as littering your code with calls like `if os.path.exists('feature_1.csv') then load('feature_1.csv') else calculate_feature_1()` or as structured as using a dedicated pipeline tool such as [Luigi](https://luigi.readthedocs.io/en/stable/index.html). I've personally found the Luigi API to be a bit of a hassle and ended up in the middle for my data preparation pipeline, however.",
    "454549": "Thank you. what about dask? It can work with parquet as well",
    "454583": "I haven't tried dask with parquet.  The best speedups I've gotten from dask over pandas in the past were from loading the data, but parquet should already be using all available cores by default.",
    "471959": "Thanks."
  },
  "source": "meta"
}