{
  "id": 329088,
  "title": "Why is there so much test data?",
  "url": "/competitions/amex-default-prediction/discussion/329088",
  "author_name": "Carl McBride Ellis",
  "post_date": "2022-06-04T17:38:16.754000",
  "votes": 40,
  "comment_count": 12,
  "views": 0,
  "content": "<p><strong>Aghhhh; …again!….</strong></p>\n<p><img src=\"https://raw.githubusercontent.com/Carl-McBride-Ellis/images_for_kaggle/main/Notebook_restarted.png\" alt=\"\"></p>\n<p>Why does this competition have so much test data?</p>\n<p>We have a lot of training data (ostensibly 16.39 GB) but maybe that is just an aspect of the challenge; exploring formats other than CSV, downcasting word lengths, reading in chunks, sub-sampling, garbage collection <em>etc</em>.</p>\n<p>But why do we need so much test data? (33.82 GB, almost twice as much as there is train data). Is not the purpose of the test data simply to provide a reliable leaderboard score to  three significant figures; and could that not  be achieved with a tenth (or indeed much less) of the test data we have been provided with? (and if not, perhaps a more \"stable\" evaluation metric?)</p>\n<p>Any feature engineering on the training data has to also be applied to the test data, and it is this process that is causing me even more red memory banners than the training phase. At least when training we can down-sample, but we need all of the <code>customer_ID</code> to create  a valid submission file. </p>\n<p>In a deployment situation we could evaluate the  model on the test data incrementally in batches whose size is commensurate with our resources, but here we have to get the job done in a 16GB CPU notebook all at once (or 13GB if one is using GPU).</p>\n<p>I am very much in favor of the recent suggestion by <a href=\"https://www.kaggle.com/jamesmcguigan\" target=\"_blank\">@jamesmcguigan</a> of <a href=\"https://www.kaggle.com/discussions/product-feedback/328720\" target=\"_blank\">32GB RAM Notebooks</a>, either that, or reducing the test dataset to a necessary and sufficient number of <code>customer_ID</code>.</p>\n<p>Anyway, thankfully other competitors have very kindly provided some excellent approaches and observations to this '<em>data engineering</em>' problem in the comments below.</p>\n<p>All the best,<br>\ncarl </p>",
  "messages": [
    {
      "id": 1811454,
      "postDate": "2022-06-04T17:38:16.753Z",
      "content": "<p><strong>Aghhhh; …again!….</strong></p>\n<p><img src=\"https://raw.githubusercontent.com/Carl-McBride-Ellis/images_for_kaggle/main/Notebook_restarted.png\" alt=\"\"></p>\n<p>Why does this competition have so much test data?</p>\n<p>We have a lot of training data (ostensibly 16.39 GB) but maybe that is just an aspect of the challenge; exploring formats other than CSV, downcasting word lengths, reading in chunks, sub-sampling, garbage collection <em>etc</em>.</p>\n<p>But why do we need so much test data? (33.82 GB, almost twice as much as there is train data). Is not the purpose of the test data simply to provide a reliable leaderboard score to  three significant figures; and could that not  be achieved with a tenth (or indeed much less) of the test data we have been provided with? (and if not, perhaps a more \"stable\" evaluation metric?)</p>\n<p>Any feature engineering on the training data has to also be applied to the test data, and it is this process that is causing me even more red memory banners than the training phase. At least when training we can down-sample, but we need all of the <code>customer_ID</code> to create  a valid submission file. </p>\n<p>In a deployment situation we could evaluate the  model on the test data incrementally in batches whose size is commensurate with our resources, but here we have to get the job done in a 16GB CPU notebook all at once (or 13GB if one is using GPU).</p>\n<p>I am very much in favor of the recent suggestion by <a href=\"https://www.kaggle.com/jamesmcguigan\" target=\"_blank\">@jamesmcguigan</a> of <a href=\"https://www.kaggle.com/discussions/product-feedback/328720\" target=\"_blank\">32GB RAM Notebooks</a>, either that, or reducing the test dataset to a necessary and sufficient number of <code>customer_ID</code>.</p>\n<p>Anyway, thankfully other competitors have very kindly provided some excellent approaches and observations to this '<em>data engineering</em>' problem in the comments below.</p>\n<p>All the best,<br>\ncarl </p>",
      "rawMarkdown": "**Aghhhh; ...again!....**\n\n![](https://raw.githubusercontent.com/Carl-McBride-Ellis/images_for_kaggle/main/Notebook_restarted.png)\n\nWhy does this competition have so much test data?\n\nWe have a lot of training data (ostensibly 16.39 GB) but maybe that is just an aspect of the challenge; exploring formats other than CSV, downcasting word lengths, reading in chunks, sub-sampling, garbage collection *etc*.\n\nBut why do we need so much test data? (33.82 GB, almost twice as much as there is train data). Is not the purpose of the test data simply to provide a reliable leaderboard score to  three significant figures; and could that not  be achieved with a tenth (or indeed much less) of the test data we have been provided with? (and if not, perhaps a more \"stable\" evaluation metric?)\n\nAny feature engineering on the training data has to also be applied to the test data, and it is this process that is causing me even more red memory banners than the training phase. At least when training we can down-sample, but we need all of the `customer_ID` to create  a valid submission file. \n\nIn a deployment situation we could evaluate the  model on the test data incrementally in batches whose size is commensurate with our resources, but here we have to get the job done in a 16GB CPU notebook all at once (or 13GB if one is using GPU).\n\nI am very much in favor of the recent suggestion by @jamesmcguigan of [32GB RAM Notebooks](https://www.kaggle.com/discussions/product-feedback/328720), either that, or reducing the test dataset to a necessary and sufficient number of `customer_ID`.\n\nAnyway, thankfully other competitors have very kindly provided some excellent approaches and observations to this '*data engineering*' problem in the comments below.\n\nAll the best,\ncarl ",
      "votes": 39
    },
    {
      "id": 1811782,
      "postDate": "2022-06-05T06:10:08.637Z",
      "content": "<p>The competition needs so much test data <strong>because recall-based metrics on unbalanced data are inherently noisy</strong>.</p>\n<p><img src=\"https://i.imgur.com/mEuAhVd.png\" alt=\"noisy_metric\"></p>\n<p>The noisiness of the metric can be shown in an experiment. I split the training data into three parts: A training set (80 %) and two validation sets (10 %, i.e. 45891 customers each). Then I trained my gradient booster for 2000 iterations and validated it on both validation sets. The diagram shows that iteration 700 gets a 0.002 higher score with validation set B than with validation set A. Iteration 1600 gets a 0.002 lower score with validation set B than with validation set A.</p>\n<p>If you take my two validation sets for public and private leaderboard, you'll get a huge shakeup - you could be at the top of the public leaderboard and end up 0.004 behind on the private leaderboard. The competition would amount to a lottery. Even with ten times more test data, we should be prepared for a shakeup.</p>\n<p>Why is the metric so noisy? Recall <em>(default rate captured at 4 %)</em> is defined as true positives divided by real positives. My validation sets have 11883 real positives each (10 % of 118828 positive customers with a stratified split). For both validation sets together, the model produces 16000 true positives. The true positives are assigned to the validation sets at random (i.e. not stratified). It is easily possible that one validation set gets 8050 true positives and the other one 7950. This distribution gives recalls of 8050/11883 = 0.677 and 7950/11883 = 0.669, respectively.</p>\n<p>With such a noisy metric, we need a lot of test data to distinguish good models from bad ones and to guarantee a fair competition.</p>\n<p>By the way, the noisiness of the metric explains why Kaggle shows only three digits of the score in the leaderboard: Any additional digit would be noise.</p>",
      "rawMarkdown": "The competition needs so much test data **because recall-based metrics on unbalanced data are inherently noisy**.\n\n![noisy_metric](https://i.imgur.com/mEuAhVd.png)\n\nThe noisiness of the metric can be shown in an experiment. I split the training data into three parts: A training set (80 %) and two validation sets (10 %, i.e. 45891 customers each). Then I trained my gradient booster for 2000 iterations and validated it on both validation sets. The diagram shows that iteration 700 gets a 0.002 higher score with validation set B than with validation set A. Iteration 1600 gets a 0.002 lower score with validation set B than with validation set A.\n\nIf you take my two validation sets for public and private leaderboard, you'll get a huge shakeup - you could be at the top of the public leaderboard and end up 0.004 behind on the private leaderboard. The competition would amount to a lottery. Even with ten times more test data, we should be prepared for a shakeup.\n\nWhy is the metric so noisy? Recall *(default rate captured at 4 %)* is defined as true positives divided by real positives. My validation sets have 11883 real positives each (10 % of 118828 positive customers with a stratified split). For both validation sets together, the model produces 16000 true positives. The true positives are assigned to the validation sets at random (i.e. not stratified). It is easily possible that one validation set gets 8050 true positives and the other one 7950. This distribution gives recalls of 8050/11883 = 0.677 and 7950/11883 = 0.669, respectively.\n\nWith such a noisy metric, we need a lot of test data to distinguish good models from bad ones and to guarantee a fair competition.\n\nBy the way, the noisiness of the metric explains why Kaggle shows only three digits of the score in the leaderboard: Any additional digit would be noise.",
      "votes": 26,
      "replies": [
        {
          "id": 1811804,
          "postDate": "2022-06-05T06:49:28.180Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>\n<p>Thank you for your magnificent insights, your work is always thought provoking, especially when read in conjunction with your topic <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">\"Graphical explanation of the competition metric\"</a>. </p>\n<p>If I understand correctly if it takes so much test data to obtain a stable LB score to three significant figures, what implication does this have for the reliability of our cross-validation scores and hold-out score, evidently obtained using substantially less data?</p>\n<p>All the best and many thanks,<br>\ncarl</p>",
          "rawMarkdown": "Dear @ambrosm \n\nThank you for your magnificent insights, your work is always thought provoking, especially when read in conjunction with your topic [\"Graphical explanation of the competition metric\"](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464). \n\nIf I understand correctly if it takes so much test data to obtain a stable LB score to three significant figures, what implication does this have for the reliability of our cross-validation scores and hold-out score, evidently obtained using substantially less data?\n\nAll the best and many thanks,\ncarl",
          "votes": 2
        },
        {
          "id": 1817737,
          "postDate": "2022-06-11T18:05:26.593Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>\n<p>PS: It is perhaps also worth mentioning that, as spotted both by Raddar, <a href=\"https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added/notebook\" target=\"_blank\">the data has random uniform noise added</a> and also by Chris Deotte in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\" target=\"_blank\">Strange Histograms</a>, so maybe to some extent the evaluation metric is not totally to blame. This uniform noise has a range of ± 0.005.</p>\n<p>This noise can be fairly easily extracted. However, it is perhaps also worth noting that having just a little added noise (or \"jitter\") is not necessarily a bad thing, and is sometimes it is even deliberately added to features as a mechanism to reduce overfitting.</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @ambrosm \n\nPS: It is perhaps also worth mentioning that, as spotted both by Raddar, [the data has random uniform noise added](https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added/notebook) and also by Chris Deotte in [Strange Histograms](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651), so maybe to some extent the evaluation metric is not totally to blame. This uniform noise has a range of ± 0.005.\n\n This noise can be fairly easily extracted. However, it is perhaps also worth noting that having just a little added noise (or \"jitter\") is not necessarily a bad thing, and is sometimes it is even deliberately added to features as a mechanism to reduce overfitting.\n\nAll the best,\ncarl",
          "votes": 2
        }
      ]
    },
    {
      "id": 1811528,
      "postDate": "2022-06-04T19:18:16.500Z",
      "content": "<p>When inferring test data in a Kaggle notebook, we can read a few million rows at a time and therefore read, process, and infer the test data in parts. </p>\n<pre><code>all_preds = []\nfor parts in range(NUM_PARTS):\n    test_part = pd.read_csv('test', nrows = ROWS, skiprows = SKIP)\n    test_part = process(test_part)\n    all_preds.append( infer(test_part) )\n</code></pre>\n<p>I provide an example in my XGB starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a> where I read test in 4 parts. When creating 3D data from the 2D provided CSV, then i infer test data in 20 chunks. I provide an example in my GRU starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "When inferring test data in a Kaggle notebook, we can read a few million rows at a time and therefore read, process, and infer the test data in parts. \n\n    all_preds = []\n    for parts in range(NUM_PARTS):\n        test_part = pd.read_csv('test', nrows = ROWS, skiprows = SKIP)\n        test_part = process(test_part)\n        all_preds.append( infer(test_part) )\n\nI provide an example in my XGB starter notebook [here][1] where I read test in 4 parts. When creating 3D data from the 2D provided CSV, then i infer test data in 20 chunks. I provide an example in my GRU starter notebook [here][2]\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790",
      "votes": 10,
      "replies": [
        {
          "id": 1811562,
          "postDate": "2022-06-04T20:28:57.170Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>Excellent suggestion. Worth being careful not to split any <code>customer_ID</code> when creating the chunks, otherwise it could create two predictions for the same <code>customer_ID</code> and  shift all the subsequent predictions incorrectly…</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @cdeotte \n\nExcellent suggestion. Worth being careful not to split any `customer_ID` when creating the chunks, otherwise it could create two predictions for the same `customer_ID` and  shift all the subsequent predictions incorrectly...\n\nAll the best,\ncarl",
          "votes": 2
        },
        {
          "id": 1811596,
          "postDate": "2022-06-04T21:32:58.387Z",
          "content": "<p>Yes correct. In my notebook first i read the entire column of <code>customer_ID</code>. Then i find which row positions we can split the list without splitting customers into two different parts. Then i read those rows from the file on disk.</p>",
          "rawMarkdown": "Yes correct. In my notebook first i read the entire column of `customer_ID`. Then i find which row positions we can split the list without splitting customers into two different parts. Then i read those rows from the file on disk.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1811543,
      "postDate": "2022-06-04T19:48:29.090Z",
      "content": "<p>I think the host has given a clear answer to this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327189\" target=\"_blank\">here</a>, emphasis mine.</p>\n<blockquote>\n  <p>We're excited to launch this competition and publish an <em>industrial scale</em> data for the community to stress test your [modelling] <em>ingenuity</em></p>\n</blockquote>\n<p>I think the constraints make this competition very interesting, there are a couple of stand-out datasets made by kagglers that can be used where precision and size is balanced allowing competitors to choose if they wish to go for less-is-more or brute-force feature engineering. </p>\n<p>Starter notebooks achieving very high LB scores also back this statement up! Splitting notebooks into training and inference is also the easiest way to expand on those already excellent notebooks.</p>",
      "rawMarkdown": "I think the host has given a clear answer to this [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327189), emphasis mine.\n>  We're excited to launch this competition and publish an *industrial scale* data for the community to stress test your [modelling] *ingenuity*\n\nI think the constraints make this competition very interesting, there are a couple of stand-out datasets made by kagglers that can be used where precision and size is balanced allowing competitors to choose if they wish to go for less-is-more or brute-force feature engineering. \n\nStarter notebooks achieving very high LB scores also back this statement up! Splitting notebooks into training and inference is also the easiest way to expand on those already excellent notebooks.",
      "votes": 3
    },
    {
      "id": 1811535,
      "postDate": "2022-06-04T19:31:30.230Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a>, on top of the usual memory tricks (float32) there is some unusual noise in the data changing some features from categorical to float (process is questionnable at best). <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> has posted a notebook to remove that noise (and an associated dataset). I use that dataset. After that, yes, the solution is to process data in chunks, for feature engineering and prediction. </p>",
      "rawMarkdown": "Hi @carlmcbrideellis, on top of the usual memory tricks (float32) there is some unusual noise in the data changing some features from categorical to float (process is questionnable at best). @raddar has posted a notebook to remove that noise (and an associated dataset). I use that dataset. After that, yes, the solution is to process data in chunks, for feature engineering and prediction. ",
      "votes": 3
    },
    {
      "id": 1811781,
      "postDate": "2022-06-05T06:07:47.423Z",
      "content": "<p><a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> Here are my notes on doing CPU / RAM optimization for large parquet files, including a pandas generator function for batch reading a large dataset (Bengali AI dataset was large but AMEX large). Unsure what needs to be changed to make it work for feather files rather than parquet.</p>\n<p>Also consider being explicit about your column dtypes, as there is a 4x difference between <code>float64</code> and <code>float16</code> when it comes to RAM optimization</p>\n<p><a href=\"https://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization/\" target=\"_blank\">https://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization/</a></p>",
      "rawMarkdown": "@carlmcbrideellis Here are my notes on doing CPU / RAM optimization for large parquet files, including a pandas generator function for batch reading a large dataset (Bengali AI dataset was large but AMEX large). Unsure what needs to be changed to make it work for feather files rather than parquet.\n\nAlso consider being explicit about your column dtypes, as there is a 4x difference between `float64` and `float16` when it comes to RAM optimization\n\nhttps://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization/",
      "votes": 2
    },
    {
      "id": 1811756,
      "postDate": "2022-06-05T05:34:34.120Z",
      "content": "<p>I agree. I mean .csv??? </p>\n<p>I wonder if they are already implementing the techniques fished out of the comp so far at AMEX. </p>\n<p>Camera pans to data engineers at AMEX reading through the comp discussions and notebooks wide eyed at these new \"parquet\" and \"feather\" techniques…</p>\n<p>Not to mention the lack of constraints on value types for many variables.</p>\n<p>I hope this data was not just scraped from their \"actual\" warehouse.</p>",
      "rawMarkdown": "I agree. I mean .csv??? \n\nI wonder if they are already implementing the techniques fished out of the comp so far at AMEX. \n\nCamera pans to data engineers at AMEX reading through the comp discussions and notebooks wide eyed at these new \"parquet\" and \"feather\" techniques...\n\nNot to mention the lack of constraints on value types for many variables.\n\nI hope this data was not just scraped from their \"actual\" warehouse.",
      "votes": 2,
      "replies": [
        {
          "id": 1811798,
          "postDate": "2022-06-05T06:37:59.713Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/krist0phersmith\" target=\"_blank\">@krist0phersmith</a> </p>\n<p><code>CSV</code>: Actually my very first notebook for this competition I set out to trying PySpark with SQL queries (having <a href=\"https://en.wikipedia.org/wiki/Lazy_evaluation\" target=\"_blank\">lazy evaluation</a>) by loading the CSV into a table. It is  worth mentioning that <a href=\"https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html\" target=\"_blank\">as of Spark 3.2 one can use pandas</a> via <code>import pyspark.pandas as ps</code> without the need for UDFs. If kaggle wants to add some data engineering flavor to competitions perhaps they could start moving away from monolithic CSV files and provide the competition data via a database? I am sure many people would enjoy trying out their SQL skills on kaggle!</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @krist0phersmith \n\n`CSV`: Actually my very first notebook for this competition I set out to trying PySpark with SQL queries (having [lazy evaluation](https://en.wikipedia.org/wiki/Lazy_evaluation)) by loading the CSV into a table. It is  worth mentioning that [as of Spark 3.2 one can use pandas](https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html) via `import pyspark.pandas as ps` without the need for UDFs. If kaggle wants to add some data engineering flavor to competitions perhaps they could start moving away from monolithic CSV files and provide the competition data via a database? I am sure many people would enjoy trying out their SQL skills on kaggle!\n\nAll the best,\ncarl\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 1815181,
      "postDate": "2022-06-08T19:03:28.027Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1811782,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2022-06-05T06:10:08.637000",
      "content": "<p>The competition needs so much test data <strong>because recall-based metrics on unbalanced data are inherently noisy</strong>.</p>\n<p><img src=\"https://i.imgur.com/mEuAhVd.png\" alt=\"noisy_metric\"></p>\n<p>The noisiness of the metric can be shown in an experiment. I split the training data into three parts: A training set (80 %) and two validation sets (10 %, i.e. 45891 customers each). Then I trained my gradient booster for 2000 iterations and validated it on both validation sets. The diagram shows that iteration 700 gets a 0.002 higher score with validation set B than with validation set A. Iteration 1600 gets a 0.002 lower score with validation set B than with validation set A.</p>\n<p>If you take my two validation sets for public and private leaderboard, you'll get a huge shakeup - you could be at the top of the public leaderboard and end up 0.004 behind on the private leaderboard. The competition would amount to a lottery. Even with ten times more test data, we should be prepared for a shakeup.</p>\n<p>Why is the metric so noisy? Recall <em>(default rate captured at 4 %)</em> is defined as true positives divided by real positives. My validation sets have 11883 real positives each (10 % of 118828 positive customers with a stratified split). For both validation sets together, the model produces 16000 true positives. The true positives are assigned to the validation sets at random (i.e. not stratified). It is easily possible that one validation set gets 8050 true positives and the other one 7950. This distribution gives recalls of 8050/11883 = 0.677 and 7950/11883 = 0.669, respectively.</p>\n<p>With such a noisy metric, we need a lot of test data to distinguish good models from bad ones and to guarantee a fair competition.</p>\n<p>By the way, the noisiness of the metric explains why Kaggle shows only three digits of the score in the leaderboard: Any additional digit would be noise.</p>",
      "votes": 26,
      "replies": [
        {
          "id": 1811804,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-06-05T06:49:28.180000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>\n<p>Thank you for your magnificent insights, your work is always thought provoking, especially when read in conjunction with your topic <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">\"Graphical explanation of the competition metric\"</a>. </p>\n<p>If I understand correctly if it takes so much test data to obtain a stable LB score to three significant figures, what implication does this have for the reliability of our cross-validation scores and hold-out score, evidently obtained using substantially less data?</p>\n<p>All the best and many thanks,<br>\ncarl</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1817737,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-06-11T18:05:26.593000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>\n<p>PS: It is perhaps also worth mentioning that, as spotted both by Raddar, <a href=\"https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added/notebook\" target=\"_blank\">the data has random uniform noise added</a> and also by Chris Deotte in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327651\" target=\"_blank\">Strange Histograms</a>, so maybe to some extent the evaluation metric is not totally to blame. This uniform noise has a range of ± 0.005.</p>\n<p>This noise can be fairly easily extracted. However, it is perhaps also worth noting that having just a little added noise (or \"jitter\") is not necessarily a bad thing, and is sometimes it is even deliberately added to features as a mechanism to reduce overfitting.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1811528,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-06-04T19:18:16.500000",
      "content": "<p>When inferring test data in a Kaggle notebook, we can read a few million rows at a time and therefore read, process, and infer the test data in parts. </p>\n<pre><code>all_preds = []\nfor parts in range(NUM_PARTS):\n    test_part = pd.read_csv('test', nrows = ROWS, skiprows = SKIP)\n    test_part = process(test_part)\n    all_preds.append( infer(test_part) )\n</code></pre>\n<p>I provide an example in my XGB starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a> where I read test in 4 parts. When creating 3D data from the 2D provided CSV, then i infer test data in 20 chunks. I provide an example in my GRU starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a></p>",
      "votes": 10,
      "replies": [
        {
          "id": 1811562,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-06-04T20:28:57.170000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>Excellent suggestion. Worth being careful not to split any <code>customer_ID</code> when creating the chunks, otherwise it could create two predictions for the same <code>customer_ID</code> and  shift all the subsequent predictions incorrectly…</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1811596,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-04T21:32:58.387000",
          "content": "<p>Yes correct. In my notebook first i read the entire column of <code>customer_ID</code>. Then i find which row positions we can split the list without splitting customers into two different parts. Then i read those rows from the file on disk.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1811543,
      "author_name": "Munum",
      "author_url": "",
      "post_date": "2022-06-04T19:48:29.090000",
      "content": "<p>I think the host has given a clear answer to this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327189\" target=\"_blank\">here</a>, emphasis mine.</p>\n<blockquote>\n  <p>We're excited to launch this competition and publish an <em>industrial scale</em> data for the community to stress test your [modelling] <em>ingenuity</em></p>\n</blockquote>\n<p>I think the constraints make this competition very interesting, there are a couple of stand-out datasets made by kagglers that can be used where precision and size is balanced allowing competitors to choose if they wish to go for less-is-more or brute-force feature engineering. </p>\n<p>Starter notebooks achieving very high LB scores also back this statement up! Splitting notebooks into training and inference is also the easiest way to expand on those already excellent notebooks.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1811535,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-06-04T19:31:30.230000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a>, on top of the usual memory tricks (float32) there is some unusual noise in the data changing some features from categorical to float (process is questionnable at best). <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> has posted a notebook to remove that noise (and an associated dataset). I use that dataset. After that, yes, the solution is to process data in chunks, for feature engineering and prediction. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1811781,
      "author_name": "James McGuigan",
      "author_url": "",
      "post_date": "2022-06-05T06:07:47.423000",
      "content": "<p><a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> Here are my notes on doing CPU / RAM optimization for large parquet files, including a pandas generator function for batch reading a large dataset (Bengali AI dataset was large but AMEX large). Unsure what needs to be changed to make it work for feather files rather than parquet.</p>\n<p>Also consider being explicit about your column dtypes, as there is a 4x difference between <code>float64</code> and <code>float16</code> when it comes to RAM optimization</p>\n<p><a href=\"https://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization/\" target=\"_blank\">https://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization/</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1811756,
      "author_name": "Kris Smith",
      "author_url": "",
      "post_date": "2022-06-05T05:34:34.120000",
      "content": "<p>I agree. I mean .csv??? </p>\n<p>I wonder if they are already implementing the techniques fished out of the comp so far at AMEX. </p>\n<p>Camera pans to data engineers at AMEX reading through the comp discussions and notebooks wide eyed at these new \"parquet\" and \"feather\" techniques…</p>\n<p>Not to mention the lack of constraints on value types for many variables.</p>\n<p>I hope this data was not just scraped from their \"actual\" warehouse.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1811798,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-06-05T06:37:59.713000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/krist0phersmith\" target=\"_blank\">@krist0phersmith</a> </p>\n<p><code>CSV</code>: Actually my very first notebook for this competition I set out to trying PySpark with SQL queries (having <a href=\"https://en.wikipedia.org/wiki/Lazy_evaluation\" target=\"_blank\">lazy evaluation</a>) by loading the CSV into a table. It is  worth mentioning that <a href=\"https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html\" target=\"_blank\">as of Spark 3.2 one can use pandas</a> via <code>import pyspark.pandas as ps</code> without the need for UDFs. If kaggle wants to add some data engineering flavor to competitions perhaps they could start moving away from monolithic CSV files and provide the competition data via a database? I am sure many people would enjoy trying out their SQL skills on kaggle!</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1815181,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-08T19:03:28.027000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1811454": "**Aghhhh; ...again!....**\n\n![](https://raw.githubusercontent.com/Carl-McBride-Ellis/images_for_kaggle/main/Notebook_restarted.png)\n\nWhy does this competition have so much test data?\n\nWe have a lot of training data (ostensibly 16.39 GB) but maybe that is just an aspect of the challenge; exploring formats other than CSV, downcasting word lengths, reading in chunks, sub-sampling, garbage collection *etc*.\n\nBut why do we need so much test data? (33.82 GB, almost twice as much as there is train data). Is not the purpose of the test data simply to provide a reliable leaderboard score to  three significant figures; and could that not  be achieved with a tenth (or indeed much less) of the test data we have been provided with? (and if not, perhaps a more \"stable\" evaluation metric?)\n\nAny feature engineering on the training data has to also be applied to the test data, and it is this process that is causing me even more red memory banners than the training phase. At least when training we can down-sample, but we need all of the `customer_ID` to create  a valid submission file. \n\nIn a deployment situation we could evaluate the  model on the test data incrementally in batches whose size is commensurate with our resources, but here we have to get the job done in a 16GB CPU notebook all at once (or 13GB if one is using GPU).\n\nI am very much in favor of the recent suggestion by @jamesmcguigan of [32GB RAM Notebooks](https://www.kaggle.com/discussions/product-feedback/328720), either that, or reducing the test dataset to a necessary and sufficient number of `customer_ID`.\n\nAnyway, thankfully other competitors have very kindly provided some excellent approaches and observations to this '*data engineering*' problem in the comments below.\n\nAll the best,\ncarl ",
    "1811782": "The competition needs so much test data **because recall-based metrics on unbalanced data are inherently noisy**.\n\n![noisy_metric](https://i.imgur.com/mEuAhVd.png)\n\nThe noisiness of the metric can be shown in an experiment. I split the training data into three parts: A training set (80 %) and two validation sets (10 %, i.e. 45891 customers each). Then I trained my gradient booster for 2000 iterations and validated it on both validation sets. The diagram shows that iteration 700 gets a 0.002 higher score with validation set B than with validation set A. Iteration 1600 gets a 0.002 lower score with validation set B than with validation set A.\n\nIf you take my two validation sets for public and private leaderboard, you'll get a huge shakeup - you could be at the top of the public leaderboard and end up 0.004 behind on the private leaderboard. The competition would amount to a lottery. Even with ten times more test data, we should be prepared for a shakeup.\n\nWhy is the metric so noisy? Recall *(default rate captured at 4 %)* is defined as true positives divided by real positives. My validation sets have 11883 real positives each (10 % of 118828 positive customers with a stratified split). For both validation sets together, the model produces 16000 true positives. The true positives are assigned to the validation sets at random (i.e. not stratified). It is easily possible that one validation set gets 8050 true positives and the other one 7950. This distribution gives recalls of 8050/11883 = 0.677 and 7950/11883 = 0.669, respectively.\n\nWith such a noisy metric, we need a lot of test data to distinguish good models from bad ones and to guarantee a fair competition.\n\nBy the way, the noisiness of the metric explains why Kaggle shows only three digits of the score in the leaderboard: Any additional digit would be noise.",
    "1811528": "When inferring test data in a Kaggle notebook, we can read a few million rows at a time and therefore read, process, and infer the test data in parts. \n\n    all_preds = []\n    for parts in range(NUM_PARTS):\n        test_part = pd.read_csv('test', nrows = ROWS, skiprows = SKIP)\n        test_part = process(test_part)\n        all_preds.append( infer(test_part) )\n\nI provide an example in my XGB starter notebook [here][1] where I read test in 4 parts. When creating 3D data from the 2D provided CSV, then i infer test data in 20 chunks. I provide an example in my GRU starter notebook [here][2]\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790",
    "1811543": "I think the host has given a clear answer to this [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327189), emphasis mine.\n>  We're excited to launch this competition and publish an *industrial scale* data for the community to stress test your [modelling] *ingenuity*\n\nI think the constraints make this competition very interesting, there are a couple of stand-out datasets made by kagglers that can be used where precision and size is balanced allowing competitors to choose if they wish to go for less-is-more or brute-force feature engineering. \n\nStarter notebooks achieving very high LB scores also back this statement up! Splitting notebooks into training and inference is also the easiest way to expand on those already excellent notebooks.",
    "1811535": "Hi @carlmcbrideellis, on top of the usual memory tricks (float32) there is some unusual noise in the data changing some features from categorical to float (process is questionnable at best). @raddar has posted a notebook to remove that noise (and an associated dataset). I use that dataset. After that, yes, the solution is to process data in chunks, for feature engineering and prediction. ",
    "1811781": "@carlmcbrideellis Here are my notes on doing CPU / RAM optimization for large parquet files, including a pandas generator function for batch reading a large dataset (Bengali AI dataset was large but AMEX large). Unsure what needs to be changed to make it work for feather files rather than parquet.\n\nAlso consider being explicit about your column dtypes, as there is a 4x difference between `float64` and `float16` when it comes to RAM optimization\n\nhttps://www.kaggle.com/code/jamesmcguigan/reading-parquet-files-ram-cpu-optimization/",
    "1811756": "I agree. I mean .csv??? \n\nI wonder if they are already implementing the techniques fished out of the comp so far at AMEX. \n\nCamera pans to data engineers at AMEX reading through the comp discussions and notebooks wide eyed at these new \"parquet\" and \"feather\" techniques...\n\nNot to mention the lack of constraints on value types for many variables.\n\nI hope this data was not just scraped from their \"actual\" warehouse.",
    "1815181": ""
  }
}