{
  "id": 52739,
  "title": "Let's save bandwidth!",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52739",
  "author_name": "NxGTR",
  "post_date": "2018-03-22T16:45:41.080000",
  "votes": 17,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hey guys, <br>\nUnless I am missing something, anything after the 8th decimal point is useless, don't do it :).</p>\n\n<p>We all can save dear bits here and there that could be used for \"better\" stuff. <br>\nAnd by \"better\" you know what I mean.</p>",
  "messages": [
    {
      "id": 301348,
      "postDate": "2018-03-22T16:45:41.080Z",
      "content": "<p>Hey guys, <br>\nUnless I am missing something, anything after the 8th decimal point is useless, don't do it :).</p>\n\n<p>We all can save dear bits here and there that could be used for \"better\" stuff. <br>\nAnd by \"better\" you know what I mean.</p>",
      "rawMarkdown": "Hey guys,  \nUnless I am missing something, anything after the 8th decimal point is useless, don't do it :).\n\nWe all can save dear bits here and there that could be used for \"better\" stuff.  \nAnd by \"better\" you know what I mean.",
      "votes": 17
    },
    {
      "id": 302546,
      "postDate": "2018-03-24T09:29:47.080Z",
      "content": "<p>For pandas users reducing bandwidth is as easy as ...</p>\n\n<pre><code>pandasdataframe = pd.read_csv('abc.csv')\npandasdataframe.to_csv('abc.csv.gz', index=False,compression='gzip')\n</code></pre>",
      "rawMarkdown": "For pandas users reducing bandwidth is as easy as ...\n\n    pandasdataframe = pd.read_csv('abc.csv')\n    pandasdataframe.to_csv('abc.csv.gz', index=False,compression='gzip')",
      "votes": 11,
      "replies": [
        {
          "id": 302576,
          "postDate": "2018-03-24T11:16:52.840Z",
          "content": "<p>Thanks for the code !</p>",
          "rawMarkdown": "Thanks for the code !",
          "votes": 2
        },
        {
          "id": 304792,
          "postDate": "2018-03-28T02:46:03.897Z",
          "content": "<p>Is that means we can upload compressed file as submission？</p>",
          "rawMarkdown": "Is that means we can upload compressed file as submission？",
          "votes": 1
        },
        {
          "id": 306043,
          "postDate": "2018-03-29T19:30:00.847Z",
          "content": "<p>Yes, you can.</p>",
          "rawMarkdown": "Yes, you can.",
          "votes": 1
        }
      ]
    },
    {
      "id": 301942,
      "postDate": "2018-03-23T13:55:13.097Z",
      "content": "<p>Guys you could also use zip or 7zip formats. It also helps reduce the bandwith and upload time ;-)</p>\n\n<p>Update : 320MB submission file becomes 64MB in zip format and 40MB with 7z \ncompression takes about 2 minutes on my PC for 7z</p>",
      "rawMarkdown": "Guys you could also use zip or 7zip formats. It also helps reduce the bandwith and upload time ;-)\n\nUpdate : 320MB submission file becomes 64MB in zip format and 40MB with 7z \ncompression takes about 2 minutes on my PC for 7z",
      "votes": 9,
      "replies": [
        {
          "id": 302293,
          "postDate": "2018-03-23T21:41:47.610Z",
          "content": "<p>Is 7zip better than gzip?</p>",
          "rawMarkdown": "Is 7zip better than gzip?",
          "votes": 1
        },
        {
          "id": 302452,
          "postDate": "2018-03-24T04:34:16.967Z",
          "content": "<p>Yes. </p>",
          "rawMarkdown": "Yes. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 302547,
      "postDate": "2018-03-24T09:33:32.920Z",
      "content": "<pre><code>library(R.utils)\nlibrary(data.table)\n\n#Write gzip file\ndf &lt;- data.table(var1='Compress me',var2=', please!')\nfwrite(df,'filename.csv',sep=',')\ngzip('filename.csv',destname='filename.csv.gz')`\n\n#Read gzip file\nfread('gzip -dc filename.csv.gz')\n</code></pre>\n\n<p>Nicked from <a href=\"https://stackoverflow.com/questions/42788401/is-possible-to-use-fwrite-from-data-table-with-gzfile\">here</a></p>",
      "rawMarkdown": "    library(R.utils)\n    library(data.table)\n\n    #Write gzip file\n    df &lt;- data.table(var1='Compress me',var2=', please!')\n    fwrite(df,'filename.csv',sep=',')\n    gzip('filename.csv',destname='filename.csv.gz')`\n\n    #Read gzip file\n    fread('gzip -dc filename.csv.gz')\n\nNicked from [here][1]\n\n\n  [1]: https://stackoverflow.com/questions/42788401/is-possible-to-use-fwrite-from-data-table-with-gzfile",
      "votes": 8
    },
    {
      "id": 302317,
      "postDate": "2018-03-23T22:38:23.890Z",
      "content": "<p>I made <a href=\"https://www.kaggle.com/aharless/smallification\">a very short kernel</a> to convert a submission file to an equivalent gzipped ranked version.  In the example I used, it saves about 70% of the space.  In general, there could be significant digits after the 8th decimal place in the input file, but once the results are expressed in rank form, there is no need for more than 8 digits.  I figure, when I remember to do so, I will cut and paste the code from it into anything new I write that creates a submission file.</p>",
      "rawMarkdown": "I made [a very short kernel][1] to convert a submission file to an equivalent gzipped ranked version.  In the example I used, it saves about 70% of the space.  In general, there could be significant digits after the 8th decimal place in the input file, but once the results are expressed in rank form, there is no need for more than 8 digits.  I figure, when I remember to do so, I will cut and paste the code from it into anything new I write that creates a submission file.\n\n [1]: https://www.kaggle.com/aharless/smallification",
      "votes": 3,
      "replies": [
        {
          "id": 316737,
          "postDate": "2018-04-19T19:01:26.817Z",
          "content": "<p>Nice idea. I think it might be compressed even further by sorting by rank prior to compression. </p>",
          "rawMarkdown": "Nice idea. I think it might be compressed even further by sorting by rank prior to compression. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 301491,
      "postDate": "2018-03-22T20:22:03.777Z",
      "content": "<p>Only the sort order of the predictions matters, so how about an alternate submission file format where only the click_id is included, in the order of our predictions? Saves the scorer having to sort, and hopefully confuses the blenders for a while :)</p>",
      "rawMarkdown": "Only the sort order of the predictions matters, so how about an alternate submission file format where only the click_id is included, in the order of our predictions? Saves the scorer having to sort, and hopefully confuses the blenders for a while :)",
      "votes": 4,
      "replies": [
        {
          "id": 301519,
          "postDate": "2018-03-22T21:32:46.850Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 301548,
          "postDate": "2018-03-22T22:38:37.320Z",
          "content": "<p>The competition <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection#evaluation\">evaluation metric</a> is AUC, short for Area Under the <a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic\">ROC</a> curve.</p>\n\n<p>Essentially, the aim is to order the click_id's by the probability that the is_attributed target is true. A perfect score of 1 would result if your ordering of test set click_id's listed all the is_attributed==0 first, and all the is_attributed==1 afterwards.</p>\n\n<p>There's a fuller <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/Fast_Computation_of_AUC_ROC_score?lang=en\">explanation here</a>, with a faster Python implementation of AUC.</p>",
          "rawMarkdown": "The competition [evaluation metric][1] is AUC, short for Area Under the [ROC][2] curve.\n\nEssentially, the aim is to order the click_id's by the probability that the is_attributed target is true. A perfect score of 1 would result if your ordering of test set click_id's listed all the is_attributed==0 first, and all the is_attributed==1 afterwards.\n\nThere's a fuller [explanation here][3], with a faster Python implementation of AUC.\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection#evaluation\n  [2]: https://en.wikipedia.org/wiki/Receiver_operating_characteristic\n  [3]: https://www.ibm.com/developerworks/community/blogs/jfp/entry/Fast_Computation_of_AUC_ROC_score?lang=en\n",
          "votes": 2
        },
        {
          "id": 301554,
          "postDate": "2018-03-22T22:52:11.893Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 301599,
          "postDate": "2018-03-23T01:39:51.420Z",
          "content": "<p>Great idea.  It took me 2 hours to upload my first sub.</p>\n\n<p>Sure, making twice the same mistake in the file did not help, but still, upload is way too long.</p>",
          "rawMarkdown": "Great idea.  It took me 2 hours to upload my first sub.\n\nSure, making twice the same mistake in the file did not help, but still, upload is way too long."
        },
        {
          "id": 301635,
          "postDate": "2018-03-23T02:44:34.663Z",
          "content": "<p>And to think: the original test set was about three times the size! That would've meant ~1Gb CSV submissions and a six hour wait...</p>\n\n<p>Thinking about it some more, it would be very easy to support either a one column or two column submission format. If the submission uploaded has only one column the implicit 'is_attributed' could be added.</p>\n\n<p>i.e. assuming sub is single column sorted click_id</p>\n\n<pre><code>if 'is_attributed' not in sub.columns:\n    sub['is_attributed'] = np.arange(sub.shape[0])\n</code></pre>\n\n<p>... which could be sent to the existing scorer. (It'd be more efficient to use a different scoring routine that no longer has to do a sort.) Either format would work so it does not complicate things (for users) at all.</p>",
          "rawMarkdown": "And to think: the original test set was about three times the size! That would've meant ~1Gb CSV submissions and a six hour wait...\n\nThinking about it some more, it would be very easy to support either a one column or two column submission format. If the submission uploaded has only one column the implicit 'is_attributed' could be added.\n\ni.e. assuming sub is single column sorted click_id\n\n    if 'is_attributed' not in sub.columns:\n        sub['is_attributed'] = np.arange(sub.shape[0])\n\n... which could be sent to the existing scorer. (It'd be more efficient to use a different scoring routine that no longer has to do a sort.) Either format would work so it does not complicate things (for users) at all.\n"
        }
      ]
    },
    {
      "id": 316839,
      "postDate": "2018-04-20T03:50:09.647Z",
      "content": "<p>I'm using pd.hdf. \ndata.to_hdf('data.h5',key='1')\nIt saves much time.</p>",
      "rawMarkdown": "I'm using pd.hdf. \ndata.to_hdf('data.h5',key='1')\nIt saves much time.",
      "votes": 1,
      "replies": [
        {
          "id": 316844,
          "postDate": "2018-04-20T04:02:14.793Z",
          "content": "<p>It looks like the first letter in key should be alpha but not numeric.</p>",
          "rawMarkdown": "It looks like the first letter in key should be alpha but not numeric."
        },
        {
          "id": 316854,
          "postDate": "2018-04-20T04:47:52.843Z",
          "content": "<p>Numeric is ok.I have done it in my code.</p>",
          "rawMarkdown": "Numeric is ok.I have done it in my code."
        }
      ]
    },
    {
      "id": 301663,
      "postDate": "2018-03-23T04:06:10.803Z",
      "content": "<p>It Took me 2 hours to understand what you mean by 8 digit.</p>",
      "rawMarkdown": "It Took me 2 hours to understand what you mean by 8 digit.",
      "votes": 1
    },
    {
      "id": 302154,
      "postDate": "2018-03-23T18:35:34.980Z",
      "content": "<p>I have same question</p>",
      "rawMarkdown": "I have same question"
    },
    {
      "id": 301622,
      "postDate": "2018-03-23T02:12:47.540Z",
      "content": "<p>Has anyone noticed that the click_id is not in a monotonically increasing order? For example. 9 is before 8. Do we need to sort the submission based on the click_id?</p>",
      "rawMarkdown": "Has anyone noticed that the click_id is not in a monotonically increasing order? For example. 9 is before 8. Do we need to sort the submission based on the click_id?",
      "replies": [
        {
          "id": 301770,
          "postDate": "2018-03-23T08:05:08.633Z",
          "content": "<p>I noticed this too. But it seems that Kaggle takes care of sorting by <code>click_id</code>. I tried to submit with different orderings and got the same score both times. So don't worry about it.</p>",
          "rawMarkdown": "I noticed this too. But it seems that Kaggle takes care of sorting by `click_id`. I tried to submit with different orderings and got the same score both times. So don't worry about it.",
          "votes": 3
        }
      ]
    },
    {
      "id": 316821,
      "postDate": "2018-04-20T02:00:22.340Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 302546,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-03-24T09:29:47.080000",
      "content": "<p>For pandas users reducing bandwidth is as easy as ...</p>\n\n<pre><code>pandasdataframe = pd.read_csv('abc.csv')\npandasdataframe.to_csv('abc.csv.gz', index=False,compression='gzip')\n</code></pre>",
      "votes": 11,
      "replies": [
        {
          "id": 302576,
          "author_name": "Nathan Lauga",
          "author_url": "",
          "post_date": "2018-03-24T11:16:52.840000",
          "content": "<p>Thanks for the code !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 304792,
          "author_name": "Paul Yang",
          "author_url": "",
          "post_date": "2018-03-28T02:46:03.897000",
          "content": "<p>Is that means we can upload compressed file as submission？</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 306043,
          "author_name": "Luis Moneda",
          "author_url": "",
          "post_date": "2018-03-29T19:30:00.847000",
          "content": "<p>Yes, you can.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 301942,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-03-23T13:55:13.097000",
      "content": "<p>Guys you could also use zip or 7zip formats. It also helps reduce the bandwith and upload time ;-)</p>\n\n<p>Update : 320MB submission file becomes 64MB in zip format and 40MB with 7z \ncompression takes about 2 minutes on my PC for 7z</p>",
      "votes": 9,
      "replies": [
        {
          "id": 302293,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-23T21:41:47.610000",
          "content": "<p>Is 7zip better than gzip?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 302452,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-03-24T04:34:16.967000",
          "content": "<p>Yes. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 302547,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-03-24T09:33:32.920000",
      "content": "<pre><code>library(R.utils)\nlibrary(data.table)\n\n#Write gzip file\ndf &lt;- data.table(var1='Compress me',var2=', please!')\nfwrite(df,'filename.csv',sep=',')\ngzip('filename.csv',destname='filename.csv.gz')`\n\n#Read gzip file\nfread('gzip -dc filename.csv.gz')\n</code></pre>\n\n<p>Nicked from <a href=\"https://stackoverflow.com/questions/42788401/is-possible-to-use-fwrite-from-data-table-with-gzfile\">here</a></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 302317,
      "author_name": "Andy Harless",
      "author_url": "",
      "post_date": "2018-03-23T22:38:23.890000",
      "content": "<p>I made <a href=\"https://www.kaggle.com/aharless/smallification\">a very short kernel</a> to convert a submission file to an equivalent gzipped ranked version.  In the example I used, it saves about 70% of the space.  In general, there could be significant digits after the 8th decimal place in the input file, but once the results are expressed in rank form, there is no need for more than 8 digits.  I figure, when I remember to do so, I will cut and paste the code from it into anything new I write that creates a submission file.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 316737,
          "author_name": "Oscar Takeshita",
          "author_url": "",
          "post_date": "2018-04-19T19:01:26.817000",
          "content": "<p>Nice idea. I think it might be compressed even further by sorting by rank prior to compression. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 301491,
      "author_name": "James Trotman",
      "author_url": "",
      "post_date": "2018-03-22T20:22:03.777000",
      "content": "<p>Only the sort order of the predictions matters, so how about an alternate submission file format where only the click_id is included, in the order of our predictions? Saves the scorer having to sort, and hopefully confuses the blenders for a while :)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 301519,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-03-22T21:32:46.850000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 301548,
          "author_name": "James Trotman",
          "author_url": "",
          "post_date": "2018-03-22T22:38:37.320000",
          "content": "<p>The competition <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection#evaluation\">evaluation metric</a> is AUC, short for Area Under the <a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic\">ROC</a> curve.</p>\n\n<p>Essentially, the aim is to order the click_id's by the probability that the is_attributed target is true. A perfect score of 1 would result if your ordering of test set click_id's listed all the is_attributed==0 first, and all the is_attributed==1 afterwards.</p>\n\n<p>There's a fuller <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/Fast_Computation_of_AUC_ROC_score?lang=en\">explanation here</a>, with a faster Python implementation of AUC.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 301554,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-03-22T22:52:11.893000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 301599,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-03-23T01:39:51.420000",
          "content": "<p>Great idea.  It took me 2 hours to upload my first sub.</p>\n\n<p>Sure, making twice the same mistake in the file did not help, but still, upload is way too long.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 301635,
          "author_name": "James Trotman",
          "author_url": "",
          "post_date": "2018-03-23T02:44:34.663000",
          "content": "<p>And to think: the original test set was about three times the size! That would've meant ~1Gb CSV submissions and a six hour wait...</p>\n\n<p>Thinking about it some more, it would be very easy to support either a one column or two column submission format. If the submission uploaded has only one column the implicit 'is_attributed' could be added.</p>\n\n<p>i.e. assuming sub is single column sorted click_id</p>\n\n<pre><code>if 'is_attributed' not in sub.columns:\n    sub['is_attributed'] = np.arange(sub.shape[0])\n</code></pre>\n\n<p>... which could be sent to the existing scorer. (It'd be more efficient to use a different scoring routine that no longer has to do a sort.) Either format would work so it does not complicate things (for users) at all.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 316839,
      "author_name": "shark",
      "author_url": "",
      "post_date": "2018-04-20T03:50:09.647000",
      "content": "<p>I'm using pd.hdf. \ndata.to_hdf('data.h5',key='1')\nIt saves much time.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 316844,
          "author_name": "Oscar Takeshita",
          "author_url": "",
          "post_date": "2018-04-20T04:02:14.793000",
          "content": "<p>It looks like the first letter in key should be alpha but not numeric.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 316854,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-20T04:47:52.843000",
          "content": "<p>Numeric is ok.I have done it in my code.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 301663,
      "author_name": "Muhammad Alfiansyah",
      "author_url": "",
      "post_date": "2018-03-23T04:06:10.803000",
      "content": "<p>It Took me 2 hours to understand what you mean by 8 digit.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 302154,
      "author_name": "sbihero",
      "author_url": "",
      "post_date": "2018-03-23T18:35:34.980000",
      "content": "<p>I have same question</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 301622,
      "author_name": "Tapioca",
      "author_url": "",
      "post_date": "2018-03-23T02:12:47.540000",
      "content": "<p>Has anyone noticed that the click_id is not in a monotonically increasing order? For example. 9 is before 8. Do we need to sort the submission based on the click_id?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 301770,
          "author_name": "Max Halford",
          "author_url": "",
          "post_date": "2018-03-23T08:05:08.633000",
          "content": "<p>I noticed this too. But it seems that Kaggle takes care of sorting by <code>click_id</code>. I tried to submit with different orderings and got the same score both times. So don't worry about it.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 316821,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-20T02:00:22.340000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "301348": "Hey guys,  \nUnless I am missing something, anything after the 8th decimal point is useless, don't do it :).\n\nWe all can save dear bits here and there that could be used for \"better\" stuff.  \nAnd by \"better\" you know what I mean.",
    "302546": "For pandas users reducing bandwidth is as easy as ...\n\n    pandasdataframe = pd.read_csv('abc.csv')\n    pandasdataframe.to_csv('abc.csv.gz', index=False,compression='gzip')",
    "301942": "Guys you could also use zip or 7zip formats. It also helps reduce the bandwith and upload time ;-)\n\nUpdate : 320MB submission file becomes 64MB in zip format and 40MB with 7z \ncompression takes about 2 minutes on my PC for 7z",
    "302547": "    library(R.utils)\n    library(data.table)\n\n    #Write gzip file\n    df &lt;- data.table(var1='Compress me',var2=', please!')\n    fwrite(df,'filename.csv',sep=',')\n    gzip('filename.csv',destname='filename.csv.gz')`\n\n    #Read gzip file\n    fread('gzip -dc filename.csv.gz')\n\nNicked from [here][1]\n\n\n  [1]: https://stackoverflow.com/questions/42788401/is-possible-to-use-fwrite-from-data-table-with-gzfile",
    "302317": "I made [a very short kernel][1] to convert a submission file to an equivalent gzipped ranked version.  In the example I used, it saves about 70% of the space.  In general, there could be significant digits after the 8th decimal place in the input file, but once the results are expressed in rank form, there is no need for more than 8 digits.  I figure, when I remember to do so, I will cut and paste the code from it into anything new I write that creates a submission file.\n\n [1]: https://www.kaggle.com/aharless/smallification",
    "301491": "Only the sort order of the predictions matters, so how about an alternate submission file format where only the click_id is included, in the order of our predictions? Saves the scorer having to sort, and hopefully confuses the blenders for a while :)",
    "316839": "I'm using pd.hdf. \ndata.to_hdf('data.h5',key='1')\nIt saves much time.",
    "301663": "It Took me 2 hours to understand what you mean by 8 digit.",
    "302154": "I have same question",
    "301622": "Has anyone noticed that the click_id is not in a monotonically increasing order? For example. 9 is before 8. Do we need to sort the submission based on the click_id?",
    "316821": ""
  }
}