{
  "id": 24272,
  "title": "creating a submission fast",
  "url": "/competitions/outbrain-click-prediction/discussion/24272",
  "author_name": "Amit",
  "post_date": "2016-10-10T23:18:29.763000",
  "votes": 0,
  "comment_count": 0,
  "views": 207,
  "content": "<p>I use R, but this time creating a submission file takes way too long with the processing needed. Here is what I do (and apologies for being lazy to make it a true &quot;kernel&quot; - feel free to do that):</p>\n\n<p>First, I take the test data given (with display_id and ad_id columns for the test.\nI then bind them with my scoring/probabilities (e.g. for each pair there is a score). I write this file as a csv file (it is a big file so it does take a few seconds, but not too many. the expensive part is the processing of the file to the right format which I do outside R.</p>\n\n<p>I go to my ubuntu/linux command line and type the following:</p>\n\n<pre><code>cat tmp.csv | sort -t&quot;,&quot; -k 1,1n -k 3,3rn | perl -aF, -ne 'chomp $F[1]; print &quot;\\n&quot; x (1 != $.), &quot;$F[0],&quot; if $l ne $F[0]; print &quot; &quot; x ($l eq $F[0]), $F[1]; $l = $F[0] }{ print &quot;\\n&quot;' &gt; output.csv\n</code></pre>\n\n<p>tmp.csv is the file wrote from R. the first part (the sort command) takes it and sorts it in the correct way, first grouping it by display_id then by decreasing probabilities within each display id.\nThe cryptic part next is a relatively simple perl script that converts it to the right format really fast, and puts it in the output.csv file.</p>\n\n<p>use at your own risk :)</p>",
  "messages": [
    {
      "id": 138806,
      "postDate": "2016-10-10T23:18:29.763Z",
      "content": "<p>I use R, but this time creating a submission file takes way too long with the processing needed. Here is what I do (and apologies for being lazy to make it a true &quot;kernel&quot; - feel free to do that):</p>\n\n<p>First, I take the test data given (with display_id and ad_id columns for the test.\nI then bind them with my scoring/probabilities (e.g. for each pair there is a score). I write this file as a csv file (it is a big file so it does take a few seconds, but not too many. the expensive part is the processing of the file to the right format which I do outside R.</p>\n\n<p>I go to my ubuntu/linux command line and type the following:</p>\n\n<pre><code>cat tmp.csv | sort -t&quot;,&quot; -k 1,1n -k 3,3rn | perl -aF, -ne 'chomp $F[1]; print &quot;\\n&quot; x (1 != $.), &quot;$F[0],&quot; if $l ne $F[0]; print &quot; &quot; x ($l eq $F[0]), $F[1]; $l = $F[0] }{ print &quot;\\n&quot;' &gt; output.csv\n</code></pre>\n\n<p>tmp.csv is the file wrote from R. the first part (the sort command) takes it and sorts it in the correct way, first grouping it by display_id then by decreasing probabilities within each display id.\nThe cryptic part next is a relatively simple perl script that converts it to the right format really fast, and puts it in the output.csv file.</p>\n\n<p>use at your own risk :)</p>",
      "rawMarkdown": "I use R, but this time creating a submission file takes way too long with the processing needed. Here is what I do (and apologies for being lazy to make it a true \"kernel\" - feel free to do that):\r\n\r\nFirst, I take the test data given (with display_id and ad_id columns for the test.\r\nI then bind them with my scoring/probabilities (e.g. for each pair there is a score). I write this file as a csv file (it is a big file so it does take a few seconds, but not too many. the expensive part is the processing of the file to the right format which I do outside R.\r\n\r\nI go to my ubuntu/linux command line and type the following:\r\n\r\n    cat tmp.csv | sort -t\",\" -k 1,1n -k 3,3rn | perl -aF, -ne 'chomp $F[1]; print \"\\n\" x (1 != $.), \"$F[0],\" if $l ne $F[0]; print \" \" x ($l eq $F[0]), $F[1]; $l = $F[0] }{ print \"\\n\"' > output.csv\r\n\r\ntmp.csv is the file wrote from R. the first part (the sort command) takes it and sorts it in the correct way, first grouping it by display_id then by decreasing probabilities within each display id.\r\nThe cryptic part next is a relatively simple perl script that converts it to the right format really fast, and puts it in the output.csv file.\r\n\r\nuse at your own risk :)\r\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "138806": "I use R, but this time creating a submission file takes way too long with the processing needed. Here is what I do (and apologies for being lazy to make it a true \"kernel\" - feel free to do that):\r\n\r\nFirst, I take the test data given (with display_id and ad_id columns for the test.\r\nI then bind them with my scoring/probabilities (e.g. for each pair there is a score). I write this file as a csv file (it is a big file so it does take a few seconds, but not too many. the expensive part is the processing of the file to the right format which I do outside R.\r\n\r\nI go to my ubuntu/linux command line and type the following:\r\n\r\n    cat tmp.csv | sort -t\",\" -k 1,1n -k 3,3rn | perl -aF, -ne 'chomp $F[1]; print \"\\n\" x (1 != $.), \"$F[0],\" if $l ne $F[0]; print \" \" x ($l eq $F[0]), $F[1]; $l = $F[0] }{ print \"\\n\"' > output.csv\r\n\r\ntmp.csv is the file wrote from R. the first part (the sort command) takes it and sorts it in the correct way, first grouping it by display_id then by decreasing probabilities within each display id.\r\nThe cryptic part next is a relatively simple perl script that converts it to the right format really fast, and puts it in the output.csv file.\r\n\r\nuse at your own risk :)\r\n"
  }
}