{
  "id": 2300,
  "title": "Submission Instructions??",
  "url": "/competitions/Raising-Money-to-Fund-an-Organizational-Mission/discussion/2300",
  "author_name": "",
  "post_date": "2012-08-02T20:22:46.120Z",
  "votes": null,
  "comment_count": 3,
  "views": 1876,
  "content": "<p>Hi,</p>\r\n<p>On the &quot;Make submission&quot; page it says:</p>\r\n<p>Your entry must:</p>\r\n<ul>\r\n<li>be in CSV format <em>(can be in a zip/gzip/rar/7z archive)</em> </li><li>have your prediction in column 3 </li><li>have exactly 5,146,737 rows </li></ul>\r\n<p>Are the first 2 columns GROUP, ID followed by PREDICTED_AMOUNT2?</p>\r\n<p>Does the submission need to be sorted in any particular order?</p>\r\n<p>&nbsp;</p>\r\n<p>Also, how can I create variables GROUP &amp; ID for the training set. I ask this since I have validation set on which I would like to test my model before making a leaderboard submission.</p>",
  "messages": [
    {
      "id": "12888",
      "postDate": "08/02/2012 20:22:46",
      "content": "<p>Hi,</p>\r\n<p>On the &quot;Make submission&quot; page it says:</p>\r\n<p>Your entry must:</p>\r\n<ul>\r\n<li>be in CSV format <em>(can be in a zip/gzip/rar/7z archive)</em> </li><li>have your prediction in column 3 </li><li>have exactly 5,146,737 rows </li></ul>\r\n<p>Are the first 2 columns GROUP, ID followed by PREDICTED_AMOUNT2?</p>\r\n<p>Does the submission need to be sorted in any particular order?</p>\r\n<p>&nbsp;</p>\r\n<p>Also, how can I create variables GROUP &amp; ID for the training set. I ask this since I have validation set on which I would like to test my model before making a leaderboard submission.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "12890",
      "postDate": "08/02/2012 23:00:56",
      "content": "<p>The solution format could be ID with PREDICTED<em>AMOUNT2, or just PREDICTED</em>AMOUNT2 in the right order.</p>\r\n<p>Group is a 1-1 mapping with &quot;jobid&quot;. I needed an integer for the back end scoring system, but you could ignore that and just use jobid.</p>\r\n<p>ID is a random number for each row of the test set. Shouldn't matter.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "12895",
      "postDate": "08/03/2012 06:19:29",
      "content": "<p>[quote=DavidC;12890]</p>\r\n<p>The solution format could be ID with PREDICTED<em>AMOUNT2, or just PREDICTED</em>AMOUNT2 in the right order.</p>\r\n<p>Group is a 1-1 mapping with &quot;jobid&quot;. I needed an integer for the back end scoring system, but you could ignore that and just use jobid.</p>\r\n<p>ID is a random number for each row of the test set. Shouldn't matter.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Hi,</p>\r\n<p>Are the submissions to be sorted by GROUP and then by Predicted AMOUNT2?</p>\r\n<p>&nbsp;</p>\r\n<p>I made my first submission with &lt;ID&gt;, &lt;Predicted Amount2&gt; and got AAT75=66339 and the following</p>\r\n<p>INFO: Assuming that column 1 with header value 'ID' maps to the required expected column 'amount2' (Line 1, Column 2)</p>\r\n<p>&nbsp;</p>\r\n<p>Then I made a second submission with just the &lt;Predicted Amount2&gt; column {note this is sorted by ID} and got AAT75 = 0.63781 and the following message</p>\r\n<p>INFO: Assuming that column 1 with header value 'DONATED_Pred' maps to the required expected column 'amount2' (Line 1, Column 2).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "12899",
      "postDate": "08/03/2012 16:37:23",
      "content": "<p>If it's just one column, then it needs to be sorted in the same way as the test data set (by &quot;id&quot;).</p>\r\n<p>If you include the ID, you should make sure the header names match what the system expects (&quot;id&quot; and &quot;amount2&quot;).</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 12890,
      "author_name": "dchudz",
      "author_url": "",
      "post_date": "08/02/2012 23:00:56",
      "content": "<p>The solution format could be ID with PREDICTED<em>AMOUNT2, or just PREDICTED</em>AMOUNT2 in the right order.</p>\r\n<p>Group is a 1-1 mapping with &quot;jobid&quot;. I needed an integer for the back end scoring system, but you could ignore that and just use jobid.</p>\r\n<p>ID is a random number for each row of the test set. Shouldn't matter.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 12895,
      "author_name": "sashikanthdareddy",
      "author_url": "",
      "post_date": "08/03/2012 06:19:29",
      "content": "<p>[quote=DavidC;12890]</p>\r\n<p>The solution format could be ID with PREDICTED<em>AMOUNT2, or just PREDICTED</em>AMOUNT2 in the right order.</p>\r\n<p>Group is a 1-1 mapping with &quot;jobid&quot;. I needed an integer for the back end scoring system, but you could ignore that and just use jobid.</p>\r\n<p>ID is a random number for each row of the test set. Shouldn't matter.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>Hi,</p>\r\n<p>Are the submissions to be sorted by GROUP and then by Predicted AMOUNT2?</p>\r\n<p>&nbsp;</p>\r\n<p>I made my first submission with &lt;ID&gt;, &lt;Predicted Amount2&gt; and got AAT75=66339 and the following</p>\r\n<p>INFO: Assuming that column 1 with header value 'ID' maps to the required expected column 'amount2' (Line 1, Column 2)</p>\r\n<p>&nbsp;</p>\r\n<p>Then I made a second submission with just the &lt;Predicted Amount2&gt; column {note this is sorted by ID} and got AAT75 = 0.63781 and the following message</p>\r\n<p>INFO: Assuming that column 1 with header value 'DONATED_Pred' maps to the required expected column 'amount2' (Line 1, Column 2).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 12899,
      "author_name": "dchudz",
      "author_url": "",
      "post_date": "08/03/2012 16:37:23",
      "content": "<p>If it's just one column, then it needs to be sorted in the same way as the test data set (by &quot;id&quot;).</p>\r\n<p>If you include the ID, you should make sure the header names match what the system expects (&quot;id&quot; and &quot;amount2&quot;).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "12888": "",
    "12890": "",
    "12895": "",
    "12899": ""
  },
  "source": "meta"
}