{
  "id": 127877,
  "title": "How do you deal with such large amount of data?",
  "url": "/competitions/deepfake-detection-challenge/discussion/127877",
  "author_name": "dagnelies",
  "post_date": "2020-01-27T11:27:44.218000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,\nI'm just wondering how others do it, especially the ones up the ladder. I'm constantly facing a dilemma:\n- when I train quickly with small data, it quickly goes into overfitting \n- ...but when I take more data, every little \"experiment\" takes ages\nIn the end, it's quite unproductive.👀 \nSo, how do you do it? Do you tend to work on small chunks? Or lots of long running jobs? Got some tips to boost productivity or execution speed?</p>",
  "messages": [
    {
      "id": 730661,
      "postDate": "2020-01-27T18:54:14.217Z",
      "content": "<p>I believe your question is not quite as old as the oldest profession - but it's been around a long time.</p>\n\n<p>When I was looking at large data set 50+ years ago it was a huge question since the analysis for a simple regression was a paper, pencil and Monroe hand crank calculator job were \"quick\" was a couple of days.</p>\n\n<p>My approach is biased towards smallest possible sample size since over most of those 50+ years my compute was small.\n1.  Try anything new with 100 rows or less - ie, execution time in minute or two.\n2.  Increase rows in log fashion - 1000, 10000, etc.  Stop when the execution time in in your happy window. <br>\n3.  Low sample sizes should guide you but never use them to carve a new law in stone.\n4. Put your computer to bed with a large number of rows - carve your new laws in stone in the morning.</p>",
      "rawMarkdown": "I believe your question is not quite as old as the oldest profession - but it's been around a long time.\n\nWhen I was looking at large data set 50+ years ago it was a huge question since the analysis for a simple regression was a paper, pencil and Monroe hand crank calculator job were \"quick\" was a couple of days.\n\nMy approach is biased towards smallest possible sample size since over most of those 50+ years my compute was small.\n1.  Try anything new with 100 rows or less - ie, execution time in minute or two.\n2.  Increase rows in log fashion - 1000, 10000, etc.  Stop when the execution time in in your happy window.  \n3.  Low sample sizes should guide you but never use them to carve a new law in stone.\n4. Put your computer to bed with a large number of rows - carve your new laws in stone in the morning.\n\n"
    },
    {
      "id": 730350,
      "postDate": "2020-01-27T11:27:44.220Z",
      "content": "<p>Hi,\nI'm just wondering how others do it, especially the ones up the ladder. I'm constantly facing a dilemma:\n- when I train quickly with small data, it quickly goes into overfitting \n- ...but when I take more data, every little \"experiment\" takes ages\nIn the end, it's quite unproductive.👀 \nSo, how do you do it? Do you tend to work on small chunks? Or lots of long running jobs? Got some tips to boost productivity or execution speed?</p>",
      "rawMarkdown": "Hi,\nI'm just wondering how others do it, especially the ones up the ladder. I'm constantly facing a dilemma:\n- when I train quickly with small data, it quickly goes into overfitting \n- ...but when I take more data, every little \"experiment\" takes ages\nIn the end, it's quite unproductive.👀 \nSo, how do you do it? Do you tend to work on small chunks? Or lots of long running jobs? Got some tips to boost productivity or execution speed?"
    },
    {
      "id": 730381,
      "postDate": "2020-01-27T12:09:54.600Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 732640,
          "postDate": "2020-01-30T03:18:50.763Z",
          "content": "<p>Thanks, very helpful</p>",
          "rawMarkdown": "Thanks, very helpful"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 730661,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2020-01-27T18:54:14.217000",
      "content": "<p>I believe your question is not quite as old as the oldest profession - but it's been around a long time.</p>\n\n<p>When I was looking at large data set 50+ years ago it was a huge question since the analysis for a simple regression was a paper, pencil and Monroe hand crank calculator job were \"quick\" was a couple of days.</p>\n\n<p>My approach is biased towards smallest possible sample size since over most of those 50+ years my compute was small.\n1.  Try anything new with 100 rows or less - ie, execution time in minute or two.\n2.  Increase rows in log fashion - 1000, 10000, etc.  Stop when the execution time in in your happy window. <br>\n3.  Low sample sizes should guide you but never use them to carve a new law in stone.\n4. Put your computer to bed with a large number of rows - carve your new laws in stone in the morning.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 730381,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-27T12:09:54.600000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 732640,
          "author_name": "beeaware (looking for job)",
          "author_url": "",
          "post_date": "2020-01-30T03:18:50.763000",
          "content": "<p>Thanks, very helpful</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "730661": "I believe your question is not quite as old as the oldest profession - but it's been around a long time.\n\nWhen I was looking at large data set 50+ years ago it was a huge question since the analysis for a simple regression was a paper, pencil and Monroe hand crank calculator job were \"quick\" was a couple of days.\n\nMy approach is biased towards smallest possible sample size since over most of those 50+ years my compute was small.\n1.  Try anything new with 100 rows or less - ie, execution time in minute or two.\n2.  Increase rows in log fashion - 1000, 10000, etc.  Stop when the execution time in in your happy window.  \n3.  Low sample sizes should guide you but never use them to carve a new law in stone.\n4. Put your computer to bed with a large number of rows - carve your new laws in stone in the morning.\n\n",
    "730350": "Hi,\nI'm just wondering how others do it, especially the ones up the ladder. I'm constantly facing a dilemma:\n- when I train quickly with small data, it quickly goes into overfitting \n- ...but when I take more data, every little \"experiment\" takes ages\nIn the end, it's quite unproductive.👀 \nSo, how do you do it? Do you tend to work on small chunks? Or lots of long running jobs? Got some tips to boost productivity or execution speed?",
    "730381": ""
  }
}