{
  "id": 263635,
  "title": "Time out concern for the private testing data",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/263635",
  "author_name": "Shanshan Yu",
  "post_date": "2021-08-09T22:10:49.293000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The rule says the private and public testing data is ~ 5x than the training dataset. Because the training dataset has 582 patients and 400k files, can we assume the private testing set has 2500+ patients and over 4 million dicom files?<br>\nI did some code analysis and find out the following:</p>\n<p>public_training_patient = 583<br>\npublic_test_patient = 87<br>\npublic_test_files = 51473<br>\npublic_training_files = 348641</p>\n<p>Then for the private sets, it will be as following:<br>\nprivate_test_patient = public_training_patient * 5 - public_test_patient<br>\nprivate_test_files = public_training_files * 5 - public_test_files</p>\n<p>Because we have 9 hours to predict the private data(maybe), for each patient and file, we can calculate the pred limited time by:<br>\ntime_per_img_file = (9 * 3600) / private_test_files<br>\ntime_per_patient = (9 * 3600) / private_test_patient</p>\n<p>Then I got:<br>\ntime_per_img_file = 0.02 seconds<br>\ntime_per_patient = 11.46 seconds</p>\n<p>It is quite not enough for using four folders' MRI to predict the result. I spend ~ 15 mins for 87 predictions, each patient cost around 10 seconds. But I get rid of half images for each patient and only trained one model for the FLAIR class.</p>",
  "messages": [
    {
      "id": 1462529,
      "postDate": "2021-08-09T22:10:49.293Z",
      "content": "<p>The rule says the private and public testing data is ~ 5x than the training dataset. Because the training dataset has 582 patients and 400k files, can we assume the private testing set has 2500+ patients and over 4 million dicom files?<br>\nI did some code analysis and find out the following:</p>\n<p>public_training_patient = 583<br>\npublic_test_patient = 87<br>\npublic_test_files = 51473<br>\npublic_training_files = 348641</p>\n<p>Then for the private sets, it will be as following:<br>\nprivate_test_patient = public_training_patient * 5 - public_test_patient<br>\nprivate_test_files = public_training_files * 5 - public_test_files</p>\n<p>Because we have 9 hours to predict the private data(maybe), for each patient and file, we can calculate the pred limited time by:<br>\ntime_per_img_file = (9 * 3600) / private_test_files<br>\ntime_per_patient = (9 * 3600) / private_test_patient</p>\n<p>Then I got:<br>\ntime_per_img_file = 0.02 seconds<br>\ntime_per_patient = 11.46 seconds</p>\n<p>It is quite not enough for using four folders' MRI to predict the result. I spend ~ 15 mins for 87 predictions, each patient cost around 10 seconds. But I get rid of half images for each patient and only trained one model for the FLAIR class.</p>",
      "rawMarkdown": "The rule says the private and public testing data is ~ 5x than the training dataset. Because the training dataset has 582 patients and 400k files, can we assume the private testing set has 2500+ patients and over 4 million dicom files?\nI did some code analysis and find out the following:\n\npublic_training_patient = 583\npublic_test_patient = 87\npublic_test_files = 51473\npublic_training_files = 348641\n\nThen for the private sets, it will be as following:\nprivate_test_patient = public_training_patient * 5 - public_test_patient\nprivate_test_files = public_training_files * 5 - public_test_files\n\nBecause we have 9 hours to predict the private data(maybe), for each patient and file, we can calculate the pred limited time by:\ntime_per_img_file = (9 * 3600) / private_test_files\ntime_per_patient = (9 * 3600) / private_test_patient\n\nThen I got:\ntime_per_img_file = 0.02 seconds\ntime_per_patient = 11.46 seconds\n\nIt is quite not enough for using four folders' MRI to predict the result. I spend ~ 15 mins for 87 predictions, each patient cost around 10 seconds. But I get rid of half images for each patient and only trained one model for the FLAIR class.",
      "votes": 3
    },
    {
      "id": 1462653,
      "postDate": "2021-08-10T00:40:15.630Z",
      "content": "<p>It's 5 times the size of the public test set, not the train set.</p>\n<blockquote>\n  <p>NOTE: the total size of the rerun test set (Public and Private) is ~5x the size of the Public test set</p>\n</blockquote>",
      "rawMarkdown": "It's 5 times the size of the public test set, not the train set.\n\n> NOTE: the total size of the rerun test set (Public and Private) is ~5x the size of the Public test set",
      "votes": 2,
      "replies": [
        {
          "id": 1462876,
          "postDate": "2021-08-10T03:01:06.557Z",
          "content": "<p>Thanks, it makes more sense. </p>",
          "rawMarkdown": "Thanks, it makes more sense. "
        },
        {
          "id": 1518267,
          "postDate": "2021-09-20T14:26:10.543Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1462653,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2021-08-10T00:40:15.630000",
      "content": "<p>It's 5 times the size of the public test set, not the train set.</p>\n<blockquote>\n  <p>NOTE: the total size of the rerun test set (Public and Private) is ~5x the size of the Public test set</p>\n</blockquote>",
      "votes": 2,
      "replies": [
        {
          "id": 1462876,
          "author_name": "Shanshan Yu",
          "author_url": "",
          "post_date": "2021-08-10T03:01:06.557000",
          "content": "<p>Thanks, it makes more sense. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1518267,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-09-20T14:26:10.543000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1462529": "The rule says the private and public testing data is ~ 5x than the training dataset. Because the training dataset has 582 patients and 400k files, can we assume the private testing set has 2500+ patients and over 4 million dicom files?\nI did some code analysis and find out the following:\n\npublic_training_patient = 583\npublic_test_patient = 87\npublic_test_files = 51473\npublic_training_files = 348641\n\nThen for the private sets, it will be as following:\nprivate_test_patient = public_training_patient * 5 - public_test_patient\nprivate_test_files = public_training_files * 5 - public_test_files\n\nBecause we have 9 hours to predict the private data(maybe), for each patient and file, we can calculate the pred limited time by:\ntime_per_img_file = (9 * 3600) / private_test_files\ntime_per_patient = (9 * 3600) / private_test_patient\n\nThen I got:\ntime_per_img_file = 0.02 seconds\ntime_per_patient = 11.46 seconds\n\nIt is quite not enough for using four folders' MRI to predict the result. I spend ~ 15 mins for 87 predictions, each patient cost around 10 seconds. But I get rid of half images for each patient and only trained one model for the FLAIR class.",
    "1462653": "It's 5 times the size of the public test set, not the train set.\n\n> NOTE: the total size of the rerun test set (Public and Private) is ~5x the size of the Public test set"
  }
}