{
  "id": 225415,
  "title": "Submissions trend to be timeout!",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/225415",
  "author_name": "朴大福",
  "post_date": "2021-03-12T06:15:13.383000",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>My submission runs about 5 hours. I guess it's due to the pre-submission running resources occupying issue. <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
  "messages": [
    {
      "id": 1235436,
      "postDate": "2021-03-12T06:15:13.383Z",
      "content": "<p>My submission runs about 5 hours. I guess it's due to the pre-submission running resources occupying issue. <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
      "rawMarkdown": "My submission runs about 5 hours. I guess it's due to the pre-submission running resources occupying issue. @addisonhoward ",
      "votes": 5
    },
    {
      "id": 1235502,
      "postDate": "2021-03-12T07:41:36.067Z",
      "content": "<p>I had the same problem, but after memory optimizing my submission code, I could submit. </p>\n<p>This is the code I have used to submit:</p>\n<pre><code>%%time\n\npreprocess_input = get_preprocessing()\n\np = Path(DATA_PATH)\n\nsubm = {}\n\nfor i, filename in tqdm(enumerate(p.glob('test/*.tiff')), total = len(list(p.glob('test/*.tiff')))):\n    print(filename)\n\n    dataset = rasterio.open(filename.as_posix(), transform = identity)\n    slices = make_grid(dataset.shape, window=WINDOW, min_overlap=MIN_OVERLAP)\n\n    preds = np.zeros(dataset.shape, dtype=np.uint8)\n    if dataset.count != 3:\n        print('Image file with subdatasets as channels')\n        layers = [rasterio.open(subd) for subd in dataset.subdatasets]\n\n    for (x1,x2,y1,y2) in tqdm(slices, total = len(slices)):\n        if dataset.count == 3:\n            image = dataset.read([1,2,3],\n                        window=Window.from_slices((x1,x2),(y1,y2)))\n            image = np.moveaxis(image, 0, -1)\n        else:\n            image = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n            for fl in range(3):\n                image[:,:,fl] = layers[fl].read(window=Window.from_slices((x1,x2),(y1,y2)))\n\n        image = preprocess_input(image = image)['image']\n        image = cv2.resize(image, (NEW_SIZE, NEW_SIZE))\n        image = np.moveaxis(image, -1, 0)\n        image = torch.from_numpy(image)\n        pred = np.zeros([WINDOW, WINDOW])\n        for fold_model in fold_models:\n            with torch.no_grad():\n                score = fold_model(image.float().to(DEVICE)[None])\n                score = score.squeeze().cpu().numpy()\n                pred += cv2.resize(score, (WINDOW, WINDOW))\n        pred = pred / len(fold_models)\n        preds[x1:x2,y1:y2] = (pred &gt; 0).astype(np.uint8)\n\n    subm[i] = {'id':filename.stem, 'predicted': rle_numba_encode(preds)}\n    del preds\n    del image\n    del dataset\n    gc.collect();\n</code></pre>\n<p>If you are using <code>rasterio</code>, please try to use the <code>read</code> method to read single slices, instead of trying to ingest the whole file into memory and then read slices from it. The latter method actually brought me timeouts.</p>",
      "rawMarkdown": "I had the same problem, but after memory optimizing my submission code, I could submit. \n\nThis is the code I have used to submit:\n\n```\n%%time\n\npreprocess_input = get_preprocessing()\n\np = Path(DATA_PATH)\n\nsubm = {}\n\nfor i, filename in tqdm(enumerate(p.glob('test/*.tiff')), total = len(list(p.glob('test/*.tiff')))):\n    print(filename)\n    \n    dataset = rasterio.open(filename.as_posix(), transform = identity)\n    slices = make_grid(dataset.shape, window=WINDOW, min_overlap=MIN_OVERLAP)\n    \n    preds = np.zeros(dataset.shape, dtype=np.uint8)\n    if dataset.count != 3:\n        print('Image file with subdatasets as channels')\n        layers = [rasterio.open(subd) for subd in dataset.subdatasets]\n    \n    for (x1,x2,y1,y2) in tqdm(slices, total = len(slices)):\n        if dataset.count == 3:\n            image = dataset.read([1,2,3],\n                        window=Window.from_slices((x1,x2),(y1,y2)))\n            image = np.moveaxis(image, 0, -1)\n        else:\n            image = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n            for fl in range(3):\n                image[:,:,fl] = layers[fl].read(window=Window.from_slices((x1,x2),(y1,y2)))\n            \n        image = preprocess_input(image = image)['image']\n        image = cv2.resize(image, (NEW_SIZE, NEW_SIZE))\n        image = np.moveaxis(image, -1, 0)\n        image = torch.from_numpy(image)\n        pred = np.zeros([WINDOW, WINDOW])\n        for fold_model in fold_models:\n            with torch.no_grad():\n                score = fold_model(image.float().to(DEVICE)[None])\n                score = score.squeeze().cpu().numpy()\n                pred += cv2.resize(score, (WINDOW, WINDOW))\n        pred = pred / len(fold_models)\n        preds[x1:x2,y1:y2] = (pred > 0).astype(np.uint8)\n        \n    subm[i] = {'id':filename.stem, 'predicted': rle_numba_encode(preds)}\n    del preds\n    del image\n    del dataset\n    gc.collect();\n```\n\nIf you are using `rasterio`, please try to use the `read` method to read single slices, instead of trying to ingest the whole file into memory and then read slices from it. The latter method actually brought me timeouts.",
      "votes": 2,
      "replies": [
        {
          "id": 1235576,
          "postDate": "2021-03-12T09:02:21.337Z",
          "content": "<p>thx, I will have a try</p>",
          "rawMarkdown": "thx, I will have a try"
        }
      ]
    },
    {
      "id": 1235675,
      "postDate": "2021-03-12T11:24:15.010Z",
      "content": "<p>My submission keep running for 10 hours without result or timeout. Waiting for the official.</p>",
      "rawMarkdown": "My submission keep running for 10 hours without result or timeout. Waiting for the official.",
      "replies": [
        {
          "id": 1235908,
          "postDate": "2021-03-12T15:22:48.470Z",
          "content": "<p>the same ：（</p>",
          "rawMarkdown": "the same ：（"
        }
      ]
    },
    {
      "id": 1235568,
      "postDate": "2021-03-12T08:55:51.920Z",
      "content": "<p>I also have issues to submit despite the fact I compute the submission file externally that way:</p>\n<pre><code>local_file = '/some_input_path/submission.csv'\ndf_submit = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv', index_col='id')\ndf_local  = pd.read_csv(local_file, index_col='id')\ndf_submit.loc[df_local.index.values] = df_local.values  \ndf_submit.to_csv('submission.csv')\n</code></pre>\n<p>There shouldn't be OOM issues that way ?</p>",
      "rawMarkdown": "I also have issues to submit despite the fact I compute the submission file externally that way:\n\n```\nlocal_file = '/some_input_path/submission.csv'\ndf_submit = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv', index_col='id')\ndf_local  = pd.read_csv(local_file, index_col='id')\ndf_submit.loc[df_local.index.values] = df_local.values  \ndf_submit.to_csv('submission.csv')\n```\n\n\nThere shouldn't be OOM issues that way ?\n\n\n",
      "replies": [
        {
          "id": 1235574,
          "postDate": "2021-03-12T09:01:45.590Z",
          "content": "<p>we did the same thing.</p>",
          "rawMarkdown": "we did the same thing."
        }
      ]
    },
    {
      "id": 1235580,
      "postDate": "2021-03-12T09:16:03.980Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1235502,
      "author_name": "gil fernandes",
      "author_url": "",
      "post_date": "2021-03-12T07:41:36.067000",
      "content": "<p>I had the same problem, but after memory optimizing my submission code, I could submit. </p>\n<p>This is the code I have used to submit:</p>\n<pre><code>%%time\n\npreprocess_input = get_preprocessing()\n\np = Path(DATA_PATH)\n\nsubm = {}\n\nfor i, filename in tqdm(enumerate(p.glob('test/*.tiff')), total = len(list(p.glob('test/*.tiff')))):\n    print(filename)\n\n    dataset = rasterio.open(filename.as_posix(), transform = identity)\n    slices = make_grid(dataset.shape, window=WINDOW, min_overlap=MIN_OVERLAP)\n\n    preds = np.zeros(dataset.shape, dtype=np.uint8)\n    if dataset.count != 3:\n        print('Image file with subdatasets as channels')\n        layers = [rasterio.open(subd) for subd in dataset.subdatasets]\n\n    for (x1,x2,y1,y2) in tqdm(slices, total = len(slices)):\n        if dataset.count == 3:\n            image = dataset.read([1,2,3],\n                        window=Window.from_slices((x1,x2),(y1,y2)))\n            image = np.moveaxis(image, 0, -1)\n        else:\n            image = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n            for fl in range(3):\n                image[:,:,fl] = layers[fl].read(window=Window.from_slices((x1,x2),(y1,y2)))\n\n        image = preprocess_input(image = image)['image']\n        image = cv2.resize(image, (NEW_SIZE, NEW_SIZE))\n        image = np.moveaxis(image, -1, 0)\n        image = torch.from_numpy(image)\n        pred = np.zeros([WINDOW, WINDOW])\n        for fold_model in fold_models:\n            with torch.no_grad():\n                score = fold_model(image.float().to(DEVICE)[None])\n                score = score.squeeze().cpu().numpy()\n                pred += cv2.resize(score, (WINDOW, WINDOW))\n        pred = pred / len(fold_models)\n        preds[x1:x2,y1:y2] = (pred &gt; 0).astype(np.uint8)\n\n    subm[i] = {'id':filename.stem, 'predicted': rle_numba_encode(preds)}\n    del preds\n    del image\n    del dataset\n    gc.collect();\n</code></pre>\n<p>If you are using <code>rasterio</code>, please try to use the <code>read</code> method to read single slices, instead of trying to ingest the whole file into memory and then read slices from it. The latter method actually brought me timeouts.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1235576,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-12T09:02:21.337000",
          "content": "<p>thx, I will have a try</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1235675,
      "author_name": "Lin Deng",
      "author_url": "",
      "post_date": "2021-03-12T11:24:15.010000",
      "content": "<p>My submission keep running for 10 hours without result or timeout. Waiting for the official.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1235908,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-12T15:22:48.470000",
          "content": "<p>the same ：（</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1235568,
      "author_name": "FabienDaniel",
      "author_url": "",
      "post_date": "2021-03-12T08:55:51.920000",
      "content": "<p>I also have issues to submit despite the fact I compute the submission file externally that way:</p>\n<pre><code>local_file = '/some_input_path/submission.csv'\ndf_submit = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv', index_col='id')\ndf_local  = pd.read_csv(local_file, index_col='id')\ndf_submit.loc[df_local.index.values] = df_local.values  \ndf_submit.to_csv('submission.csv')\n</code></pre>\n<p>There shouldn't be OOM issues that way ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1235574,
          "author_name": "朴大福",
          "author_url": "",
          "post_date": "2021-03-12T09:01:45.590000",
          "content": "<p>we did the same thing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1235580,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-12T09:16:03.980000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1235436": "My submission runs about 5 hours. I guess it's due to the pre-submission running resources occupying issue. @addisonhoward ",
    "1235502": "I had the same problem, but after memory optimizing my submission code, I could submit. \n\nThis is the code I have used to submit:\n\n```\n%%time\n\npreprocess_input = get_preprocessing()\n\np = Path(DATA_PATH)\n\nsubm = {}\n\nfor i, filename in tqdm(enumerate(p.glob('test/*.tiff')), total = len(list(p.glob('test/*.tiff')))):\n    print(filename)\n    \n    dataset = rasterio.open(filename.as_posix(), transform = identity)\n    slices = make_grid(dataset.shape, window=WINDOW, min_overlap=MIN_OVERLAP)\n    \n    preds = np.zeros(dataset.shape, dtype=np.uint8)\n    if dataset.count != 3:\n        print('Image file with subdatasets as channels')\n        layers = [rasterio.open(subd) for subd in dataset.subdatasets]\n    \n    for (x1,x2,y1,y2) in tqdm(slices, total = len(slices)):\n        if dataset.count == 3:\n            image = dataset.read([1,2,3],\n                        window=Window.from_slices((x1,x2),(y1,y2)))\n            image = np.moveaxis(image, 0, -1)\n        else:\n            image = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n            for fl in range(3):\n                image[:,:,fl] = layers[fl].read(window=Window.from_slices((x1,x2),(y1,y2)))\n            \n        image = preprocess_input(image = image)['image']\n        image = cv2.resize(image, (NEW_SIZE, NEW_SIZE))\n        image = np.moveaxis(image, -1, 0)\n        image = torch.from_numpy(image)\n        pred = np.zeros([WINDOW, WINDOW])\n        for fold_model in fold_models:\n            with torch.no_grad():\n                score = fold_model(image.float().to(DEVICE)[None])\n                score = score.squeeze().cpu().numpy()\n                pred += cv2.resize(score, (WINDOW, WINDOW))\n        pred = pred / len(fold_models)\n        preds[x1:x2,y1:y2] = (pred > 0).astype(np.uint8)\n        \n    subm[i] = {'id':filename.stem, 'predicted': rle_numba_encode(preds)}\n    del preds\n    del image\n    del dataset\n    gc.collect();\n```\n\nIf you are using `rasterio`, please try to use the `read` method to read single slices, instead of trying to ingest the whole file into memory and then read slices from it. The latter method actually brought me timeouts.",
    "1235675": "My submission keep running for 10 hours without result or timeout. Waiting for the official.",
    "1235568": "I also have issues to submit despite the fact I compute the submission file externally that way:\n\n```\nlocal_file = '/some_input_path/submission.csv'\ndf_submit = pd.read_csv('../input/hubmap-kidney-segmentation/sample_submission.csv', index_col='id')\ndf_local  = pd.read_csv(local_file, index_col='id')\ndf_submit.loc[df_local.index.values] = df_local.values  \ndf_submit.to_csv('submission.csv')\n```\n\n\nThere shouldn't be OOM issues that way ?\n\n\n",
    "1235580": ""
  }
}