{
  "id": 62895,
  "title": "Solution Journal [LB 0.37185]",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/62895",
  "author_name": "",
  "post_date": "2018-08-08T14:54:46.051224900Z",
  "votes": 16,
  "comment_count": 86,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>We would like to start sharing our results in this competition.</p>\n\n<h3>The Open Solution approach</h3>\n\n<p>It means that we are going to open:</p>\n\n<ol>\n<li>the <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection\">code on GitHub</a> - <em>(no worries competitive guys -&gt; we only publish code that scores below bronze medal)</em></li>\n<li>our <a href=\"https://app.neptune.ml/neptune-ml/Google-AI-Object-Detection-Challenge\">experiments results</a></li>\n<li>our approach, that is what we have tried, what worked well, etc.</li>\n</ol>\n\n<h3>Goals</h3>\n\n<p>These are pretty straightforward:</p>\n\n<ol>\n<li>Learning from the process.</li>\n<li>Encourage more Kagglers to start working on this competition.</li>\n<li>Share open source solution with no strings attached, so that less experienced Kagglers can join competition.</li>\n</ol>\n\n<h3>What can you find here?</h3>\n\n<p>In this thread we will discuss our approach to this solution, techniques used, network architectures and all other deep learning related stuff! We want this place to be good address for people who want to share knowledge or gain knowledge :)</p>\n\n<p>Happy Training!</p>\n\n<p>Kamil &amp; Kuba</p>",
  "messages": [
    {
      "id": "367792",
      "postDate": "08/08/2018 14:54:46",
      "content": "<p>Hi,</p>\n\n<p>We would like to start sharing our results in this competition.</p>\n\n<h3>The Open Solution approach</h3>\n\n<p>It means that we are going to open:</p>\n\n<ol>\n<li>the <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection\">code on GitHub</a> - <em>(no worries competitive guys -&gt; we only publish code that scores below bronze medal)</em></li>\n<li>our <a href=\"https://app.neptune.ml/neptune-ml/Google-AI-Object-Detection-Challenge\">experiments results</a></li>\n<li>our approach, that is what we have tried, what worked well, etc.</li>\n</ol>\n\n<h3>Goals</h3>\n\n<p>These are pretty straightforward:</p>\n\n<ol>\n<li>Learning from the process.</li>\n<li>Encourage more Kagglers to start working on this competition.</li>\n<li>Share open source solution with no strings attached, so that less experienced Kagglers can join competition.</li>\n</ol>\n\n<h3>What can you find here?</h3>\n\n<p>In this thread we will discuss our approach to this solution, techniques used, network architectures and all other deep learning related stuff! We want this place to be good address for people who want to share knowledge or gain knowledge :)</p>\n\n<p>Happy Training!</p>\n\n<p>Kamil &amp; Kuba</p>",
      "rawMarkdown": "Hi,\n\nWe would like to start sharing our results in this competition.\n\n### The Open Solution approach\nIt means that we are going to open:\n\n1. the [code on GitHub](https://github.com/neptune-ml/open-solution-googleai-object-detection) - *(no worries competitive guys -&gt; we only publish code that scores below bronze medal)*\n2. our [experiments results](https://app.neptune.ml/neptune-ml/Google-AI-Object-Detection-Challenge)\n3. our approach, that is what we have tried, what worked well, etc.\n\n### Goals\nThese are pretty straightforward:\n\n1. Learning from the process.\n2. Encourage more Kagglers to start working on this competition.\n3. Share open source solution with no strings attached, so that less experienced Kagglers can join competition.\n\n### What can you find here?\nIn this thread we will discuss our approach to this solution, techniques used, network architectures and all other deep learning related stuff! We want this place to be good address for people who want to share knowledge or gain knowledge :)\n\nHappy Training!\n\nKamil &amp; Kuba",
      "votes": null
    },
    {
      "id": "367849",
      "postDate": "08/08/2018 17:11:27",
      "content": "<p>Gotta say that our story in this comp is a constant battle with making retinanetwork work. </p>\n\n<p>First the performance and anchor setup, then working with 500 classes, batching by aspect ratios and dividing the problem into smaller subproblems.</p>\n\n<p>Now we are brewing some experiments (check on neptune.ml) and expecting improvements\nbut obviously still strugling. At this point the biggest debacle is, how to balance classes for the images with multiple bboxes. </p>\n\n<p>If you have any ideas feel free to chime in!</p>",
      "rawMarkdown": "Gotta say that our story in this comp is a constant battle with making retinanetwork work. \n\nFirst the performance and anchor setup, then working with 500 classes, batching by aspect ratios and dividing the problem into smaller subproblems.\n\nNow we are brewing some experiments (check on neptune.ml) and expecting improvements\nbut obviously still strugling. At this point the biggest debacle is, how to balance classes for the images with multiple bboxes. \n\nIf you have any ideas feel free to chime in!",
      "votes": null
    },
    {
      "id": "368079",
      "postDate": "08/09/2018 08:09:54",
      "content": "<p>LB 0.21694 score falls in the silver range. In a discussion at TGS Salt ID competition, you promised <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/62162#365436\">here</a> that you won't publish code that'll f**k up the leaderboard. Come on dude, it takes like a billion years to train on this dataset. Not cool. </p>",
      "rawMarkdown": "LB 0.21694 score falls in the silver range. In a discussion at TGS Salt ID competition, you promised [here][1] that you won't publish code that'll f**k up the leaderboard. Come on dude, it takes like a billion years to train on this dataset. Not cool. \n\n\n  [1]: https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/62162#365436",
      "votes": null
    },
    {
      "id": "368100",
      "postDate": "08/09/2018 09:01:26",
      "content": "<p>Hi,\ndon't worry, the code from the public repository won't take you to the medal zone. As stated in the post:</p>\n\n<p>&gt; (no worries competitive guys -&gt; we only publish code that scores below bronze medal)</p>",
      "rawMarkdown": "Hi,\ndon't worry, the code from the public repository won't take you to the medal zone. As stated in the post:\n\n&gt; (no worries competitive guys -&gt; we only publish code that scores below bronze medal)",
      "votes": null
    },
    {
      "id": "368108",
      "postDate": "08/09/2018 09:11:26",
      "content": "<p>Hi <a href=\"/yk1598\">@yk1598</a>,</p>\n\n<blockquote>\n  <p>the code on GitHub - (no worries competitive guys -&gt; we only publish code that scores below bronze medal)</p>\n</blockquote>\n\n<p>Code in the repository will not give your silver medal. No worries, we know what we are publishing.</p>\n\n<p>In the next post I will describe our work in more detail. Also, this starter gives you good starting point for your own work :)</p>\n\n<p>Best,</p>\n\n<p>Kamil</p>\n\n<p>BTW: nice avatar!</p>",
      "rawMarkdown": "Hi @yk1598,\n\n&gt; the code on GitHub - (no worries competitive guys -&gt; we only publish code that scores below bronze medal)\n\nCode in the repository will not give your silver medal. No worries, we know what we are publishing.\n\nIn the next post I will describe our work in more detail. Also, this starter gives you good starting point for your own work :)\n\nBest,\n\nKamil\n\nBTW: nice avatar!",
      "votes": null
    },
    {
      "id": "368110",
      "postDate": "08/09/2018 09:15:54",
      "content": "<p>Ah I see! I got confused by the title. Apologies.</p>",
      "rawMarkdown": "Ah I see! I got confused by the title. Apologies.",
      "votes": null
    },
    {
      "id": "368142",
      "postDate": "08/09/2018 10:32:34",
      "content": "<h2>Competition update</h2>\n\n<p>It took us quite a lot of time to develop reasonable solution to this competition.</p>\n\n<p>We have decided to work with <strong>RetinaNet</strong> and focal loss, described in this paper: <a href=\"https://arxiv.org/abs/1708.02002\">Focal Loss for Dense Object Detection</a>. If you are new to RetinaNet - I recommend to skim through blog post that describes <a href=\"https://medium.com/@14prakash/the-intuition-behind-retinanet-eb636755607d\">The intuition behind RetinaNet</a>.</p>\n\n<h2>Code</h2>\n\n<p>What you can see on <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection\">master branch</a> is training procedure on 10 classes -&gt; you can extend it to work on the entire dataset. We worked with this code to quickly iterate over various ideas.</p>\n\n<h2>Our analysis of the problem</h2>\n\n<p>We quickly decided that this competition consist of two subproblems, each to be approached separately:</p>\n\n<ul>\n<li><em>First subproblem</em> is classes related to <em>people</em> and <em>clothing</em>, because the bboxes overlap a lot and there are multiple bboxes per image. Here, we have approximately 80 classes.</li>\n<li><em>Second subproblem</em> is remaining classes. Here, we take all these classes and divide it into 7 bins. Each bin is occupied by classes with similar frequency in the dataset. We need such bins to prepare proper epoch as described below. </li>\n</ul>\n\n<h2>Preprocessing for training</h2>\n\n<ul>\n<li>When we run training for the remaining classes, we make sure that each class (within an epoch) has similar number of occurrences -&gt; we implemented sampler to do this work. Thanks to this we have more balanced problem. In practice we oversample rare classes and subsample frequent classes. 7 bins mentioned above are utilized here.</li>\n<li>Next, we calculate <strong>aspect ratio</strong> and we prepare batches only for images with similar <strong>aspect ratio</strong>. We need this in the next step - resize. After resize all images are similarly squeezed - training signal is better balanced.</li>\n<li>At this point we are ready to feed batch to the network. Images are with similar aspect ratio, classes within the epoch are balanced, so training signal is stronger.</li>\n<li>Resulting experiment is like <a href=\"https://app.neptune.ml/-/dashboard/experiment/f945da64-6dd3-459b-94c5-58bc6a83f590\">this one</a>.</li>\n</ul>\n\n<h2>Other remarks</h2>\n\n<ul>\n<li>It is good to start experimenting with few classes (like 10) and get better feel of the problem. We also run <a href=\"https://app.neptune.ml/-/dashboard/experiment/c779468e-d3f7-44b8-a3a4-43a012315708\">training on 10 classes</a>.</li>\n<li>We noticed, that for rare classes augmentations are necessary :)</li>\n</ul>\n\n<h2>Open Questions</h2>\n\n<ul>\n<li>What to do with highly overlapping bboxes (<em>people</em> and <em>clothing</em> subproblem).</li>\n</ul>",
      "rawMarkdown": "## Competition update\nIt took us quite a lot of time to develop reasonable solution to this competition.\n\nWe have decided to work with **RetinaNet** and focal loss, described in this paper: [Focal Loss for Dense Object Detection](https://arxiv.org/abs/1708.02002). If you are new to RetinaNet - I recommend to skim through blog post that describes [The intuition behind RetinaNet](https://medium.com/@14prakash/the-intuition-behind-retinanet-eb636755607d).\n\n## Code\nWhat you can see on [master branch](https://github.com/neptune-ml/open-solution-googleai-object-detection) is training procedure on 10 classes -&gt; you can extend it to work on the entire dataset. We worked with this code to quickly iterate over various ideas.\n\n## Our analysis of the problem\nWe quickly decided that this competition consist of two subproblems, each to be approached separately:\n\n- *First subproblem* is classes related to *people* and *clothing*, because the bboxes overlap a lot and there are multiple bboxes per image. Here, we have approximately 80 classes.\n- *Second subproblem* is remaining classes. Here, we take all these classes and divide it into 7 bins. Each bin is occupied by classes with similar frequency in the dataset. We need such bins to prepare proper epoch as described below. \n\n## Preprocessing for training\n* When we run training for the remaining classes, we make sure that each class (within an epoch) has similar number of occurrences -&gt; we implemented sampler to do this work. Thanks to this we have more balanced problem. In practice we oversample rare classes and subsample frequent classes. 7 bins mentioned above are utilized here.\n* Next, we calculate **aspect ratio** and we prepare batches only for images with similar **aspect ratio**. We need this in the next step - resize. After resize all images are similarly squeezed - training signal is better balanced.\n* At this point we are ready to feed batch to the network. Images are with similar aspect ratio, classes within the epoch are balanced, so training signal is stronger.\n* Resulting experiment is like [this one](https://app.neptune.ml/-/dashboard/experiment/f945da64-6dd3-459b-94c5-58bc6a83f590).\n\n## Other remarks\n- It is good to start experimenting with few classes (like 10) and get better feel of the problem. We also run [training on 10 classes](https://app.neptune.ml/-/dashboard/experiment/c779468e-d3f7-44b8-a3a4-43a012315708).\n- We noticed, that for rare classes augmentations are necessary :)\n\n## Open Questions\n* What to do with highly overlapping bboxes (*people* and *clothing* subproblem).",
      "votes": null
    },
    {
      "id": "368312",
      "postDate": "08/09/2018 17:18:39",
      "content": "<p>@Kamil I think this may answer your questions <a href=\"https://stackoverflow.com/questions/49951422/get-rid-of-overlapping-bounding-boxes-across-different-classes-in-tensorflow-obj\">non_max_suppression over all classes</a></p>",
      "rawMarkdown": "Kamil I think this may answer your questions [non_max_suppression over all classes][1]\n\n\n  [1]: https://stackoverflow.com/questions/49951422/get-rid-of-overlapping-bounding-boxes-across-different-classes-in-tensorflow-obj \"StackOverflow\"",
      "votes": null
    },
    {
      "id": "368397",
      "postDate": "08/09/2018 20:18:04",
      "content": "<p>Hi <a href=\"/dskswu\">@dskswu</a>, thanks I'll check it :)</p>",
      "rawMarkdown": "Hi @dskswu, thanks I'll check it :)",
      "votes": null
    },
    {
      "id": "368500",
      "postDate": "08/10/2018 03:06:49",
      "content": "<p>I think you have a typo <code>python main.py train --pipeline_name retinanet</code> in your instructions.</p>",
      "rawMarkdown": "I think you have a typo `python main.py train --pipeline_name retinanet` in your instructions.",
      "votes": null
    },
    {
      "id": "368909",
      "postDate": "08/11/2018 09:00:40",
      "content": "<p>Hi, it should be <code>python main.py -- train --pipeline_name retinanet</code></p>\n\n<p>thank you!</p>",
      "rawMarkdown": "Hi, it should be `python main.py -- train --pipeline_name retinanet`\n\nthank you!",
      "votes": null
    },
    {
      "id": "369813",
      "postDate": "08/13/2018 20:09:14",
      "content": "<p>Is this the correct data path parameters? </p>\n\n<pre><code>   train_imgs_dir: train_03/\n   test_imgs_dir: test/\n   annotations_filepath: annotations/\n   annotations_human_labels_filepath: annotations_human_labels/\n   bbox_hierarchy_filepath: bbox_hierarchy/\n   valid_ids_filepath: valid_ids/\n   sample_submission: sample_submission.csv\n   experiment_dir:  experiment\n   class_mappings_filepath: class_mappings/\n</code></pre>\n\n<p>I am trying to troubleshoot this error: </p>\n\n<pre><code>neptune: Executing in Offline Mode.\nneptune: Executing in Offline Mode.\n2018-08-13 20-04-33 google-ai-odt &gt;&gt;&gt; training\nTraceback (most recent call last):\n File \"main.py\", line 78, in &lt;module&gt;\nmain()\n File \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 16, in train\npipeline_manager.train(pipeline_name, dev_mode)\nFile \"/floyd/home/src/pipeline_manager.py\", line 21, in train\ntrain(pipeline_name, dev_mode)\nFile \"/floyd/home/src/pipeline_manager.py\", line 38, in train\nannotations = pd.read_csv(PARAMS.annotations_filepath)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 655, in parser_f\nreturn _read(filepath_or_buffer, kwds)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 405, in _read\nparser = TextFileReader(filepath_or_buffer, **kwds)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 764, in __init__\nself._make_engine(self.engine)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 985, in _make_engine\nself._engine = CParserWrapper(self.f, **self.options)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 1605, in __init__\nself._reader = parsers.TextReader(src, **kwds)\nFile \"pandas/_libs/parsers.pyx\", line 394, in pandas._libs.parsers.TextReader.__cinit__ (pandas/_libs/parsers.c:4209)\nFile \"pandas/_libs/parsers.pyx\", line 710, in pandas._libs.parsers.TextReader._setup_parser_source \n(pandas/_libs/parsers.c:8873)\n FileNotFoundError: File b'' does not exist\n</code></pre>",
      "rawMarkdown": "Is this the correct data path parameters? \n\n       train_imgs_dir: train_03/\n       test_imgs_dir: test/\n       annotations_filepath: annotations/\n       annotations_human_labels_filepath: annotations_human_labels/\n       bbox_hierarchy_filepath: bbox_hierarchy/\n       valid_ids_filepath: valid_ids/\n       sample_submission: sample_submission.csv\n       experiment_dir:  experiment\n       class_mappings_filepath: class_mappings/\n\nI am trying to troubleshoot this error: \n\n    neptune: Executing in Offline Mode.\n    neptune: Executing in Offline Mode.\n    2018-08-13 20-04-33 google-ai-odt &gt;&gt;&gt; training\n    Traceback (most recent call last):\n     File \"main.py\", line 78, in",
      "votes": null
    },
    {
      "id": "370045",
      "postDate": "08/14/2018 06:20:18",
      "content": "<p>Not quite. Mine looks something like this:</p>\n\n<pre><code>parameters:\n # Data Paths\n\n train_imgs_dir: .../open-images-v4/bounding-boxes/train\n\n test_imgs_dir: .../open-images-v4/bounding-boxes/test_challenge_2018\n\n annotations_filepath: .../googleai-object-detection/data/annotations/challenge-2018-train-annotations-bbox.csv\n\n annotations_human_labels_filepath:.../googleai-object-detection/data/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n\n bbox_hierarchy_filepath: .../googleai-object-detection/data/metadata/bbox_labels_500_hierarchy.json\n\n class_mappings_filepath: .../googleai-object-detection/data/metadata/challenge-2018-class-descriptions-500.csv\n\n valid_ids_filepath: .../googleai-object-detection/data/metadata/challenge-2018-image-ids-valset-od.csv\n\nsample_submission: .../googleai-object-detection/data/sample_submission.csv\n\nmetadata_filepath: .../googleai-object-detection/files/metadata.csv\n\nexperiment_dir:  .../googleai-object-detection/experiments/10_classes\n</code></pre>",
      "rawMarkdown": "Not quite. Mine looks something like this:\n\n    parameters:\n     # Data Paths\n\n     train_imgs_dir: .../open-images-v4/bounding-boxes/train\n\n     test_imgs_dir: .../open-images-v4/bounding-boxes/test_challenge_2018\n\n     annotations_filepath: .../googleai-object-detection/data/annotations/challenge-2018-train-annotations-bbox.csv\n\n     annotations_human_labels_filepath:.../googleai-object-detection/data/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n\n     bbox_hierarchy_filepath: .../googleai-object-detection/data/metadata/bbox_labels_500_hierarchy.json\n\n     class_mappings_filepath: .../googleai-object-detection/data/metadata/challenge-2018-class-descriptions-500.csv\n\n     valid_ids_filepath: .../googleai-object-detection/data/metadata/challenge-2018-image-ids-valset-od.csv\n     \n    sample_submission: .../googleai-object-detection/data/sample_submission.csv\n     \n    metadata_filepath: .../googleai-object-detection/files/metadata.csv\n    \n    experiment_dir:  .../googleai-object-detection/experiments/10_classes",
      "votes": null
    },
    {
      "id": "370325",
      "postDate": "08/14/2018 16:15:23",
      "content": "<p>@Jakub, Thank you. I totally missed this on the download page. </p>",
      "rawMarkdown": "Jakub, Thank you. I totally missed this on the download page.",
      "votes": null
    },
    {
      "id": "370944",
      "postDate": "08/15/2018 17:55:24",
      "content": "<p>I am running the latest master branch (offline), and when the code gets to the training point it crashes when trying to forward() the model and evaluate the loss function:</p>\n\n<pre><code>neptune: Executing in Offline Mode.\n2018-08-15 18-34-40 google-ai-odt &gt;&gt;&gt; training\n2018-08-15 18-35-03 google-ai-odt &gt;&gt;&gt; Training on a reduced class subset: ['Person', 'Car', 'Dress', 'Footwear']\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing Step label_encoder...\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing Step label_encoder...\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing experiment directories under experiments\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing experiment directories under experiments\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; done: initializing experiment directories\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; done: initializing experiment directories\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; Step label_encoder initialized\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; Step label_encoder initialized\n\n[skipped]\n\n2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, fitting and transforming...\n2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, fitting and transforming...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; starting training...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; starting training...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; initial lr: 1e-05\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; initial lr: 1e-05\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 ...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 ...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 batch 0 ...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 batch 0 ...\nTraceback (most recent call last):\n  File \"main.py\", line 78, in &lt;module&gt;\n    main()\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 722, in __call__\n    return self.main(*args, **kwargs)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 697, in main\n    rv = self.invoke(ctx)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 1066, in invoke\n    return _process_result(sub_ctx.command.invoke(sub_ctx))\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 895, in invoke\n    return ctx.invoke(self.callback, **ctx.params)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 535, in invoke\n    return callback(*args, **kwargs)\n  File \"main.py\", line 16, in train\n    pipeline_manager.train(pipeline_name, dev_mode)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 21, in train\n    train(pipeline_name, dev_mode)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 85, in train\n    pipeline.fit_transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 280, in fit_transform\n    step_output_data = self._cached_fit_transform(step_inputs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 390, in _cached_fit_transform\n    step_output_data = self.transformer.fit_transform(**step_inputs)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/steppy/base.py\", line 605, in fit_transform\n    self.fit(*args, **kwargs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/models.py\", line 32, in fit\n    metrics = self._fit_loop(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/models.py\", line 63, in _fit_loop\n    batch_loss = loss_function(outputs_batch, target) * weight\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 357, in __call__\n    result = self.forward(*input, **kwargs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/parallel.py\", line 137, in forward\n    outputs = _criterion_parallel_apply(replicas, inputs, targets, kwargs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/parallel.py\", line 192, in _criterion_parallel_apply\n    raise output\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/parallel.py\", line 167, in _worker\n    output = module(*(input + target), **kwargs)\nTypeError: can only concatenate tuple (not \"dict\") to tuple\n</code></pre>\n\n<p>It looks like the \"target\" variable used for the loss function is supposed to be a tuple, but instead it is a dictionary. I have to admit I'm not sure what exactly's causing this, but I wanted to see if you have any immediate ideas before I spend time going through the code line by line. Execution command is just: <code>python main.py -- train --pipeline_name retinanet</code>, and the whole config has been filled out with (supposedly) the correct files.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "I am running the latest master branch (offline), and when the code gets to the training point it crashes when trying to forward() the model and evaluate the loss function:\n\n    neptune: Executing in Offline Mode.\n    2018-08-15 18-34-40 google-ai-odt &gt;&gt;&gt; training\n    2018-08-15 18-35-03 google-ai-odt &gt;&gt;&gt; Training on a reduced class subset: ['Person', 'Car', 'Dress', 'Footwear']\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing Step label_encoder...\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing Step label_encoder...\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing experiment directories under experiments\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing experiment directories under experiments\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; done: initializing experiment directories\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; done: initializing experiment directories\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; Step label_encoder initialized\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; Step label_encoder initialized\n\n    [skipped]\n\n    2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n    2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n    2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, fitting and transforming...\n    2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, fitting and transforming...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; starting training...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; starting training...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; initial lr: 1e-05\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; initial lr: 1e-05\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 ...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 ...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 batch 0 ...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 batch 0 ...\n    Traceback (most recent call last):\n      File \"main.py\", line 78, in",
      "votes": null
    },
    {
      "id": "370976",
      "postDate": "08/15/2018 18:52:11",
      "content": "<p>Hi there <a href=\"/anokas\">@anokas</a>.</p>\n\n<p>I think the problem is we have worked with, and tested it on multigpu. The loss is also parallelized (extension of pytorch data parallelization so that it better utilized many cards). </p>\n\n<p>So there are 2 options going forward.</p>\n\n<ul>\n<li><p>If you have multigpu change batch_size_train and batch_size_inference to something larger than 1. For example I am training some batches of classes right now with <code>batch_size_train: 8</code> and <code>batch_size_inference: 8</code> on 4 gpus <a href=\"https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce\">https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce</a> . If you do so, remember to change the batch_size_inference to 1 when running evaluate_predict pipeline. For some unknown reason it can crash the memory if you don't do that.</p></li>\n<li><p>Second option if you are training on just 1 gpu is to substitute all DataParallel stuff in the <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py\">https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py</a> file. To be more specific those are <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py#L19\">model</a> and <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py#L105\">loss</a></p></li>\n</ul>\n\n<p>I will add those as issues and try and fix it as soon as I can (likely tomorrow).\nJust out of curiousity what is your hardware setup?</p>\n\n<p>Best,\nJakub</p>",
      "rawMarkdown": "Hi there @anokas.\n\nI think the problem is we have worked with, and tested it on multigpu. The loss is also parallelized (extension of pytorch data parallelization so that it better utilized many cards). \n\nSo there are 2 options going forward.\n\n - If you have multigpu change batch_size_train and batch_size_inference to something larger than 1. For example I am training some batches of classes right now with `batch_size_train: 8` and `batch_size_inference: 8` on 4 gpus https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce . If you do so, remember to change the batch_size_inference to 1 when running evaluate_predict pipeline. For some unknown reason it can crash the memory if you don't do that.\n\n - Second option if you are training on just 1 gpu is to substitute all DataParallel stuff in the https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py file. To be more specific those are [model](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py#L19) and [loss](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py#L105)\n\nI will add those as issues and try and fix it as soon as I can (likely tomorrow).\nJust out of curiousity what is your hardware setup?\n\nBest,\nJakub",
      "votes": null
    },
    {
      "id": "370984",
      "postDate": "08/15/2018 18:57:02",
      "content": "<p>Hi Jakub,</p>\n\n<p>I am running on 4x 1080Ti, the issue was that both batch_size_train and batch_size_inference were set to 1 (the default in the repo config file). I fixed it now, thank you :)</p>",
      "rawMarkdown": "Hi Jakub,\n\nI am running on 4x 1080Ti, the issue was that both batch_size_train and batch_size_inference were set to 1 (the default in the repo config file). I fixed it now, thank you :)",
      "votes": null
    },
    {
      "id": "370985",
      "postDate": "08/15/2018 19:01:26",
      "content": "<p>My apologies for that.</p>\n\n<p>I hope it will run smoothly from now on. I am curious to hear your thoughts on this project both during and after the competition. So feel free to drop a comment whenever you feel like it.</p>",
      "rawMarkdown": "My apologies for that.\n\nI hope it will run smoothly from now on. I am curious to hear your thoughts on this project both during and after the competition. So feel free to drop a comment whenever you feel like it.",
      "votes": null
    },
    {
      "id": "371008",
      "postDate": "08/15/2018 19:41:35",
      "content": "<p>Hey Jakub, </p>\n\n<p>Is there any additional setup for running the open solution in neptune-ml? I think I want to test it out in neptune. </p>",
      "rawMarkdown": "Hey Jakub, \n\nIs there any additional setup for running the open solution in neptune-ml? I think I want to test it out in neptune.",
      "votes": null
    },
    {
      "id": "371009",
      "postDate": "08/15/2018 19:46:14",
      "content": "<p>Sure, will do. I ran into one other issue when calling evaluate or predict:</p>\n\n<pre><code>2018-08-15 20:28:49 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n\nTraceback (most recent call last):\n  File \"main.py\", line 78, in &lt;module&gt;\n    main()\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 722, in __call__\n    return self.main(*args, **kwargs)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 697, in main\n    rv = self.invoke(ctx)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 1066, in invoke\n    return _process_result(sub_ctx.command.invoke(sub_ctx))\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 895, in invoke\n    return ctx.invoke(self.callback, **ctx.params)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 535, in invoke\n    return callback(*args, **kwargs)\n  File \"main.py\", line 25, in evaluate\n    pipeline_manager.evaluate(pipeline_name, dev_mode, chunk_size)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 24, in evaluate\n    evaluate(pipeline_name, dev_mode, chunk_size)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 117, in evaluate\n    prediction = generate_prediction(valid_img_ids, pipeline, chunk_size)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 182, in generate_prediction\n    return _generate_prediction_in_chunks(img_ids, pipeline, chunk_size)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 214, in _generate_prediction_in_chunks\n    output = pipeline.transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 315, in transform\n    step_inputs[input_step.name] = input_step.transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 315, in transform\n    step_inputs[input_step.name] = input_step.transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 315, in transform\n    step_inputs[input_step.name] = input_step.transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 321, in transform\n    step_output_data = self._cached_transform(step_inputs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 424, in _cached_transform\n    raise ValueError('No transformer cached {}'.format(self.name))\nValueError: No transformer cached retinanet\n</code></pre>\n\n<p>It appears to be looking for <code>experiments/transformers/retinanet</code> - I only have the file <code>experiments/transformers/label_encoder</code> - I have only run the model for one epoch but just wanted to test evaluation/prediction as a sanity check (I can see it saved the weight checkpoint too). Is this a file generated when the script completely finishes? Or is there some way to evaluate/predict an intermediate model?</p>\n\n<p>Thanks again.</p>",
      "rawMarkdown": "Sure, will do. I ran into one other issue when calling evaluate or predict:\n\n    2018-08-15 20:28:49 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n    \n    Traceback (most recent call last):\n      File \"main.py\", line 78, in",
      "votes": null
    },
    {
      "id": "371011",
      "postDate": "08/15/2018 19:58:40",
      "content": "<p>@Jakub what is the configs to set up in neptune for this comp? </p>",
      "rawMarkdown": "Jakub what is the configs to set up in neptune for this comp?",
      "votes": null
    },
    {
      "id": "371012",
      "postDate": "08/15/2018 20:04:45",
      "content": "<p>Update: I copied the <code>best.torch</code> weights file to experiments/transformers/retinanet, and it appears to work!</p>",
      "rawMarkdown": "Update: I copied the `best.torch` weights file to experiments/transformers/retinanet, and it appears to work!",
      "votes": null
    },
    {
      "id": "371207",
      "postDate": "08/16/2018 07:48:23",
      "content": "<p>Hi <a href=\"/anokas\">@anokas</a> @Jakub,\nWhen I try to predict (python main.py -- predict --pipeline_name retinanet) for a sanity check, I am getting the following error:\n<strong>\"RuntimeError: $ Torch: not enough memory: you tried to allocate 0GB. Buy new RAM! at /pytorch/torch/lib/TH/THGeneral.c:253\".</strong>\nI am using 3 x TITAN Xp and I think I have enough memory (12GB for each GPU). I wonder if you can help me to overcome this issue. Thanks in advance.</p>",
      "rawMarkdown": "Hi @anokas @Jakub,\nWhen I try to predict (python main.py -- predict --pipeline_name retinanet) for a sanity check, I am getting the following error:\n**\"RuntimeError: $ Torch: not enough memory: you tried to allocate 0GB. Buy new RAM! at /pytorch/torch/lib/TH/THGeneral.c:253\".**\nI am using 3 x TITAN Xp and I think I have enough memory (12GB for each GPU). I wonder if you can help me to overcome this issue. Thanks in advance.",
      "votes": null
    },
    {
      "id": "371221",
      "postDate": "08/16/2018 08:57:16",
      "content": "<p>Probably you need to predict in chunks by adding <code>.. --chunk_size 100</code> at the end of your command. </p>",
      "rawMarkdown": "Probably you need to predict in chunks by adding `.. --chunk_size 100` at the end of your command.",
      "votes": null
    },
    {
      "id": "371238",
      "postDate": "08/16/2018 10:24:49",
      "content": "<p><a href=\"/anokas\">@anokas</a> Steppy saves the transformer to /transformers/NAME after said transformer has been fitted. In case training was stopped mid way you have to manually copy the checkpoint as you have. </p>",
      "rawMarkdown": "anokas Steppy saves the transformer to /transformers/NAME after said transformer has been fitted. In case training was stopped mid way you have to manually copy the checkpoint as you have.",
      "votes": null
    },
    {
      "id": "371283",
      "postDate": "08/16/2018 12:56:25",
      "content": "<p>Hi @William Green ,</p>\n\n<p>All you need to do is to change data paths in the very same <code>neptune.yaml</code> you are using.</p>\n\n<p>I will update the readme.md today with all the instructions needed to run it on multigpu with <code>neptune send</code></p>",
      "rawMarkdown": "Hi @William Green ,\n\nAll you need to do is to change data paths in the very same `neptune.yaml` you are using.\n\nI will update the readme.md today with all the instructions needed to run it on multigpu with `neptune send`",
      "votes": null
    },
    {
      "id": "371317",
      "postDate": "08/16/2018 14:31:16",
      "content": "<p><strong>Update</strong>\nI updated the master branch of our <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection\">repo</a></p>\n\n<ol>\n<li><p>Added submission merge notebook for those of you who want to train batches of classes and join the predictions </p></li>\n<li><p>Updated the <code>neptune.yaml</code> config file with filepaths for those of you that may want to run in the cloud.\nIn that case pretty much all you need to do is run</p>\n\n<pre><code>neptune send --worker m-4p100 \\\n--environment pytorch-0.3.1-gpu-py3 \\\n--config configs/neptune.yaml \\\nmain.py train --pipeline_name retinanet\n</code></pre>\n\n<p>Yup there are multigpu workers on\nneptune now.</p></li>\n<li><p>Dropped some redundant steppy bits since the steppy 0.1.6 already has all we need here.</p></li>\n</ol>",
      "rawMarkdown": "**Update**\nI updated the master branch of our [repo](https://github.com/neptune-ml/open-solution-googleai-object-detection)\n\n1. Added submission merge notebook for those of you who want to train batches of classes and join the predictions \n\n1. Updated the `neptune.yaml` config file with filepaths for those of you that may want to run in the cloud.\n   In that case pretty much all you need to do is run\n\n        neptune send --worker m-4p100 \\\n        --environment pytorch-0.3.1-gpu-py3 \\\n        --config configs/neptune.yaml \\\n        main.py train --pipeline_name retinanet\n\n   Yup there are multigpu workers on\n   neptune now.\n\n1. Dropped some redundant steppy bits since the steppy 0.1.6 already has all we need here.",
      "votes": null
    },
    {
      "id": "371510",
      "postDate": "08/17/2018 01:01:12",
      "content": "<p>Hi There,\n   I'm using neptune.ml for my experiments using your base code. Just getting my feet wet at the moment.\n  I did a train run based on defaults but with --dev_mode flag on\n  However running into trouble when trying to run the eval part. The suggestion below (from the documentation):</p>\n\n<p>\"With cloud environment you need to change the experiment directory to the one that you have just trained. Let's assume that your experiment id was GAI-14. You should go to neptune.yaml and change:</p>\n\n<p>experiment_dir:  ../GAI-14/output/experiment\n\"\nis not working for me. Please note that I am indeed changing the \"GAI-14\" to the name of my experimental (training) run. But ../my_expt_id/.. does not seem to exist.\nPlease help!</p>",
      "rawMarkdown": "Hi There,\n   I'm using neptune.ml for my experiments using your base code. Just getting my feet wet at the moment.\n  I did a train run based on defaults but with --dev_mode flag on\n  However running into trouble when trying to run the eval part. The suggestion below (from the documentation):\n\n\"With cloud environment you need to change the experiment directory to the one that you have just trained. Let's assume that your experiment id was GAI-14. You should go to neptune.yaml and change:\n\n  experiment_dir:  ../GAI-14/output/experiment\n\"\nis not working for me. Please note that I am indeed changing the \"GAI-14\" to the name of my experimental (training) run. But ../my_expt_id/.. does not seem to exist.\nPlease help!",
      "votes": null
    },
    {
      "id": "371625",
      "postDate": "08/17/2018 08:54:58",
      "content": "<p>Hi @Giri Gopalan. </p>\n\n<p>I forgot to add a snippet that copies the experiment from your protected (read only) <code>/input</code> directory to the <code>/output</code> directory before running anything. Also when running neptune in the cloud you need to specify <code>--input my_dir</code> directories if you want to use something that is outside of your experiment. It is all nicely explained in the docs <a href=\"https://docs.neptune.ml/advanced-topics/storage/\">https://docs.neptune.ml/advanced-topics/storage/</a> .</p>\n\n<p>Anyways sorry for the trouble.\nBoth the code and instructions are updated so please get the newest master and follow the instructions.\nIn case of any trouble let me know. </p>",
      "rawMarkdown": "Hi @Giri Gopalan. \n\nI forgot to add a snippet that copies the experiment from your protected (read only) `/input` directory to the `/output` directory before running anything. Also when running neptune in the cloud you need to specify `--input my_dir` directories if you want to use something that is outside of your experiment. It is all nicely explained in the docs https://docs.neptune.ml/advanced-topics/storage/ .\n\nAnyways sorry for the trouble.\nBoth the code and instructions are updated so please get the newest master and follow the instructions.\nIn case of any trouble let me know.",
      "votes": null
    },
    {
      "id": "371801",
      "postDate": "08/17/2018 16:20:11",
      "content": "<p>I recieve this error when running in the cloud: </p>\n\n<pre><code>/usr/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary \nincompatibility. Expected 96, got 88\nreturn f(*args, **kwds)\n/usr/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary \nincompatibility. Expected 96, got 88\n return f(*args, **kwds)\n 2018-08-17 16:16:58,612 google-ai-odt WARNING  pipeline_manager.py:58 - train() Validation sample-size is smaller \n then desired validation sample size ... clipping\n 2018-08-17 16:16:58,612 google-ai-odt WARNING  pipeline_manager.py:58 - train() Validation sample-size is smaller \n then desired validation sample size ... clipping\n Traceback (most recent call last):\n File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in &lt;module&gt;\nexecute()\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 103, in execute\nexecfile(job_filepath, job_globals)\nFile \"/usr/local/lib/python3.6/dist-packages/past/builtins/misc.py\", line 82, in execfile\nexec_(code, myglobals, mylocals)\nFile \"main.py\", line 78, in &lt;module&gt;\nmain()\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 16, in train\npipeline_manager.train(pipeline_name, dev_mode)\nFile \"/neptune/src/pipeline_manager.py\", line 21, in train\ntrain(pipeline_name, dev_mode)\nFile \"/neptune/src/pipeline_manager.py\", line 60, in train\na_max=valid_ids_data.shape[0])\nTypeError: clip() missing 1 required positional argument: 'a_min'\n</code></pre>",
      "rawMarkdown": "I recieve this error when running in the cloud: \n\n    /usr/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary \n    incompatibility. Expected 96, got 88\n    return f(*args, **kwds)\n    /usr/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary \n    incompatibility. Expected 96, got 88\n     return f(*args, **kwds)\n     2018-08-17 16:16:58,612 google-ai-odt WARNING  pipeline_manager.py:58 - train() Validation sample-size is smaller \n     then desired validation sample size ... clipping\n     2018-08-17 16:16:58,612 google-ai-odt WARNING  pipeline_manager.py:58 - train() Validation sample-size is smaller \n     then desired validation sample size ... clipping\n     Traceback (most recent call last):\n     File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in",
      "votes": null
    },
    {
      "id": "371837",
      "postDate": "08/17/2018 17:06:28",
      "content": "<p>Hi @William Green.</p>\n\n<p>Are you sure you are running the newest master?</p>\n\n<p>I ran it this morning in the cloud and it worked just fine.\nThis is a link to the experiment <a href=\"https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca\">https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca</a></p>",
      "rawMarkdown": "Hi @William Green.\n\nAre you sure you are running the newest master?\n\nI ran it this morning in the cloud and it worked just fine.\nThis is a link to the experiment https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca",
      "votes": null
    },
    {
      "id": "371887",
      "postDate": "08/17/2018 19:11:58",
      "content": "<p>Hey @Jakub,</p>\n\n<p>I guess I didn't have the newest master. I have it working now. Thank you. </p>",
      "rawMarkdown": "Hey @Jakub,\n\nI guess I didn't have the newest master. I have it working now. Thank you.",
      "votes": null
    },
    {
      "id": "371936",
      "postDate": "08/17/2018 21:28:46",
      "content": "<p>Hi Jakub,</p>\n\n<p>Thanks. This brings me one step closer, but I am encountering the error below when I follow new instructions. Not too sure what is wrong.</p>\n\n<p>Unexpected end of /proc/mounts line `overlay / overlay rw,relatime,lowerdir=/var/lib/docker/overlay2/l/PKZCZTI7LLZ3N7KN675YAWDO2A:/var/lib/docker/overlay2/l/BZKXUKDO5K62J4XGTM2YWPDFWB:/var/lib/docker/overlay2/l/TWQXPQNTQQCLJMMPWR5GK7XWEV:/var/lib/docker/overlay2/l/4SULQLH2FKKFFP5YH6ST4JUC7W:/var/lib/docker/overlay2/l/QT7PKBTZCF652T2A5ZSXX56EHN:/var/lib/docker/overlay2/l/YEFVDSP2J4P6AKUSB7IPZCHYBA:/var/lib/docker/overlay2/l/LRAWBPHZID7W7RKBL43V3FJSZX:/var/lib/docker/overlay2/l/DG3MICQYVNJOX6YXDPHTIQU64W:/var/lib/docker/overlay2/l/MNIQIBWYDKSM5'</p>\n\n<p>Unexpected end of /proc/mounts line `J3EWKYNO536PV:/var/lib/docker/overlay2/l/UGWX4WBW3IQHMO5R5BDV27FMQE:/var/lib/docker/overlay2/l/NWR22WBG66ZRI7EAVALSIQZUDW:/var/lib/docker/overlay2/l/WF564LYDXGQWBGDQC3VPAO74KP:/var/lib/docker/overlay2/l/KWSW4LTKEPLIQX5V5NDA6HATPF:/var/lib/docker/overlay2/l/5D6ZQBN2PVELJCG5L4B34NVXEV:/var/lib/docker/overlay2/l/OZR3TE5HSEHSK6X2DM62OINVH5:/var/lib/docker/overlay2/l/I4NS2OYZDHG2YHUYVGVNVKN7KC:/var/lib/docker/overlay2/l/5UBGWBX2QJ7L4I45XICUS7W5WZ:/var/lib/docker/overlay2/l/X7YX7LUJRJ7NOVYUMVYJ7ER5WG:/var/lib/do'</p>\n\n<p>Unexpected end of /proc/mounts line `cker/overlay2/l/5VUUIAN3JGDVEFUKTLBV3BUJ64:/var/lib/docker/overlay2/l/IHUECGATV342F57ZGCTVIGG7CN:/var/lib/docker/overlay2/l/BTK3WTDIK4C6AUJEKR7I557WHA:/var/lib/docker/overlay2/l/IO5CIAELD3QELFU4N7MCI7FOT3:/var/lib/docker/overlay2/l/YGDBE4IOHQXR5ECOI4ZCBHAIIT:/var/lib/docker/overlay2/l/XGGPCHWJVDANU2B4ISJZ7JX2KI:/var/lib/docker/overlay2/l/VJROJ5LPTOXKSVYKJK463A3SZB:/var/lib/docker/overlay2/l/ZIWL5PNUTTEWS23HPG44QDSMNN,upperdir=/var/lib/docker/overlay2/7e427cf00268b04985194be27ee928c391b82a3eba75b9b18da9e9b0'</p>",
      "rawMarkdown": "Hi Jakub,\n\n  Thanks. This brings me one step closer, but I am encountering the error below when I follow new instructions. Not too sure what is wrong.\n\nUnexpected end of /proc/mounts line `overlay / overlay rw,relatime,lowerdir=/var/lib/docker/overlay2/l/PKZCZTI7LLZ3N7KN675YAWDO2A:/var/lib/docker/overlay2/l/BZKXUKDO5K62J4XGTM2YWPDFWB:/var/lib/docker/overlay2/l/TWQXPQNTQQCLJMMPWR5GK7XWEV:/var/lib/docker/overlay2/l/4SULQLH2FKKFFP5YH6ST4JUC7W:/var/lib/docker/overlay2/l/QT7PKBTZCF652T2A5ZSXX56EHN:/var/lib/docker/overlay2/l/YEFVDSP2J4P6AKUSB7IPZCHYBA:/var/lib/docker/overlay2/l/LRAWBPHZID7W7RKBL43V3FJSZX:/var/lib/docker/overlay2/l/DG3MICQYVNJOX6YXDPHTIQU64W:/var/lib/docker/overlay2/l/MNIQIBWYDKSM5'\n\nUnexpected end of /proc/mounts line `J3EWKYNO536PV:/var/lib/docker/overlay2/l/UGWX4WBW3IQHMO5R5BDV27FMQE:/var/lib/docker/overlay2/l/NWR22WBG66ZRI7EAVALSIQZUDW:/var/lib/docker/overlay2/l/WF564LYDXGQWBGDQC3VPAO74KP:/var/lib/docker/overlay2/l/KWSW4LTKEPLIQX5V5NDA6HATPF:/var/lib/docker/overlay2/l/5D6ZQBN2PVELJCG5L4B34NVXEV:/var/lib/docker/overlay2/l/OZR3TE5HSEHSK6X2DM62OINVH5:/var/lib/docker/overlay2/l/I4NS2OYZDHG2YHUYVGVNVKN7KC:/var/lib/docker/overlay2/l/5UBGWBX2QJ7L4I45XICUS7W5WZ:/var/lib/docker/overlay2/l/X7YX7LUJRJ7NOVYUMVYJ7ER5WG:/var/lib/do'\n\nUnexpected end of /proc/mounts line `cker/overlay2/l/5VUUIAN3JGDVEFUKTLBV3BUJ64:/var/lib/docker/overlay2/l/IHUECGATV342F57ZGCTVIGG7CN:/var/lib/docker/overlay2/l/BTK3WTDIK4C6AUJEKR7I557WHA:/var/lib/docker/overlay2/l/IO5CIAELD3QELFU4N7MCI7FOT3:/var/lib/docker/overlay2/l/YGDBE4IOHQXR5ECOI4ZCBHAIIT:/var/lib/docker/overlay2/l/XGGPCHWJVDANU2B4ISJZ7JX2KI:/var/lib/docker/overlay2/l/VJROJ5LPTOXKSVYKJK463A3SZB:/var/lib/docker/overlay2/l/ZIWL5PNUTTEWS23HPG44QDSMNN,upperdir=/var/lib/docker/overlay2/7e427cf00268b04985194be27ee928c391b82a3eba75b9b18da9e9b0'",
      "votes": null
    },
    {
      "id": "372057",
      "postDate": "08/18/2018 06:47:21",
      "content": "<p>It looks like some random system trouble.</p>\n\n<p>Can you try again?</p>",
      "rawMarkdown": "It looks like some random system trouble.\n\nCan you try again?",
      "votes": null
    },
    {
      "id": "372402",
      "postDate": "08/19/2018 09:05:58",
      "content": "<p><a href=\"/kkaczmarek\">@kkaczmarek</a> - Thanks for your starter solution.. I have 2 questions here..</p>\n\n<ol>\n<li>What exactly is this metadata_filepath: /mnt/ml-team/minerva/open-solutions/googleai-object-detection/files/metadata.csv. I cannot find it on the Open Images download page.</li>\n<li>For train, do I need to download all the files from 00 to 08? I'm just trying to get some hands-on on this problem. So just the train_00.zip would be enough?</li>\n</ol>\n\n<p>TIA..</p>",
      "rawMarkdown": "kkaczmarek - Thanks for your starter solution.. I have 2 questions here..\n\n1. What exactly is this metadata_filepath: /mnt/ml-team/minerva/open-solutions/googleai-object-detection/files/metadata.csv. I cannot find it on the Open Images download page.\n2. For train, do I need to download all the files from 00 to 08? I'm just trying to get some hands-on on this problem. So just the train_00.zip would be enough?\n\nTIA..",
      "votes": null
    },
    {
      "id": "372525",
      "postDate": "08/19/2018 15:35:01",
      "content": "<p>I get this error when executing step 4. Evaluate/Predict RetinaNet:</p>\n\n<pre><code>Traceback (most recent call last):\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in &lt;module&gt;\nexecute()\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 103, in execute\nexecfile(job_filepath, job_globals)\nFile \"/usr/local/lib/python3.6/dist-packages/past/builtins/misc.py\", line 82, in execfile\nexec_(code, myglobals, mylocals)\n File \"main.py\", line 78, in &lt;module&gt;\nmain()\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 67, in evaluate_predict\npipeline_manager.evaluate(pipeline_name, dev_mode, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 24, in evaluate\nevaluate(pipeline_name, dev_mode, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 117, in evaluate\nprediction = generate_prediction(valid_img_ids, pipeline, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 182, in generate_prediction\nreturn _generate_prediction_in_chunks(img_ids, pipeline, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 214, in _generate_prediction_in_chunks\noutput = pipeline.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 370, in transform\nstep_output_data = self._cached_transform(step_inputs)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 477, in _cached_transform\nraise ValueError('No transformer cached {}'.format(self.name))\nValueError: No transformer cached label_encoder\n</code></pre>",
      "rawMarkdown": "I get this error when executing step 4. Evaluate/Predict RetinaNet:\n\n    Traceback (most recent call last):\n    File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in",
      "votes": null
    },
    {
      "id": "372551",
      "postDate": "08/19/2018 17:11:18",
      "content": "<p>Tried 3 different times over the weekend.  All 3 failed.</p>\n\n<p>My local (on my machine) works fine, but I would love to get more experiments running on different platforms. This is a beast of a dataset. So would be nice if I can run on neptune too.</p>",
      "rawMarkdown": "Tried 3 different times over the weekend.  All 3 failed.\n\nMy local (on my machine) works fine, but I would love to get more experiments running on different platforms. This is a beast of a dataset. So would be nice if I can run on neptune too.",
      "votes": null
    },
    {
      "id": "372789",
      "postDate": "08/20/2018 10:38:26",
      "content": "<p>It looks as if you didn't have your model/transformers saved.</p>\n\n<p>Check whether you specified your train folder correctly:</p>\n\n<pre><code> experiment_dir:  /output/experiment\n clone_experiment_dir_from:  /input/GAI-14/output/experiment\n</code></pre>\n\n<p>And pointed to it in the cli :</p>\n\n<pre><code> neptune send --worker m-4p100 \\\n --environment pytorch-0.3.1-gpu-py3 \\\n --config configs/neptune.yaml \\\n --input /GAI-14 \\\n main.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n</code></pre>\n\n<p>So in the example GAI-14 experiment above your <code>label_encoder</code> should be in the  <code>GAI-14/output/experiment/transformers/label_encoder</code> . Can you confirm that?</p>",
      "rawMarkdown": "It looks as if you didn't have your model/transformers saved.\n\nCheck whether you specified your train folder correctly:\n\n     experiment_dir:  /output/experiment\n     clone_experiment_dir_from:  /input/GAI-14/output/experiment\n\nAnd pointed to it in the cli :\n\n     neptune send --worker m-4p100 \\\n     --environment pytorch-0.3.1-gpu-py3 \\\n     --config configs/neptune.yaml \\\n     --input /GAI-14 \\\n     main.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n\nSo in the example GAI-14 experiment above your `label_encoder` should be in the  `GAI-14/output/experiment/transformers/label_encoder` . Can you confirm that?",
      "votes": null
    },
    {
      "id": "372791",
      "postDate": "08/20/2018 10:42:20",
      "content": "<p>Hi there @Samrat P .</p>\n\n<p>Are you talking about the latest master?\nBy accident it was left there in the <code>neptune.yaml</code> before.</p>\n\n<p>ad 1.\nBut we have a version of code that uses <code>metadata.csv</code> where we calculate aspect ratio for all images and later batch the train so that batches have give or take the same aspect ratio.</p>\n\n<p>ad. 2\nIf you want to train locally then yes you need to download it all (we downloaded in one go not in chunks).\nBut if you wanna run it in the cloud we have uploaded it all to neptune.\nSo if you simply follow the instructions in the repo you can start getting your hands dirty in no time!</p>",
      "rawMarkdown": "Hi there @Samrat P .\n\nAre you talking about the latest master?\nBy accident it was left there in the `neptune.yaml` before.\n\nad 1.\nBut we have a version of code that uses `metadata.csv` where we calculate aspect ratio for all images and later batch the train so that batches have give or take the same aspect ratio.\n\nad. 2\nIf you want to train locally then yes you need to download it all (we downloaded in one go not in chunks).\nBut if you wanna run it in the cloud we have uploaded it all to neptune.\nSo if you simply follow the instructions in the repo you can start getting your hands dirty in no time!",
      "votes": null
    },
    {
      "id": "372792",
      "postDate": "08/20/2018 10:48:36",
      "content": "<p><strong>Update</strong></p>\n\n<p>Hi all. </p>\n\n<p>We continue our work on retinanet. \nWe are training everything in batches of classes making sure that the classes that fall into one batch are give or take of similar prevalence in the dataset.\nSo far our solution contains as we call them batch_{1-8} . Those classes do not contain human related classes (we should add them todayish).</p>\n\n<p>Surprisingly (to me) augmentation did help. If you look at some of those classes there are very few examples so it isn't that surprising I guess it's just with the dataset of 1.7M images you kinda feel like those problems are not really important anymore.</p>\n\n<p>We are also working on weighted classification loss in retinanet where we weigh the loss per class based on the train distribution (and clip it to something like 10).</p>\n\n<p>Let's see what happens with that!</p>",
      "rawMarkdown": "**Update**\n\nHi all. \n\nWe continue our work on retinanet. \nWe are training everything in batches of classes making sure that the classes that fall into one batch are give or take of similar prevalence in the dataset.\nSo far our solution contains as we call them batch_{1-8} . Those classes do not contain human related classes (we should add them todayish).\n\nSurprisingly (to me) augmentation did help. If you look at some of those classes there are very few examples so it isn't that surprising I guess it's just with the dataset of 1.7M images you kinda feel like those problems are not really important anymore.\n\nWe are also working on weighted classification loss in retinanet where we weigh the loss per class based on the train distribution (and clip it to something like 10).\n\nLet's see what happens with that!",
      "votes": null
    },
    {
      "id": "372838",
      "postDate": "08/20/2018 13:02:09",
      "content": "<p>@Jakub Thank you. I check on it later today. </p>",
      "rawMarkdown": "Jakub Thank you. I check on it later today.",
      "votes": null
    },
    {
      "id": "372850",
      "postDate": "08/20/2018 13:19:01",
      "content": "<p>@Jakub,  What augmentation techniques did you use?</p>",
      "rawMarkdown": "Jakub,  What augmentation techniques did you use?",
      "votes": null
    },
    {
      "id": "372891",
      "postDate": "08/20/2018 15:00:56",
      "content": "<ul>\n<li>very small rotation -5:5 degrees</li>\n<li>small scaling 0.8 : 1.2</li>\n<li>left/right flip</li>\n<li>color augmentations (bluring and stuff)</li>\n</ul>",
      "rawMarkdown": "very small rotation -5:5 degrees\n - small scaling 0.8 : 1.2\n - left/right flip\n - color augmentations (bluring and stuff)",
      "votes": null
    },
    {
      "id": "372938",
      "postDate": "08/20/2018 17:36:42",
      "content": "<p><strong>Update</strong></p>\n\n<p>Added class batches with human labels.\nSo training on all classes we got to CV 515 LB 368 . Quite a large gap if u ask me. </p>\n\n<p>I wonder how much false negatives our batch &amp; merge approach generates.</p>",
      "rawMarkdown": "**Update**\n\nAdded class batches with human labels.\nSo training on all classes we got to CV 515 LB 368 . Quite a large gap if u ask me. \n\nI wonder how much false negatives our batch &amp; merge approach generates.",
      "votes": null
    },
    {
      "id": "372959",
      "postDate": "08/20/2018 18:48:10",
      "content": "<p>@Jakub, </p>\n\n<p>Here is what I inserted: </p>\n\n<pre><code>neptune send --worker m-4p100 \\\n--environment pytorch-0.3.1-gpu-py3 \\\n--config configs/neptune.yaml \\\n--input /KAG-18 \\\nmain.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n</code></pre>\n\n<p>I still get the same error message: </p>\n\n<pre><code>  experiment_dir:  KAG-18/output/experiment/transformers/label_encoder\n  clone_experiment_dir_from:  /input/KAG-18/output/experiment\n</code></pre>\n\n<p>I also tried : </p>\n\n<pre><code>experiment_dir:  /output/experiment\nclone_experiment_dir_from:  /input/KAG-18/output/experiment\n</code></pre>\n\n<p>Error:</p>\n\n<pre><code>Traceback (most recent call last):\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in \n&lt;module&gt;\nexecute()\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 103, in execute\nexecfile(job_filepath, job_globals)\nFile \"/usr/local/lib/python3.6/dist-packages/past/builtins/misc.py\", line 82, in execfile\nexec_(code, myglobals, mylocals)\nFile \"main.py\", line 78, in &lt;module&gt;\nmain()\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 67, in evaluate_predict\npipeline_manager.evaluate(pipeline_name, dev_mode, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 24, in evaluate\nevaluate(pipeline_name, dev_mode, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 117, in evaluate\nprediction = generate_prediction(valid_img_ids, pipeline, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 182, in generate_prediction\nreturn _generate_prediction_in_chunks(img_ids, pipeline, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 214, in _generate_prediction_in_chunks\noutput = pipeline.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 370, in transform\nstep_output_data = self._cached_transform(step_inputs)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 477, in _cached_transform\nraise ValueError('No transformer cached {}'.format(self.name))\nValueError: No transformer cached label_encoder\n</code></pre>",
      "rawMarkdown": "Jakub, \n\nHere is what I inserted: \n\n    neptune send --worker m-4p100 \\\n    --environment pytorch-0.3.1-gpu-py3 \\\n    --config configs/neptune.yaml \\\n    --input /KAG-18 \\\n    main.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n\nI still get the same error message: \n\n      experiment_dir:  KAG-18/output/experiment/transformers/label_encoder\n      clone_experiment_dir_from:  /input/KAG-18/output/experiment\n\nI also tried : \n\n    experiment_dir:  /output/experiment\n    clone_experiment_dir_from:  /input/KAG-18/output/experiment\n\nError:\n\n    Traceback (most recent call last):\n    File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in",
      "votes": null
    },
    {
      "id": "373279",
      "postDate": "08/21/2018 05:31:27",
      "content": "<p>@Jakub,\nDisregard, I was able to get the evaluation to run. I had to clone the github since I did not do after the update. So far everything is running okay. I will know for sure once the prediction is complete. </p>",
      "rawMarkdown": "Jakub,\nDisregard, I was able to get the evaluation to run. I had to clone the github since I did not do after the update. So far everything is running okay. I will know for sure once the prediction is complete.",
      "votes": null
    },
    {
      "id": "373324",
      "postDate": "08/21/2018 07:07:43",
      "content": "<p>I was able to get both the train and predict to run. However, evaluation_prediction file was empty. </p>\n\n<pre><code>Traceback (most recent call last):\n File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in &lt;module&gt;\nexecute()\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 103, in execute\nexecfile(job_filepath, job_globals)\nFile \"/usr/local/lib/python3.6/dist-packages/past/builtins/misc.py\", line 82, in execfile\nexec_(code, myglobals, mylocals)\nFile \"main.py\", line 78, in &lt;module&gt;\nmain()\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 68, in evaluate_predict\npipeline_manager.predict(pipeline_name, dev_mode, submit_predictions, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 27, in predict\npredict(pipeline_name, dev_mode, submit_predictions, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 161, in predict\nshutil.copytree(PARAMS.clone_experiment_dir_from, PARAMS.experiment_dir)\nFile \"/usr/lib/python3.6/shutil.py\", line 315, in copytree\nos.makedirs(dst)\nFile \"/usr/lib/python3.6/os.py\", line 220, in makedirs\nmkdir(name, mode)\nFileExistsError: [Errno 17] File exists: '/output/experiment'\n</code></pre>\n\n<p>Code was hang up here:</p>\n\n<pre><code> &gt;&gt;&gt; Formatting prediction\n &gt;&gt;&gt; Background predicted for all the images. Metric cannot be calculated\n &gt;&gt;&gt; predicting\n</code></pre>",
      "rawMarkdown": "I was able to get both the train and predict to run. However, evaluation_prediction file was empty. \n\n    Traceback (most recent call last):\n     File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in",
      "votes": null
    },
    {
      "id": "373326",
      "postDate": "08/21/2018 07:08:15",
      "content": "<p>Interesting, which classes did you group as \"human\"? </p>",
      "rawMarkdown": "Interesting, which classes did you group as \"human\"?",
      "votes": null
    },
    {
      "id": "373382",
      "postDate": "08/21/2018 09:12:57",
      "content": "<p>So we have 4 \"human\" batches:</p>\n\n<pre><code>[\"Footwear\", \"Human hair\",\"Person\", \"Human arm\", \"Human eye\", \"Human face\", \"Suit\", \"Human hand\", \"Human leg\", \"Human nose\", \"Dress\", \"Human mouth\"]\n\n[\"Jacket\", \"Sports uniform\", \"Human ear\", \"Shorts\", \"Helmet\", \"Hat\", \"Tie\", \"Swimwear\"]\n\n[\"Trousers\", \"Shirt\", \"Umbrella\", \"Coat\", \"Human beard\", \"Necklace\", \"Human head\"]\n\n[\"Scarf\", \"Human foot\", \"Luggage and bags\", \"Watch\", \"Brassiere\",\"Earrings\",\"Sock\",\"Skirt\",\"Glove\",\"Crown\",\"Swim cap\",\"Belt\",\"Tiara\"]\n</code></pre>\n\n<p>I guess they are more fashion/human after all.</p>",
      "rawMarkdown": "So we have 4 \"human\" batches:\n\n    [\"Footwear\", \"Human hair\",\"Person\", \"Human arm\", \"Human eye\", \"Human face\", \"Suit\", \"Human hand\", \"Human leg\", \"Human nose\", \"Dress\", \"Human mouth\"]\n    \n    [\"Jacket\", \"Sports uniform\", \"Human ear\", \"Shorts\", \"Helmet\", \"Hat\", \"Tie\", \"Swimwear\"]\n    \n    [\"Trousers\", \"Shirt\", \"Umbrella\", \"Coat\", \"Human beard\", \"Necklace\", \"Human head\"]\n    \n    [\"Scarf\", \"Human foot\", \"Luggage and bags\", \"Watch\", \"Brassiere\",\"Earrings\",\"Sock\",\"Skirt\",\"Glove\",\"Crown\",\"Swim cap\",\"Belt\",\"Tiara\"]\n\nI guess they are more fashion/human after all.",
      "votes": null
    },
    {
      "id": "373476",
      "postDate": "08/21/2018 12:19:07",
      "content": "<p>I like the idea of the approach, but is there a reason you decided to make the \"batch size\" so small? I looked at your other batches and it seems like they have about 100 classes each. Is it because of sampling reasons (meaning there are a lot more examples of human/fashion items in the train set)?</p>",
      "rawMarkdown": "I like the idea of the approach, but is there a reason you decided to make the \"batch size\" so small? I looked at your other batches and it seems like they have about 100 classes each. Is it because of sampling reasons (meaning there are a lot more examples of human/fashion items in the train set)?",
      "votes": null
    },
    {
      "id": "373506",
      "postDate": "08/21/2018 13:11:19",
      "content": "<p>Hi <a href=\"/muhammedazamkhan\">@muhammedazamkhan</a> <a href=\"/jakubczakon\">@jakubczakon</a>. I ran the predict command, but got back an empty submission file. I feel like my model didn't train long enough, especially given that I trained on all classes. Did you run your model on all 500 classes <a href=\"/muhammedazamkhan\">@muhammedazamkhan</a>? And also for 1000 epochs like the Neptune Team?</p>",
      "rawMarkdown": "Hi @muhammedazamkhan @jakubczakon. I ran the predict command, but got back an empty submission file. I feel like my model didn't train long enough, especially given that I trained on all classes. Did you run your model on all 500 classes @muhammedazamkhan? And also for 1000 epochs like the Neptune Team?",
      "votes": null
    },
    {
      "id": "373632",
      "postDate": "08/21/2018 17:13:03",
      "content": "<p>Hi Jakub,</p>\n\n<p>I am just wondering how long it takes you to train your whole solution on 4x P100? What about for just one of your class batches? Apologies if this info is in the neptune.ml experiments, I am a little confused by which run is which there. Also, do you train on the whole dataset?</p>",
      "rawMarkdown": "Hi Jakub,\n\nI am just wondering how long it takes you to train your whole solution on 4x P100? What about for just one of your class batches? Apologies if this info is in the neptune.ml experiments, I am a little confused by which run is which there. Also, do you train on the whole dataset?",
      "votes": null
    },
    {
      "id": "373644",
      "postDate": "08/21/2018 17:31:06",
      "content": "<p>For instance I trained this one \"batch_8\" <a href=\"https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce\">https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce</a> .</p>\n\n<p>I ran it local on 4 gtx 1070 though. Anyhow it was around 20 epochs/day  (120 total). Usually the batches would train for give or take 60 epochs and the smaller ones can actually be put on 2 gpus. </p>\n\n<p>I train on samples of 50000k per epoch for training and 10k for validation.\nThe samplers aren't a part of the public repo but they will be released right after the competition. \nOn the other hand we could release the code with the samplers but I am afraid it will make it to easy to get a medal with it.</p>",
      "rawMarkdown": "For instance I trained this one \"batch_8\" https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce .\n\n I ran it local on 4 gtx 1070 though. Anyhow it was around 20 epochs/day  (120 total). Usually the batches would train for give or take 60 epochs and the smaller ones can actually be put on 2 gpus. \n\nI train on samples of 50000k per epoch for training and 10k for validation.\nThe samplers aren't a part of the public repo but they will be released right after the competition. \nOn the other hand we could release the code with the samplers but I am afraid it will make it to easy to get a medal with it.",
      "votes": null
    },
    {
      "id": "373647",
      "postDate": "08/21/2018 17:34:04",
      "content": "<p>So does the <code>training_sample_size</code> parameter in the config mean number of samples per epoch but it still trains on the full dataset? Or does it mean that the dataset is truncated to N samples and those samples are reused for each epoch.</p>\n\n<p>I also assume by sampler you are referring to the method of which images are drawn from the dataset for training?</p>",
      "rawMarkdown": "So does the `training_sample_size` parameter in the config mean number of samples per epoch but it still trains on the full dataset? Or does it mean that the dataset is truncated to N samples and those samples are reused for each epoch.\n\nI also assume by sampler you are referring to the method of which images are drawn from the dataset for training?",
      "votes": null
    },
    {
      "id": "373653",
      "postDate": "08/21/2018 17:45:53",
      "content": "<p>Sorry for the confusion <a href=\"/anokas\">@anokas</a>, my bad. I wasn't sure which version was on public github.</p>\n\n<p>Anyways, yes <code>training_sample_size</code> is the size of the sample that is drawn every epoch. Sampler refers to the pytorch sampler which is later passed to the pytorch loader in this <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/loaders.py\">file</a>.</p>\n\n<p>And we do sample from the entire dataset every epoch (just different samples).</p>",
      "rawMarkdown": "Sorry for the confusion @anokas, my bad. I wasn't sure which version was on public github.\n\nAnyways, yes `training_sample_size` is the size of the sample that is drawn every epoch. Sampler refers to the pytorch sampler which is later passed to the pytorch loader in this [file](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/loaders.py).\n\nAnd we do sample from the entire dataset every epoch (just different samples).",
      "votes": null
    },
    {
      "id": "373725",
      "postDate": "08/21/2018 20:02:30",
      "content": "<p>I keep getting this error when evaluating:</p>\n\n<pre><code>FileExistsError: [Errno 17] File exists: '/output/experiment'\n</code></pre>\n\n<p>Any idea how to fix?</p>",
      "rawMarkdown": "I keep getting this error when evaluating:\n\n    FileExistsError: [Errno 17] File exists: '/output/experiment'\n\nAny idea how to fix?",
      "votes": null
    },
    {
      "id": "373768",
      "postDate": "08/21/2018 21:47:27",
      "content": "<p>I'm a little curious. Would we call that function in the pipeline_config and config files? </p>",
      "rawMarkdown": "I'm a little curious. Would we call that function in the pipeline_config and config files?",
      "votes": null
    },
    {
      "id": "373794",
      "postDate": "08/21/2018 23:32:33",
      "content": "<p>Thanks Jakub, I understand now.</p>\n\n<p>One thing I am curious about:\nThe model seems to spend a long time waiting between epochs. I understand the 6 minutes for validation, but does it take an additional 6 minutes to save the model, and then another 12 minutes to start the next epoch? Do you have any insights into what is going on in between these steps that might be making it slow?</p>\n\n<pre><code>2018-08-21 23:16:42 steppy &gt;&gt;&gt;; epoch 1 sum:     0.58112\n2018-08-21 23:22:44 steppy &gt;&gt;&gt;; epoch 1 validation sum:     0.53714\n2018-08-21 23:28:45 steppy &gt;&gt;&gt;; epoch 1 model persisted to experiment2/checkpoints/retinanet/best.torch\n2018-08-21 23:28:45 steppy &gt;&gt;&gt;; epoch 2 current lr: 1e-05\n2018-08-21 23:40:48 steppy &gt;&gt;&gt;; epoch 1 time 0:54:58\n2018-08-21 23:40:48 steppy &gt;&gt;&gt;; epoch 2 ...\n</code></pre>",
      "rawMarkdown": "Thanks Jakub, I understand now.\n\nOne thing I am curious about:\nThe model seems to spend a long time waiting between epochs. I understand the 6 minutes for validation, but does it take an additional 6 minutes to save the model, and then another 12 minutes to start the next epoch? Do you have any insights into what is going on in between these steps that might be making it slow?\n\n    2018-08-21 23:16:42 steppy &gt;&gt;&gt;; epoch 1 sum:     0.58112\n    2018-08-21 23:22:44 steppy &gt;&gt;&gt;; epoch 1 validation sum:     0.53714\n    2018-08-21 23:28:45 steppy &gt;&gt;&gt;; epoch 1 model persisted to experiment2/checkpoints/retinanet/best.torch\n    2018-08-21 23:28:45 steppy &gt;&gt;&gt;; epoch 2 current lr: 1e-05\n    2018-08-21 23:40:48 steppy &gt;&gt;&gt;; epoch 1 time 0:54:58\n    2018-08-21 23:40:48 steppy &gt;&gt;&gt;; epoch 2 ...",
      "votes": null
    },
    {
      "id": "373925",
      "postDate": "08/22/2018 05:39:05",
      "content": "<p>Thanks for spotting that <a href=\"/anokas\">@anokas</a> I haven't noticed that before.\nI checked the latest experiment and the same thing happened.</p>\n\n<p>My guess is that it could be that by mistake the loss on validation is calculated multiple times. Looking at the callback list:</p>\n\n<pre><code>return CallbackList(\n    callbacks=[experiment_timing, training_monitor, validation_monitor,\n               model_checkpoints, lr_scheduler, early_stopping, neptune_monitor,\n               ]) \n</code></pre>\n\n<p>after the validation_monitor the validation loss is needed in model_checkpoints, early_stopping and neptune_monitor. I wonder if turning some of those off speeds things up considerably.</p>\n\n<p>Unfortunately, I am not sure that I will have the time to fix that till Monday. So if you would like to fix/check that yourself. The vast majority of the logic is in the <code>src/callbacks.py</code> the rest sits in the <a href=\"https://github.com/neptune-ml/steppy-toolkit/blob/master/toolkit/pytorch_transformers/callbacks.py\">steppy-toolkit package</a>. </p>",
      "rawMarkdown": "Thanks for spotting that @anokas I haven't noticed that before.\nI checked the latest experiment and the same thing happened.\n\nMy guess is that it could be that by mistake the loss on validation is calculated multiple times. Looking at the callback list:\n\n    return CallbackList(\n        callbacks=[experiment_timing, training_monitor, validation_monitor,\n                   model_checkpoints, lr_scheduler, early_stopping, neptune_monitor,\n                   ]) \n\nafter the validation_monitor the validation loss is needed in model_checkpoints, early_stopping and neptune_monitor. I wonder if turning some of those off speeds things up considerably.\n\nUnfortunately, I am not sure that I will have the time to fix that till Monday. So if you would like to fix/check that yourself. The vast majority of the logic is in the `src/callbacks.py` the rest sits in the [steppy-toolkit package](https://github.com/neptune-ml/steppy-toolkit/blob/master/toolkit/pytorch_transformers/callbacks.py).",
      "votes": null
    },
    {
      "id": "373971",
      "postDate": "08/22/2018 07:38:58",
      "content": "<p>Can you paste your data paths? Does it fail after evaluation ( map value has been calculated) or straight away? What os the exact command you are using?</p>",
      "rawMarkdown": "Can you paste your data paths? Does it fail after evaluation ( map value has been calculated) or straight away? What os the exact command you are using?",
      "votes": null
    },
    {
      "id": "374073",
      "postDate": "08/22/2018 11:52:33",
      "content": "<p>Hi @Giri Gopalan.</p>\n\n<p>I am running my evaluate_predict experiment right now and it seems to be working just fine.\nSo in my case I:</p>\n\n<ul>\n<li><p>trained the model with this <a href=\"https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca\">neptune\nexperiment GAI-473</a> running the following command</p>\n\n<pre><code> neptune send --worker m-4p100 \\\n --environment pytorch-0.3.1-gpu-py3 \\\n --config configs/neptune.yaml \\\n main.py train --pipeline_name retinanet\n</code></pre></li>\n</ul>\n\n<p>and my <code>neptune.yaml</code> looked like this</p>\n\n<pre><code>parameters:\n# Data Paths\n  train_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/train\n  test_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/test_challenge_2018\n  annotations_filepath: /public/challenges/google-ai-open-images-object-detection-track/annotations/challenge-2018-train-annotations-bbox.csv\n  annotations_human_labels_filepath: /public/challenges/google-ai-open-images-object-detection-track/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n  bbox_hierarchy_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/bbox_labels_500_hierarchy.json\n  class_mappings_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/challenge-2018-class-descriptions-500.csv\n  valid_ids_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/challenge-2018-image-ids-valset-od.csv\n  sample_submission: /public/challenges/google-ai-open-images-object-detection-track/sample_submission.csv\n  experiment_dir:  /output/experiment\n  clone_experiment_dir_from: ''\n</code></pre>\n\n<p>Notice that the <code>clone_experiment_dir_from: ''</code> during training. </p>\n\n<ul>\n<li><p>To evaluate I changed the <code>neptune.yaml</code> to</p>\n\n<pre><code>clone_experiment_dir_from: /input/GAI-473/output/experiment \n</code></pre>\n\n<p>and ran the following command</p>\n\n<pre><code> neptune send --worker m-p100 \\\n--environment pytorch-0.3.1-gpu-py3 \\\n--config configs/neptune.yaml \\\n--input /GAI-473 \\\nmain.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n</code></pre></li>\n</ul>\n\n<p>Notice that during prediction I am using just one p100 machine with the command <code>m-p100</code> but it is of course not necessary to change that (just cheaper).</p>\n\n<p>Have you done it exactly this way as well ?</p>",
      "rawMarkdown": "Hi @Giri Gopalan.\n\nI am running my evaluate_predict experiment right now and it seems to be working just fine.\nSo in my case I:\n\n - trained the model with this [neptune\n   experiment GAI-473](https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca) running the following command\n\n         neptune send --worker m-4p100 \\\n         --environment pytorch-0.3.1-gpu-py3 \\\n         --config configs/neptune.yaml \\\n         main.py train --pipeline_name retinanet\n\nand my `neptune.yaml` looked like this\n\n    parameters:\n    # Data Paths\n      train_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/train\n      test_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/test_challenge_2018\n      annotations_filepath: /public/challenges/google-ai-open-images-object-detection-track/annotations/challenge-2018-train-annotations-bbox.csv\n      annotations_human_labels_filepath: /public/challenges/google-ai-open-images-object-detection-track/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n      bbox_hierarchy_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/bbox_labels_500_hierarchy.json\n      class_mappings_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/challenge-2018-class-descriptions-500.csv\n      valid_ids_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/challenge-2018-image-ids-valset-od.csv\n      sample_submission: /public/challenges/google-ai-open-images-object-detection-track/sample_submission.csv\n      experiment_dir:  /output/experiment\n      clone_experiment_dir_from: ''\n\nNotice that the `clone_experiment_dir_from: '' ` during training. \n\n  - To evaluate I changed the `neptune.yaml` to\n\n        clone_experiment_dir_from: /input/GAI-473/output/experiment \n    \n    and ran the following command\n\n         neptune send --worker m-p100 \\\n        --environment pytorch-0.3.1-gpu-py3 \\\n        --config configs/neptune.yaml \\\n        --input /GAI-473 \\\n        main.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n\nNotice that during prediction I am using just one p100 machine with the command `m-p100` but it is of course not necessary to change that (just cheaper).\n\nHave you done it exactly this way as well ?",
      "votes": null
    },
    {
      "id": "374080",
      "postDate": "08/22/2018 12:05:51",
      "content": "<pre><code># Data Paths\ntrain_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/train\ntest_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/test_challenge_2018\nannotations_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/annotations/challenge-2018-train-annotations-bbox.csv\nannotations_human_labels_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/annotations/challenge-2018-train-annotations-human-imagelabels.csv\nbbox_hierarchy_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/metadata/bbox_labels_500_hierarchy.json\nclass_mappings_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/metadata/challenge-2018-class-descriptions-500.csv\nvalid_ids_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/metadata/challenge-2018-image-ids-valset-od.csv\nsample_submission: /public/challenges/google-ai-open-images-object-detection- \ntrack/sample_submission.csv\nexperiment_dir:  /output/experiment\nclone_experiment_dir_from:  /input/GOOG-2/output/experiment\n</code></pre>\n\n<p>I am trying this setup per your recommendation:</p>\n\n<pre><code>experiment_dir:  /output/experiment\nclone_experiment_dir_from:  GOOG-2/output/experiment/transformers/label_encoder\n</code></pre>",
      "rawMarkdown": "# Data Paths\n    train_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/train\n    test_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/test_challenge_2018\n    annotations_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/annotations/challenge-2018-train-annotations-bbox.csv\n    annotations_human_labels_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n    bbox_hierarchy_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/metadata/bbox_labels_500_hierarchy.json\n    class_mappings_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/metadata/challenge-2018-class-descriptions-500.csv\n    valid_ids_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/metadata/challenge-2018-image-ids-valset-od.csv\n    sample_submission: /public/challenges/google-ai-open-images-object-detection- \n    track/sample_submission.csv\n    experiment_dir:  /output/experiment\n    clone_experiment_dir_from:  /input/GOOG-2/output/experiment\n\nI am trying this setup per your recommendation:\n\n    experiment_dir:  /output/experiment\n    clone_experiment_dir_from:  GOOG-2/output/experiment/transformers/label_encoder",
      "votes": null
    },
    {
      "id": "374089",
      "postDate": "08/22/2018 12:23:36",
      "content": "<p>@William Greene I think it should be:</p>\n\n<pre><code>clone_experiment_dir_from:  GOOG-2/output/experiment\n</code></pre>\n\n<p>not </p>\n\n<pre><code>clone_experiment_dir_from:  GOOG-2/output/experiment/transformers/label_encoder\n</code></pre>",
      "rawMarkdown": "William Greene I think it should be:\n\n    clone_experiment_dir_from:  GOOG-2/output/experiment\n\nnot \n\n    clone_experiment_dir_from:  GOOG-2/output/experiment/transformers/label_encoder",
      "votes": null
    },
    {
      "id": "374101",
      "postDate": "08/22/2018 12:44:40",
      "content": "<p>Ok I think I found the culprit should have fix in no time.</p>\n\n<p>If you don't want to wait for it you can simply run 'evaluate' pipeline and 'predict' pipeline separately in two consecutive experiments and it will work just fine. </p>",
      "rawMarkdown": "Ok I think I found the culprit should have fix in no time.\n\nIf you don't want to wait for it you can simply run 'evaluate' pipeline and 'predict' pipeline separately in two consecutive experiments and it will work just fine.",
      "votes": null
    },
    {
      "id": "374103",
      "postDate": "08/22/2018 12:51:21",
      "content": "<p>Can you try running the <code>evaluate</code> pipeline and not 'evaluate_predict' pipeline first </p>",
      "rawMarkdown": "Can you try running the `evaluate` pipeline and not 'evaluate_predict' pipeline first",
      "votes": null
    },
    {
      "id": "374106",
      "postDate": "08/22/2018 12:59:06",
      "content": "<p>@William Green please check the newest master</p>",
      "rawMarkdown": "William Green please check the newest master",
      "votes": null
    },
    {
      "id": "374110",
      "postDate": "08/22/2018 13:04:04",
      "content": "<p>Thanks for all your help so far. Is there any way I can resume training from a checkpoint, if the training script is stopped?</p>\n\n<p>I am unfamiliar with steppy so I am not sure how to change the code to do this.</p>",
      "rawMarkdown": "Thanks for all your help so far. Is there any way I can resume training from a checkpoint, if the training script is stopped?\n\nI am unfamiliar with steppy so I am not sure how to change the code to do this.",
      "votes": null
    },
    {
      "id": "374143",
      "postDate": "08/22/2018 13:49:15",
      "content": "<p>Sure <a href=\"/anokas\">@anokas</a>. Glad I could be of help.</p>\n\n<p>If you go to <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py\">src/models.py</a> :</p>\n\n<p>You have something like this:</p>\n\n<pre><code>class ModelParallel(Model):\n    def fit(self, datagen, validation_datagen=None):\n        self._initialize_model_weights()\n\n        self.model = DataParallel(self.model)\n\n        if torch.cuda.is_available():\n            self.model = self.model.cuda() \n</code></pre>\n\n<p>simply add loading:</p>\n\n<pre><code>class ModelParallel(Model):\n    def fit(self, datagen, validation_datagen=None):\n        self._initialize_model_weights()\n\n        self.model = DataParallel(self.model)\n        self.load(YOUR_FILEPATH)\n\n        if torch.cuda.is_available():\n            self.model = self.model.cuda() \n</code></pre>\n\n<p>If you want to do some custom loading of say some backbone layers or something I would suggest that you go to the <code>load</code> method and check how it's done there and probably add a 'custom_load' method for your use case.</p>",
      "rawMarkdown": "Sure @anokas. Glad I could be of help.\n\n If you go to [src/models.py](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py) :\n\nYou have something like this:\n\n    class ModelParallel(Model):\n        def fit(self, datagen, validation_datagen=None):\n            self._initialize_model_weights()\n    \n            self.model = DataParallel(self.model)\n    \n            if torch.cuda.is_available():\n                self.model = self.model.cuda() \n\nsimply add loading:\n\n    class ModelParallel(Model):\n        def fit(self, datagen, validation_datagen=None):\n            self._initialize_model_weights()\n    \n            self.model = DataParallel(self.model)\n            self.load(YOUR_FILEPATH)\n\n            if torch.cuda.is_available():\n                self.model = self.model.cuda() \n\nIf you want to do some custom loading of say some backbone layers or something I would suggest that you go to the `load` method and check how it's done there and probably add a 'custom_load' method for your use case.",
      "votes": null
    },
    {
      "id": "374159",
      "postDate": "08/22/2018 14:07:26",
      "content": "<p>@Jakub Thank you. I'm checking it out right now. </p>",
      "rawMarkdown": "Jakub Thank you. I'm checking it out right now.",
      "votes": null
    },
    {
      "id": "374601",
      "postDate": "08/23/2018 11:20:59",
      "content": "<p>Hi @Jonas \nI've just checked the implementation with default 10 classes and have run for 5 epochs.</p>",
      "rawMarkdown": "Hi @Jonas \nI've just checked the implementation with default 10 classes and have run for 5 epochs.",
      "votes": null
    },
    {
      "id": "374672",
      "postDate": "08/23/2018 14:43:53",
      "content": "<p>OK, I found out what was causing my crashes. In the evaluate runs that crashed, I did NOT use --chunk-size. When I set it per your response above, it went through properly.</p>\n\n<p>I think things should not crash if an option is not set correctly.  Some sort of error message would have been easier for me to debug.</p>\n\n<p>In any case, thanks much for the help and all of the work you guys have done.</p>",
      "rawMarkdown": "OK, I found out what was causing my crashes. In the evaluate runs that crashed, I did NOT use --chunk-size. When I set it per your response above, it went through properly.\n\nI think things should not crash if an option is not set correctly.  Some sort of error message would have been easier for me to debug.\n\nIn any case, thanks much for the help and all of the work you guys have done.",
      "votes": null
    },
    {
      "id": "374682",
      "postDate": "08/23/2018 15:08:18",
      "content": "<p>What was the baseline score for solution 1? </p>",
      "rawMarkdown": "What was the baseline score for solution 1?",
      "votes": null
    },
    {
      "id": "375068",
      "postDate": "08/24/2018 14:10:11",
      "content": "<p>@ Jakub, How many batches are there total?</p>",
      "rawMarkdown": "Jakub, How many batches are there total?",
      "votes": null
    },
    {
      "id": "375070",
      "postDate": "08/24/2018 14:20:15",
      "content": "<p>We have 12 batches right now but the experiments with some of them joined together are in progress.</p>",
      "rawMarkdown": "We have 12 batches right now but the experiments with some of them joined together are in progress.",
      "votes": null
    },
    {
      "id": "375072",
      "postDate": "08/24/2018 14:26:33",
      "content": "<p>@Jakub, </p>\n\n<p>How do you apply the method for training from a checkpoint in the cloud? </p>",
      "rawMarkdown": "Jakub, \n\nHow do you apply the method for training from a checkpoint in the cloud?",
      "votes": null
    },
    {
      "id": "375091",
      "postDate": "08/24/2018 15:25:29",
      "content": "<p>Simply set data paths as you would with evaluate/predict and add --input YOUREXP-1 in the command and your checkpoint will be available in the /output/experiment/checkpoints/retinanet/best.torch .</p>",
      "rawMarkdown": "Simply set data paths as you would with evaluate/predict and add --input YOUREXP-1 in the command and your checkpoint will be available in the /output/experiment/checkpoints/retinanet/best.torch .",
      "votes": null
    },
    {
      "id": "376754",
      "postDate": "08/28/2018 02:57:00",
      "content": "<p>Missing experiment folder.  For some reason, the experiment folder was not created after training was complete. </p>",
      "rawMarkdown": "Missing experiment folder.  For some reason, the experiment folder was not created after training was complete.",
      "votes": null
    },
    {
      "id": "376857",
      "postDate": "08/28/2018 07:13:47",
      "content": "<p>Local or cloud?</p>",
      "rawMarkdown": "Local or cloud?",
      "votes": null
    },
    {
      "id": "376948",
      "postDate": "08/28/2018 11:12:43",
      "content": "<p>Cloud.</p>",
      "rawMarkdown": "Cloud.",
      "votes": null
    },
    {
      "id": "376995",
      "postDate": "08/28/2018 12:48:21",
      "content": "<p>So output/experiment is empty or non-existent? </p>",
      "rawMarkdown": "So output/experiment is empty or non-existent?",
      "votes": null
    },
    {
      "id": "377064",
      "postDate": "08/28/2018 14:36:51",
      "content": "<p>It's non-existent. I think I may have figured out why.</p>\n\n<p>I had <code>output/experiment/</code> versus <code>/output/experiment</code></p>",
      "rawMarkdown": "It's non-existent. I think I may have figured out why.\n\nI had `output/experiment/` versus ` /output/experiment`",
      "votes": null
    },
    {
      "id": "377230",
      "postDate": "08/28/2018 19:17:08",
      "content": "<p>@Jskub can I use the same cmd line for evaluation?</p>\n\n<p>clone:   /output/experiment/checkpoints/retinanet/best.torch</p>",
      "rawMarkdown": "Jskub can I use the same cmd line for evaluation?\n\n  clone:   /output/experiment/checkpoints/retinanet/best.torch",
      "votes": null
    },
    {
      "id": "377251",
      "postDate": "08/28/2018 20:29:24",
      "content": "<p>Is this the correct command to train from last checkpoint</p>\n\n<pre><code>neptune send --worker m-4p100 \\\n--environment pytorch-0.3.1-gpu-py3 \\\n--config configs/neptune.yaml \\\n--input /GOOG-50 \\\nmain.py train --pipeline_name retinanet\n</code></pre>\n\n<p>config: </p>\n\n<pre><code>      experiment_dir:  /output/experiment/checkpoints/retinanet/best.torch\n      clone_experiment_dir_from: ''\n</code></pre>\n\n<p>I get the following error when I either try to run from last know checkpoint or evaluate:</p>\n\n<pre><code>00:00   /usr/bin/python: can't open file '/usr/local/lib/python3.6/dist-packages/neptune/job_wrapper.py': [Errno 2] No \n such file or directory\n</code></pre>",
      "rawMarkdown": "Is this the correct command to train from last checkpoint\n\n    neptune send --worker m-4p100 \\\n    --environment pytorch-0.3.1-gpu-py3 \\\n    --config configs/neptune.yaml \\\n    --input /GOOG-50 \\\n    main.py train --pipeline_name retinanet\n\nconfig: \n\n          experiment_dir:  /output/experiment/checkpoints/retinanet/best.torch\n          clone_experiment_dir_from: ''\n\n\nI get the following error when I either try to run from last know checkpoint or evaluate:\n\n    00:00\t/usr/bin/python: can't open file '/usr/local/lib/python3.6/dist-packages/neptune/job_wrapper.py': [Errno 2] No \n     such file or directory",
      "votes": null
    },
    {
      "id": "377291",
      "postDate": "08/28/2018 22:12:29",
      "content": "<p>config should be:</p>\n\n<pre><code>  experiment_dir:  /output/experiment/checkpoints/retinanet/best.torch\n  clone_experiment_dir_from: /input/GOOG-50/output/experiment\n</code></pre>\n\n<p>and you should insert the following line:</p>\n\n<pre><code>    self.load('/input/GOOG-50/output/experiment/checkpoints/retinanet/best.torch')\n</code></pre>\n\n<p>In the <code>models.py</code> as explained to <a href=\"/anokas\">@anokas</a> below.</p>",
      "rawMarkdown": "config should be:\n\n      experiment_dir:  /output/experiment/checkpoints/retinanet/best.torch\n      clone_experiment_dir_from: /input/GOOG-50/output/experiment\n\nand you should insert the following line:\n\n        self.load('/input/GOOG-50/output/experiment/checkpoints/retinanet/best.torch')\n\n\nIn the `models.py` as explained to @anokas below.",
      "votes": null
    },
    {
      "id": "377629",
      "postDate": "08/29/2018 13:33:44",
      "content": "<p>I'm still getting this error </p>\n\n<pre><code>0:00    /usr/bin/python: can't open file '/usr/local/lib/python3.6/dist-packages/neptune/job_wrapper.py': [Errno 2] \nNo such file or directory\n</code></pre>",
      "rawMarkdown": "I'm still getting this error \n\n    0:00\t/usr/bin/python: can't open file '/usr/local/lib/python3.6/dist-packages/neptune/job_wrapper.py': [Errno 2] \n    No such file or directory",
      "votes": null
    },
    {
      "id": "377697",
      "postDate": "08/29/2018 16:00:12",
      "content": "<p>I am pretty sure that the problem was with the wrong neptune-cli version in the requirements.tx . </p>\n\n<p>It is fixed now so it should run with no problems. Just go to the repo and check.</p>",
      "rawMarkdown": "I am pretty sure that the problem was with the wrong neptune-cli version in the requirements.tx . \n\nIt is fixed now so it should run with no problems. Just go to the repo and check.",
      "votes": null
    },
    {
      "id": "377746",
      "postDate": "08/29/2018 17:17:24",
      "content": "<p>@Jakub </p>\n\n<p>Thank you :) I finally was able to get it to run. It works :) </p>",
      "rawMarkdown": "Jakub \n\nThank you :) I finally was able to get it to run. It works :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 367849,
      "author_name": "jakubczakon",
      "author_url": "",
      "post_date": "08/08/2018 17:11:27",
      "content": "<p>Gotta say that our story in this comp is a constant battle with making retinanetwork work. </p>\n\n<p>First the performance and anchor setup, then working with 500 classes, batching by aspect ratios and dividing the problem into smaller subproblems.</p>\n\n<p>Now we are brewing some experiments (check on neptune.ml) and expecting improvements\nbut obviously still strugling. At this point the biggest debacle is, how to balance classes for the images with multiple bboxes. </p>\n\n<p>If you have any ideas feel free to chime in!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 368079,
      "author_name": "yk1598",
      "author_url": "",
      "post_date": "08/09/2018 08:09:54",
      "content": "<p>LB 0.21694 score falls in the silver range. In a discussion at TGS Salt ID competition, you promised <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/62162#365436\">here</a> that you won't publish code that'll f**k up the leaderboard. Come on dude, it takes like a billion years to train on this dataset. Not cool. </p>",
      "votes": null,
      "replies": [
        {
          "id": 368100,
          "author_name": "taraspiotr",
          "author_url": "",
          "post_date": "08/09/2018 09:01:26",
          "content": "<p>Hi,\ndon't worry, the code from the public repository won't take you to the medal zone. As stated in the post:</p>\n\n<p>&gt; (no worries competitive guys -&gt; we only publish code that scores below bronze medal)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368108,
          "author_name": "kkaczmarek",
          "author_url": "",
          "post_date": "08/09/2018 09:11:26",
          "content": "<p>Hi <a href=\"/yk1598\">@yk1598</a>,</p>\n\n<blockquote>\n  <p>the code on GitHub - (no worries competitive guys -&gt; we only publish code that scores below bronze medal)</p>\n</blockquote>\n\n<p>Code in the repository will not give your silver medal. No worries, we know what we are publishing.</p>\n\n<p>In the next post I will describe our work in more detail. Also, this starter gives you good starting point for your own work :)</p>\n\n<p>Best,</p>\n\n<p>Kamil</p>\n\n<p>BTW: nice avatar!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368110,
          "author_name": "yk1598",
          "author_url": "",
          "post_date": "08/09/2018 09:15:54",
          "content": "<p>Ah I see! I got confused by the title. Apologies.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 368142,
      "author_name": "kkaczmarek",
      "author_url": "",
      "post_date": "08/09/2018 10:32:34",
      "content": "<h2>Competition update</h2>\n\n<p>It took us quite a lot of time to develop reasonable solution to this competition.</p>\n\n<p>We have decided to work with <strong>RetinaNet</strong> and focal loss, described in this paper: <a href=\"https://arxiv.org/abs/1708.02002\">Focal Loss for Dense Object Detection</a>. If you are new to RetinaNet - I recommend to skim through blog post that describes <a href=\"https://medium.com/@14prakash/the-intuition-behind-retinanet-eb636755607d\">The intuition behind RetinaNet</a>.</p>\n\n<h2>Code</h2>\n\n<p>What you can see on <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection\">master branch</a> is training procedure on 10 classes -&gt; you can extend it to work on the entire dataset. We worked with this code to quickly iterate over various ideas.</p>\n\n<h2>Our analysis of the problem</h2>\n\n<p>We quickly decided that this competition consist of two subproblems, each to be approached separately:</p>\n\n<ul>\n<li><em>First subproblem</em> is classes related to <em>people</em> and <em>clothing</em>, because the bboxes overlap a lot and there are multiple bboxes per image. Here, we have approximately 80 classes.</li>\n<li><em>Second subproblem</em> is remaining classes. Here, we take all these classes and divide it into 7 bins. Each bin is occupied by classes with similar frequency in the dataset. We need such bins to prepare proper epoch as described below. </li>\n</ul>\n\n<h2>Preprocessing for training</h2>\n\n<ul>\n<li>When we run training for the remaining classes, we make sure that each class (within an epoch) has similar number of occurrences -&gt; we implemented sampler to do this work. Thanks to this we have more balanced problem. In practice we oversample rare classes and subsample frequent classes. 7 bins mentioned above are utilized here.</li>\n<li>Next, we calculate <strong>aspect ratio</strong> and we prepare batches only for images with similar <strong>aspect ratio</strong>. We need this in the next step - resize. After resize all images are similarly squeezed - training signal is better balanced.</li>\n<li>At this point we are ready to feed batch to the network. Images are with similar aspect ratio, classes within the epoch are balanced, so training signal is stronger.</li>\n<li>Resulting experiment is like <a href=\"https://app.neptune.ml/-/dashboard/experiment/f945da64-6dd3-459b-94c5-58bc6a83f590\">this one</a>.</li>\n</ul>\n\n<h2>Other remarks</h2>\n\n<ul>\n<li>It is good to start experimenting with few classes (like 10) and get better feel of the problem. We also run <a href=\"https://app.neptune.ml/-/dashboard/experiment/c779468e-d3f7-44b8-a3a4-43a012315708\">training on 10 classes</a>.</li>\n<li>We noticed, that for rare classes augmentations are necessary :)</li>\n</ul>\n\n<h2>Open Questions</h2>\n\n<ul>\n<li>What to do with highly overlapping bboxes (<em>people</em> and <em>clothing</em> subproblem).</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 368312,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/09/2018 17:18:39",
      "content": "<p>@Kamil I think this may answer your questions <a href=\"https://stackoverflow.com/questions/49951422/get-rid-of-overlapping-bounding-boxes-across-different-classes-in-tensorflow-obj\">non_max_suppression over all classes</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 368397,
          "author_name": "kkaczmarek",
          "author_url": "",
          "post_date": "08/09/2018 20:18:04",
          "content": "<p>Hi <a href=\"/dskswu\">@dskswu</a>, thanks I'll check it :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 368500,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/10/2018 03:06:49",
      "content": "<p>I think you have a typo <code>python main.py train --pipeline_name retinanet</code> in your instructions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 368909,
          "author_name": "kkaczmarek",
          "author_url": "",
          "post_date": "08/11/2018 09:00:40",
          "content": "<p>Hi, it should be <code>python main.py -- train --pipeline_name retinanet</code></p>\n\n<p>thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 369813,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/13/2018 20:09:14",
      "content": "<p>Is this the correct data path parameters? </p>\n\n<pre><code>   train_imgs_dir: train_03/\n   test_imgs_dir: test/\n   annotations_filepath: annotations/\n   annotations_human_labels_filepath: annotations_human_labels/\n   bbox_hierarchy_filepath: bbox_hierarchy/\n   valid_ids_filepath: valid_ids/\n   sample_submission: sample_submission.csv\n   experiment_dir:  experiment\n   class_mappings_filepath: class_mappings/\n</code></pre>\n\n<p>I am trying to troubleshoot this error: </p>\n\n<pre><code>neptune: Executing in Offline Mode.\nneptune: Executing in Offline Mode.\n2018-08-13 20-04-33 google-ai-odt &gt;&gt;&gt; training\nTraceback (most recent call last):\n File \"main.py\", line 78, in &lt;module&gt;\nmain()\n File \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/site-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 16, in train\npipeline_manager.train(pipeline_name, dev_mode)\nFile \"/floyd/home/src/pipeline_manager.py\", line 21, in train\ntrain(pipeline_name, dev_mode)\nFile \"/floyd/home/src/pipeline_manager.py\", line 38, in train\nannotations = pd.read_csv(PARAMS.annotations_filepath)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 655, in parser_f\nreturn _read(filepath_or_buffer, kwds)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 405, in _read\nparser = TextFileReader(filepath_or_buffer, **kwds)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 764, in __init__\nself._make_engine(self.engine)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 985, in _make_engine\nself._engine = CParserWrapper(self.f, **self.options)\nFile \"/usr/local/lib/python3.6/site-packages/pandas/io/parsers.py\", line 1605, in __init__\nself._reader = parsers.TextReader(src, **kwds)\nFile \"pandas/_libs/parsers.pyx\", line 394, in pandas._libs.parsers.TextReader.__cinit__ (pandas/_libs/parsers.c:4209)\nFile \"pandas/_libs/parsers.pyx\", line 710, in pandas._libs.parsers.TextReader._setup_parser_source \n(pandas/_libs/parsers.c:8873)\n FileNotFoundError: File b'' does not exist\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 370045,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/14/2018 06:20:18",
          "content": "<p>Not quite. Mine looks something like this:</p>\n\n<pre><code>parameters:\n # Data Paths\n\n train_imgs_dir: .../open-images-v4/bounding-boxes/train\n\n test_imgs_dir: .../open-images-v4/bounding-boxes/test_challenge_2018\n\n annotations_filepath: .../googleai-object-detection/data/annotations/challenge-2018-train-annotations-bbox.csv\n\n annotations_human_labels_filepath:.../googleai-object-detection/data/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n\n bbox_hierarchy_filepath: .../googleai-object-detection/data/metadata/bbox_labels_500_hierarchy.json\n\n class_mappings_filepath: .../googleai-object-detection/data/metadata/challenge-2018-class-descriptions-500.csv\n\n valid_ids_filepath: .../googleai-object-detection/data/metadata/challenge-2018-image-ids-valset-od.csv\n\nsample_submission: .../googleai-object-detection/data/sample_submission.csv\n\nmetadata_filepath: .../googleai-object-detection/files/metadata.csv\n\nexperiment_dir:  .../googleai-object-detection/experiments/10_classes\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370325,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/14/2018 16:15:23",
          "content": "<p>@Jakub, Thank you. I totally missed this on the download page. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 370944,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "08/15/2018 17:55:24",
      "content": "<p>I am running the latest master branch (offline), and when the code gets to the training point it crashes when trying to forward() the model and evaluate the loss function:</p>\n\n<pre><code>neptune: Executing in Offline Mode.\n2018-08-15 18-34-40 google-ai-odt &gt;&gt;&gt; training\n2018-08-15 18-35-03 google-ai-odt &gt;&gt;&gt; Training on a reduced class subset: ['Person', 'Car', 'Dress', 'Footwear']\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing Step label_encoder...\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing Step label_encoder...\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing experiment directories under experiments\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing experiment directories under experiments\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; done: initializing experiment directories\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; done: initializing experiment directories\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; Step label_encoder initialized\n2018-08-15 18:35:05 steppy &gt;&gt;&gt; Step label_encoder initialized\n\n[skipped]\n\n2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, fitting and transforming...\n2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, fitting and transforming...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; starting training...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; starting training...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; initial lr: 1e-05\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; initial lr: 1e-05\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 ...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 ...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 batch 0 ...\n2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 batch 0 ...\nTraceback (most recent call last):\n  File \"main.py\", line 78, in &lt;module&gt;\n    main()\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 722, in __call__\n    return self.main(*args, **kwargs)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 697, in main\n    rv = self.invoke(ctx)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 1066, in invoke\n    return _process_result(sub_ctx.command.invoke(sub_ctx))\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 895, in invoke\n    return ctx.invoke(self.callback, **ctx.params)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 535, in invoke\n    return callback(*args, **kwargs)\n  File \"main.py\", line 16, in train\n    pipeline_manager.train(pipeline_name, dev_mode)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 21, in train\n    train(pipeline_name, dev_mode)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 85, in train\n    pipeline.fit_transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 280, in fit_transform\n    step_output_data = self._cached_fit_transform(step_inputs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 390, in _cached_fit_transform\n    step_output_data = self.transformer.fit_transform(**step_inputs)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/steppy/base.py\", line 605, in fit_transform\n    self.fit(*args, **kwargs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/models.py\", line 32, in fit\n    metrics = self._fit_loop(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/models.py\", line 63, in _fit_loop\n    batch_loss = loss_function(outputs_batch, target) * weight\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/torch/nn/modules/module.py\", line 357, in __call__\n    result = self.forward(*input, **kwargs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/parallel.py\", line 137, in forward\n    outputs = _criterion_parallel_apply(replicas, inputs, targets, kwargs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/parallel.py\", line 192, in _criterion_parallel_apply\n    raise output\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/parallel.py\", line 167, in _worker\n    output = module(*(input + target), **kwargs)\nTypeError: can only concatenate tuple (not \"dict\") to tuple\n</code></pre>\n\n<p>It looks like the \"target\" variable used for the loss function is supposed to be a tuple, but instead it is a dictionary. I have to admit I'm not sure what exactly's causing this, but I wanted to see if you have any immediate ideas before I spend time going through the code line by line. Execution command is just: <code>python main.py -- train --pipeline_name retinanet</code>, and the whole config has been filled out with (supposedly) the correct files.</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 370976,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/15/2018 18:52:11",
          "content": "<p>Hi there <a href=\"/anokas\">@anokas</a>.</p>\n\n<p>I think the problem is we have worked with, and tested it on multigpu. The loss is also parallelized (extension of pytorch data parallelization so that it better utilized many cards). </p>\n\n<p>So there are 2 options going forward.</p>\n\n<ul>\n<li><p>If you have multigpu change batch_size_train and batch_size_inference to something larger than 1. For example I am training some batches of classes right now with <code>batch_size_train: 8</code> and <code>batch_size_inference: 8</code> on 4 gpus <a href=\"https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce\">https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce</a> . If you do so, remember to change the batch_size_inference to 1 when running evaluate_predict pipeline. For some unknown reason it can crash the memory if you don't do that.</p></li>\n<li><p>Second option if you are training on just 1 gpu is to substitute all DataParallel stuff in the <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py\">https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py</a> file. To be more specific those are <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py#L19\">model</a> and <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py#L105\">loss</a></p></li>\n</ul>\n\n<p>I will add those as issues and try and fix it as soon as I can (likely tomorrow).\nJust out of curiousity what is your hardware setup?</p>\n\n<p>Best,\nJakub</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370984,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/15/2018 18:57:02",
          "content": "<p>Hi Jakub,</p>\n\n<p>I am running on 4x 1080Ti, the issue was that both batch_size_train and batch_size_inference were set to 1 (the default in the repo config file). I fixed it now, thank you :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370985,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/15/2018 19:01:26",
          "content": "<p>My apologies for that.</p>\n\n<p>I hope it will run smoothly from now on. I am curious to hear your thoughts on this project both during and after the competition. So feel free to drop a comment whenever you feel like it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371009,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/15/2018 19:46:14",
          "content": "<p>Sure, will do. I ran into one other issue when calling evaluate or predict:</p>\n\n<pre><code>2018-08-15 20:28:49 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n\nTraceback (most recent call last):\n  File \"main.py\", line 78, in &lt;module&gt;\n    main()\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 722, in __call__\n    return self.main(*args, **kwargs)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 697, in main\n    rv = self.invoke(ctx)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 1066, in invoke\n    return _process_result(sub_ctx.command.invoke(sub_ctx))\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 895, in invoke\n    return ctx.invoke(self.callback, **ctx.params)\n  File \"/home/m09170/anaconda3/lib/python3.6/site-packages/click/core.py\", line 535, in invoke\n    return callback(*args, **kwargs)\n  File \"main.py\", line 25, in evaluate\n    pipeline_manager.evaluate(pipeline_name, dev_mode, chunk_size)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 24, in evaluate\n    evaluate(pipeline_name, dev_mode, chunk_size)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 117, in evaluate\n    prediction = generate_prediction(valid_img_ids, pipeline, chunk_size)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 182, in generate_prediction\n    return _generate_prediction_in_chunks(img_ids, pipeline, chunk_size)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/pipeline_manager.py\", line 214, in _generate_prediction_in_chunks\n    output = pipeline.transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 315, in transform\n    step_inputs[input_step.name] = input_step.transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 315, in transform\n    step_inputs[input_step.name] = input_step.transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 315, in transform\n    step_inputs[input_step.name] = input_step.transform(data)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 321, in transform\n    step_output_data = self._cached_transform(step_inputs)\n  File \"/media/nvme1/kaggle-openimages/src/open-solution-googleai-object-detection/src/steppy_dev/base.py\", line 424, in _cached_transform\n    raise ValueError('No transformer cached {}'.format(self.name))\nValueError: No transformer cached retinanet\n</code></pre>\n\n<p>It appears to be looking for <code>experiments/transformers/retinanet</code> - I only have the file <code>experiments/transformers/label_encoder</code> - I have only run the model for one epoch but just wanted to test evaluation/prediction as a sanity check (I can see it saved the weight checkpoint too). Is this a file generated when the script completely finishes? Or is there some way to evaluate/predict an intermediate model?</p>\n\n<p>Thanks again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371012,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/15/2018 20:04:45",
          "content": "<p>Update: I copied the <code>best.torch</code> weights file to experiments/transformers/retinanet, and it appears to work!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371207,
          "author_name": "muhammedazamkhan",
          "author_url": "",
          "post_date": "08/16/2018 07:48:23",
          "content": "<p>Hi <a href=\"/anokas\">@anokas</a> @Jakub,\nWhen I try to predict (python main.py -- predict --pipeline_name retinanet) for a sanity check, I am getting the following error:\n<strong>\"RuntimeError: $ Torch: not enough memory: you tried to allocate 0GB. Buy new RAM! at /pytorch/torch/lib/TH/THGeneral.c:253\".</strong>\nI am using 3 x TITAN Xp and I think I have enough memory (12GB for each GPU). I wonder if you can help me to overcome this issue. Thanks in advance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371221,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/16/2018 08:57:16",
          "content": "<p>Probably you need to predict in chunks by adding <code>.. --chunk_size 100</code> at the end of your command. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371238,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/16/2018 10:24:49",
          "content": "<p><a href=\"/anokas\">@anokas</a> Steppy saves the transformer to /transformers/NAME after said transformer has been fitted. In case training was stopped mid way you have to manually copy the checkpoint as you have. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373506,
          "author_name": "hellojonas",
          "author_url": "",
          "post_date": "08/21/2018 13:11:19",
          "content": "<p>Hi <a href=\"/muhammedazamkhan\">@muhammedazamkhan</a> <a href=\"/jakubczakon\">@jakubczakon</a>. I ran the predict command, but got back an empty submission file. I feel like my model didn't train long enough, especially given that I trained on all classes. Did you run your model on all 500 classes <a href=\"/muhammedazamkhan\">@muhammedazamkhan</a>? And also for 1000 epochs like the Neptune Team?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374601,
          "author_name": "muhammedazamkhan",
          "author_url": "",
          "post_date": "08/23/2018 11:20:59",
          "content": "<p>Hi @Jonas \nI've just checked the implementation with default 10 classes and have run for 5 epochs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 371008,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/15/2018 19:41:35",
      "content": "<p>Hey Jakub, </p>\n\n<p>Is there any additional setup for running the open solution in neptune-ml? I think I want to test it out in neptune. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 371011,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/15/2018 19:58:40",
      "content": "<p>@Jakub what is the configs to set up in neptune for this comp? </p>",
      "votes": null,
      "replies": [
        {
          "id": 371283,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/16/2018 12:56:25",
          "content": "<p>Hi @William Green ,</p>\n\n<p>All you need to do is to change data paths in the very same <code>neptune.yaml</code> you are using.</p>\n\n<p>I will update the readme.md today with all the instructions needed to run it on multigpu with <code>neptune send</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 371317,
      "author_name": "jakubczakon",
      "author_url": "",
      "post_date": "08/16/2018 14:31:16",
      "content": "<p><strong>Update</strong>\nI updated the master branch of our <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection\">repo</a></p>\n\n<ol>\n<li><p>Added submission merge notebook for those of you who want to train batches of classes and join the predictions </p></li>\n<li><p>Updated the <code>neptune.yaml</code> config file with filepaths for those of you that may want to run in the cloud.\nIn that case pretty much all you need to do is run</p>\n\n<pre><code>neptune send --worker m-4p100 \\\n--environment pytorch-0.3.1-gpu-py3 \\\n--config configs/neptune.yaml \\\nmain.py train --pipeline_name retinanet\n</code></pre>\n\n<p>Yup there are multigpu workers on\nneptune now.</p></li>\n<li><p>Dropped some redundant steppy bits since the steppy 0.1.6 already has all we need here.</p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 371510,
      "author_name": "ggopalan",
      "author_url": "",
      "post_date": "08/17/2018 01:01:12",
      "content": "<p>Hi There,\n   I'm using neptune.ml for my experiments using your base code. Just getting my feet wet at the moment.\n  I did a train run based on defaults but with --dev_mode flag on\n  However running into trouble when trying to run the eval part. The suggestion below (from the documentation):</p>\n\n<p>\"With cloud environment you need to change the experiment directory to the one that you have just trained. Let's assume that your experiment id was GAI-14. You should go to neptune.yaml and change:</p>\n\n<p>experiment_dir:  ../GAI-14/output/experiment\n\"\nis not working for me. Please note that I am indeed changing the \"GAI-14\" to the name of my experimental (training) run. But ../my_expt_id/.. does not seem to exist.\nPlease help!</p>",
      "votes": null,
      "replies": [
        {
          "id": 371625,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/17/2018 08:54:58",
          "content": "<p>Hi @Giri Gopalan. </p>\n\n<p>I forgot to add a snippet that copies the experiment from your protected (read only) <code>/input</code> directory to the <code>/output</code> directory before running anything. Also when running neptune in the cloud you need to specify <code>--input my_dir</code> directories if you want to use something that is outside of your experiment. It is all nicely explained in the docs <a href=\"https://docs.neptune.ml/advanced-topics/storage/\">https://docs.neptune.ml/advanced-topics/storage/</a> .</p>\n\n<p>Anyways sorry for the trouble.\nBoth the code and instructions are updated so please get the newest master and follow the instructions.\nIn case of any trouble let me know. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371936,
          "author_name": "ggopalan",
          "author_url": "",
          "post_date": "08/17/2018 21:28:46",
          "content": "<p>Hi Jakub,</p>\n\n<p>Thanks. This brings me one step closer, but I am encountering the error below when I follow new instructions. Not too sure what is wrong.</p>\n\n<p>Unexpected end of /proc/mounts line `overlay / overlay rw,relatime,lowerdir=/var/lib/docker/overlay2/l/PKZCZTI7LLZ3N7KN675YAWDO2A:/var/lib/docker/overlay2/l/BZKXUKDO5K62J4XGTM2YWPDFWB:/var/lib/docker/overlay2/l/TWQXPQNTQQCLJMMPWR5GK7XWEV:/var/lib/docker/overlay2/l/4SULQLH2FKKFFP5YH6ST4JUC7W:/var/lib/docker/overlay2/l/QT7PKBTZCF652T2A5ZSXX56EHN:/var/lib/docker/overlay2/l/YEFVDSP2J4P6AKUSB7IPZCHYBA:/var/lib/docker/overlay2/l/LRAWBPHZID7W7RKBL43V3FJSZX:/var/lib/docker/overlay2/l/DG3MICQYVNJOX6YXDPHTIQU64W:/var/lib/docker/overlay2/l/MNIQIBWYDKSM5'</p>\n\n<p>Unexpected end of /proc/mounts line `J3EWKYNO536PV:/var/lib/docker/overlay2/l/UGWX4WBW3IQHMO5R5BDV27FMQE:/var/lib/docker/overlay2/l/NWR22WBG66ZRI7EAVALSIQZUDW:/var/lib/docker/overlay2/l/WF564LYDXGQWBGDQC3VPAO74KP:/var/lib/docker/overlay2/l/KWSW4LTKEPLIQX5V5NDA6HATPF:/var/lib/docker/overlay2/l/5D6ZQBN2PVELJCG5L4B34NVXEV:/var/lib/docker/overlay2/l/OZR3TE5HSEHSK6X2DM62OINVH5:/var/lib/docker/overlay2/l/I4NS2OYZDHG2YHUYVGVNVKN7KC:/var/lib/docker/overlay2/l/5UBGWBX2QJ7L4I45XICUS7W5WZ:/var/lib/docker/overlay2/l/X7YX7LUJRJ7NOVYUMVYJ7ER5WG:/var/lib/do'</p>\n\n<p>Unexpected end of /proc/mounts line `cker/overlay2/l/5VUUIAN3JGDVEFUKTLBV3BUJ64:/var/lib/docker/overlay2/l/IHUECGATV342F57ZGCTVIGG7CN:/var/lib/docker/overlay2/l/BTK3WTDIK4C6AUJEKR7I557WHA:/var/lib/docker/overlay2/l/IO5CIAELD3QELFU4N7MCI7FOT3:/var/lib/docker/overlay2/l/YGDBE4IOHQXR5ECOI4ZCBHAIIT:/var/lib/docker/overlay2/l/XGGPCHWJVDANU2B4ISJZ7JX2KI:/var/lib/docker/overlay2/l/VJROJ5LPTOXKSVYKJK463A3SZB:/var/lib/docker/overlay2/l/ZIWL5PNUTTEWS23HPG44QDSMNN,upperdir=/var/lib/docker/overlay2/7e427cf00268b04985194be27ee928c391b82a3eba75b9b18da9e9b0'</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372057,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/18/2018 06:47:21",
          "content": "<p>It looks like some random system trouble.</p>\n\n<p>Can you try again?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372551,
          "author_name": "ggopalan",
          "author_url": "",
          "post_date": "08/19/2018 17:11:18",
          "content": "<p>Tried 3 different times over the weekend.  All 3 failed.</p>\n\n<p>My local (on my machine) works fine, but I would love to get more experiments running on different platforms. This is a beast of a dataset. So would be nice if I can run on neptune too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374073,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/22/2018 11:52:33",
          "content": "<p>Hi @Giri Gopalan.</p>\n\n<p>I am running my evaluate_predict experiment right now and it seems to be working just fine.\nSo in my case I:</p>\n\n<ul>\n<li><p>trained the model with this <a href=\"https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca\">neptune\nexperiment GAI-473</a> running the following command</p>\n\n<pre><code> neptune send --worker m-4p100 \\\n --environment pytorch-0.3.1-gpu-py3 \\\n --config configs/neptune.yaml \\\n main.py train --pipeline_name retinanet\n</code></pre></li>\n</ul>\n\n<p>and my <code>neptune.yaml</code> looked like this</p>\n\n<pre><code>parameters:\n# Data Paths\n  train_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/train\n  test_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/test_challenge_2018\n  annotations_filepath: /public/challenges/google-ai-open-images-object-detection-track/annotations/challenge-2018-train-annotations-bbox.csv\n  annotations_human_labels_filepath: /public/challenges/google-ai-open-images-object-detection-track/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n  bbox_hierarchy_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/bbox_labels_500_hierarchy.json\n  class_mappings_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/challenge-2018-class-descriptions-500.csv\n  valid_ids_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/challenge-2018-image-ids-valset-od.csv\n  sample_submission: /public/challenges/google-ai-open-images-object-detection-track/sample_submission.csv\n  experiment_dir:  /output/experiment\n  clone_experiment_dir_from: ''\n</code></pre>\n\n<p>Notice that the <code>clone_experiment_dir_from: ''</code> during training. </p>\n\n<ul>\n<li><p>To evaluate I changed the <code>neptune.yaml</code> to</p>\n\n<pre><code>clone_experiment_dir_from: /input/GAI-473/output/experiment \n</code></pre>\n\n<p>and ran the following command</p>\n\n<pre><code> neptune send --worker m-p100 \\\n--environment pytorch-0.3.1-gpu-py3 \\\n--config configs/neptune.yaml \\\n--input /GAI-473 \\\nmain.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n</code></pre></li>\n</ul>\n\n<p>Notice that during prediction I am using just one p100 machine with the command <code>m-p100</code> but it is of course not necessary to change that (just cheaper).</p>\n\n<p>Have you done it exactly this way as well ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374103,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/22/2018 12:51:21",
          "content": "<p>Can you try running the <code>evaluate</code> pipeline and not 'evaluate_predict' pipeline first </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374106,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/22/2018 12:59:06",
          "content": "<p>@William Green please check the newest master</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374159,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/22/2018 14:07:26",
          "content": "<p>@Jakub Thank you. I'm checking it out right now. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374672,
          "author_name": "ggopalan",
          "author_url": "",
          "post_date": "08/23/2018 14:43:53",
          "content": "<p>OK, I found out what was causing my crashes. In the evaluate runs that crashed, I did NOT use --chunk-size. When I set it per your response above, it went through properly.</p>\n\n<p>I think things should not crash if an option is not set correctly.  Some sort of error message would have been easier for me to debug.</p>\n\n<p>In any case, thanks much for the help and all of the work you guys have done.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 371801,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/17/2018 16:20:11",
      "content": "<p>I recieve this error when running in the cloud: </p>\n\n<pre><code>/usr/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary \nincompatibility. Expected 96, got 88\nreturn f(*args, **kwds)\n/usr/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary \nincompatibility. Expected 96, got 88\n return f(*args, **kwds)\n 2018-08-17 16:16:58,612 google-ai-odt WARNING  pipeline_manager.py:58 - train() Validation sample-size is smaller \n then desired validation sample size ... clipping\n 2018-08-17 16:16:58,612 google-ai-odt WARNING  pipeline_manager.py:58 - train() Validation sample-size is smaller \n then desired validation sample size ... clipping\n Traceback (most recent call last):\n File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in &lt;module&gt;\nexecute()\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 103, in execute\nexecfile(job_filepath, job_globals)\nFile \"/usr/local/lib/python3.6/dist-packages/past/builtins/misc.py\", line 82, in execfile\nexec_(code, myglobals, mylocals)\nFile \"main.py\", line 78, in &lt;module&gt;\nmain()\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 16, in train\npipeline_manager.train(pipeline_name, dev_mode)\nFile \"/neptune/src/pipeline_manager.py\", line 21, in train\ntrain(pipeline_name, dev_mode)\nFile \"/neptune/src/pipeline_manager.py\", line 60, in train\na_max=valid_ids_data.shape[0])\nTypeError: clip() missing 1 required positional argument: 'a_min'\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 371837,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/17/2018 17:06:28",
          "content": "<p>Hi @William Green.</p>\n\n<p>Are you sure you are running the newest master?</p>\n\n<p>I ran it this morning in the cloud and it worked just fine.\nThis is a link to the experiment <a href=\"https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca\">https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371887,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/17/2018 19:11:58",
          "content": "<p>Hey @Jakub,</p>\n\n<p>I guess I didn't have the newest master. I have it working now. Thank you. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 372402,
      "author_name": "samratp",
      "author_url": "",
      "post_date": "08/19/2018 09:05:58",
      "content": "<p><a href=\"/kkaczmarek\">@kkaczmarek</a> - Thanks for your starter solution.. I have 2 questions here..</p>\n\n<ol>\n<li>What exactly is this metadata_filepath: /mnt/ml-team/minerva/open-solutions/googleai-object-detection/files/metadata.csv. I cannot find it on the Open Images download page.</li>\n<li>For train, do I need to download all the files from 00 to 08? I'm just trying to get some hands-on on this problem. So just the train_00.zip would be enough?</li>\n</ol>\n\n<p>TIA..</p>",
      "votes": null,
      "replies": [
        {
          "id": 372791,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/20/2018 10:42:20",
          "content": "<p>Hi there @Samrat P .</p>\n\n<p>Are you talking about the latest master?\nBy accident it was left there in the <code>neptune.yaml</code> before.</p>\n\n<p>ad 1.\nBut we have a version of code that uses <code>metadata.csv</code> where we calculate aspect ratio for all images and later batch the train so that batches have give or take the same aspect ratio.</p>\n\n<p>ad. 2\nIf you want to train locally then yes you need to download it all (we downloaded in one go not in chunks).\nBut if you wanna run it in the cloud we have uploaded it all to neptune.\nSo if you simply follow the instructions in the repo you can start getting your hands dirty in no time!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 372525,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/19/2018 15:35:01",
      "content": "<p>I get this error when executing step 4. Evaluate/Predict RetinaNet:</p>\n\n<pre><code>Traceback (most recent call last):\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in &lt;module&gt;\nexecute()\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 103, in execute\nexecfile(job_filepath, job_globals)\nFile \"/usr/local/lib/python3.6/dist-packages/past/builtins/misc.py\", line 82, in execfile\nexec_(code, myglobals, mylocals)\n File \"main.py\", line 78, in &lt;module&gt;\nmain()\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 67, in evaluate_predict\npipeline_manager.evaluate(pipeline_name, dev_mode, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 24, in evaluate\nevaluate(pipeline_name, dev_mode, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 117, in evaluate\nprediction = generate_prediction(valid_img_ids, pipeline, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 182, in generate_prediction\nreturn _generate_prediction_in_chunks(img_ids, pipeline, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 214, in _generate_prediction_in_chunks\noutput = pipeline.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 370, in transform\nstep_output_data = self._cached_transform(step_inputs)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 477, in _cached_transform\nraise ValueError('No transformer cached {}'.format(self.name))\nValueError: No transformer cached label_encoder\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 372789,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/20/2018 10:38:26",
          "content": "<p>It looks as if you didn't have your model/transformers saved.</p>\n\n<p>Check whether you specified your train folder correctly:</p>\n\n<pre><code> experiment_dir:  /output/experiment\n clone_experiment_dir_from:  /input/GAI-14/output/experiment\n</code></pre>\n\n<p>And pointed to it in the cli :</p>\n\n<pre><code> neptune send --worker m-4p100 \\\n --environment pytorch-0.3.1-gpu-py3 \\\n --config configs/neptune.yaml \\\n --input /GAI-14 \\\n main.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n</code></pre>\n\n<p>So in the example GAI-14 experiment above your <code>label_encoder</code> should be in the  <code>GAI-14/output/experiment/transformers/label_encoder</code> . Can you confirm that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372838,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/20/2018 13:02:09",
          "content": "<p>@Jakub Thank you. I check on it later today. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372959,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/20/2018 18:48:10",
          "content": "<p>@Jakub, </p>\n\n<p>Here is what I inserted: </p>\n\n<pre><code>neptune send --worker m-4p100 \\\n--environment pytorch-0.3.1-gpu-py3 \\\n--config configs/neptune.yaml \\\n--input /KAG-18 \\\nmain.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n</code></pre>\n\n<p>I still get the same error message: </p>\n\n<pre><code>  experiment_dir:  KAG-18/output/experiment/transformers/label_encoder\n  clone_experiment_dir_from:  /input/KAG-18/output/experiment\n</code></pre>\n\n<p>I also tried : </p>\n\n<pre><code>experiment_dir:  /output/experiment\nclone_experiment_dir_from:  /input/KAG-18/output/experiment\n</code></pre>\n\n<p>Error:</p>\n\n<pre><code>Traceback (most recent call last):\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in \n&lt;module&gt;\nexecute()\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 103, in execute\nexecfile(job_filepath, job_globals)\nFile \"/usr/local/lib/python3.6/dist-packages/past/builtins/misc.py\", line 82, in execfile\nexec_(code, myglobals, mylocals)\nFile \"main.py\", line 78, in &lt;module&gt;\nmain()\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 67, in evaluate_predict\npipeline_manager.evaluate(pipeline_name, dev_mode, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 24, in evaluate\nevaluate(pipeline_name, dev_mode, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 117, in evaluate\nprediction = generate_prediction(valid_img_ids, pipeline, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 182, in generate_prediction\nreturn _generate_prediction_in_chunks(img_ids, pipeline, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 214, in _generate_prediction_in_chunks\noutput = pipeline.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 364, in transform\nstep_inputs[input_step.name] = input_step.transform(data)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 370, in transform\nstep_output_data = self._cached_transform(step_inputs)\nFile \"/usr/local/lib/python3.6/dist-packages/steppy/base.py\", line 477, in _cached_transform\nraise ValueError('No transformer cached {}'.format(self.name))\nValueError: No transformer cached label_encoder\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373279,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/21/2018 05:31:27",
          "content": "<p>@Jakub,\nDisregard, I was able to get the evaluation to run. I had to clone the github since I did not do after the update. So far everything is running okay. I will know for sure once the prediction is complete. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 372792,
      "author_name": "jakubczakon",
      "author_url": "",
      "post_date": "08/20/2018 10:48:36",
      "content": "<p><strong>Update</strong></p>\n\n<p>Hi all. </p>\n\n<p>We continue our work on retinanet. \nWe are training everything in batches of classes making sure that the classes that fall into one batch are give or take of similar prevalence in the dataset.\nSo far our solution contains as we call them batch_{1-8} . Those classes do not contain human related classes (we should add them todayish).</p>\n\n<p>Surprisingly (to me) augmentation did help. If you look at some of those classes there are very few examples so it isn't that surprising I guess it's just with the dataset of 1.7M images you kinda feel like those problems are not really important anymore.</p>\n\n<p>We are also working on weighted classification loss in retinanet where we weigh the loss per class based on the train distribution (and clip it to something like 10).</p>\n\n<p>Let's see what happens with that!</p>",
      "votes": null,
      "replies": [
        {
          "id": 372850,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/20/2018 13:19:01",
          "content": "<p>@Jakub,  What augmentation techniques did you use?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372891,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/20/2018 15:00:56",
          "content": "<ul>\n<li>very small rotation -5:5 degrees</li>\n<li>small scaling 0.8 : 1.2</li>\n<li>left/right flip</li>\n<li>color augmentations (bluring and stuff)</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 372938,
      "author_name": "jakubczakon",
      "author_url": "",
      "post_date": "08/20/2018 17:36:42",
      "content": "<p><strong>Update</strong></p>\n\n<p>Added class batches with human labels.\nSo training on all classes we got to CV 515 LB 368 . Quite a large gap if u ask me. </p>\n\n<p>I wonder how much false negatives our batch &amp; merge approach generates.</p>",
      "votes": null,
      "replies": [
        {
          "id": 373326,
          "author_name": "differentialmind",
          "author_url": "",
          "post_date": "08/21/2018 07:08:15",
          "content": "<p>Interesting, which classes did you group as \"human\"? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373382,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/21/2018 09:12:57",
          "content": "<p>So we have 4 \"human\" batches:</p>\n\n<pre><code>[\"Footwear\", \"Human hair\",\"Person\", \"Human arm\", \"Human eye\", \"Human face\", \"Suit\", \"Human hand\", \"Human leg\", \"Human nose\", \"Dress\", \"Human mouth\"]\n\n[\"Jacket\", \"Sports uniform\", \"Human ear\", \"Shorts\", \"Helmet\", \"Hat\", \"Tie\", \"Swimwear\"]\n\n[\"Trousers\", \"Shirt\", \"Umbrella\", \"Coat\", \"Human beard\", \"Necklace\", \"Human head\"]\n\n[\"Scarf\", \"Human foot\", \"Luggage and bags\", \"Watch\", \"Brassiere\",\"Earrings\",\"Sock\",\"Skirt\",\"Glove\",\"Crown\",\"Swim cap\",\"Belt\",\"Tiara\"]\n</code></pre>\n\n<p>I guess they are more fashion/human after all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373476,
          "author_name": "differentialmind",
          "author_url": "",
          "post_date": "08/21/2018 12:19:07",
          "content": "<p>I like the idea of the approach, but is there a reason you decided to make the \"batch size\" so small? I looked at your other batches and it seems like they have about 100 classes each. Is it because of sampling reasons (meaning there are a lot more examples of human/fashion items in the train set)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373632,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/21/2018 17:13:03",
          "content": "<p>Hi Jakub,</p>\n\n<p>I am just wondering how long it takes you to train your whole solution on 4x P100? What about for just one of your class batches? Apologies if this info is in the neptune.ml experiments, I am a little confused by which run is which there. Also, do you train on the whole dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373644,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/21/2018 17:31:06",
          "content": "<p>For instance I trained this one \"batch_8\" <a href=\"https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce\">https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce</a> .</p>\n\n<p>I ran it local on 4 gtx 1070 though. Anyhow it was around 20 epochs/day  (120 total). Usually the batches would train for give or take 60 epochs and the smaller ones can actually be put on 2 gpus. </p>\n\n<p>I train on samples of 50000k per epoch for training and 10k for validation.\nThe samplers aren't a part of the public repo but they will be released right after the competition. \nOn the other hand we could release the code with the samplers but I am afraid it will make it to easy to get a medal with it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373647,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/21/2018 17:34:04",
          "content": "<p>So does the <code>training_sample_size</code> parameter in the config mean number of samples per epoch but it still trains on the full dataset? Or does it mean that the dataset is truncated to N samples and those samples are reused for each epoch.</p>\n\n<p>I also assume by sampler you are referring to the method of which images are drawn from the dataset for training?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373653,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/21/2018 17:45:53",
          "content": "<p>Sorry for the confusion <a href=\"/anokas\">@anokas</a>, my bad. I wasn't sure which version was on public github.</p>\n\n<p>Anyways, yes <code>training_sample_size</code> is the size of the sample that is drawn every epoch. Sampler refers to the pytorch sampler which is later passed to the pytorch loader in this <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/loaders.py\">file</a>.</p>\n\n<p>And we do sample from the entire dataset every epoch (just different samples).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373768,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/21/2018 21:47:27",
          "content": "<p>I'm a little curious. Would we call that function in the pipeline_config and config files? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373794,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/21/2018 23:32:33",
          "content": "<p>Thanks Jakub, I understand now.</p>\n\n<p>One thing I am curious about:\nThe model seems to spend a long time waiting between epochs. I understand the 6 minutes for validation, but does it take an additional 6 minutes to save the model, and then another 12 minutes to start the next epoch? Do you have any insights into what is going on in between these steps that might be making it slow?</p>\n\n<pre><code>2018-08-21 23:16:42 steppy &gt;&gt;&gt;; epoch 1 sum:     0.58112\n2018-08-21 23:22:44 steppy &gt;&gt;&gt;; epoch 1 validation sum:     0.53714\n2018-08-21 23:28:45 steppy &gt;&gt;&gt;; epoch 1 model persisted to experiment2/checkpoints/retinanet/best.torch\n2018-08-21 23:28:45 steppy &gt;&gt;&gt;; epoch 2 current lr: 1e-05\n2018-08-21 23:40:48 steppy &gt;&gt;&gt;; epoch 1 time 0:54:58\n2018-08-21 23:40:48 steppy &gt;&gt;&gt;; epoch 2 ...\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 373925,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/22/2018 05:39:05",
          "content": "<p>Thanks for spotting that <a href=\"/anokas\">@anokas</a> I haven't noticed that before.\nI checked the latest experiment and the same thing happened.</p>\n\n<p>My guess is that it could be that by mistake the loss on validation is calculated multiple times. Looking at the callback list:</p>\n\n<pre><code>return CallbackList(\n    callbacks=[experiment_timing, training_monitor, validation_monitor,\n               model_checkpoints, lr_scheduler, early_stopping, neptune_monitor,\n               ]) \n</code></pre>\n\n<p>after the validation_monitor the validation loss is needed in model_checkpoints, early_stopping and neptune_monitor. I wonder if turning some of those off speeds things up considerably.</p>\n\n<p>Unfortunately, I am not sure that I will have the time to fix that till Monday. So if you would like to fix/check that yourself. The vast majority of the logic is in the <code>src/callbacks.py</code> the rest sits in the <a href=\"https://github.com/neptune-ml/steppy-toolkit/blob/master/toolkit/pytorch_transformers/callbacks.py\">steppy-toolkit package</a>. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374110,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/22/2018 13:04:04",
          "content": "<p>Thanks for all your help so far. Is there any way I can resume training from a checkpoint, if the training script is stopped?</p>\n\n<p>I am unfamiliar with steppy so I am not sure how to change the code to do this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374143,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/22/2018 13:49:15",
          "content": "<p>Sure <a href=\"/anokas\">@anokas</a>. Glad I could be of help.</p>\n\n<p>If you go to <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py\">src/models.py</a> :</p>\n\n<p>You have something like this:</p>\n\n<pre><code>class ModelParallel(Model):\n    def fit(self, datagen, validation_datagen=None):\n        self._initialize_model_weights()\n\n        self.model = DataParallel(self.model)\n\n        if torch.cuda.is_available():\n            self.model = self.model.cuda() \n</code></pre>\n\n<p>simply add loading:</p>\n\n<pre><code>class ModelParallel(Model):\n    def fit(self, datagen, validation_datagen=None):\n        self._initialize_model_weights()\n\n        self.model = DataParallel(self.model)\n        self.load(YOUR_FILEPATH)\n\n        if torch.cuda.is_available():\n            self.model = self.model.cuda() \n</code></pre>\n\n<p>If you want to do some custom loading of say some backbone layers or something I would suggest that you go to the <code>load</code> method and check how it's done there and probably add a 'custom_load' method for your use case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 375068,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/24/2018 14:10:11",
          "content": "<p>@ Jakub, How many batches are there total?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 375070,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/24/2018 14:20:15",
          "content": "<p>We have 12 batches right now but the experiments with some of them joined together are in progress.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 375072,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/24/2018 14:26:33",
          "content": "<p>@Jakub, </p>\n\n<p>How do you apply the method for training from a checkpoint in the cloud? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 375091,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/24/2018 15:25:29",
          "content": "<p>Simply set data paths as you would with evaluate/predict and add --input YOUREXP-1 in the command and your checkpoint will be available in the /output/experiment/checkpoints/retinanet/best.torch .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 377230,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/28/2018 19:17:08",
          "content": "<p>@Jskub can I use the same cmd line for evaluation?</p>\n\n<p>clone:   /output/experiment/checkpoints/retinanet/best.torch</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 373324,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/21/2018 07:07:43",
      "content": "<p>I was able to get both the train and predict to run. However, evaluation_prediction file was empty. </p>\n\n<pre><code>Traceback (most recent call last):\n File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in &lt;module&gt;\nexecute()\nFile \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 103, in execute\nexecfile(job_filepath, job_globals)\nFile \"/usr/local/lib/python3.6/dist-packages/past/builtins/misc.py\", line 82, in execfile\nexec_(code, myglobals, mylocals)\nFile \"main.py\", line 78, in &lt;module&gt;\nmain()\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 722, in __call__\nreturn self.main(*args, **kwargs)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 697, in main\nrv = self.invoke(ctx)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 1066, in invoke\nreturn _process_result(sub_ctx.command.invoke(sub_ctx))\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 895, in invoke\nreturn ctx.invoke(self.callback, **ctx.params)\nFile \"/usr/local/lib/python3.6/dist-packages/click/core.py\", line 535, in invoke\nreturn callback(*args, **kwargs)\nFile \"main.py\", line 68, in evaluate_predict\npipeline_manager.predict(pipeline_name, dev_mode, submit_predictions, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 27, in predict\npredict(pipeline_name, dev_mode, submit_predictions, chunk_size)\nFile \"/neptune/src/pipeline_manager.py\", line 161, in predict\nshutil.copytree(PARAMS.clone_experiment_dir_from, PARAMS.experiment_dir)\nFile \"/usr/lib/python3.6/shutil.py\", line 315, in copytree\nos.makedirs(dst)\nFile \"/usr/lib/python3.6/os.py\", line 220, in makedirs\nmkdir(name, mode)\nFileExistsError: [Errno 17] File exists: '/output/experiment'\n</code></pre>\n\n<p>Code was hang up here:</p>\n\n<pre><code> &gt;&gt;&gt; Formatting prediction\n &gt;&gt;&gt; Background predicted for all the images. Metric cannot be calculated\n &gt;&gt;&gt; predicting\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 373725,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/21/2018 20:02:30",
      "content": "<p>I keep getting this error when evaluating:</p>\n\n<pre><code>FileExistsError: [Errno 17] File exists: '/output/experiment'\n</code></pre>\n\n<p>Any idea how to fix?</p>",
      "votes": null,
      "replies": [
        {
          "id": 373971,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/22/2018 07:38:58",
          "content": "<p>Can you paste your data paths? Does it fail after evaluation ( map value has been calculated) or straight away? What os the exact command you are using?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374080,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/22/2018 12:05:51",
          "content": "<pre><code># Data Paths\ntrain_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/train\ntest_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/test_challenge_2018\nannotations_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/annotations/challenge-2018-train-annotations-bbox.csv\nannotations_human_labels_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/annotations/challenge-2018-train-annotations-human-imagelabels.csv\nbbox_hierarchy_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/metadata/bbox_labels_500_hierarchy.json\nclass_mappings_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/metadata/challenge-2018-class-descriptions-500.csv\nvalid_ids_filepath: /public/challenges/google-ai-open-images-object-detection- \ntrack/metadata/challenge-2018-image-ids-valset-od.csv\nsample_submission: /public/challenges/google-ai-open-images-object-detection- \ntrack/sample_submission.csv\nexperiment_dir:  /output/experiment\nclone_experiment_dir_from:  /input/GOOG-2/output/experiment\n</code></pre>\n\n<p>I am trying this setup per your recommendation:</p>\n\n<pre><code>experiment_dir:  /output/experiment\nclone_experiment_dir_from:  GOOG-2/output/experiment/transformers/label_encoder\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374089,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/22/2018 12:23:36",
          "content": "<p>@William Greene I think it should be:</p>\n\n<pre><code>clone_experiment_dir_from:  GOOG-2/output/experiment\n</code></pre>\n\n<p>not </p>\n\n<pre><code>clone_experiment_dir_from:  GOOG-2/output/experiment/transformers/label_encoder\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374101,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/22/2018 12:44:40",
          "content": "<p>Ok I think I found the culprit should have fix in no time.</p>\n\n<p>If you don't want to wait for it you can simply run 'evaluate' pipeline and 'predict' pipeline separately in two consecutive experiments and it will work just fine. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 374682,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/23/2018 15:08:18",
      "content": "<p>What was the baseline score for solution 1? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 376754,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/28/2018 02:57:00",
      "content": "<p>Missing experiment folder.  For some reason, the experiment folder was not created after training was complete. </p>",
      "votes": null,
      "replies": [
        {
          "id": 376857,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/28/2018 07:13:47",
          "content": "<p>Local or cloud?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 376948,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/28/2018 11:12:43",
          "content": "<p>Cloud.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 376995,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/28/2018 12:48:21",
          "content": "<p>So output/experiment is empty or non-existent? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 377064,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/28/2018 14:36:51",
          "content": "<p>It's non-existent. I think I may have figured out why.</p>\n\n<p>I had <code>output/experiment/</code> versus <code>/output/experiment</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 377251,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/28/2018 20:29:24",
      "content": "<p>Is this the correct command to train from last checkpoint</p>\n\n<pre><code>neptune send --worker m-4p100 \\\n--environment pytorch-0.3.1-gpu-py3 \\\n--config configs/neptune.yaml \\\n--input /GOOG-50 \\\nmain.py train --pipeline_name retinanet\n</code></pre>\n\n<p>config: </p>\n\n<pre><code>      experiment_dir:  /output/experiment/checkpoints/retinanet/best.torch\n      clone_experiment_dir_from: ''\n</code></pre>\n\n<p>I get the following error when I either try to run from last know checkpoint or evaluate:</p>\n\n<pre><code>00:00   /usr/bin/python: can't open file '/usr/local/lib/python3.6/dist-packages/neptune/job_wrapper.py': [Errno 2] No \n such file or directory\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 377291,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/28/2018 22:12:29",
          "content": "<p>config should be:</p>\n\n<pre><code>  experiment_dir:  /output/experiment/checkpoints/retinanet/best.torch\n  clone_experiment_dir_from: /input/GOOG-50/output/experiment\n</code></pre>\n\n<p>and you should insert the following line:</p>\n\n<pre><code>    self.load('/input/GOOG-50/output/experiment/checkpoints/retinanet/best.torch')\n</code></pre>\n\n<p>In the <code>models.py</code> as explained to <a href=\"/anokas\">@anokas</a> below.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 377629,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "08/29/2018 13:33:44",
      "content": "<p>I'm still getting this error </p>\n\n<pre><code>0:00    /usr/bin/python: can't open file '/usr/local/lib/python3.6/dist-packages/neptune/job_wrapper.py': [Errno 2] \nNo such file or directory\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 377697,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "08/29/2018 16:00:12",
          "content": "<p>I am pretty sure that the problem was with the wrong neptune-cli version in the requirements.tx . </p>\n\n<p>It is fixed now so it should run with no problems. Just go to the repo and check.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 377746,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "08/29/2018 17:17:24",
          "content": "<p>@Jakub </p>\n\n<p>Thank you :) I finally was able to get it to run. It works :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "367792": "Hi,\n\nWe would like to start sharing our results in this competition.\n\n### The Open Solution approach\nIt means that we are going to open:\n\n1. the [code on GitHub](https://github.com/neptune-ml/open-solution-googleai-object-detection) - *(no worries competitive guys -&gt; we only publish code that scores below bronze medal)*\n2. our [experiments results](https://app.neptune.ml/neptune-ml/Google-AI-Object-Detection-Challenge)\n3. our approach, that is what we have tried, what worked well, etc.\n\n### Goals\nThese are pretty straightforward:\n\n1. Learning from the process.\n2. Encourage more Kagglers to start working on this competition.\n3. Share open source solution with no strings attached, so that less experienced Kagglers can join competition.\n\n### What can you find here?\nIn this thread we will discuss our approach to this solution, techniques used, network architectures and all other deep learning related stuff! We want this place to be good address for people who want to share knowledge or gain knowledge :)\n\nHappy Training!\n\nKamil &amp; Kuba",
    "367849": "Gotta say that our story in this comp is a constant battle with making retinanetwork work. \n\nFirst the performance and anchor setup, then working with 500 classes, batching by aspect ratios and dividing the problem into smaller subproblems.\n\nNow we are brewing some experiments (check on neptune.ml) and expecting improvements\nbut obviously still strugling. At this point the biggest debacle is, how to balance classes for the images with multiple bboxes. \n\nIf you have any ideas feel free to chime in!",
    "368079": "LB 0.21694 score falls in the silver range. In a discussion at TGS Salt ID competition, you promised [here][1] that you won't publish code that'll f**k up the leaderboard. Come on dude, it takes like a billion years to train on this dataset. Not cool. \n\n\n  [1]: https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/62162#365436",
    "368100": "Hi,\ndon't worry, the code from the public repository won't take you to the medal zone. As stated in the post:\n\n&gt; (no worries competitive guys -&gt; we only publish code that scores below bronze medal)",
    "368108": "Hi @yk1598,\n\n&gt; the code on GitHub - (no worries competitive guys -&gt; we only publish code that scores below bronze medal)\n\nCode in the repository will not give your silver medal. No worries, we know what we are publishing.\n\nIn the next post I will describe our work in more detail. Also, this starter gives you good starting point for your own work :)\n\nBest,\n\nKamil\n\nBTW: nice avatar!",
    "368110": "Ah I see! I got confused by the title. Apologies.",
    "368142": "## Competition update\nIt took us quite a lot of time to develop reasonable solution to this competition.\n\nWe have decided to work with **RetinaNet** and focal loss, described in this paper: [Focal Loss for Dense Object Detection](https://arxiv.org/abs/1708.02002). If you are new to RetinaNet - I recommend to skim through blog post that describes [The intuition behind RetinaNet](https://medium.com/@14prakash/the-intuition-behind-retinanet-eb636755607d).\n\n## Code\nWhat you can see on [master branch](https://github.com/neptune-ml/open-solution-googleai-object-detection) is training procedure on 10 classes -&gt; you can extend it to work on the entire dataset. We worked with this code to quickly iterate over various ideas.\n\n## Our analysis of the problem\nWe quickly decided that this competition consist of two subproblems, each to be approached separately:\n\n- *First subproblem* is classes related to *people* and *clothing*, because the bboxes overlap a lot and there are multiple bboxes per image. Here, we have approximately 80 classes.\n- *Second subproblem* is remaining classes. Here, we take all these classes and divide it into 7 bins. Each bin is occupied by classes with similar frequency in the dataset. We need such bins to prepare proper epoch as described below. \n\n## Preprocessing for training\n* When we run training for the remaining classes, we make sure that each class (within an epoch) has similar number of occurrences -&gt; we implemented sampler to do this work. Thanks to this we have more balanced problem. In practice we oversample rare classes and subsample frequent classes. 7 bins mentioned above are utilized here.\n* Next, we calculate **aspect ratio** and we prepare batches only for images with similar **aspect ratio**. We need this in the next step - resize. After resize all images are similarly squeezed - training signal is better balanced.\n* At this point we are ready to feed batch to the network. Images are with similar aspect ratio, classes within the epoch are balanced, so training signal is stronger.\n* Resulting experiment is like [this one](https://app.neptune.ml/-/dashboard/experiment/f945da64-6dd3-459b-94c5-58bc6a83f590).\n\n## Other remarks\n- It is good to start experimenting with few classes (like 10) and get better feel of the problem. We also run [training on 10 classes](https://app.neptune.ml/-/dashboard/experiment/c779468e-d3f7-44b8-a3a4-43a012315708).\n- We noticed, that for rare classes augmentations are necessary :)\n\n## Open Questions\n* What to do with highly overlapping bboxes (*people* and *clothing* subproblem).",
    "368312": "Kamil I think this may answer your questions [non_max_suppression over all classes][1]\n\n\n  [1]: https://stackoverflow.com/questions/49951422/get-rid-of-overlapping-bounding-boxes-across-different-classes-in-tensorflow-obj \"StackOverflow\"",
    "368397": "Hi @dskswu, thanks I'll check it :)",
    "368500": "I think you have a typo `python main.py train --pipeline_name retinanet` in your instructions.",
    "368909": "Hi, it should be `python main.py -- train --pipeline_name retinanet`\n\nthank you!",
    "369813": "Is this the correct data path parameters? \n\n       train_imgs_dir: train_03/\n       test_imgs_dir: test/\n       annotations_filepath: annotations/\n       annotations_human_labels_filepath: annotations_human_labels/\n       bbox_hierarchy_filepath: bbox_hierarchy/\n       valid_ids_filepath: valid_ids/\n       sample_submission: sample_submission.csv\n       experiment_dir:  experiment\n       class_mappings_filepath: class_mappings/\n\nI am trying to troubleshoot this error: \n\n    neptune: Executing in Offline Mode.\n    neptune: Executing in Offline Mode.\n    2018-08-13 20-04-33 google-ai-odt &gt;&gt;&gt; training\n    Traceback (most recent call last):\n     File \"main.py\", line 78, in",
    "370045": "Not quite. Mine looks something like this:\n\n    parameters:\n     # Data Paths\n\n     train_imgs_dir: .../open-images-v4/bounding-boxes/train\n\n     test_imgs_dir: .../open-images-v4/bounding-boxes/test_challenge_2018\n\n     annotations_filepath: .../googleai-object-detection/data/annotations/challenge-2018-train-annotations-bbox.csv\n\n     annotations_human_labels_filepath:.../googleai-object-detection/data/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n\n     bbox_hierarchy_filepath: .../googleai-object-detection/data/metadata/bbox_labels_500_hierarchy.json\n\n     class_mappings_filepath: .../googleai-object-detection/data/metadata/challenge-2018-class-descriptions-500.csv\n\n     valid_ids_filepath: .../googleai-object-detection/data/metadata/challenge-2018-image-ids-valset-od.csv\n     \n    sample_submission: .../googleai-object-detection/data/sample_submission.csv\n     \n    metadata_filepath: .../googleai-object-detection/files/metadata.csv\n    \n    experiment_dir:  .../googleai-object-detection/experiments/10_classes",
    "370325": "Jakub, Thank you. I totally missed this on the download page.",
    "370944": "I am running the latest master branch (offline), and when the code gets to the training point it crashes when trying to forward() the model and evaluate the loss function:\n\n    neptune: Executing in Offline Mode.\n    2018-08-15 18-34-40 google-ai-odt &gt;&gt;&gt; training\n    2018-08-15 18-35-03 google-ai-odt &gt;&gt;&gt; Training on a reduced class subset: ['Person', 'Car', 'Dress', 'Footwear']\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing Step label_encoder...\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing Step label_encoder...\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing experiment directories under experiments\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; initializing experiment directories under experiments\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; done: initializing experiment directories\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; done: initializing experiment directories\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; Step label_encoder initialized\n    2018-08-15 18:35:05 steppy &gt;&gt;&gt; Step label_encoder initialized\n\n    [skipped]\n\n    2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n    2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n    2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, fitting and transforming...\n    2018-08-15 18:35:10 steppy &gt;&gt;&gt; Step retinanet, fitting and transforming...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; starting training...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; starting training...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; initial lr: 1e-05\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; initial lr: 1e-05\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 ...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 ...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 batch 0 ...\n    2018-08-15 18:35:13 steppy &gt;&gt;&gt; epoch 0 batch 0 ...\n    Traceback (most recent call last):\n      File \"main.py\", line 78, in",
    "370976": "Hi there @anokas.\n\nI think the problem is we have worked with, and tested it on multigpu. The loss is also parallelized (extension of pytorch data parallelization so that it better utilized many cards). \n\nSo there are 2 options going forward.\n\n - If you have multigpu change batch_size_train and batch_size_inference to something larger than 1. For example I am training some batches of classes right now with `batch_size_train: 8` and `batch_size_inference: 8` on 4 gpus https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce . If you do so, remember to change the batch_size_inference to 1 when running evaluate_predict pipeline. For some unknown reason it can crash the memory if you don't do that.\n\n - Second option if you are training on just 1 gpu is to substitute all DataParallel stuff in the https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py file. To be more specific those are [model](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py#L19) and [loss](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py#L105)\n\nI will add those as issues and try and fix it as soon as I can (likely tomorrow).\nJust out of curiousity what is your hardware setup?\n\nBest,\nJakub",
    "370984": "Hi Jakub,\n\nI am running on 4x 1080Ti, the issue was that both batch_size_train and batch_size_inference were set to 1 (the default in the repo config file). I fixed it now, thank you :)",
    "370985": "My apologies for that.\n\nI hope it will run smoothly from now on. I am curious to hear your thoughts on this project both during and after the competition. So feel free to drop a comment whenever you feel like it.",
    "371008": "Hey Jakub, \n\nIs there any additional setup for running the open solution in neptune-ml? I think I want to test it out in neptune.",
    "371009": "Sure, will do. I ran into one other issue when calling evaluate or predict:\n\n    2018-08-15 20:28:49 steppy &gt;&gt;&gt; Step retinanet, unpacking inputs...\n    \n    Traceback (most recent call last):\n      File \"main.py\", line 78, in",
    "371011": "Jakub what is the configs to set up in neptune for this comp?",
    "371012": "Update: I copied the `best.torch` weights file to experiments/transformers/retinanet, and it appears to work!",
    "371207": "Hi @anokas @Jakub,\nWhen I try to predict (python main.py -- predict --pipeline_name retinanet) for a sanity check, I am getting the following error:\n**\"RuntimeError: $ Torch: not enough memory: you tried to allocate 0GB. Buy new RAM! at /pytorch/torch/lib/TH/THGeneral.c:253\".**\nI am using 3 x TITAN Xp and I think I have enough memory (12GB for each GPU). I wonder if you can help me to overcome this issue. Thanks in advance.",
    "371221": "Probably you need to predict in chunks by adding `.. --chunk_size 100` at the end of your command.",
    "371238": "anokas Steppy saves the transformer to /transformers/NAME after said transformer has been fitted. In case training was stopped mid way you have to manually copy the checkpoint as you have.",
    "371283": "Hi @William Green ,\n\nAll you need to do is to change data paths in the very same `neptune.yaml` you are using.\n\nI will update the readme.md today with all the instructions needed to run it on multigpu with `neptune send`",
    "371317": "**Update**\nI updated the master branch of our [repo](https://github.com/neptune-ml/open-solution-googleai-object-detection)\n\n1. Added submission merge notebook for those of you who want to train batches of classes and join the predictions \n\n1. Updated the `neptune.yaml` config file with filepaths for those of you that may want to run in the cloud.\n   In that case pretty much all you need to do is run\n\n        neptune send --worker m-4p100 \\\n        --environment pytorch-0.3.1-gpu-py3 \\\n        --config configs/neptune.yaml \\\n        main.py train --pipeline_name retinanet\n\n   Yup there are multigpu workers on\n   neptune now.\n\n1. Dropped some redundant steppy bits since the steppy 0.1.6 already has all we need here.",
    "371510": "Hi There,\n   I'm using neptune.ml for my experiments using your base code. Just getting my feet wet at the moment.\n  I did a train run based on defaults but with --dev_mode flag on\n  However running into trouble when trying to run the eval part. The suggestion below (from the documentation):\n\n\"With cloud environment you need to change the experiment directory to the one that you have just trained. Let's assume that your experiment id was GAI-14. You should go to neptune.yaml and change:\n\n  experiment_dir:  ../GAI-14/output/experiment\n\"\nis not working for me. Please note that I am indeed changing the \"GAI-14\" to the name of my experimental (training) run. But ../my_expt_id/.. does not seem to exist.\nPlease help!",
    "371625": "Hi @Giri Gopalan. \n\nI forgot to add a snippet that copies the experiment from your protected (read only) `/input` directory to the `/output` directory before running anything. Also when running neptune in the cloud you need to specify `--input my_dir` directories if you want to use something that is outside of your experiment. It is all nicely explained in the docs https://docs.neptune.ml/advanced-topics/storage/ .\n\nAnyways sorry for the trouble.\nBoth the code and instructions are updated so please get the newest master and follow the instructions.\nIn case of any trouble let me know.",
    "371801": "I recieve this error when running in the cloud: \n\n    /usr/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary \n    incompatibility. Expected 96, got 88\n    return f(*args, **kwds)\n    /usr/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary \n    incompatibility. Expected 96, got 88\n     return f(*args, **kwds)\n     2018-08-17 16:16:58,612 google-ai-odt WARNING  pipeline_manager.py:58 - train() Validation sample-size is smaller \n     then desired validation sample size ... clipping\n     2018-08-17 16:16:58,612 google-ai-odt WARNING  pipeline_manager.py:58 - train() Validation sample-size is smaller \n     then desired validation sample size ... clipping\n     Traceback (most recent call last):\n     File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in",
    "371837": "Hi @William Green.\n\nAre you sure you are running the newest master?\n\nI ran it this morning in the cloud and it worked just fine.\nThis is a link to the experiment https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca",
    "371887": "Hey @Jakub,\n\nI guess I didn't have the newest master. I have it working now. Thank you.",
    "371936": "Hi Jakub,\n\n  Thanks. This brings me one step closer, but I am encountering the error below when I follow new instructions. Not too sure what is wrong.\n\nUnexpected end of /proc/mounts line `overlay / overlay rw,relatime,lowerdir=/var/lib/docker/overlay2/l/PKZCZTI7LLZ3N7KN675YAWDO2A:/var/lib/docker/overlay2/l/BZKXUKDO5K62J4XGTM2YWPDFWB:/var/lib/docker/overlay2/l/TWQXPQNTQQCLJMMPWR5GK7XWEV:/var/lib/docker/overlay2/l/4SULQLH2FKKFFP5YH6ST4JUC7W:/var/lib/docker/overlay2/l/QT7PKBTZCF652T2A5ZSXX56EHN:/var/lib/docker/overlay2/l/YEFVDSP2J4P6AKUSB7IPZCHYBA:/var/lib/docker/overlay2/l/LRAWBPHZID7W7RKBL43V3FJSZX:/var/lib/docker/overlay2/l/DG3MICQYVNJOX6YXDPHTIQU64W:/var/lib/docker/overlay2/l/MNIQIBWYDKSM5'\n\nUnexpected end of /proc/mounts line `J3EWKYNO536PV:/var/lib/docker/overlay2/l/UGWX4WBW3IQHMO5R5BDV27FMQE:/var/lib/docker/overlay2/l/NWR22WBG66ZRI7EAVALSIQZUDW:/var/lib/docker/overlay2/l/WF564LYDXGQWBGDQC3VPAO74KP:/var/lib/docker/overlay2/l/KWSW4LTKEPLIQX5V5NDA6HATPF:/var/lib/docker/overlay2/l/5D6ZQBN2PVELJCG5L4B34NVXEV:/var/lib/docker/overlay2/l/OZR3TE5HSEHSK6X2DM62OINVH5:/var/lib/docker/overlay2/l/I4NS2OYZDHG2YHUYVGVNVKN7KC:/var/lib/docker/overlay2/l/5UBGWBX2QJ7L4I45XICUS7W5WZ:/var/lib/docker/overlay2/l/X7YX7LUJRJ7NOVYUMVYJ7ER5WG:/var/lib/do'\n\nUnexpected end of /proc/mounts line `cker/overlay2/l/5VUUIAN3JGDVEFUKTLBV3BUJ64:/var/lib/docker/overlay2/l/IHUECGATV342F57ZGCTVIGG7CN:/var/lib/docker/overlay2/l/BTK3WTDIK4C6AUJEKR7I557WHA:/var/lib/docker/overlay2/l/IO5CIAELD3QELFU4N7MCI7FOT3:/var/lib/docker/overlay2/l/YGDBE4IOHQXR5ECOI4ZCBHAIIT:/var/lib/docker/overlay2/l/XGGPCHWJVDANU2B4ISJZ7JX2KI:/var/lib/docker/overlay2/l/VJROJ5LPTOXKSVYKJK463A3SZB:/var/lib/docker/overlay2/l/ZIWL5PNUTTEWS23HPG44QDSMNN,upperdir=/var/lib/docker/overlay2/7e427cf00268b04985194be27ee928c391b82a3eba75b9b18da9e9b0'",
    "372057": "It looks like some random system trouble.\n\nCan you try again?",
    "372402": "kkaczmarek - Thanks for your starter solution.. I have 2 questions here..\n\n1. What exactly is this metadata_filepath: /mnt/ml-team/minerva/open-solutions/googleai-object-detection/files/metadata.csv. I cannot find it on the Open Images download page.\n2. For train, do I need to download all the files from 00 to 08? I'm just trying to get some hands-on on this problem. So just the train_00.zip would be enough?\n\nTIA..",
    "372525": "I get this error when executing step 4. Evaluate/Predict RetinaNet:\n\n    Traceback (most recent call last):\n    File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in",
    "372551": "Tried 3 different times over the weekend.  All 3 failed.\n\nMy local (on my machine) works fine, but I would love to get more experiments running on different platforms. This is a beast of a dataset. So would be nice if I can run on neptune too.",
    "372789": "It looks as if you didn't have your model/transformers saved.\n\nCheck whether you specified your train folder correctly:\n\n     experiment_dir:  /output/experiment\n     clone_experiment_dir_from:  /input/GAI-14/output/experiment\n\nAnd pointed to it in the cli :\n\n     neptune send --worker m-4p100 \\\n     --environment pytorch-0.3.1-gpu-py3 \\\n     --config configs/neptune.yaml \\\n     --input /GAI-14 \\\n     main.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n\nSo in the example GAI-14 experiment above your `label_encoder` should be in the  `GAI-14/output/experiment/transformers/label_encoder` . Can you confirm that?",
    "372791": "Hi there @Samrat P .\n\nAre you talking about the latest master?\nBy accident it was left there in the `neptune.yaml` before.\n\nad 1.\nBut we have a version of code that uses `metadata.csv` where we calculate aspect ratio for all images and later batch the train so that batches have give or take the same aspect ratio.\n\nad. 2\nIf you want to train locally then yes you need to download it all (we downloaded in one go not in chunks).\nBut if you wanna run it in the cloud we have uploaded it all to neptune.\nSo if you simply follow the instructions in the repo you can start getting your hands dirty in no time!",
    "372792": "**Update**\n\nHi all. \n\nWe continue our work on retinanet. \nWe are training everything in batches of classes making sure that the classes that fall into one batch are give or take of similar prevalence in the dataset.\nSo far our solution contains as we call them batch_{1-8} . Those classes do not contain human related classes (we should add them todayish).\n\nSurprisingly (to me) augmentation did help. If you look at some of those classes there are very few examples so it isn't that surprising I guess it's just with the dataset of 1.7M images you kinda feel like those problems are not really important anymore.\n\nWe are also working on weighted classification loss in retinanet where we weigh the loss per class based on the train distribution (and clip it to something like 10).\n\nLet's see what happens with that!",
    "372838": "Jakub Thank you. I check on it later today.",
    "372850": "Jakub,  What augmentation techniques did you use?",
    "372891": "very small rotation -5:5 degrees\n - small scaling 0.8 : 1.2\n - left/right flip\n - color augmentations (bluring and stuff)",
    "372938": "**Update**\n\nAdded class batches with human labels.\nSo training on all classes we got to CV 515 LB 368 . Quite a large gap if u ask me. \n\nI wonder how much false negatives our batch &amp; merge approach generates.",
    "372959": "Jakub, \n\nHere is what I inserted: \n\n    neptune send --worker m-4p100 \\\n    --environment pytorch-0.3.1-gpu-py3 \\\n    --config configs/neptune.yaml \\\n    --input /KAG-18 \\\n    main.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n\nI still get the same error message: \n\n      experiment_dir:  KAG-18/output/experiment/transformers/label_encoder\n      clone_experiment_dir_from:  /input/KAG-18/output/experiment\n\nI also tried : \n\n    experiment_dir:  /output/experiment\n    clone_experiment_dir_from:  /input/KAG-18/output/experiment\n\nError:\n\n    Traceback (most recent call last):\n    File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in",
    "373279": "Jakub,\nDisregard, I was able to get the evaluation to run. I had to clone the github since I did not do after the update. So far everything is running okay. I will know for sure once the prediction is complete.",
    "373324": "I was able to get both the train and predict to run. However, evaluation_prediction file was empty. \n\n    Traceback (most recent call last):\n     File \"/usr/local/lib/python3.6/dist-packages/deepsense/neptune/job_wrapper.py\", line 107, in",
    "373326": "Interesting, which classes did you group as \"human\"?",
    "373382": "So we have 4 \"human\" batches:\n\n    [\"Footwear\", \"Human hair\",\"Person\", \"Human arm\", \"Human eye\", \"Human face\", \"Suit\", \"Human hand\", \"Human leg\", \"Human nose\", \"Dress\", \"Human mouth\"]\n    \n    [\"Jacket\", \"Sports uniform\", \"Human ear\", \"Shorts\", \"Helmet\", \"Hat\", \"Tie\", \"Swimwear\"]\n    \n    [\"Trousers\", \"Shirt\", \"Umbrella\", \"Coat\", \"Human beard\", \"Necklace\", \"Human head\"]\n    \n    [\"Scarf\", \"Human foot\", \"Luggage and bags\", \"Watch\", \"Brassiere\",\"Earrings\",\"Sock\",\"Skirt\",\"Glove\",\"Crown\",\"Swim cap\",\"Belt\",\"Tiara\"]\n\nI guess they are more fashion/human after all.",
    "373476": "I like the idea of the approach, but is there a reason you decided to make the \"batch size\" so small? I looked at your other batches and it seems like they have about 100 classes each. Is it because of sampling reasons (meaning there are a lot more examples of human/fashion items in the train set)?",
    "373506": "Hi @muhammedazamkhan @jakubczakon. I ran the predict command, but got back an empty submission file. I feel like my model didn't train long enough, especially given that I trained on all classes. Did you run your model on all 500 classes @muhammedazamkhan? And also for 1000 epochs like the Neptune Team?",
    "373632": "Hi Jakub,\n\nI am just wondering how long it takes you to train your whole solution on 4x P100? What about for just one of your class batches? Apologies if this info is in the neptune.ml experiments, I am a little confused by which run is which there. Also, do you train on the whole dataset?",
    "373644": "For instance I trained this one \"batch_8\" https://app.neptune.ml/-/dashboard/experiment/6fbb78f8-67e1-4edf-816a-6ce2234504ce .\n\n I ran it local on 4 gtx 1070 though. Anyhow it was around 20 epochs/day  (120 total). Usually the batches would train for give or take 60 epochs and the smaller ones can actually be put on 2 gpus. \n\nI train on samples of 50000k per epoch for training and 10k for validation.\nThe samplers aren't a part of the public repo but they will be released right after the competition. \nOn the other hand we could release the code with the samplers but I am afraid it will make it to easy to get a medal with it.",
    "373647": "So does the `training_sample_size` parameter in the config mean number of samples per epoch but it still trains on the full dataset? Or does it mean that the dataset is truncated to N samples and those samples are reused for each epoch.\n\nI also assume by sampler you are referring to the method of which images are drawn from the dataset for training?",
    "373653": "Sorry for the confusion @anokas, my bad. I wasn't sure which version was on public github.\n\nAnyways, yes `training_sample_size` is the size of the sample that is drawn every epoch. Sampler refers to the pytorch sampler which is later passed to the pytorch loader in this [file](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/loaders.py).\n\nAnd we do sample from the entire dataset every epoch (just different samples).",
    "373725": "I keep getting this error when evaluating:\n\n    FileExistsError: [Errno 17] File exists: '/output/experiment'\n\nAny idea how to fix?",
    "373768": "I'm a little curious. Would we call that function in the pipeline_config and config files?",
    "373794": "Thanks Jakub, I understand now.\n\nOne thing I am curious about:\nThe model seems to spend a long time waiting between epochs. I understand the 6 minutes for validation, but does it take an additional 6 minutes to save the model, and then another 12 minutes to start the next epoch? Do you have any insights into what is going on in between these steps that might be making it slow?\n\n    2018-08-21 23:16:42 steppy &gt;&gt;&gt;; epoch 1 sum:     0.58112\n    2018-08-21 23:22:44 steppy &gt;&gt;&gt;; epoch 1 validation sum:     0.53714\n    2018-08-21 23:28:45 steppy &gt;&gt;&gt;; epoch 1 model persisted to experiment2/checkpoints/retinanet/best.torch\n    2018-08-21 23:28:45 steppy &gt;&gt;&gt;; epoch 2 current lr: 1e-05\n    2018-08-21 23:40:48 steppy &gt;&gt;&gt;; epoch 1 time 0:54:58\n    2018-08-21 23:40:48 steppy &gt;&gt;&gt;; epoch 2 ...",
    "373925": "Thanks for spotting that @anokas I haven't noticed that before.\nI checked the latest experiment and the same thing happened.\n\nMy guess is that it could be that by mistake the loss on validation is calculated multiple times. Looking at the callback list:\n\n    return CallbackList(\n        callbacks=[experiment_timing, training_monitor, validation_monitor,\n                   model_checkpoints, lr_scheduler, early_stopping, neptune_monitor,\n                   ]) \n\nafter the validation_monitor the validation loss is needed in model_checkpoints, early_stopping and neptune_monitor. I wonder if turning some of those off speeds things up considerably.\n\nUnfortunately, I am not sure that I will have the time to fix that till Monday. So if you would like to fix/check that yourself. The vast majority of the logic is in the `src/callbacks.py` the rest sits in the [steppy-toolkit package](https://github.com/neptune-ml/steppy-toolkit/blob/master/toolkit/pytorch_transformers/callbacks.py).",
    "373971": "Can you paste your data paths? Does it fail after evaluation ( map value has been calculated) or straight away? What os the exact command you are using?",
    "374073": "Hi @Giri Gopalan.\n\nI am running my evaluate_predict experiment right now and it seems to be working just fine.\nSo in my case I:\n\n - trained the model with this [neptune\n   experiment GAI-473](https://app.neptune.ml/-/dashboard/experiment/ca63378d-eaef-4992-a181-1582ca31e5ca) running the following command\n\n         neptune send --worker m-4p100 \\\n         --environment pytorch-0.3.1-gpu-py3 \\\n         --config configs/neptune.yaml \\\n         main.py train --pipeline_name retinanet\n\nand my `neptune.yaml` looked like this\n\n    parameters:\n    # Data Paths\n      train_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/train\n      test_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/test_challenge_2018\n      annotations_filepath: /public/challenges/google-ai-open-images-object-detection-track/annotations/challenge-2018-train-annotations-bbox.csv\n      annotations_human_labels_filepath: /public/challenges/google-ai-open-images-object-detection-track/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n      bbox_hierarchy_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/bbox_labels_500_hierarchy.json\n      class_mappings_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/challenge-2018-class-descriptions-500.csv\n      valid_ids_filepath: /public/challenges/google-ai-open-images-object-detection-track/metadata/challenge-2018-image-ids-valset-od.csv\n      sample_submission: /public/challenges/google-ai-open-images-object-detection-track/sample_submission.csv\n      experiment_dir:  /output/experiment\n      clone_experiment_dir_from: ''\n\nNotice that the `clone_experiment_dir_from: '' ` during training. \n\n  - To evaluate I changed the `neptune.yaml` to\n\n        clone_experiment_dir_from: /input/GAI-473/output/experiment \n    \n    and ran the following command\n\n         neptune send --worker m-p100 \\\n        --environment pytorch-0.3.1-gpu-py3 \\\n        --config configs/neptune.yaml \\\n        --input /GAI-473 \\\n        main.py evaluate_predict --pipeline_name retinanet --chunk_size 100\n\nNotice that during prediction I am using just one p100 machine with the command `m-p100` but it is of course not necessary to change that (just cheaper).\n\nHave you done it exactly this way as well ?",
    "374080": "# Data Paths\n    train_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/train\n    test_imgs_dir: /public/datasets/open-images-dataset-v4/bounding-boxes/test_challenge_2018\n    annotations_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/annotations/challenge-2018-train-annotations-bbox.csv\n    annotations_human_labels_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/annotations/challenge-2018-train-annotations-human-imagelabels.csv\n    bbox_hierarchy_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/metadata/bbox_labels_500_hierarchy.json\n    class_mappings_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/metadata/challenge-2018-class-descriptions-500.csv\n    valid_ids_filepath: /public/challenges/google-ai-open-images-object-detection- \n    track/metadata/challenge-2018-image-ids-valset-od.csv\n    sample_submission: /public/challenges/google-ai-open-images-object-detection- \n    track/sample_submission.csv\n    experiment_dir:  /output/experiment\n    clone_experiment_dir_from:  /input/GOOG-2/output/experiment\n\nI am trying this setup per your recommendation:\n\n    experiment_dir:  /output/experiment\n    clone_experiment_dir_from:  GOOG-2/output/experiment/transformers/label_encoder",
    "374089": "William Greene I think it should be:\n\n    clone_experiment_dir_from:  GOOG-2/output/experiment\n\nnot \n\n    clone_experiment_dir_from:  GOOG-2/output/experiment/transformers/label_encoder",
    "374101": "Ok I think I found the culprit should have fix in no time.\n\nIf you don't want to wait for it you can simply run 'evaluate' pipeline and 'predict' pipeline separately in two consecutive experiments and it will work just fine.",
    "374103": "Can you try running the `evaluate` pipeline and not 'evaluate_predict' pipeline first",
    "374106": "William Green please check the newest master",
    "374110": "Thanks for all your help so far. Is there any way I can resume training from a checkpoint, if the training script is stopped?\n\nI am unfamiliar with steppy so I am not sure how to change the code to do this.",
    "374143": "Sure @anokas. Glad I could be of help.\n\n If you go to [src/models.py](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/models.py) :\n\nYou have something like this:\n\n    class ModelParallel(Model):\n        def fit(self, datagen, validation_datagen=None):\n            self._initialize_model_weights()\n    \n            self.model = DataParallel(self.model)\n    \n            if torch.cuda.is_available():\n                self.model = self.model.cuda() \n\nsimply add loading:\n\n    class ModelParallel(Model):\n        def fit(self, datagen, validation_datagen=None):\n            self._initialize_model_weights()\n    \n            self.model = DataParallel(self.model)\n            self.load(YOUR_FILEPATH)\n\n            if torch.cuda.is_available():\n                self.model = self.model.cuda() \n\nIf you want to do some custom loading of say some backbone layers or something I would suggest that you go to the `load` method and check how it's done there and probably add a 'custom_load' method for your use case.",
    "374159": "Jakub Thank you. I'm checking it out right now.",
    "374601": "Hi @Jonas \nI've just checked the implementation with default 10 classes and have run for 5 epochs.",
    "374672": "OK, I found out what was causing my crashes. In the evaluate runs that crashed, I did NOT use --chunk-size. When I set it per your response above, it went through properly.\n\nI think things should not crash if an option is not set correctly.  Some sort of error message would have been easier for me to debug.\n\nIn any case, thanks much for the help and all of the work you guys have done.",
    "374682": "What was the baseline score for solution 1?",
    "375068": "Jakub, How many batches are there total?",
    "375070": "We have 12 batches right now but the experiments with some of them joined together are in progress.",
    "375072": "Jakub, \n\nHow do you apply the method for training from a checkpoint in the cloud?",
    "375091": "Simply set data paths as you would with evaluate/predict and add --input YOUREXP-1 in the command and your checkpoint will be available in the /output/experiment/checkpoints/retinanet/best.torch .",
    "376754": "Missing experiment folder.  For some reason, the experiment folder was not created after training was complete.",
    "376857": "Local or cloud?",
    "376948": "Cloud.",
    "376995": "So output/experiment is empty or non-existent?",
    "377064": "It's non-existent. I think I may have figured out why.\n\nI had `output/experiment/` versus ` /output/experiment`",
    "377230": "Jskub can I use the same cmd line for evaluation?\n\n  clone:   /output/experiment/checkpoints/retinanet/best.torch",
    "377251": "Is this the correct command to train from last checkpoint\n\n    neptune send --worker m-4p100 \\\n    --environment pytorch-0.3.1-gpu-py3 \\\n    --config configs/neptune.yaml \\\n    --input /GOOG-50 \\\n    main.py train --pipeline_name retinanet\n\nconfig: \n\n          experiment_dir:  /output/experiment/checkpoints/retinanet/best.torch\n          clone_experiment_dir_from: ''\n\n\nI get the following error when I either try to run from last know checkpoint or evaluate:\n\n    00:00\t/usr/bin/python: can't open file '/usr/local/lib/python3.6/dist-packages/neptune/job_wrapper.py': [Errno 2] No \n     such file or directory",
    "377291": "config should be:\n\n      experiment_dir:  /output/experiment/checkpoints/retinanet/best.torch\n      clone_experiment_dir_from: /input/GOOG-50/output/experiment\n\nand you should insert the following line:\n\n        self.load('/input/GOOG-50/output/experiment/checkpoints/retinanet/best.torch')\n\n\nIn the `models.py` as explained to @anokas below.",
    "377629": "I'm still getting this error \n\n    0:00\t/usr/bin/python: can't open file '/usr/local/lib/python3.6/dist-packages/neptune/job_wrapper.py': [Errno 2] \n    No such file or directory",
    "377697": "I am pretty sure that the problem was with the wrong neptune-cli version in the requirements.tx . \n\nIt is fixed now so it should run with no problems. Just go to the repo and check.",
    "377746": "Jakub \n\nThank you :) I finally was able to get it to run. It works :)"
  },
  "source": "meta"
}