{
  "id": 30462,
  "title": "Fun with qsub on Colfax",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30462",
  "author_name": "",
  "post_date": "2017-03-21T16:44:49.456513100Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>If you are using Colfax Cluster provided, probably you use <code>qsub</code> tool to submit your <em>huge</em> job to the cluster.  Then you wait until the task is terminated and you can look to STDIN.oXXXX or STDIN.eXXXX (or your customized output name). </p>\n\n<p>However, there is an option to launch <code>qsub</code> in the interactive mode and execute the code on node directly. </p>\n\n<p><strong>Terminal</strong> </p>\n\n<p>Here is the magic command that you can run once connected to the cluster : </p>\n\n<p><code>\nqsub -V -I interactive_qsub\n</code></p>\n\n<p>where <code>interactive_qsub</code> file is</p>\n\n<p><code>#PBS -l nodes=1:knl7210:ram96gb -N interactive_qsub -d /path/to/your/home</code></p>\n\n<p><strong>Jupyter-Notebook</strong> </p>\n\n<p>Another option is to run a code from <em>Jupyter-Notebook</em> once you are running jupyter on the cluster without browser and visualize the interface on your local machine. Follow <a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/how-to-start-with-python-on-colfax-cluster\">this</a> for more details. </p>\n\n<p>Thus, you can, for example run caffe training on a dataset and follow the progress. It can be done like this. Write the following in a jupyter cell and execute:</p>\n\n<p><code>args_str = \" \".join(['train', '-solver', solver_filename, '-weights', weights_filename])</code></p>\n\n<p><code>qsub_script = \"\\\"#PBS -V -I -x -N caffe_train -l nodes=1:knl7210:ram96gb\\ncaffe %s\\\"\" % args_str</code></p>\n\n<p><code>!echo {qsub_script} &gt; qsub_script</code></p>\n\n<p><code>!chmod +x qsub_script</code></p>\n\n<p><code>!qsub $PWD/qsub_script</code></p>\n\n<p>This script submits a job to the cluster and runs created script file <code>qsub_script</code> in the interactive mode. </p>\n\n<p>See <a href=\"http://docs.adaptivecomputing.com/torque/5-1-3/Content/topics/torque/commands/qsub.htm\">here</a> for more details and have fun !</p>\n\n<p>PS. Thank you Intel for providing freely such a thing</p>",
  "messages": [
    {
      "id": "169587",
      "postDate": "03/21/2017 16:44:49",
      "content": "<p>If you are using Colfax Cluster provided, probably you use <code>qsub</code> tool to submit your <em>huge</em> job to the cluster.  Then you wait until the task is terminated and you can look to STDIN.oXXXX or STDIN.eXXXX (or your customized output name). </p>\n\n<p>However, there is an option to launch <code>qsub</code> in the interactive mode and execute the code on node directly. </p>\n\n<p><strong>Terminal</strong> </p>\n\n<p>Here is the magic command that you can run once connected to the cluster : </p>\n\n<p><code>\nqsub -V -I interactive_qsub\n</code></p>\n\n<p>where <code>interactive_qsub</code> file is</p>\n\n<p><code>#PBS -l nodes=1:knl7210:ram96gb -N interactive_qsub -d /path/to/your/home</code></p>\n\n<p><strong>Jupyter-Notebook</strong> </p>\n\n<p>Another option is to run a code from <em>Jupyter-Notebook</em> once you are running jupyter on the cluster without browser and visualize the interface on your local machine. Follow <a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/how-to-start-with-python-on-colfax-cluster\">this</a> for more details. </p>\n\n<p>Thus, you can, for example run caffe training on a dataset and follow the progress. It can be done like this. Write the following in a jupyter cell and execute:</p>\n\n<p><code>args_str = \" \".join(['train', '-solver', solver_filename, '-weights', weights_filename])</code></p>\n\n<p><code>qsub_script = \"\\\"#PBS -V -I -x -N caffe_train -l nodes=1:knl7210:ram96gb\\ncaffe %s\\\"\" % args_str</code></p>\n\n<p><code>!echo {qsub_script} &gt; qsub_script</code></p>\n\n<p><code>!chmod +x qsub_script</code></p>\n\n<p><code>!qsub $PWD/qsub_script</code></p>\n\n<p>This script submits a job to the cluster and runs created script file <code>qsub_script</code> in the interactive mode. </p>\n\n<p>See <a href=\"http://docs.adaptivecomputing.com/torque/5-1-3/Content/topics/torque/commands/qsub.htm\">here</a> for more details and have fun !</p>\n\n<p>PS. Thank you Intel for providing freely such a thing</p>",
      "rawMarkdown": "If you are using Colfax Cluster provided, probably you use `qsub` tool to submit your *huge* job to the cluster.  Then you wait until the task is terminated and you can look to STDIN.oXXXX or STDIN.eXXXX (or your customized output name). \n\nHowever, there is an option to launch `qsub` in the interactive mode and execute the code on node directly. \n\n\n**Terminal** \n\nHere is the magic command that you can run once connected to the cluster : \n\n```\nqsub -V -I interactive_qsub\n```\n\nwhere `interactive_qsub` file is\n\n```#PBS -l nodes=1:knl7210:ram96gb -N interactive_qsub -d /path/to/your/home```\n\n**Jupyter-Notebook** \n\nAnother option is to run a code from *Jupyter-Notebook* once you are running jupyter on the cluster without browser and visualize the interface on your local machine. Follow [this][2] for more details. \n\nThus, you can, for example run caffe training on a dataset and follow the progress. It can be done like this. Write the following in a jupyter cell and execute:\n\n```args_str = \" \".join(['train', '-solver', solver_filename, '-weights', weights_filename])```\n\n```qsub_script = \"\\\"#PBS -V -I -x -N caffe_train -l nodes=1:knl7210:ram96gb\\ncaffe %s\\\"\" % args_str```\n\n```!echo {qsub_script} > qsub_script```\n\n```!chmod +x qsub_script```\n\n```!qsub $PWD/qsub_script```\n\nThis script submits a job to the cluster and runs created script file `qsub_script` in the interactive mode. \n\n\nSee [here][1] for more details and have fun !\n\nPS. Thank you Intel for providing freely such a thing\n\n\n\n  [1]: http://docs.adaptivecomputing.com/torque/5-1-3/Content/topics/torque/commands/qsub.htm\n\n  [2]: https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/how-to-start-with-python-on-colfax-cluster",
      "votes": null
    },
    {
      "id": "169590",
      "postDate": "03/21/2017 16:53:18",
      "content": "<p>Each node has 64*4 cores. If you request 4 nodes you're requesting a total of 1024 cores distributed through 4 worker nodes. These resources are all dedicated exclusively to you which is probably not adequate for interactive use through Jupyter, etc.</p>",
      "rawMarkdown": "Each node has 64*4 cores. If you request 4 nodes you're requesting a total of 1024 cores distributed through 4 worker nodes. These resources are all dedicated exclusively to you which is probably not adequate for interactive use through Jupyter, etc.",
      "votes": null
    },
    {
      "id": "169591",
      "postDate": "03/21/2017 16:58:36",
      "content": "<p>I thought about directly launching scripts and observing the stdout without waiting. </p>\n\n<p>I'm not sure about the possibility of launching Jupyter inside the job and port forwarding up to ones local machine. </p>",
      "rawMarkdown": "I thought about directly launching scripts and observing the stdout without waiting. \n\nI'm not sure about the possibility of launching Jupyter inside the job and port forwarding up to ones local machine.",
      "votes": null
    },
    {
      "id": "169594",
      "postDate": "03/21/2017 17:02:40",
      "content": "<p>If you request multiple nodes you need to make your job use all nodes (that doesn't work transparently). For example, by accessing $PBSNODES and making your code distributed. If you don't do that you're running on a single node and having 3 idle nodes. If your code is not prepared to run on multiple nodes use <strong>#PBS -l nodes=1</strong>. If everyone runs the command above without knowing what's going on then I'm afraid there won't be resources for everybody.</p>",
      "rawMarkdown": "If you request multiple nodes you need to make your job use all nodes (that doesn't work transparently). For example, by accessing $PBSNODES and making your code distributed. If you don't do that you're running on a single node and having 3 idle nodes. If your code is not prepared to run on multiple nodes use **#PBS -l nodes=1**. If everyone runs the command above without knowing what's going on then I'm afraid there won't be resources for everybody.",
      "votes": null
    },
    {
      "id": "169596",
      "postDate": "03/21/2017 17:07:57",
      "content": "<p>Right, it is not evident. I changed it in the text to prevent this overload. But anyway it is up to all of us to use properly the resources. And hopefully, they provide Intel flavoured Caffe which should be adapted to the cluster and the nodes.</p>",
      "rawMarkdown": "Right, it is not evident. I changed it in the text to prevent this overload. But anyway it is up to all of us to use properly the resources. And hopefully, they provide Intel flavoured Caffe which should be adapted to the cluster and the nodes.",
      "votes": null
    },
    {
      "id": "169891",
      "postDate": "03/23/2017 06:16:45",
      "content": "<blockquote>\n  <p>I think, we can use only 1 node on Colfax. <br>\n  If you request 4 node, you can use only 1 node.</p>\n</blockquote>\n\n<p>I'm very sorry, I was wrong! <br>\nIt was something wrong that I did try before:(  </p>\n\n<p>However, so that, it may be a problem of nothing of the guidelines to use nodes.  </p>\n\n<p>Anyways, thank you so much to make this discussion!!</p>",
      "rawMarkdown": "> I think, we can use only 1 node on Colfax.  \n> If you request 4 node, you can use only 1 node.\n\nI'm very sorry, I was wrong!  \nIt was something wrong that I did try before:(  \n\nHowever, so that, it may be a problem of nothing of the guidelines to use nodes.  \n\nAnyways, thank you so much to make this discussion!!",
      "votes": null
    },
    {
      "id": "169928",
      "postDate": "03/23/2017 09:13:53",
      "content": "<p>Again, it depends on the code you executing. Basic example of Colfax documentation shows that you can use all requested nodes:</p>\n\n<pre><code>$ qsub -V -I -x  ~/tmp/distrJob \nqsub: waiting for job 4171.c001 to start\nqsub: job 4171.c001 ready\n###################################################################\n# Colfax Cluster - https://colfaxresearch.com/\n#      Date:           Thu Mar 23 02:00:20 PDT 2017\n#    Job ID:           4171.c001\n#      User:           USER\n# Resources:           neednodes=4:knl,nodes=4:knl,walltime=24:00:00\n###################################################################\nLaunching the parallel job from mother superior c001-n038...\n[1] DAPL startup: RLIMIT_MEMLOCK too small\n[2] DAPL startup: RLIMIT_MEMLOCK too small\n[3] DAPL startup: RLIMIT_MEMLOCK too small\nHello world from host c001-n040 (rank 2)!\nHello world from host c001-n038 (rank 0)!\nHello world from host c001-n041 (rank 3)!\nHello world from host c001-n039 (rank 1)!\nqsub: job 4171.c001 completed\n</code></pre>\n\n<p>They provide also caffe with <a href=\"https://github.com/intel/caffe/wiki/Multinode-guide\">mlsl options</a> at <code>/opt/caffe-mlsl</code>, so caffe funs have a huge advantage :)</p>",
      "rawMarkdown": "Again, it depends on the code you executing. Basic example of Colfax documentation shows that you can use all requested nodes:\n\n    $ qsub -V -I -x  ~/tmp/distrJob \n    qsub: waiting for job 4171.c001 to start\n    qsub: job 4171.c001 ready\n    ###################################################################\n    # Colfax Cluster - https://colfaxresearch.com/\n    #      Date:           Thu Mar 23 02:00:20 PDT 2017\n    #    Job ID:           4171.c001\n    #      User:           USER\n    # Resources:           neednodes=4:knl,nodes=4:knl,walltime=24:00:00\n    ###################################################################\n    Launching the parallel job from mother superior c001-n038...\n    [1] DAPL startup: RLIMIT_MEMLOCK too small\n    [2] DAPL startup: RLIMIT_MEMLOCK too small\n    [3] DAPL startup: RLIMIT_MEMLOCK too small\n    Hello world from host c001-n040 (rank 2)!\n    Hello world from host c001-n038 (rank 0)!\n    Hello world from host c001-n041 (rank 3)!\n    Hello world from host c001-n039 (rank 1)!\n    qsub: job 4171.c001 completed\n\n\nThey provide also caffe with [mlsl options](https://github.com/intel/caffe/wiki/Multinode-guide) at `/opt/caffe-mlsl`, so caffe funs have a huge advantage :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 169590,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "03/21/2017 16:53:18",
      "content": "<p>Each node has 64*4 cores. If you request 4 nodes you're requesting a total of 1024 cores distributed through 4 worker nodes. These resources are all dedicated exclusively to you which is probably not adequate for interactive use through Jupyter, etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 169591,
          "author_name": "vfdev5",
          "author_url": "",
          "post_date": "03/21/2017 16:58:36",
          "content": "<p>I thought about directly launching scripts and observing the stdout without waiting. </p>\n\n<p>I'm not sure about the possibility of launching Jupyter inside the job and port forwarding up to ones local machine. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169594,
          "author_name": "aamaia",
          "author_url": "",
          "post_date": "03/21/2017 17:02:40",
          "content": "<p>If you request multiple nodes you need to make your job use all nodes (that doesn't work transparently). For example, by accessing $PBSNODES and making your code distributed. If you don't do that you're running on a single node and having 3 idle nodes. If your code is not prepared to run on multiple nodes use <strong>#PBS -l nodes=1</strong>. If everyone runs the command above without knowing what's going on then I'm afraid there won't be resources for everybody.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169596,
          "author_name": "vfdev5",
          "author_url": "",
          "post_date": "03/21/2017 17:07:57",
          "content": "<p>Right, it is not evident. I changed it in the text to prevent this overload. But anyway it is up to all of us to use properly the resources. And hopefully, they provide Intel flavoured Caffe which should be adapted to the cluster and the nodes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169891,
          "author_name": "kambarakun",
          "author_url": "",
          "post_date": "03/23/2017 06:16:45",
          "content": "<blockquote>\n  <p>I think, we can use only 1 node on Colfax. <br>\n  If you request 4 node, you can use only 1 node.</p>\n</blockquote>\n\n<p>I'm very sorry, I was wrong! <br>\nIt was something wrong that I did try before:(  </p>\n\n<p>However, so that, it may be a problem of nothing of the guidelines to use nodes.  </p>\n\n<p>Anyways, thank you so much to make this discussion!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169928,
          "author_name": "vfdev5",
          "author_url": "",
          "post_date": "03/23/2017 09:13:53",
          "content": "<p>Again, it depends on the code you executing. Basic example of Colfax documentation shows that you can use all requested nodes:</p>\n\n<pre><code>$ qsub -V -I -x  ~/tmp/distrJob \nqsub: waiting for job 4171.c001 to start\nqsub: job 4171.c001 ready\n###################################################################\n# Colfax Cluster - https://colfaxresearch.com/\n#      Date:           Thu Mar 23 02:00:20 PDT 2017\n#    Job ID:           4171.c001\n#      User:           USER\n# Resources:           neednodes=4:knl,nodes=4:knl,walltime=24:00:00\n###################################################################\nLaunching the parallel job from mother superior c001-n038...\n[1] DAPL startup: RLIMIT_MEMLOCK too small\n[2] DAPL startup: RLIMIT_MEMLOCK too small\n[3] DAPL startup: RLIMIT_MEMLOCK too small\nHello world from host c001-n040 (rank 2)!\nHello world from host c001-n038 (rank 0)!\nHello world from host c001-n041 (rank 3)!\nHello world from host c001-n039 (rank 1)!\nqsub: job 4171.c001 completed\n</code></pre>\n\n<p>They provide also caffe with <a href=\"https://github.com/intel/caffe/wiki/Multinode-guide\">mlsl options</a> at <code>/opt/caffe-mlsl</code>, so caffe funs have a huge advantage :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "169587": "If you are using Colfax Cluster provided, probably you use `qsub` tool to submit your *huge* job to the cluster.  Then you wait until the task is terminated and you can look to STDIN.oXXXX or STDIN.eXXXX (or your customized output name). \n\nHowever, there is an option to launch `qsub` in the interactive mode and execute the code on node directly. \n\n\n**Terminal** \n\nHere is the magic command that you can run once connected to the cluster : \n\n```\nqsub -V -I interactive_qsub\n```\n\nwhere `interactive_qsub` file is\n\n```#PBS -l nodes=1:knl7210:ram96gb -N interactive_qsub -d /path/to/your/home```\n\n**Jupyter-Notebook** \n\nAnother option is to run a code from *Jupyter-Notebook* once you are running jupyter on the cluster without browser and visualize the interface on your local machine. Follow [this][2] for more details. \n\nThus, you can, for example run caffe training on a dataset and follow the progress. It can be done like this. Write the following in a jupyter cell and execute:\n\n```args_str = \" \".join(['train', '-solver', solver_filename, '-weights', weights_filename])```\n\n```qsub_script = \"\\\"#PBS -V -I -x -N caffe_train -l nodes=1:knl7210:ram96gb\\ncaffe %s\\\"\" % args_str```\n\n```!echo {qsub_script} > qsub_script```\n\n```!chmod +x qsub_script```\n\n```!qsub $PWD/qsub_script```\n\nThis script submits a job to the cluster and runs created script file `qsub_script` in the interactive mode. \n\n\nSee [here][1] for more details and have fun !\n\nPS. Thank you Intel for providing freely such a thing\n\n\n\n  [1]: http://docs.adaptivecomputing.com/torque/5-1-3/Content/topics/torque/commands/qsub.htm\n\n  [2]: https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/how-to-start-with-python-on-colfax-cluster",
    "169590": "Each node has 64*4 cores. If you request 4 nodes you're requesting a total of 1024 cores distributed through 4 worker nodes. These resources are all dedicated exclusively to you which is probably not adequate for interactive use through Jupyter, etc.",
    "169591": "I thought about directly launching scripts and observing the stdout without waiting. \n\nI'm not sure about the possibility of launching Jupyter inside the job and port forwarding up to ones local machine.",
    "169594": "If you request multiple nodes you need to make your job use all nodes (that doesn't work transparently). For example, by accessing $PBSNODES and making your code distributed. If you don't do that you're running on a single node and having 3 idle nodes. If your code is not prepared to run on multiple nodes use **#PBS -l nodes=1**. If everyone runs the command above without knowing what's going on then I'm afraid there won't be resources for everybody.",
    "169596": "Right, it is not evident. I changed it in the text to prevent this overload. But anyway it is up to all of us to use properly the resources. And hopefully, they provide Intel flavoured Caffe which should be adapted to the cluster and the nodes.",
    "169891": "> I think, we can use only 1 node on Colfax.  \n> If you request 4 node, you can use only 1 node.\n\nI'm very sorry, I was wrong!  \nIt was something wrong that I did try before:(  \n\nHowever, so that, it may be a problem of nothing of the guidelines to use nodes.  \n\nAnyways, thank you so much to make this discussion!!",
    "169928": "Again, it depends on the code you executing. Basic example of Colfax documentation shows that you can use all requested nodes:\n\n    $ qsub -V -I -x  ~/tmp/distrJob \n    qsub: waiting for job 4171.c001 to start\n    qsub: job 4171.c001 ready\n    ###################################################################\n    # Colfax Cluster - https://colfaxresearch.com/\n    #      Date:           Thu Mar 23 02:00:20 PDT 2017\n    #    Job ID:           4171.c001\n    #      User:           USER\n    # Resources:           neednodes=4:knl,nodes=4:knl,walltime=24:00:00\n    ###################################################################\n    Launching the parallel job from mother superior c001-n038...\n    [1] DAPL startup: RLIMIT_MEMLOCK too small\n    [2] DAPL startup: RLIMIT_MEMLOCK too small\n    [3] DAPL startup: RLIMIT_MEMLOCK too small\n    Hello world from host c001-n040 (rank 2)!\n    Hello world from host c001-n038 (rank 0)!\n    Hello world from host c001-n041 (rank 3)!\n    Hello world from host c001-n039 (rank 1)!\n    qsub: job 4171.c001 completed\n\n\nThey provide also caffe with [mlsl options](https://github.com/intel/caffe/wiki/Multinode-guide) at `/opt/caffe-mlsl`, so caffe funs have a huge advantage :)"
  },
  "source": "meta"
}