{
  "id": 30272,
  "title": "Colfax -- Caffe, Keras, etc.",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30272",
  "author_name": "",
  "post_date": "2017-03-17T15:45:22.591947Z",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<ul>\n<li><p>Caffe: my first attempt failed becase lmdb-dev is not installed on the cluster (when you try to \"import caffe\" from python it fails). I compiled lmdb locally on my account and made sure to put <code>LD_LIBRARY_PATH=~/tmp/lib:$LD_LIBRARY_PATH</code> in my PBS script but  had no luck. Also, we need to compile our own caffe version from <a href=\"https://github.com/intel/caffe\">https://github.com/intel/caffe</a> to make use of multiple nodes. See <a href=\"https://github.com/intel/caffe/wiki/Multinode-guide\">https://github.com/intel/caffe/wiki/Multinode-guide</a>. Maybe it didn't work because of the interaction between conda environment and LD_LIBRARY_PATH. I put this on hold.</p></li>\n<li><p>Keras:  theano backend failed in my first attempt. Afterwards I discovered this: <a href=\"https://colfaxresearch.com/discussion/topic/keras-with-theano-backend/\">https://colfaxresearch.com/discussion/topic/keras-with-theano-backend/</a></p></li>\n<li><p>Keras: to use tensorflow backend make sure to configure ~/.keras/keras.jon and <strong>ALSO</strong> write in your submission script <code>export env KERAS_BACKEND=tensorflow</code> (to override the default). I'm using keras 2.0 with tensorflow 1.0 from conda-forge. Running on a single node and a single CPU is extremely slow. More than on my local machine. I'm not sure why. Slow I/O only explains part of the picture, and I don't have the tools to investigate remotely properly . Adding threads to tensorflow session via <code>intra_op_parallelism_threads</code> does not help. </p></li>\n<li><p>Next steps for me: trying to find the cause of the slowness, then trying to make use of the compute power, probably will need Tensorflow compiled with OpenBlas OR using  <a href=\"https://www.tensorflow.org/deploy/distributed\">distributed tensorflow</a> even on a single node to make use of multiple CPUs.</p></li>\n</ul>\n\n<p>I realize all these problems are temporary while we figure out how to work on Colfax in the initial stages of the competition and I hope that  by sharing we can overcome these problems faster.</p>",
  "messages": [
    {
      "id": "168614",
      "postDate": "03/17/2017 15:45:22",
      "content": "<p>Hi,</p>\n\n<ul>\n<li><p>Caffe: my first attempt failed becase lmdb-dev is not installed on the cluster (when you try to \"import caffe\" from python it fails). I compiled lmdb locally on my account and made sure to put <code>LD_LIBRARY_PATH=~/tmp/lib:$LD_LIBRARY_PATH</code> in my PBS script but  had no luck. Also, we need to compile our own caffe version from <a href=\"https://github.com/intel/caffe\">https://github.com/intel/caffe</a> to make use of multiple nodes. See <a href=\"https://github.com/intel/caffe/wiki/Multinode-guide\">https://github.com/intel/caffe/wiki/Multinode-guide</a>. Maybe it didn't work because of the interaction between conda environment and LD_LIBRARY_PATH. I put this on hold.</p></li>\n<li><p>Keras:  theano backend failed in my first attempt. Afterwards I discovered this: <a href=\"https://colfaxresearch.com/discussion/topic/keras-with-theano-backend/\">https://colfaxresearch.com/discussion/topic/keras-with-theano-backend/</a></p></li>\n<li><p>Keras: to use tensorflow backend make sure to configure ~/.keras/keras.jon and <strong>ALSO</strong> write in your submission script <code>export env KERAS_BACKEND=tensorflow</code> (to override the default). I'm using keras 2.0 with tensorflow 1.0 from conda-forge. Running on a single node and a single CPU is extremely slow. More than on my local machine. I'm not sure why. Slow I/O only explains part of the picture, and I don't have the tools to investigate remotely properly . Adding threads to tensorflow session via <code>intra_op_parallelism_threads</code> does not help. </p></li>\n<li><p>Next steps for me: trying to find the cause of the slowness, then trying to make use of the compute power, probably will need Tensorflow compiled with OpenBlas OR using  <a href=\"https://www.tensorflow.org/deploy/distributed\">distributed tensorflow</a> even on a single node to make use of multiple CPUs.</p></li>\n</ul>\n\n<p>I realize all these problems are temporary while we figure out how to work on Colfax in the initial stages of the competition and I hope that  by sharing we can overcome these problems faster.</p>",
      "rawMarkdown": "Hi,\n\n* Caffe: my first attempt failed becase lmdb-dev is not installed on the cluster (when you try to \"import caffe\" from python it fails). I compiled lmdb locally on my account and made sure to put `LD_LIBRARY_PATH=~/tmp/lib:$LD_LIBRARY_PATH` in my PBS script but  had no luck. Also, we need to compile our own caffe version from [https://github.com/intel/caffe][1] to make use of multiple nodes. See [https://github.com/intel/caffe/wiki/Multinode-guide][2]. Maybe it didn't work because of the interaction between conda environment and LD_LIBRARY_PATH. I put this on hold.\n\n* Keras:  theano backend failed in my first attempt. Afterwards I discovered this: [https://colfaxresearch.com/discussion/topic/keras-with-theano-backend/][3]\n\n* Keras: to use tensorflow backend make sure to configure ~/.keras/keras.jon and **ALSO** write in your submission script `export env KERAS_BACKEND=tensorflow` (to override the default). I'm using keras 2.0 with tensorflow 1.0 from conda-forge. Running on a single node and a single CPU is extremely slow. More than on my local machine. I'm not sure why. Slow I/O only explains part of the picture, and I don't have the tools to investigate remotely properly . Adding threads to tensorflow session via `intra_op_parallelism_threads` does not help. \n\n* Next steps for me: trying to find the cause of the slowness, then trying to make use of the compute power, probably will need Tensorflow compiled with OpenBlas OR using  [distributed tensorflow][4] even on a single node to make use of multiple CPUs.\n \nI realize all these problems are temporary while we figure out how to work on Colfax in the initial stages of the competition and I hope that  by sharing we can overcome these problems faster.\n\n\n  [1]: https://github.com/intel/caffe\n  [2]: https://github.com/intel/caffe/wiki/Multinode-guide\n  [3]: https://colfaxresearch.com/discussion/topic/keras-with-theano-backend/\n  [4]: https://www.tensorflow.org/deploy/distributed",
      "votes": null
    },
    {
      "id": "168887",
      "postDate": "03/18/2017 14:16:21",
      "content": "<p>I have a similar impressions about Colfax cluster with Keras. I use Theano as backend. However, if you specify <code>export OMP_NUM_THREADS=16</code> and use the option <code>openmp = True</code> in .theanorc, when you run a theano testing script </p>\n\n<p><code>#PBS -N test -l nodes=4:knl</code> </p>\n\n<p><code>python /opt/theano/theano/misc/elemwise_openmp_speedup.py</code></p>\n\n<p>you can see a theano performance report.</p>",
      "rawMarkdown": "I have a similar impressions about Colfax cluster with Keras. I use Theano as backend. However, if you specify `export OMP_NUM_THREADS=16` and use the option `openmp = True` in .theanorc, when you run a theano testing script \n\n```#PBS -N test -l nodes=4:knl``` \n\n```python /opt/theano/theano/misc/elemwise_openmp_speedup.py```\n\n you can see a theano performance report.",
      "votes": null
    },
    {
      "id": "168986",
      "postDate": "03/18/2017 23:47:09",
      "content": "<p>@amaia Could you share how you created a conda environment with tensorflow 1.0.1 and keras 2.0.1 on the colfax cluster? </p>\n\n<p>I created a conda python 3.5 environment and installed tf and keras from conda-forge but kept getting errors as the conda Python reverted to the 'system' Intel python35. Many thanks!</p>",
      "rawMarkdown": "amaia Could you share how you created a conda environment with tensorflow 1.0.1 and keras 2.0.1 on the colfax cluster? \n\nI created a conda python 3.5 environment and installed tf and keras from conda-forge but kept getting errors as the conda Python reverted to the 'system' Intel python35. Many thanks!",
      "votes": null
    },
    {
      "id": "168991",
      "postDate": "03/19/2017 00:22:36",
      "content": "<p>If you're submitting a job make sure to \"source activate\" inside the PBS script. </p>\n\n<p>Other than this, suppose you create a conda environment, activate and then call ipython. If you didn't install ipython in the environment the one being called is the other one in $PATH, which is <code>/opt/intel/intelpython35/bin/ipython</code>. This is confusing. You can also use the program <code>which</code> to see which one is being called.</p>",
      "rawMarkdown": "If you're submitting a job make sure to \"source activate\" inside the PBS script. \n\nOther than this, suppose you create a conda environment, activate and then call ipython. If you didn't install ipython in the environment the one being called is the other one in $PATH, which is ```/opt/intel/intelpython35/bin/ipython```. This is confusing. You can also use the program ```which``` to see which one is being called.",
      "votes": null
    },
    {
      "id": "169005",
      "postDate": "03/19/2017 03:32:35",
      "content": "<p>UPDATE: I'm trying MXNet compiled with MKL now. The version in /opt/intel/mkl is not recent enough, a more recent version of MKLML is needed by mxnet. I'm running benchmarks and it's fast but from time to time it segfaults mysteriously while training.</p>",
      "rawMarkdown": "UPDATE: I'm trying MXNet compiled with MKL now. The version in /opt/intel/mkl is not recent enough, a more recent version of MKLML is needed by mxnet. I'm running benchmarks and it's fast but from time to time it segfaults mysteriously while training.",
      "votes": null
    },
    {
      "id": "169060",
      "postDate": "03/19/2017 11:43:59",
      "content": "<p>I created a conda env using:</p>\n\n<p>$ conda create -n my_root --clone=/opt/intel/intelpython35</p>\n\n<p>And, then installed TF 1.0.0 and keras 2.0.1 using conda-forge.</p>\n\n<p>When running myprogram.py, the error is:</p>\n\n<p>Traceback (most recent call last):\n  File \"/home/u2606/VM/dev/myprogram.py\", line 498, in \n    main()\n  File \"/home/u2606/VM/dev/myprogram.py\", line 355, in main\n    base_model = InceptionV3(weights='imagenet')\n  File \"/opt/intel/intelpython35/lib/python3.5/site-packages/Keras-1.1.0-py3.5.egg/keras/applications/inception_v3.py\", line 296, in InceptionV3\n    md5_hash='fe114b3ff2ea4bf891e9353d1bbfb32f')\n  File \"/opt/intel/intelpython35/lib/python3.5/site-packages/Keras-1.1.0-py3.5.egg/keras/utils/data_utils.py\", line 98, in get_file\n    raise Exception(error_msg.format(origin, e.errno, e.reason))\nException: URL fetch failure on <a href=\"https://github.com/fchollet/deep-learning-models/releases/download/v0.2/inception_v3_weights_tf_dim_ordering_tf_kernels.h5\">https://github.com/fchollet/deep-learning-models/releases/download/v0.2/inception_v3_weights_tf_dim_ordering_tf_kernels.h5</a>: None -- [Errno -3] Temporary failure in name resolution</p>\n\n<p>The newly created conda env is reverting to the 'system' python and the older versions of TF and keras. How can this be fixed?</p>",
      "rawMarkdown": "I created a conda env using:\n\n$ conda create -n my_root --clone=/opt/intel/intelpython35\n\nAnd, then installed TF 1.0.0 and keras 2.0.1 using conda-forge.\n\nWhen running myprogram.py, the error is:\n\nTraceback (most recent call last):\n  File \"/home/u2606/VM/dev/myprogram.py\", line 498, in <module>\n    main()\n  File \"/home/u2606/VM/dev/myprogram.py\", line 355, in main\n    base_model = InceptionV3(weights='imagenet')\n  File \"/opt/intel/intelpython35/lib/python3.5/site-packages/Keras-1.1.0-py3.5.egg/keras/applications/inception_v3.py\", line 296, in InceptionV3\n    md5_hash='fe114b3ff2ea4bf891e9353d1bbfb32f')\n  File \"/opt/intel/intelpython35/lib/python3.5/site-packages/Keras-1.1.0-py3.5.egg/keras/utils/data_utils.py\", line 98, in get_file\n    raise Exception(error_msg.format(origin, e.errno, e.reason))\nException: URL fetch failure on https://github.com/fchollet/deep-learning-models/releases/download/v0.2/inception_v3_weights_tf_dim_ordering_tf_kernels.h5: None -- [Errno -3] Temporary failure in name resolution\n\nThe newly created conda env is reverting to the 'system' python and the older versions of TF and keras. How can this be fixed?",
      "votes": null
    },
    {
      "id": "169108",
      "postDate": "03/19/2017 15:52:23",
      "content": "<p>So, you \"cloned\" the python in /opt when creating the environment and it's using the one you specified, what's wrong? if you create the environment with \"conda create -n my_root2 python=3.5\" it will use another version.</p>\n\n<p>The second error (\"temporary failure in name resolution\") happens because the work nodes don't have internet connectivity. You need to download weights from the master node.</p>",
      "rawMarkdown": "So, you \"cloned\" the python in /opt when creating the environment and it's using the one you specified, what's wrong? if you create the environment with \"conda create -n my_root2 python=3.5\" it will use another version.\n\nThe second error (\"temporary failure in name resolution\") happens because the work nodes don't have internet connectivity. You need to download weights from the master node.",
      "votes": null
    },
    {
      "id": "169778",
      "postDate": "03/22/2017 14:57:37",
      "content": "<p>I am training a CNN with Keras (theano) using the  parameters in vfdev's post. I am seeing a ~50x slower performance compared with my GeForce.  </p>",
      "rawMarkdown": "I am training a CNN with Keras (theano) using the  parameters in vfdev's post. I am seeing a ~50x slower performance compared with my GeForce.",
      "votes": null
    },
    {
      "id": "173363",
      "postDate": "04/06/2017 16:37:37",
      "content": "<p>UPDATE2:</p>\n\n<ul>\n<li><p>I/O speed is not consistent as it's a shared resource between all users. Removing this bottleneck helps improving overall performance. If you try naively to read from disk using 255 threads at the same time performance will be terrible.</p></li>\n<li><p><strong>Keras with Theano (using MKL)</strong> backend: this was the fastest I could achieve so far using Colfax, but still not acceptable performance.</p></li>\n</ul>\n\n<p>(Tensorflow is slower, even more on CPU).</p>\n\n<ul>\n<li><p>Caffe with MLSL as provided in <strong>/opt/caffe-mlsl</strong>. Performance worse than running on a single CPU. I think it's a configuration problem, only setting OMP_NUM_THREADS is not enough. I tried changing layer engine to MKL2017 but couldn't figure out (Caffe wasn't even trying to parallelize -- running \"top\" shows 100% CPU usage, i.e. only one CPU being used).</p></li>\n<li><p>Custom Caffe compiled with OpenBLAS and GCC instead of Intel compiler. In this setup setting OMP_NUM_THREADS makes \"top\" show multiple cpus are being used but it's still slow.</p></li>\n</ul>\n\n<p>Anyone had a better experience? All my attempts so far are trying to use more than one CPU within a single node. </p>",
      "rawMarkdown": "UPDATE2:\n\n* I/O speed is not consistent as it's a shared resource between all users. Removing this bottleneck helps improving overall performance. If you try naively to read from disk using 255 threads at the same time performance will be terrible.\n\n* **Keras with Theano (using MKL)** backend: this was the fastest I could achieve so far using Colfax, but still not acceptable performance.\n\n(Tensorflow is slower, even more on CPU).\n\n* Caffe with MLSL as provided in **/opt/caffe-mlsl**. Performance worse than running on a single CPU. I think it's a configuration problem, only setting OMP_NUM_THREADS is not enough. I tried changing layer engine to MKL2017 but couldn't figure out (Caffe wasn't even trying to parallelize -- running \"top\" shows 100% CPU usage, i.e. only one CPU being used).\n\n* Custom Caffe compiled with OpenBLAS and GCC instead of Intel compiler. In this setup setting OMP_NUM_THREADS makes \"top\" show multiple cpus are being used but it's still slow.\n\n\nAnyone had a better experience? All my attempts so far are trying to use more than one CPU within a single node.",
      "votes": null
    },
    {
      "id": "174256",
      "postDate": "04/10/2017 20:22:54",
      "content": "<p>I tried caffe installation in /opt/intel/intelpython_update_2/intelpython2.\nWith engine=\"MKL2017\", it's about 1.5x faster than without that.  But still, the speed is only about 1/40 of GTX1060.  By setting OMP_NUM_THREADS=32, I'm gaining about 30% speedup (network too small?).  I think something is seriously wrong.  </p>",
      "rawMarkdown": "I tried caffe installation in /opt/intel/intelpython_update_2/intelpython2.\nWith engine=\"MKL2017\", it's about 1.5x faster than without that.  But still, the speed is only about 1/40 of GTX1060.  By setting OMP_NUM_THREADS=32, I'm gaining about 30% speedup (network too small?).  I think something is seriously wrong.",
      "votes": null
    },
    {
      "id": "186987",
      "postDate": "05/30/2017 03:50:31",
      "content": "<p>I am using tflearn / tensorflow. I set the following parameters. (64 as intel processors hae 64 cores)</p>\n\n<pre><code>export OMP_NUM_THREADS=64\n</code></pre>\n\n<p>in model code</p>\n\n<pre><code>tflearn.init_graph(num_cores=64)\n</code></pre>\n\n<p>still training on colfax is slower than my laptop cpu :-(\nThis works perfectly on my laptop but it doesnt work on colfax. \nany tips on how can we train faster on colfax using tflearn / tensorflow? </p>",
      "rawMarkdown": "I am using tflearn / tensorflow. I set the following parameters. (64 as intel processors hae 64 cores)\n\n    export OMP_NUM_THREADS=64\n\nin model code\n\n    tflearn.init_graph(num_cores=64)\n\nstill training on colfax is slower than my laptop cpu :-(\nThis works perfectly on my laptop but it doesnt work on colfax. \nany tips on how can we train faster on colfax using tflearn / tensorflow?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 168887,
      "author_name": "vfdev5",
      "author_url": "",
      "post_date": "03/18/2017 14:16:21",
      "content": "<p>I have a similar impressions about Colfax cluster with Keras. I use Theano as backend. However, if you specify <code>export OMP_NUM_THREADS=16</code> and use the option <code>openmp = True</code> in .theanorc, when you run a theano testing script </p>\n\n<p><code>#PBS -N test -l nodes=4:knl</code> </p>\n\n<p><code>python /opt/theano/theano/misc/elemwise_openmp_speedup.py</code></p>\n\n<p>you can see a theano performance report.</p>",
      "votes": null,
      "replies": [
        {
          "id": 169778,
          "author_name": "chuijh",
          "author_url": "",
          "post_date": "03/22/2017 14:57:37",
          "content": "<p>I am training a CNN with Keras (theano) using the  parameters in vfdev's post. I am seeing a ~50x slower performance compared with my GeForce.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 168986,
      "author_name": "intaka",
      "author_url": "",
      "post_date": "03/18/2017 23:47:09",
      "content": "<p>@amaia Could you share how you created a conda environment with tensorflow 1.0.1 and keras 2.0.1 on the colfax cluster? </p>\n\n<p>I created a conda python 3.5 environment and installed tf and keras from conda-forge but kept getting errors as the conda Python reverted to the 'system' Intel python35. Many thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 168991,
          "author_name": "aamaia",
          "author_url": "",
          "post_date": "03/19/2017 00:22:36",
          "content": "<p>If you're submitting a job make sure to \"source activate\" inside the PBS script. </p>\n\n<p>Other than this, suppose you create a conda environment, activate and then call ipython. If you didn't install ipython in the environment the one being called is the other one in $PATH, which is <code>/opt/intel/intelpython35/bin/ipython</code>. This is confusing. You can also use the program <code>which</code> to see which one is being called.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169060,
          "author_name": "intaka",
          "author_url": "",
          "post_date": "03/19/2017 11:43:59",
          "content": "<p>I created a conda env using:</p>\n\n<p>$ conda create -n my_root --clone=/opt/intel/intelpython35</p>\n\n<p>And, then installed TF 1.0.0 and keras 2.0.1 using conda-forge.</p>\n\n<p>When running myprogram.py, the error is:</p>\n\n<p>Traceback (most recent call last):\n  File \"/home/u2606/VM/dev/myprogram.py\", line 498, in \n    main()\n  File \"/home/u2606/VM/dev/myprogram.py\", line 355, in main\n    base_model = InceptionV3(weights='imagenet')\n  File \"/opt/intel/intelpython35/lib/python3.5/site-packages/Keras-1.1.0-py3.5.egg/keras/applications/inception_v3.py\", line 296, in InceptionV3\n    md5_hash='fe114b3ff2ea4bf891e9353d1bbfb32f')\n  File \"/opt/intel/intelpython35/lib/python3.5/site-packages/Keras-1.1.0-py3.5.egg/keras/utils/data_utils.py\", line 98, in get_file\n    raise Exception(error_msg.format(origin, e.errno, e.reason))\nException: URL fetch failure on <a href=\"https://github.com/fchollet/deep-learning-models/releases/download/v0.2/inception_v3_weights_tf_dim_ordering_tf_kernels.h5\">https://github.com/fchollet/deep-learning-models/releases/download/v0.2/inception_v3_weights_tf_dim_ordering_tf_kernels.h5</a>: None -- [Errno -3] Temporary failure in name resolution</p>\n\n<p>The newly created conda env is reverting to the 'system' python and the older versions of TF and keras. How can this be fixed?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169108,
          "author_name": "aamaia",
          "author_url": "",
          "post_date": "03/19/2017 15:52:23",
          "content": "<p>So, you \"cloned\" the python in /opt when creating the environment and it's using the one you specified, what's wrong? if you create the environment with \"conda create -n my_root2 python=3.5\" it will use another version.</p>\n\n<p>The second error (\"temporary failure in name resolution\") happens because the work nodes don't have internet connectivity. You need to download weights from the master node.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 169005,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "03/19/2017 03:32:35",
      "content": "<p>UPDATE: I'm trying MXNet compiled with MKL now. The version in /opt/intel/mkl is not recent enough, a more recent version of MKLML is needed by mxnet. I'm running benchmarks and it's fast but from time to time it segfaults mysteriously while training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 173363,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "04/06/2017 16:37:37",
      "content": "<p>UPDATE2:</p>\n\n<ul>\n<li><p>I/O speed is not consistent as it's a shared resource between all users. Removing this bottleneck helps improving overall performance. If you try naively to read from disk using 255 threads at the same time performance will be terrible.</p></li>\n<li><p><strong>Keras with Theano (using MKL)</strong> backend: this was the fastest I could achieve so far using Colfax, but still not acceptable performance.</p></li>\n</ul>\n\n<p>(Tensorflow is slower, even more on CPU).</p>\n\n<ul>\n<li><p>Caffe with MLSL as provided in <strong>/opt/caffe-mlsl</strong>. Performance worse than running on a single CPU. I think it's a configuration problem, only setting OMP_NUM_THREADS is not enough. I tried changing layer engine to MKL2017 but couldn't figure out (Caffe wasn't even trying to parallelize -- running \"top\" shows 100% CPU usage, i.e. only one CPU being used).</p></li>\n<li><p>Custom Caffe compiled with OpenBLAS and GCC instead of Intel compiler. In this setup setting OMP_NUM_THREADS makes \"top\" show multiple cpus are being used but it's still slow.</p></li>\n</ul>\n\n<p>Anyone had a better experience? All my attempts so far are trying to use more than one CPU within a single node. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 174256,
      "author_name": "aaalgo",
      "author_url": "",
      "post_date": "04/10/2017 20:22:54",
      "content": "<p>I tried caffe installation in /opt/intel/intelpython_update_2/intelpython2.\nWith engine=\"MKL2017\", it's about 1.5x faster than without that.  But still, the speed is only about 1/40 of GTX1060.  By setting OMP_NUM_THREADS=32, I'm gaining about 30% speedup (network too small?).  I think something is seriously wrong.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 186987,
      "author_name": "sbharti",
      "author_url": "",
      "post_date": "05/30/2017 03:50:31",
      "content": "<p>I am using tflearn / tensorflow. I set the following parameters. (64 as intel processors hae 64 cores)</p>\n\n<pre><code>export OMP_NUM_THREADS=64\n</code></pre>\n\n<p>in model code</p>\n\n<pre><code>tflearn.init_graph(num_cores=64)\n</code></pre>\n\n<p>still training on colfax is slower than my laptop cpu :-(\nThis works perfectly on my laptop but it doesnt work on colfax. \nany tips on how can we train faster on colfax using tflearn / tensorflow? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "168614": "Hi,\n\n* Caffe: my first attempt failed becase lmdb-dev is not installed on the cluster (when you try to \"import caffe\" from python it fails). I compiled lmdb locally on my account and made sure to put `LD_LIBRARY_PATH=~/tmp/lib:$LD_LIBRARY_PATH` in my PBS script but  had no luck. Also, we need to compile our own caffe version from [https://github.com/intel/caffe][1] to make use of multiple nodes. See [https://github.com/intel/caffe/wiki/Multinode-guide][2]. Maybe it didn't work because of the interaction between conda environment and LD_LIBRARY_PATH. I put this on hold.\n\n* Keras:  theano backend failed in my first attempt. Afterwards I discovered this: [https://colfaxresearch.com/discussion/topic/keras-with-theano-backend/][3]\n\n* Keras: to use tensorflow backend make sure to configure ~/.keras/keras.jon and **ALSO** write in your submission script `export env KERAS_BACKEND=tensorflow` (to override the default). I'm using keras 2.0 with tensorflow 1.0 from conda-forge. Running on a single node and a single CPU is extremely slow. More than on my local machine. I'm not sure why. Slow I/O only explains part of the picture, and I don't have the tools to investigate remotely properly . Adding threads to tensorflow session via `intra_op_parallelism_threads` does not help. \n\n* Next steps for me: trying to find the cause of the slowness, then trying to make use of the compute power, probably will need Tensorflow compiled with OpenBlas OR using  [distributed tensorflow][4] even on a single node to make use of multiple CPUs.\n \nI realize all these problems are temporary while we figure out how to work on Colfax in the initial stages of the competition and I hope that  by sharing we can overcome these problems faster.\n\n\n  [1]: https://github.com/intel/caffe\n  [2]: https://github.com/intel/caffe/wiki/Multinode-guide\n  [3]: https://colfaxresearch.com/discussion/topic/keras-with-theano-backend/\n  [4]: https://www.tensorflow.org/deploy/distributed",
    "168887": "I have a similar impressions about Colfax cluster with Keras. I use Theano as backend. However, if you specify `export OMP_NUM_THREADS=16` and use the option `openmp = True` in .theanorc, when you run a theano testing script \n\n```#PBS -N test -l nodes=4:knl``` \n\n```python /opt/theano/theano/misc/elemwise_openmp_speedup.py```\n\n you can see a theano performance report.",
    "168986": "amaia Could you share how you created a conda environment with tensorflow 1.0.1 and keras 2.0.1 on the colfax cluster? \n\nI created a conda python 3.5 environment and installed tf and keras from conda-forge but kept getting errors as the conda Python reverted to the 'system' Intel python35. Many thanks!",
    "168991": "If you're submitting a job make sure to \"source activate\" inside the PBS script. \n\nOther than this, suppose you create a conda environment, activate and then call ipython. If you didn't install ipython in the environment the one being called is the other one in $PATH, which is ```/opt/intel/intelpython35/bin/ipython```. This is confusing. You can also use the program ```which``` to see which one is being called.",
    "169005": "UPDATE: I'm trying MXNet compiled with MKL now. The version in /opt/intel/mkl is not recent enough, a more recent version of MKLML is needed by mxnet. I'm running benchmarks and it's fast but from time to time it segfaults mysteriously while training.",
    "169060": "I created a conda env using:\n\n$ conda create -n my_root --clone=/opt/intel/intelpython35\n\nAnd, then installed TF 1.0.0 and keras 2.0.1 using conda-forge.\n\nWhen running myprogram.py, the error is:\n\nTraceback (most recent call last):\n  File \"/home/u2606/VM/dev/myprogram.py\", line 498, in <module>\n    main()\n  File \"/home/u2606/VM/dev/myprogram.py\", line 355, in main\n    base_model = InceptionV3(weights='imagenet')\n  File \"/opt/intel/intelpython35/lib/python3.5/site-packages/Keras-1.1.0-py3.5.egg/keras/applications/inception_v3.py\", line 296, in InceptionV3\n    md5_hash='fe114b3ff2ea4bf891e9353d1bbfb32f')\n  File \"/opt/intel/intelpython35/lib/python3.5/site-packages/Keras-1.1.0-py3.5.egg/keras/utils/data_utils.py\", line 98, in get_file\n    raise Exception(error_msg.format(origin, e.errno, e.reason))\nException: URL fetch failure on https://github.com/fchollet/deep-learning-models/releases/download/v0.2/inception_v3_weights_tf_dim_ordering_tf_kernels.h5: None -- [Errno -3] Temporary failure in name resolution\n\nThe newly created conda env is reverting to the 'system' python and the older versions of TF and keras. How can this be fixed?",
    "169108": "So, you \"cloned\" the python in /opt when creating the environment and it's using the one you specified, what's wrong? if you create the environment with \"conda create -n my_root2 python=3.5\" it will use another version.\n\nThe second error (\"temporary failure in name resolution\") happens because the work nodes don't have internet connectivity. You need to download weights from the master node.",
    "169778": "I am training a CNN with Keras (theano) using the  parameters in vfdev's post. I am seeing a ~50x slower performance compared with my GeForce.",
    "173363": "UPDATE2:\n\n* I/O speed is not consistent as it's a shared resource between all users. Removing this bottleneck helps improving overall performance. If you try naively to read from disk using 255 threads at the same time performance will be terrible.\n\n* **Keras with Theano (using MKL)** backend: this was the fastest I could achieve so far using Colfax, but still not acceptable performance.\n\n(Tensorflow is slower, even more on CPU).\n\n* Caffe with MLSL as provided in **/opt/caffe-mlsl**. Performance worse than running on a single CPU. I think it's a configuration problem, only setting OMP_NUM_THREADS is not enough. I tried changing layer engine to MKL2017 but couldn't figure out (Caffe wasn't even trying to parallelize -- running \"top\" shows 100% CPU usage, i.e. only one CPU being used).\n\n* Custom Caffe compiled with OpenBLAS and GCC instead of Intel compiler. In this setup setting OMP_NUM_THREADS makes \"top\" show multiple cpus are being used but it's still slow.\n\n\nAnyone had a better experience? All my attempts so far are trying to use more than one CPU within a single node.",
    "174256": "I tried caffe installation in /opt/intel/intelpython_update_2/intelpython2.\nWith engine=\"MKL2017\", it's about 1.5x faster than without that.  But still, the speed is only about 1/40 of GTX1060.  By setting OMP_NUM_THREADS=32, I'm gaining about 30% speedup (network too small?).  I think something is seriously wrong.",
    "186987": "I am using tflearn / tensorflow. I set the following parameters. (64 as intel processors hae 64 cores)\n\n    export OMP_NUM_THREADS=64\n\nin model code\n\n    tflearn.init_graph(num_cores=64)\n\nstill training on colfax is slower than my laptop cpu :-(\nThis works perfectly on my laptop but it doesnt work on colfax. \nany tips on how can we train faster on colfax using tflearn / tensorflow?"
  },
  "source": "meta"
}