{
  "id": 30186,
  "title": "Nvidia+TF vs Intel+Caffe",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30186",
  "author_name": "",
  "post_date": "2017-03-16T10:55:36.819217300Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi!</p>\n\n<p>Nice and interesting competition!</p>\n\n<p>I want to see your opinions (or better, facts) about the Intel infrastructure. Intel Phi is a newbie in the field. BUT even 16G for your models is more than you can get on the latest Titan X, and more than you can get in your regular home computer powered by some older nvidia card. In 90G of memory you can load all the data and cross validate the hell out of your model!</p>\n\n<p>So, my question is: From a holistic perspective, what is the fastest platform to work on? </p>\n\n<ul>\n<li>Nvidia powered ami instances? (costs $ but you can run your existing tf+keras or caffe) at known speeds and with known restrictions. </li>\n<li>Intel platform (free (!) ) but with high unknowns wrt to brute ML performance (64 cores can't beat 1000 cores) and overall performance gains (90G  RAM vs batch loading data) and operation schedule.  I saw in DSTL competition that on the last days the submission server couldn't handle all the requests. What about here, is there a risk of running out of resources at the end of the competition? And wait forever for your latest model to be run?</li>\n</ul>\n\n<p>Another thing to consider:\n -  According to tf github, they merged support for AVX512 architecture. Have smb tried it? Can you seamlessly take your (python) code from home and run it on Intel Phi?\n - What is the learning curve to deploy on the intel free instances? There are A LOT of videos and documentation there. </p>\n\n<p>I am asking because working remotely on an always on, free instance is VERY appealing but it takes time until you setup/run all your pipeline.</p>\n\n<p>Thanks!</p>\n\n<p><strong>EDIT</strong></p>\n\n<p>The conda setup on colfax HAS both tensorflow and theano. IMHO tensorflow is optimized for hi architecture. At least Anaconda have paid distribution for Intel platform. Downside is that there is only tensorflow 0.12 but that's ok.</p>",
  "messages": [
    {
      "id": "168135",
      "postDate": "03/16/2017 10:55:36",
      "content": "<p>Hi!</p>\n\n<p>Nice and interesting competition!</p>\n\n<p>I want to see your opinions (or better, facts) about the Intel infrastructure. Intel Phi is a newbie in the field. BUT even 16G for your models is more than you can get on the latest Titan X, and more than you can get in your regular home computer powered by some older nvidia card. In 90G of memory you can load all the data and cross validate the hell out of your model!</p>\n\n<p>So, my question is: From a holistic perspective, what is the fastest platform to work on? </p>\n\n<ul>\n<li>Nvidia powered ami instances? (costs $ but you can run your existing tf+keras or caffe) at known speeds and with known restrictions. </li>\n<li>Intel platform (free (!) ) but with high unknowns wrt to brute ML performance (64 cores can't beat 1000 cores) and overall performance gains (90G  RAM vs batch loading data) and operation schedule.  I saw in DSTL competition that on the last days the submission server couldn't handle all the requests. What about here, is there a risk of running out of resources at the end of the competition? And wait forever for your latest model to be run?</li>\n</ul>\n\n<p>Another thing to consider:\n -  According to tf github, they merged support for AVX512 architecture. Have smb tried it? Can you seamlessly take your (python) code from home and run it on Intel Phi?\n - What is the learning curve to deploy on the intel free instances? There are A LOT of videos and documentation there. </p>\n\n<p>I am asking because working remotely on an always on, free instance is VERY appealing but it takes time until you setup/run all your pipeline.</p>\n\n<p>Thanks!</p>\n\n<p><strong>EDIT</strong></p>\n\n<p>The conda setup on colfax HAS both tensorflow and theano. IMHO tensorflow is optimized for hi architecture. At least Anaconda have paid distribution for Intel platform. Downside is that there is only tensorflow 0.12 but that's ok.</p>",
      "rawMarkdown": "Hi!\n\nNice and interesting competition!\n\nI want to see your opinions (or better, facts) about the Intel infrastructure. Intel Phi is a newbie in the field. BUT even 16G for your models is more than you can get on the latest Titan X, and more than you can get in your regular home computer powered by some older nvidia card. In 90G of memory you can load all the data and cross validate the hell out of your model!\n\nSo, my question is: From a holistic perspective, what is the fastest platform to work on? \n\n - Nvidia powered ami instances? (costs $ but you can run your existing tf+keras or caffe) at known speeds and with known restrictions. \n - Intel platform (free (!) ) but with high unknowns wrt to brute ML performance (64 cores can't beat 1000 cores) and overall performance gains (90G  RAM vs batch loading data) and operation schedule.  I saw in DSTL competition that on the last days the submission server couldn't handle all the requests. What about here, is there a risk of running out of resources at the end of the competition? And wait forever for your latest model to be run?\n\nAnother thing to consider:\n -  According to tf github, they merged support for AVX512 architecture. Have smb tried it? Can you seamlessly take your (python) code from home and run it on Intel Phi?\n - What is the learning curve to deploy on the intel free instances? There are A LOT of videos and documentation there. \n\nI am asking because working remotely on an always on, free instance is VERY appealing but it takes time until you setup/run all your pipeline.\n\nThanks!\n\n **EDIT**\n\nThe conda setup on colfax HAS both tensorflow and theano. IMHO tensorflow is optimized for hi architecture. At least Anaconda have paid distribution for Intel platform. Downside is that there is only tensorflow 0.12 but that's ok.",
      "votes": null
    },
    {
      "id": "168149",
      "postDate": "03/16/2017 11:56:53",
      "content": "<pre><code>Queue batch\n    queue_type = Execution\n    total_jobs = 1\n    state_count = Transit:0 Queued:0 Held:0 Waiting:0 Running:1 Exiting:0 Complete:0 \n    resources_default.nodes = 1\n    resources_default.walltime = 24:00:00\n    mtime = Wed Mar 15 12:53:44 2017\n    resources_assigned.nodect = 1\n    enabled = True\n    started = True\n</code></pre>\n\n<p>24 hrs per job. Ok, that should be enough in the beginning. And there are 44 nodes if I am not wrong.</p>",
      "rawMarkdown": "Queue batch\n    \tqueue_type = Execution\n    \ttotal_jobs = 1\n    \tstate_count = Transit:0 Queued:0 Held:0 Waiting:0 Running:1 Exiting:0 Complete:0 \n    \tresources_default.nodes = 1\n    \tresources_default.walltime = 24:00:00\n    \tmtime = Wed Mar 15 12:53:44 2017\n    \tresources_assigned.nodect = 1\n    \tenabled = True\n    \tstarted = True\n\n24 hrs per job. Ok, that should be enough in the beginning. And there are 44 nodes if I am not wrong.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 168149,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "03/16/2017 11:56:53",
      "content": "<pre><code>Queue batch\n    queue_type = Execution\n    total_jobs = 1\n    state_count = Transit:0 Queued:0 Held:0 Waiting:0 Running:1 Exiting:0 Complete:0 \n    resources_default.nodes = 1\n    resources_default.walltime = 24:00:00\n    mtime = Wed Mar 15 12:53:44 2017\n    resources_assigned.nodect = 1\n    enabled = True\n    started = True\n</code></pre>\n\n<p>24 hrs per job. Ok, that should be enough in the beginning. And there are 44 nodes if I am not wrong.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "168135": "Hi!\n\nNice and interesting competition!\n\nI want to see your opinions (or better, facts) about the Intel infrastructure. Intel Phi is a newbie in the field. BUT even 16G for your models is more than you can get on the latest Titan X, and more than you can get in your regular home computer powered by some older nvidia card. In 90G of memory you can load all the data and cross validate the hell out of your model!\n\nSo, my question is: From a holistic perspective, what is the fastest platform to work on? \n\n - Nvidia powered ami instances? (costs $ but you can run your existing tf+keras or caffe) at known speeds and with known restrictions. \n - Intel platform (free (!) ) but with high unknowns wrt to brute ML performance (64 cores can't beat 1000 cores) and overall performance gains (90G  RAM vs batch loading data) and operation schedule.  I saw in DSTL competition that on the last days the submission server couldn't handle all the requests. What about here, is there a risk of running out of resources at the end of the competition? And wait forever for your latest model to be run?\n\nAnother thing to consider:\n -  According to tf github, they merged support for AVX512 architecture. Have smb tried it? Can you seamlessly take your (python) code from home and run it on Intel Phi?\n - What is the learning curve to deploy on the intel free instances? There are A LOT of videos and documentation there. \n\nI am asking because working remotely on an always on, free instance is VERY appealing but it takes time until you setup/run all your pipeline.\n\nThanks!\n\n **EDIT**\n\nThe conda setup on colfax HAS both tensorflow and theano. IMHO tensorflow is optimized for hi architecture. At least Anaconda have paid distribution for Intel platform. Downside is that there is only tensorflow 0.12 but that's ok.",
    "168149": "Queue batch\n    \tqueue_type = Execution\n    \ttotal_jobs = 1\n    \tstate_count = Transit:0 Queued:0 Held:0 Waiting:0 Running:1 Exiting:0 Complete:0 \n    \tresources_default.nodes = 1\n    \tresources_default.walltime = 24:00:00\n    \tmtime = Wed Mar 15 12:53:44 2017\n    \tresources_assigned.nodect = 1\n    \tenabled = True\n    \tstarted = True\n\n24 hrs per job. Ok, that should be enough in the beginning. And there are 44 nodes if I am not wrong."
  },
  "source": "meta"
}