{
  "id": 4459,
  "title": "PYLEARN2_DATA_PATH ERROR",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4459",
  "author_name": "",
  "post_date": "2013-04-28T16:37:01.960Z",
  "votes": null,
  "comment_count": 27,
  "views": 10692,
  "content": "<p>Hi All,</p>\r\n<p>I get this error while trying out the sample submission python script:</p>\r\n<p>&quot;pylearn2.utils.string_utils.EnvironmentVariableError: You need to define your PYLEARN2_DATA_PATH environment variable. If you are using a computer at LISA, this should be set to /data/lisa/data&quot;</p>\r\n<p>I am using Ubuntu and I already added below line to my .bashrc file: export PYLEARN2_DATA_PATH=/home/kirana/myDocs/blackbox/icml_2013_black_box</p>\r\n<p>&nbsp;</p>\r\n<p>Our score so far is based on R. Trying to use python and must say pylearn2 is terrible. So non intuitive. Any tips on how to get round this pylearn2_data_path error is much appreciated</p>\r\n<p>Thanks<br>\r\nKiran</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p><br>\r\n<br>\r\n</p>",
  "messages": [
    {
      "id": "23572",
      "postDate": "04/28/2013 16:37:01",
      "content": "<p>Hi All,</p>\r\n<p>I get this error while trying out the sample submission python script:</p>\r\n<p>&quot;pylearn2.utils.string_utils.EnvironmentVariableError: You need to define your PYLEARN2_DATA_PATH environment variable. If you are using a computer at LISA, this should be set to /data/lisa/data&quot;</p>\r\n<p>I am using Ubuntu and I already added below line to my .bashrc file: export PYLEARN2_DATA_PATH=/home/kirana/myDocs/blackbox/icml_2013_black_box</p>\r\n<p>&nbsp;</p>\r\n<p>Our score so far is based on R. Trying to use python and must say pylearn2 is terrible. So non intuitive. Any tips on how to get round this pylearn2_data_path error is much appreciated</p>\r\n<p>Thanks<br>\r\nKiran</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p><br>\r\n<br>\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23581",
      "postDate": "04/28/2013 19:16:29",
      "content": "<p>You have two problems.</p>\r\n<p>One, as the README explains, it's going to look for ${PYLEARN2_DATA_PATH}/icml_2013_black_box. You've linked directly to icml_2013_black_box, instead of its parent. PYLEARN2_DATA_PATH is meant to be a general subdirectory where all of the datasets you want\r\n to use with pylearn2 can be found. (If you're doing some of the other kaggle contests, you should put the data for those in there too). If you don't want to actually move all of the datasets for use with pylearn2 into one directory, you can just make the directory\r\n and put symlinks to the various datasets in it.</p>\r\n<p>The other problem is that your environment variable isn't getting set correctly. There are a lot of rookie mistakes people make when trying to set environment variables:</p>\r\n<p>-Just editing your .bashrc doesn't change the environment variable. You have to open a new terminal window, or run &quot;source ~/.bashrc&quot;</p>\r\n<p>-Only terminal windows that have been opened, or had source ~/.bashrc run in them, will have the right environment variable. So if you're running pylearn2 in a terminal window that's been open for a while, it won't see the updated value.</p>\r\n<p>-If you're running pylearn2 out of an IDE or an ipython notebook, it probably needs to be restarted. An IDE might also have some settings that override your .bashrc's environment variables.</p>\r\n<p>-If you run pylearn2 using sudo, it will see the root's environment variables, not your users. Running pylearn2 with sudo isn't recommended, but someone else on the Kaggle forum was having trouble with environment variables, and that was the cause.</p>\r\n<p>You can check to see if the environment variable has been set in your shell by running:</p>\r\n<p>echo ${PYLEARN2_DATA_PATH}</p>\r\n<p>You can check to see if it's visible to the python interpreter by running:</p>\r\n<p>python -c &quot;import os; print os.environ['PYLEARN2_DATA_PATH']&quot;</p>\r\n<p>If the python - c command works, make sure you're running pylearn2 in exactly the same way (same terminal window, not using sudo, same version of python, etc)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23609",
      "postDate": "04/29/2013 03:52:15",
      "content": "<p>My data is in a subdirectory icml_2013_black_box within the&nbsp;/home/kirana/myDocs/blackbox/ folder. So we are good.</p>\r\n<p><span>export PYLEARN2_DATA_PATH=/home/kirana/myDocs/blackbox/icml_2013_black_box</span></p>\r\n<p>&nbsp;</p>\r\n<p>I not only edited .bashrc but also restarted my system. wehn I did $PYLEARN2_DATA_PATH on my system, it still worked</p>\r\n<p>&nbsp;</p>\r\n<p>Even then it is not working.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23610",
      "postDate": "04/29/2013 04:23:44",
      "content": "<p>I don't think you understood what I was saying. If you want it to work where the files are now your command needs to be&nbsp;</p>\r\n<p><span>export PYLEARN2_DATA_PATH=/home/kirana/myDocs/blackbox</span></p>\r\n<p>not</p>\r\n<p><span>export PYLEARN2_DATA_PATH=/home/kirana/myDocs/blackbox/icml_2013_black_box</span></p>\r\n<p>Restarting your system is overkill. You don't need to do that. What do you mean by &quot;<span>wehn I did $PYLEARN2_DATA_PATH on my system, it still worked&quot;? Did you use echo to print the environment variable, or python -c to see if it's visible to the python\r\n interpreter?</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23611",
      "postDate": "04/29/2013 05:24:28",
      "content": "<p>I am doing the same thing:</p>\r\n<p>&nbsp;</p>\r\n<p>I have an icml_2013_black_box folder under icml_2013_blackbox where I keep the data. So it is the same thing:</p>\r\n<p>echo $PYLEARN2_DATA_PATH is working.</p>\r\n<p>&nbsp;</p>\r\n<p>Here is the error I am getting. I have the variable set correctly</p>\r\n<p>&nbsp;&nbsp; raise EnvironmentVariableError(&quot;You need to define your PYLEARN2_DATA_PATH environment variable. If you are using a computer at LISA, this should be set to /data/lisa/data&quot;)<br>\r\npylearn2.utils.string_utils.EnvironmentVariableError: You need to define your PYLEARN2_DATA_PATH environment variable. If you are using a computer at LISA, this should be set to /data/lisa/data<br>\r\n<br>\r\n<br>\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23612",
      "postDate": "04/29/2013 05:28:17",
      "content": "<p>it gives an assertion error without sudo:</p>\r\n<p>&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/scripts/icml_2013_wrepl/black_box/black_box_dataset.py&quot;, line 79, in __init__<br>\r\n&nbsp;&nbsp;&nbsp; assert stop &lt;= X.shape[0]<br>\r\nAssertionError</p>\r\n<p>&nbsp;</p>\r\n<p>and when I use sudo it gives the error above.</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>I have followed all instructions correctly<br>\r\n<br>\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23620",
      "postDate": "04/29/2013 14:14:38",
      "content": "<p>Don't use sudo.</p>\r\n<p>Can you explain why you were using sudo in the first place? You're the second Kaggle competittor to do that, and I have no idea why anyone would want to. It's a very risky thing to do, and I can't see what you'd hope to gain my doing it.</p>\r\n<p>As I said above, if you use sudo, python will get the environment variables for your root account, not your current user. So your ~/.bashrc has nothing to do with what gets passed to pylearn2.</p>\r\n<p>Can you give the full backtrace for the AssertionError? It looks like either you've modified the yaml file to use a different &quot;stop&quot; argument, or your data files have gotten truncated somehow.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23623",
      "postDate": "04/29/2013 14:45:47",
      "content": "<p>Ok - here is the backtrace (without sudo)</p>\r\n<p>&nbsp;</p>\r\n<p>/pylearn2/scripts/icml_2013_wrepl/black_box/mlp.yaml')<br>\r\n/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/models/mlp.py:40: UserWarning: MLP changing the recursion limit.<br>\r\n&nbsp; warnings.warn(&quot;MLP changing the recursion limit.&quot;)<br>\r\nTraceback (most recent call last):<br>\r\n&nbsp; File &quot;/home/kirana/pylearn2/pylearn2/scripts/train.py&quot;, line 95, in &lt;module&gt;<br>\r\n&nbsp;&nbsp;&nbsp; train_obj = serial.load_train_file(args.config)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/utils/serial.py&quot;, line 435, in load_train_file<br>\r\n&nbsp;&nbsp;&nbsp; return yaml_parse.load_path(config_file_path)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 87, in load_path<br>\r\n&nbsp;&nbsp;&nbsp; return load(content, **kwargs)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 54, in load<br>\r\n&nbsp;&nbsp;&nbsp; return instantiate_all(proxy_graph)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 141, in instantiate_all<br>\r\n&nbsp;&nbsp;&nbsp; graph[key] = instantiate_all(graph[key])<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 141, in instantiate_all<br>\r\n&nbsp;&nbsp;&nbsp; graph[key] = instantiate_all(graph[key])<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 141, in instantiate_all<br>\r\n&nbsp;&nbsp;&nbsp; graph[key] = instantiate_all(graph[key])<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 150, in instantiate_all<br>\r\n&nbsp;&nbsp;&nbsp; graph = graph.instantiate()<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 192, in instantiate<br>\r\n&nbsp;&nbsp;&nbsp; self.instance = checked_call(self.cls, self.kwds)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/utils/call_check.py&quot;, line 98, in checked_call<br>\r\n&nbsp;&nbsp;&nbsp; return to_call(**kwargs)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/scripts/icml_2013_wrepl/black_box/black_box_dataset.py&quot;, line 79, in __init__<br>\r\n&nbsp;&nbsp;&nbsp; assert stop &lt;= X.shape[0]<br>\r\nAssertionError<br>\r\n<br>\r\n<br>\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23631",
      "postDate": "04/29/2013 16:07:36",
      "content": "<pre>What happens if you run md5sum on your train.csv file? Here's what I get:</pre>\r\n<pre>md5sum train.csv <br>3a2f6711b34a96d54762db3ea785c66c  train.csv</pre>\r\n<pre>If you don't get the same thing, you should download it again. If you do get the same thing, write back and I'll think of something else to check.</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23649",
      "postDate": "04/29/2013 17:54:49",
      "content": "<p>yes it is the same:</p>\r\n<pre>kirana@kiran-pc:~/myDocs/blackbox/icml_2013_black_box$ md5sum train.csv<br>3a2f6711b34a96d54762db3ea785c66c  train.csv</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23650",
      "postDate": "04/29/2013 17:57:51",
      "content": "<p>Try using &quot;git pull&quot; to update pylearn2 and run it again. I've added some checking that should give a more informative error message. It will still crash, but paste the new error message back and I'll have a better idea of what's going on.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23652",
      "postDate": "04/29/2013 18:37:15",
      "content": "<p>Thanks Ian - working now.</p>\r\n<p>You are a rockstar. Many thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23653",
      "postDate": "04/29/2013 19:03:12",
      "content": "<p>Glad it's working.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23672",
      "postDate": "04/29/2013 22:14:32",
      "content": "<p>how do i extract the predicted &quot;probabilities&quot; for each class in the test set? How do i specify the test set (in case i want to predict for another test set)?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23681",
      "postDate": "04/30/2013 04:05:03",
      "content": "<p>Look at the make_submission.py script.</p>\r\n<p>Line 52,&nbsp;<span class=\"x_n\">Y</span><span class=\"x_o\">=</span><span class=\"x_n\">model</span><span class=\"x_o\">.</span><span class=\"x_n\">fprop</span><span class=\"x_p\">(</span><span class=\"x_n\">X</span><span class=\"x_p\">), computes the probabilities. Y[i,j]\r\n gives the probability that example i (in row i of X) belongs to class j.</span></p>\r\n<p>Line 34 makes the dataset. If you have a different dataset object you want to use here, just call its constructor instead of using the get_test_set method.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23702",
      "postDate": "04/30/2013 15:44:57",
      "content": "<p>How do i covert the Y matrix to float?</p>\r\n<p>I getting the following error:</p>\r\n<pre> out.write('%f' % (Y[i,j]))<br>TypeError: float argument required, not TensorVariable</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23703",
      "postDate": "04/30/2013 15:54:53",
      "content": "<p>Y is an algebraic variable and doesn't actually have a numerical value at this point in time. Read through the Theano basic tutorial to get some idea of how it works: http://deeplearning.net/software/theano/tutorial/</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23705",
      "postDate": "04/30/2013 15:59:51",
      "content": "<p>Ian,</p>\r\n<p>&nbsp; &nbsp; I don't know python, and i just want to convert that valua to a float in order to print it. I read the code to convert it to class and couldn't understand...</p>\r\n<p>&nbsp; &nbsp; How do i convert that matrix to float? Already looked at that tutorial and i found it very confusing...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23706",
      "postDate": "04/30/2013 16:04:43",
      "content": "<p>What you are asking doesn't make any sense.</p>\r\n<p>Suppose we have an equation:</p>\r\n<p>2x &#43; 3 = 7y</p>\r\n<p>You can't convert y to a float, because y doesn't have a specific floating point value. Its value depends on x. If you don't know x, you don't know y.</p>\r\n<p>That's what a TensorVariable is--it represents an algebraic variable of unknown value. You can't just &quot;convert&quot; it. You can compute a specific value of the variable given a set of specific values of the variables that it depends on.</p>\r\n<p>Can you tell me more about what you're trying to do? I can probably tell you a better way of going about it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23707",
      "postDate": "04/30/2013 16:08:55",
      "content": "<p>do the make submission.py output an matrix with nx9 (10000x9 for the test set) with the probabilities instead of the major problable class.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23708",
      "postDate": "04/30/2013 16:13:27",
      "content": "<p>I think this is it (Thanks!):</p>\r\n<p>&nbsp;</p>\r\n<pre>X = model.get_input_space().make_batch_theano()<br>Y = model.fprop(X)<br><br>from theano import tensor as T<br>from theano import function<br>f = function([X], Y)<br><br>y = []<br>for i in xrange(dataset.X.shape[0] / batch_size):<br>    x_arg = dataset.X[i*batch_size:(i&#43;1)*batch_size,:]<br>    if X.ndim &gt; 2:<br>        x_arg = dataset.get_topological_view(x_arg)<br>    y.append(f(x_arg.astype(X.dtype)))<br><br>y = np.concatenate(y)<br><br>#assert y.ndim == 9<br>#assert y.shape[0] == dataset.X.shape[0]<br><br>y = y[:m,:]</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23710",
      "postDate": "04/30/2013 16:21:30",
      "content": "<p>Are you just doing that for debugging purposes? If you upload a file that says anything but &quot;1.0&quot;, &quot;2.0&quot;, etc. Kaggle will give you 0 points.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23714",
      "postDate": "04/30/2013 16:30:49",
      "content": "<p>Nope. I just wanna run some analysis on the predictions. And i don't know enough python (i don't know python at all to be honest) to do this directly.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23719",
      "postDate": "04/30/2013 18:09:37",
      "content": "<p>By the way, can i run more than one training in paralell? Tried it but theano complained. Are the training proccess using gpus?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23721",
      "postDate": "04/30/2013 18:30:03",
      "content": "<p>If you have your THEANO_FLAGS or .theanorc set to use GPU, then it will use GPU. pylearn2 doesn't control whether you use GPU or not. pylearn2 doesn't impose any limitations on what pylearn2 jobs you can run in parallel. Theano will only let you run as many\r\n GPU jobs as you have GPUs for.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23722",
      "postDate": "04/30/2013 18:31:57",
      "content": "<p>when i try to run in parallel, i get this:&nbsp;<span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\">INFO (theano.gof.compilelock): Waiting for existing lock by process '12975' (I am process '12966')\r\n</span><span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\">INFO (theano.gof.compilelock): To manually release the lock, delete&nbsp;</span></p>\r\n<p><span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\">I just want to know if it is a limitation, or a configuration. Can that mutual exclusiviness be disabled? is it safe?</span></p>\r\n<p><span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\">those are my questions...</span></p>\r\n<p><span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\"><br>\r\n</span></p>\r\n<pre class=\"x_GNVMTOMCABB\">&nbsp;</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23723",
      "postDate": "04/30/2013 18:41:46",
      "content": "<p>Theano actually will execute the learning experiment in parallel. It's only the initial compilation that's getting serialized. If you wait for a while they will both run.</p>\r\n<p>What's going on is theano keeps a directory on disk where it caches the output of its calls to gcc. The problem is that whoever wrote theano's compilation mechanism doesn't seem to have understood locks and parallelism at all, so it is hopelessly inefficient.\r\n Whenever theano tries to acquire the lock and fails, it sleeps for about 5 seconds before trying again. This means that running two jobs in parallel can more than double the compilation time.</p>\r\n<p>For the long term, I've talked with the theano developers about switching to a lock-free design that should eliminate these synchronization issues. I have no idea how long it will be before anyone has time to implement that.</p>\r\n<p>In the short term, you can check the theano documentation to see how to control the directory that's used for the compilation cache. If you use an environment variable to set it to a different directory for job, then you'll lose the benefit of the cache,\r\n but you also won't have to pay the cost of all your jobs sleeping for no good reason.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "135753",
      "postDate": "09/16/2016 16:45:34",
      "content": "<p>Hi all \nI dont know whether it is the right forum or not. But i have been struggling to set this python data path variable in windows and when i search for this thing this forums is one of the few places where such things are mentioned.\nany help would be appreciated i am student of bachelors and i am trying to do a project from kaggle  as my final year project.\nregards\nwaleed sial</p>",
      "rawMarkdown": "Hi all \r\nI dont know whether it is the right forum or not. But i have been struggling to set this python data path variable in windows and when i search for this thing this forums is one of the few places where such things are mentioned.\r\nany help would be appreciated i am student of bachelors and i am trying to do a project from kaggle  as my final year project.\r\nregards\r\nwaleed sial",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 23581,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/28/2013 19:16:29",
      "content": "<p>You have two problems.</p>\r\n<p>One, as the README explains, it's going to look for ${PYLEARN2_DATA_PATH}/icml_2013_black_box. You've linked directly to icml_2013_black_box, instead of its parent. PYLEARN2_DATA_PATH is meant to be a general subdirectory where all of the datasets you want\r\n to use with pylearn2 can be found. (If you're doing some of the other kaggle contests, you should put the data for those in there too). If you don't want to actually move all of the datasets for use with pylearn2 into one directory, you can just make the directory\r\n and put symlinks to the various datasets in it.</p>\r\n<p>The other problem is that your environment variable isn't getting set correctly. There are a lot of rookie mistakes people make when trying to set environment variables:</p>\r\n<p>-Just editing your .bashrc doesn't change the environment variable. You have to open a new terminal window, or run &quot;source ~/.bashrc&quot;</p>\r\n<p>-Only terminal windows that have been opened, or had source ~/.bashrc run in them, will have the right environment variable. So if you're running pylearn2 in a terminal window that's been open for a while, it won't see the updated value.</p>\r\n<p>-If you're running pylearn2 out of an IDE or an ipython notebook, it probably needs to be restarted. An IDE might also have some settings that override your .bashrc's environment variables.</p>\r\n<p>-If you run pylearn2 using sudo, it will see the root's environment variables, not your users. Running pylearn2 with sudo isn't recommended, but someone else on the Kaggle forum was having trouble with environment variables, and that was the cause.</p>\r\n<p>You can check to see if the environment variable has been set in your shell by running:</p>\r\n<p>echo ${PYLEARN2_DATA_PATH}</p>\r\n<p>You can check to see if it's visible to the python interpreter by running:</p>\r\n<p>python -c &quot;import os; print os.environ['PYLEARN2_DATA_PATH']&quot;</p>\r\n<p>If the python - c command works, make sure you're running pylearn2 in exactly the same way (same terminal window, not using sudo, same version of python, etc)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23609,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/29/2013 03:52:15",
      "content": "<p>My data is in a subdirectory icml_2013_black_box within the&nbsp;/home/kirana/myDocs/blackbox/ folder. So we are good.</p>\r\n<p><span>export PYLEARN2_DATA_PATH=/home/kirana/myDocs/blackbox/icml_2013_black_box</span></p>\r\n<p>&nbsp;</p>\r\n<p>I not only edited .bashrc but also restarted my system. wehn I did $PYLEARN2_DATA_PATH on my system, it still worked</p>\r\n<p>&nbsp;</p>\r\n<p>Even then it is not working.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23610,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 04:23:44",
      "content": "<p>I don't think you understood what I was saying. If you want it to work where the files are now your command needs to be&nbsp;</p>\r\n<p><span>export PYLEARN2_DATA_PATH=/home/kirana/myDocs/blackbox</span></p>\r\n<p>not</p>\r\n<p><span>export PYLEARN2_DATA_PATH=/home/kirana/myDocs/blackbox/icml_2013_black_box</span></p>\r\n<p>Restarting your system is overkill. You don't need to do that. What do you mean by &quot;<span>wehn I did $PYLEARN2_DATA_PATH on my system, it still worked&quot;? Did you use echo to print the environment variable, or python -c to see if it's visible to the python\r\n interpreter?</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23611,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/29/2013 05:24:28",
      "content": "<p>I am doing the same thing:</p>\r\n<p>&nbsp;</p>\r\n<p>I have an icml_2013_black_box folder under icml_2013_blackbox where I keep the data. So it is the same thing:</p>\r\n<p>echo $PYLEARN2_DATA_PATH is working.</p>\r\n<p>&nbsp;</p>\r\n<p>Here is the error I am getting. I have the variable set correctly</p>\r\n<p>&nbsp;&nbsp; raise EnvironmentVariableError(&quot;You need to define your PYLEARN2_DATA_PATH environment variable. If you are using a computer at LISA, this should be set to /data/lisa/data&quot;)<br>\r\npylearn2.utils.string_utils.EnvironmentVariableError: You need to define your PYLEARN2_DATA_PATH environment variable. If you are using a computer at LISA, this should be set to /data/lisa/data<br>\r\n<br>\r\n<br>\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23612,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/29/2013 05:28:17",
      "content": "<p>it gives an assertion error without sudo:</p>\r\n<p>&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/scripts/icml_2013_wrepl/black_box/black_box_dataset.py&quot;, line 79, in __init__<br>\r\n&nbsp;&nbsp;&nbsp; assert stop &lt;= X.shape[0]<br>\r\nAssertionError</p>\r\n<p>&nbsp;</p>\r\n<p>and when I use sudo it gives the error above.</p>\r\n<p>&nbsp;</p>\r\n<p>&nbsp;</p>\r\n<p>I have followed all instructions correctly<br>\r\n<br>\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23620,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 14:14:38",
      "content": "<p>Don't use sudo.</p>\r\n<p>Can you explain why you were using sudo in the first place? You're the second Kaggle competittor to do that, and I have no idea why anyone would want to. It's a very risky thing to do, and I can't see what you'd hope to gain my doing it.</p>\r\n<p>As I said above, if you use sudo, python will get the environment variables for your root account, not your current user. So your ~/.bashrc has nothing to do with what gets passed to pylearn2.</p>\r\n<p>Can you give the full backtrace for the AssertionError? It looks like either you've modified the yaml file to use a different &quot;stop&quot; argument, or your data files have gotten truncated somehow.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23623,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/29/2013 14:45:47",
      "content": "<p>Ok - here is the backtrace (without sudo)</p>\r\n<p>&nbsp;</p>\r\n<p>/pylearn2/scripts/icml_2013_wrepl/black_box/mlp.yaml')<br>\r\n/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/models/mlp.py:40: UserWarning: MLP changing the recursion limit.<br>\r\n&nbsp; warnings.warn(&quot;MLP changing the recursion limit.&quot;)<br>\r\nTraceback (most recent call last):<br>\r\n&nbsp; File &quot;/home/kirana/pylearn2/pylearn2/scripts/train.py&quot;, line 95, in &lt;module&gt;<br>\r\n&nbsp;&nbsp;&nbsp; train_obj = serial.load_train_file(args.config)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/utils/serial.py&quot;, line 435, in load_train_file<br>\r\n&nbsp;&nbsp;&nbsp; return yaml_parse.load_path(config_file_path)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 87, in load_path<br>\r\n&nbsp;&nbsp;&nbsp; return load(content, **kwargs)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 54, in load<br>\r\n&nbsp;&nbsp;&nbsp; return instantiate_all(proxy_graph)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 141, in instantiate_all<br>\r\n&nbsp;&nbsp;&nbsp; graph[key] = instantiate_all(graph[key])<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 141, in instantiate_all<br>\r\n&nbsp;&nbsp;&nbsp; graph[key] = instantiate_all(graph[key])<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 141, in instantiate_all<br>\r\n&nbsp;&nbsp;&nbsp; graph[key] = instantiate_all(graph[key])<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 150, in instantiate_all<br>\r\n&nbsp;&nbsp;&nbsp; graph = graph.instantiate()<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/config/yaml_parse.py&quot;, line 192, in instantiate<br>\r\n&nbsp;&nbsp;&nbsp; self.instance = checked_call(self.cls, self.kwds)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/utils/call_check.py&quot;, line 98, in checked_call<br>\r\n&nbsp;&nbsp;&nbsp; return to_call(**kwargs)<br>\r\n&nbsp; File &quot;/usr/local/lib/python2.7/dist-packages/pylearn2-0.1dev-py2.7.egg/pylearn2/scripts/icml_2013_wrepl/black_box/black_box_dataset.py&quot;, line 79, in __init__<br>\r\n&nbsp;&nbsp;&nbsp; assert stop &lt;= X.shape[0]<br>\r\nAssertionError<br>\r\n<br>\r\n<br>\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23631,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 16:07:36",
      "content": "<pre>What happens if you run md5sum on your train.csv file? Here's what I get:</pre>\r\n<pre>md5sum train.csv <br>3a2f6711b34a96d54762db3ea785c66c  train.csv</pre>\r\n<pre>If you don't get the same thing, you should download it again. If you do get the same thing, write back and I'll think of something else to check.</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23649,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/29/2013 17:54:49",
      "content": "<p>yes it is the same:</p>\r\n<pre>kirana@kiran-pc:~/myDocs/blackbox/icml_2013_black_box$ md5sum train.csv<br>3a2f6711b34a96d54762db3ea785c66c  train.csv</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23650,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 17:57:51",
      "content": "<p>Try using &quot;git pull&quot; to update pylearn2 and run it again. I've added some checking that should give a more informative error message. It will still crash, but paste the new error message back and I'll have a better idea of what's going on.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23652,
      "author_name": "rkirana",
      "author_url": "",
      "post_date": "04/29/2013 18:37:15",
      "content": "<p>Thanks Ian - working now.</p>\r\n<p>You are a rockstar. Many thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23653,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 19:03:12",
      "content": "<p>Glad it's working.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23672,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/29/2013 22:14:32",
      "content": "<p>how do i extract the predicted &quot;probabilities&quot; for each class in the test set? How do i specify the test set (in case i want to predict for another test set)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23681,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 04:05:03",
      "content": "<p>Look at the make_submission.py script.</p>\r\n<p>Line 52,&nbsp;<span class=\"x_n\">Y</span><span class=\"x_o\">=</span><span class=\"x_n\">model</span><span class=\"x_o\">.</span><span class=\"x_n\">fprop</span><span class=\"x_p\">(</span><span class=\"x_n\">X</span><span class=\"x_p\">), computes the probabilities. Y[i,j]\r\n gives the probability that example i (in row i of X) belongs to class j.</span></p>\r\n<p>Line 34 makes the dataset. If you have a different dataset object you want to use here, just call its constructor instead of using the get_test_set method.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23702,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/30/2013 15:44:57",
      "content": "<p>How do i covert the Y matrix to float?</p>\r\n<p>I getting the following error:</p>\r\n<pre> out.write('%f' % (Y[i,j]))<br>TypeError: float argument required, not TensorVariable</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23703,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 15:54:53",
      "content": "<p>Y is an algebraic variable and doesn't actually have a numerical value at this point in time. Read through the Theano basic tutorial to get some idea of how it works: http://deeplearning.net/software/theano/tutorial/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23705,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/30/2013 15:59:51",
      "content": "<p>Ian,</p>\r\n<p>&nbsp; &nbsp; I don't know python, and i just want to convert that valua to a float in order to print it. I read the code to convert it to class and couldn't understand...</p>\r\n<p>&nbsp; &nbsp; How do i convert that matrix to float? Already looked at that tutorial and i found it very confusing...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23706,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 16:04:43",
      "content": "<p>What you are asking doesn't make any sense.</p>\r\n<p>Suppose we have an equation:</p>\r\n<p>2x &#43; 3 = 7y</p>\r\n<p>You can't convert y to a float, because y doesn't have a specific floating point value. Its value depends on x. If you don't know x, you don't know y.</p>\r\n<p>That's what a TensorVariable is--it represents an algebraic variable of unknown value. You can't just &quot;convert&quot; it. You can compute a specific value of the variable given a set of specific values of the variables that it depends on.</p>\r\n<p>Can you tell me more about what you're trying to do? I can probably tell you a better way of going about it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23707,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/30/2013 16:08:55",
      "content": "<p>do the make submission.py output an matrix with nx9 (10000x9 for the test set) with the probabilities instead of the major problable class.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23708,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/30/2013 16:13:27",
      "content": "<p>I think this is it (Thanks!):</p>\r\n<p>&nbsp;</p>\r\n<pre>X = model.get_input_space().make_batch_theano()<br>Y = model.fprop(X)<br><br>from theano import tensor as T<br>from theano import function<br>f = function([X], Y)<br><br>y = []<br>for i in xrange(dataset.X.shape[0] / batch_size):<br>    x_arg = dataset.X[i*batch_size:(i&#43;1)*batch_size,:]<br>    if X.ndim &gt; 2:<br>        x_arg = dataset.get_topological_view(x_arg)<br>    y.append(f(x_arg.astype(X.dtype)))<br><br>y = np.concatenate(y)<br><br>#assert y.ndim == 9<br>#assert y.shape[0] == dataset.X.shape[0]<br><br>y = y[:m,:]</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23710,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 16:21:30",
      "content": "<p>Are you just doing that for debugging purposes? If you upload a file that says anything but &quot;1.0&quot;, &quot;2.0&quot;, etc. Kaggle will give you 0 points.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23714,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/30/2013 16:30:49",
      "content": "<p>Nope. I just wanna run some analysis on the predictions. And i don't know enough python (i don't know python at all to be honest) to do this directly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23719,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/30/2013 18:09:37",
      "content": "<p>By the way, can i run more than one training in paralell? Tried it but theano complained. Are the training proccess using gpus?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23721,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 18:30:03",
      "content": "<p>If you have your THEANO_FLAGS or .theanorc set to use GPU, then it will use GPU. pylearn2 doesn't control whether you use GPU or not. pylearn2 doesn't impose any limitations on what pylearn2 jobs you can run in parallel. Theano will only let you run as many\r\n GPU jobs as you have GPUs for.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23722,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/30/2013 18:31:57",
      "content": "<p>when i try to run in parallel, i get this:&nbsp;<span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\">INFO (theano.gof.compilelock): Waiting for existing lock by process '12975' (I am process '12966')\r\n</span><span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\">INFO (theano.gof.compilelock): To manually release the lock, delete&nbsp;</span></p>\r\n<p><span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\">I just want to know if it is a limitation, or a configuration. Can that mutual exclusiviness be disabled? is it safe?</span></p>\r\n<p><span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\">those are my questions...</span></p>\r\n<p><span class=\"x_GNVMTOMCLAB x_ace_constant x_ace_language\" style=\"line-height:1.4em\"><br>\r\n</span></p>\r\n<pre class=\"x_GNVMTOMCABB\">&nbsp;</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23723,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/30/2013 18:41:46",
      "content": "<p>Theano actually will execute the learning experiment in parallel. It's only the initial compilation that's getting serialized. If you wait for a while they will both run.</p>\r\n<p>What's going on is theano keeps a directory on disk where it caches the output of its calls to gcc. The problem is that whoever wrote theano's compilation mechanism doesn't seem to have understood locks and parallelism at all, so it is hopelessly inefficient.\r\n Whenever theano tries to acquire the lock and fails, it sleeps for about 5 seconds before trying again. This means that running two jobs in parallel can more than double the compilation time.</p>\r\n<p>For the long term, I've talked with the theano developers about switching to a lock-free design that should eliminate these synchronization issues. I have no idea how long it will be before anyone has time to implement that.</p>\r\n<p>In the short term, you can check the theano documentation to see how to control the directory that's used for the compilation cache. If you use an environment variable to set it to a different directory for job, then you'll lose the benefit of the cache,\r\n but you also won't have to pay the cost of all your jobs sleeping for no good reason.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 135753,
      "author_name": "waleedsial",
      "author_url": "",
      "post_date": "09/16/2016 16:45:34",
      "content": "<p>Hi all \nI dont know whether it is the right forum or not. But i have been struggling to set this python data path variable in windows and when i search for this thing this forums is one of the few places where such things are mentioned.\nany help would be appreciated i am student of bachelors and i am trying to do a project from kaggle  as my final year project.\nregards\nwaleed sial</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "23572": "",
    "23581": "",
    "23609": "",
    "23610": "",
    "23611": "",
    "23612": "",
    "23620": "",
    "23623": "",
    "23631": "",
    "23649": "",
    "23650": "",
    "23652": "",
    "23653": "",
    "23672": "",
    "23681": "",
    "23702": "",
    "23703": "",
    "23705": "",
    "23706": "",
    "23707": "",
    "23708": "",
    "23710": "",
    "23714": "",
    "23719": "",
    "23721": "",
    "23722": "",
    "23723": "",
    "135753": "Hi all \r\nI dont know whether it is the right forum or not. But i have been struggling to set this python data path variable in windows and when i search for this thing this forums is one of the few places where such things are mentioned.\r\nany help would be appreciated i am student of bachelors and i am trying to do a project from kaggle  as my final year project.\r\nregards\r\nwaleed sial"
  },
  "source": "meta"
}