{
  "id": 171675,
  "title": "please explain me the role of partition in scoring script! ",
  "url": "/competitions/landmark-retrieval-2020/discussion/171675",
  "author_name": "Uday Kumar Gurugubelli",
  "post_date": "2020-08-02T02:25:34.145000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": 954699,
      "postDate": "2020-08-02T02:25:34.147Z",
      "rawMarkdown": "",
      "votes": 1
    },
    {
      "id": 955233,
      "postDate": "2020-08-02T13:26:34.263Z",
      "content": "<p>If you're talking about:</p>\n\n<p><code>predicted_positions = np.argpartition(distances, K, axis=1)[:, :K]\n</code></p>\n\n<p>It's equivalent of finding nearest K point. argpartition is more efficient than classical sorting to find top K</p>",
      "rawMarkdown": "If you're talking about:\n\n`        predicted_positions = np.argpartition(distances, K, axis=1)[:, :K]\n`\n\nIt's equivalent of finding nearest K point. argpartition is more efficient than classical sorting to find top K",
      "votes": 2,
      "replies": [
        {
          "id": 955243,
          "postDate": "2020-08-02T13:34:04.350Z",
          "content": "<p>thanks for reply. but what is the value of K? and how do you know the value of the array at index K. \nfor example say K=100. it sorts the values based on the value at 100 position, which seems like a random number. explain me if i am wrong.</p>",
          "rawMarkdown": "thanks for reply. but what is the value of K? and how do you know the value of the array at index K. \nfor example say K=100. it sorts the values based on the value at 100 position, which seems like a random number. explain me if i am wrong."
        },
        {
          "id": 955391,
          "postDate": "2020-08-02T15:44:03.147Z",
          "content": "<p>import numpy as np\nx = np.array([.3, .4, .2, .1])\ny=np.argpartition(x, 2)\nprint(y)\nprint(x[y])\nx = np.array([.3, .4, .2, .1])\ny=np.argpartition(x, 3)\nprint(y)\nprint(x[y])</p>\n\n<h1>[3 2 0 1]</h1>\n\n<h1>[0.1 0.2 0.3 0.4]</h1>\n\n<h1>[0 3 2 1]</h1>\n\n<h1>[0.3 0.1 0.2 0.4]</h1>\n\n<p>the result is not matching with the description given in numpy argpartition docs.</p>",
          "rawMarkdown": "import numpy as np\nx = np.array([.3, .4, .2, .1])\ny=np.argpartition(x, 2)\nprint(y)\nprint(x[y])\nx = np.array([.3, .4, .2, .1])\ny=np.argpartition(x, 3)\nprint(y)\nprint(x[y])\n\n#[3 2 0 1]\n#[0.1 0.2 0.3 0.4]\n#[0 3 2 1]\n#[0.3 0.1 0.2 0.4]\n\nthe result is not matching with the description given in numpy argpartition docs."
        },
        {
          "id": 956004,
          "postDate": "2020-08-03T06:35:29.550Z",
          "content": "<p>Argpartition will take the first K element to the left and the other to the right but non necessary in the right order. For example:</p>\n\n<p><code>\nimport numpy as np\nx = np.arange(10)\nnp.random.shuffle(x)\nprint(x, '\\n')`\nfor i in  np.arange(10):\n    y = np.argpartition(x, i)\n    print(f'K: {i}  -  Argpartition: {x[y]}  -  Top K: {x[y][:i]}\\n')\n</code></p>\n\n<p>If you see the output you get:</p>\n\n<hr>\n\n<p>Original vector   [3 4 7 6 9 2 8 5 0 1] </p>\n\n<p>K: 0  -  Argpartition: [0 4 7 6 9 2 8 5 3 1]  -  Top K: []</p>\n\n<p>K: 1  -  Argpartition: [0 1 7 6 9 2 8 5 3 4]  -  Top K: [0]</p>\n\n<p>K: 2  -  Argpartition: [0 1 2 6 9 7 8 5 3 4]  -  Top K: [0 1]</p>\n\n<p>K: 3  -  Argpartition: [2 1 0 3 4 6 8 5 7 9]  -  Top K: [2 1 0]</p>\n\n<p>K: 4  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3]</p>\n\n<p>K: 5  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4]</p>\n\n<p>K: 6  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4 5]</p>\n\n<p>K: 7  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4 5 6]</p>\n\n<p>K: 8  -  Argpartition: [2 1 0 3 7 4 6 5 8 9]  -  Top K: [2 1 0 3 7 4 6 5]</p>\n\n<p>K: 9  -  Argpartition: [2 1 0 3 7 4 6 5 8 9]  -  Top K: [2 1 0 3 7 4 6 5 8]</p>\n\n<hr>\n\n<p>Each time you will get top K (in this case as sorting in increasing order and selecting top K) but not necessary in right order as sorting.</p>\n\n<p>If you look at k = 6 --&gt; you have [2 1 0 3 4 5] which are the lower 6 element inside the vector </p>\n\n<p>This method is more efficient in fact you can look at some evaluation here:</p>\n\n<p><a href=\"https://stackoverflow.com/questions/42184499/cannot-understand-numpy-argpartition-output\">https://stackoverflow.com/questions/42184499/cannot-understand-numpy-argpartition-output</a></p>\n\n<p>```\nIn [51]: a = np.random.rand(10000)*100</p>\n\n<p>In [52]: %timeit np.argpartition(a,range(a.size-1))[:5]\n10 loops, best of 3: 105 ms per loop</p>\n\n<p>In [53]: %timeit a.argsort()\n1000 loops, best of 3: 893 µs per loop\n```</p>",
          "rawMarkdown": "Argpartition will take the first K element to the left and the other to the right but non necessary in the right order. For example:\n\n```\nimport numpy as np\nx = np.arange(10)\nnp.random.shuffle(x)\nprint(x, '\\n')`\nfor i in  np.arange(10):\n    y = np.argpartition(x, i)\n    print(f'K: {i}  -  Argpartition: {x[y]}  -  Top K: {x[y][:i]}\\n')\n```\n\nIf you see the output you get:\n\n---------------------------------------\nOriginal vector   [3 4 7 6 9 2 8 5 0 1] \n\nK: 0  -  Argpartition: [0 4 7 6 9 2 8 5 3 1]  -  Top K: []\n\nK: 1  -  Argpartition: [0 1 7 6 9 2 8 5 3 4]  -  Top K: [0]\n\nK: 2  -  Argpartition: [0 1 2 6 9 7 8 5 3 4]  -  Top K: [0 1]\n\nK: 3  -  Argpartition: [2 1 0 3 4 6 8 5 7 9]  -  Top K: [2 1 0]\n\nK: 4  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3]\n\nK: 5  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4]\n\nK: 6  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4 5]\n\nK: 7  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4 5 6]\n\nK: 8  -  Argpartition: [2 1 0 3 7 4 6 5 8 9]  -  Top K: [2 1 0 3 7 4 6 5]\n\nK: 9  -  Argpartition: [2 1 0 3 7 4 6 5 8 9]  -  Top K: [2 1 0 3 7 4 6 5 8]\n\n\n\n---------------------------------------\n\nEach time you will get top K (in this case as sorting in increasing order and selecting top K) but not necessary in right order as sorting.\n\nIf you look at k = 6 --&gt; you have [2 1 0 3 4 5] which are the lower 6 element inside the vector \n\nThis method is more efficient in fact you can look at some evaluation here:\n\nhttps://stackoverflow.com/questions/42184499/cannot-understand-numpy-argpartition-output\n\n```\nIn [51]: a = np.random.rand(10000)*100\n\nIn [52]: %timeit np.argpartition(a,range(a.size-1))[:5]\n10 loops, best of 3: 105 ms per loop\n\nIn [53]: %timeit a.argsort()\n1000 loops, best of 3: 893 µs per loop\n```\n"
        },
        {
          "id": 956649,
          "postDate": "2020-08-03T17:03:46.470Z",
          "content": "<p>so it doesn't depend on the value at index K. Thanks for the explanation.</p>",
          "rawMarkdown": " so it doesn't depend on the value at index K. Thanks for the explanation."
        }
      ]
    },
    {
      "id": 954845,
      "postDate": "2020-08-02T06:19:20.180Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 955233,
      "author_name": "Davide Stenner",
      "author_url": "",
      "post_date": "2020-08-02T13:26:34.263000",
      "content": "<p>If you're talking about:</p>\n\n<p><code>predicted_positions = np.argpartition(distances, K, axis=1)[:, :K]\n</code></p>\n\n<p>It's equivalent of finding nearest K point. argpartition is more efficient than classical sorting to find top K</p>",
      "votes": 2,
      "replies": [
        {
          "id": 955243,
          "author_name": "Uday Kumar Gurugubelli",
          "author_url": "",
          "post_date": "2020-08-02T13:34:04.350000",
          "content": "<p>thanks for reply. but what is the value of K? and how do you know the value of the array at index K. \nfor example say K=100. it sorts the values based on the value at 100 position, which seems like a random number. explain me if i am wrong.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 955391,
          "author_name": "Uday Kumar Gurugubelli",
          "author_url": "",
          "post_date": "2020-08-02T15:44:03.147000",
          "content": "<p>import numpy as np\nx = np.array([.3, .4, .2, .1])\ny=np.argpartition(x, 2)\nprint(y)\nprint(x[y])\nx = np.array([.3, .4, .2, .1])\ny=np.argpartition(x, 3)\nprint(y)\nprint(x[y])</p>\n\n<h1>[3 2 0 1]</h1>\n\n<h1>[0.1 0.2 0.3 0.4]</h1>\n\n<h1>[0 3 2 1]</h1>\n\n<h1>[0.3 0.1 0.2 0.4]</h1>\n\n<p>the result is not matching with the description given in numpy argpartition docs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 956004,
          "author_name": "Davide Stenner",
          "author_url": "",
          "post_date": "2020-08-03T06:35:29.550000",
          "content": "<p>Argpartition will take the first K element to the left and the other to the right but non necessary in the right order. For example:</p>\n\n<p><code>\nimport numpy as np\nx = np.arange(10)\nnp.random.shuffle(x)\nprint(x, '\\n')`\nfor i in  np.arange(10):\n    y = np.argpartition(x, i)\n    print(f'K: {i}  -  Argpartition: {x[y]}  -  Top K: {x[y][:i]}\\n')\n</code></p>\n\n<p>If you see the output you get:</p>\n\n<hr>\n\n<p>Original vector   [3 4 7 6 9 2 8 5 0 1] </p>\n\n<p>K: 0  -  Argpartition: [0 4 7 6 9 2 8 5 3 1]  -  Top K: []</p>\n\n<p>K: 1  -  Argpartition: [0 1 7 6 9 2 8 5 3 4]  -  Top K: [0]</p>\n\n<p>K: 2  -  Argpartition: [0 1 2 6 9 7 8 5 3 4]  -  Top K: [0 1]</p>\n\n<p>K: 3  -  Argpartition: [2 1 0 3 4 6 8 5 7 9]  -  Top K: [2 1 0]</p>\n\n<p>K: 4  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3]</p>\n\n<p>K: 5  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4]</p>\n\n<p>K: 6  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4 5]</p>\n\n<p>K: 7  -  Argpartition: [2 1 0 3 4 5 6 7 8 9]  -  Top K: [2 1 0 3 4 5 6]</p>\n\n<p>K: 8  -  Argpartition: [2 1 0 3 7 4 6 5 8 9]  -  Top K: [2 1 0 3 7 4 6 5]</p>\n\n<p>K: 9  -  Argpartition: [2 1 0 3 7 4 6 5 8 9]  -  Top K: [2 1 0 3 7 4 6 5 8]</p>\n\n<hr>\n\n<p>Each time you will get top K (in this case as sorting in increasing order and selecting top K) but not necessary in right order as sorting.</p>\n\n<p>If you look at k = 6 --&gt; you have [2 1 0 3 4 5] which are the lower 6 element inside the vector </p>\n\n<p>This method is more efficient in fact you can look at some evaluation here:</p>\n\n<p><a href=\"https://stackoverflow.com/questions/42184499/cannot-understand-numpy-argpartition-output\">https://stackoverflow.com/questions/42184499/cannot-understand-numpy-argpartition-output</a></p>\n\n<p>```\nIn [51]: a = np.random.rand(10000)*100</p>\n\n<p>In [52]: %timeit np.argpartition(a,range(a.size-1))[:5]\n10 loops, best of 3: 105 ms per loop</p>\n\n<p>In [53]: %timeit a.argsort()\n1000 loops, best of 3: 893 µs per loop\n```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 956649,
          "author_name": "Uday Kumar Gurugubelli",
          "author_url": "",
          "post_date": "2020-08-03T17:03:46.470000",
          "content": "<p>so it doesn't depend on the value at index K. Thanks for the explanation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 954845,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-02T06:19:20.180000",
      "content": "",
      "votes": -3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "954699": "",
    "955233": "If you're talking about:\n\n`        predicted_positions = np.argpartition(distances, K, axis=1)[:, :K]\n`\n\nIt's equivalent of finding nearest K point. argpartition is more efficient than classical sorting to find top K",
    "954845": ""
  }
}