{
  "id": 612928,
  "title": "Solution notebook: 3-stage Fusion Split-Ensemble",
  "url": "/competitions/acm-icaif-25-ai-agentic-retrieval-grand-challenge/discussion/612928",
  "author_name": "",
  "post_date": "2025-10-22T23:26:27.461456400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Full reproduction</strong>: please refer to the <a href=\"https://github.com/Dan-T24n/Agentic-Retrieval-Challenge-ICAIF25-Public/blob/main/notebooks/fusion_se_databricks.ipynb\" target=\"_blank\">notebook in Github repo</a></p>\n<h2>Notes &amp; disclaimers</h2>\n<p>First, a big thanks to the organizers for this interesting challenge. I've enjoyed working on this project. The baseline notebook provides a solid starting point.</p>\n<p>Disclaimers:</p>\n<ul>\n<li>To maintain fairness, I only used open-source and cheapest models (mainly gpt-oss-120b, llama 3-405b re-ranking) for submission.</li>\n<li>Some design choices are determined by cost/performance trade-offs to replicate real-world usage.</li>\n<li>Total runtime ~2h with Databricks, cost ~$2 USD.</li>\n<li>Using this solution, one can <em>crank up the scores</em> quite a bit (+5-8pp), by forced retries many attempts -&gt;cost 5-10x more (time/$$): not realistic</li>\n</ul>\n<h2>Model Design: Fusion Split-Ensemble</h2>\n<p><strong>Principles:</strong></p>\n<ul>\n<li>maximize robustness and explanability: easy to iterate and reproduce</li>\n<li>modular: swapable models, quickly adjust latency-performance ratio</li>\n</ul>\n<h3>Solution Overview</h3>\n<p><strong>Core ideas</strong>: 3-stage architecture that progressively filters candidates.</p>\n<ul>\n<li>Manage attention allocation, avoid long context, avoid crowded pools</li>\n<li>Stage-specific custom strategy: prompting + model choice + ensemble</li>\n</ul>\n<p><strong>Key mechanism:</strong></p>\n<ul>\n<li>Agentic retry with RRF ensemble, any task/stage -&gt; emulate multiple judges<ul>\n<li>Adaptive retries: based on response quality, remember and fuse multiple partial answers</li>\n<li>Forced retries: redo same query again to get more opinions, then RRF fusion</li></ul></li>\n</ul>\n<p><strong>Tasks allocation:</strong></p>\n<ul>\n<li>Chunk ranking: 3-stage Fusion SE described below. Runtime ~2h</li>\n<li>Document ranking: use same design, only final stage -&gt; run first, not in parallel to avoid QPS throttling. Runtime ~5m</li>\n</ul>\n<p><strong>Results</strong> with  using dev-sets, <em>zero retry</em>: S1: oss-120b /S2: gpt-5-mini  /S3: gpt-5</p>\n<ul>\n<li>average document score: <strong>0.92, MAP@5: 0.905</strong></li>\n<li>average chunk score: <strong>0.74, MAP@5: 0.769</strong></li>\n</ul>\n<p><a href=\"https://drive.google.com/open?id=1xHqayJIs4iXFMGBhib3du-ezPFFQ7B4d&amp;usp=drive_fs\" target=\"_blank\">Overall scores</a></p>\n<p><a href=\"https://drive.google.com/file/d/1VJDaDGirQD73eYomquNiiH_rFLgPhHva/view?usp=sharing\" target=\"_blank\">Breakdown scores at each stage</a></p>\n<p><em>If you want higher scores</em>: just crank up <code>min_retry_attempts</code>, swap bigger models in S2/S3 -&gt; higher cost &amp; latency</p>\n<ul>\n<li>not reproducible in reality</li>\n<li>easy to climb <code>leaderboard</code></li>\n</ul>\n<p>You can check the log and json results in the repo.</p>\n<h2>Data flow</h2>\n<pre><code># using 90-percentile stats\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n</code></pre>\n<h2>Architecture Details</h2>\n<h3>Stage 1: Local Ranking (Recall Focus)</h3>\n<p><strong>Purpose</strong>: Cast wide net to capture all potentially relevant chunks</p>\n<p><strong>Strategy</strong>:</p>\n<ul>\n<li>Split initial pools into equal-chunk parts (&lt;=5 parts per query)</li>\n<li>Manageable parts: with &gt;90%, max 50-70 chunks per part</li>\n<li>BM25 + LLM Fusion with frequencies (consensus) &amp; normalized scores (confidence)</li>\n</ul>\n<p><strong>Key Feature</strong>: Hybrid scoring prevents early loss</p>\n<ul>\n<li>Achieve 100% recall (0% loss) -&gt; validated with n=100 x5 experiments</li>\n<li>If use pure LLM loses 20-30%% of relevant chunks</li>\n</ul>\n<h3>Stage 2: Split Rescore</h3>\n<p><strong>Purpose</strong>: Reduce context overload while maintaining candidate diversity</p>\n<p><strong>Strategy</strong>:</p>\n<ul>\n<li>Split 50 candidates into 2 parts (~25 each)</li>\n<li>Rank each part independently with pure LLM</li>\n<li>Extract top-10 from each part (~20 total)</li>\n</ul>\n<p><strong>Key Feature</strong>: Context reduction improves instruction-following and quality of ranking</p>\n<ul>\n<li>50 chunks in one call: Context Rot &amp; loss-in-the-middle kick in (20k-60k tokens)</li>\n<li>2x25 chunks: better ranking quality, side-effect: less API throttling (TPM rate limits)</li>\n</ul>\n<h3>Stage 3: Final Rescore (Precision Focus)</h3>\n<p><strong>Purpose</strong>: Precise global ranking on reduced candidate pool</p>\n<p><strong>Strategy</strong>:</p>\n<ul>\n<li>Single LLM-call with powerful model on ~20 candidates</li>\n<li>Enable Forced Retry (if time/budget allow): min_attempts &gt; 1 -&gt; independant opinions</li>\n<li>Handle edge-cases: stuck in reasoning loop, correct unfinished answer, one-chunk answer</li>\n</ul>\n<hr>\n<h2>Q&amp;A</h2>\n<h3>Why Lexical Fusion Only at Stage 1?</h3>\n<p>Hypothesis: lexical signals help disambiguate when semantic signals are weak</p>\n<p><strong>Example</strong>:</p>\n<pre><code>Stage : Recall Ranking ( scale)\n\n\nQuery: \"What investor views emerged on geographic expansion?\"\n\nContent Reality:\n  Chunk A: \"State mortgage data: Ohio $2,725, California $2,322...\"\n           (lexically  \"expansion\" context but semantically weak)\n\n  Chunk B: \"Investors expressed optimism about expansion plans...\"\n           (semantically stronger but lower keyword )\n\n chunks  same LLM score (), but Chunk A has higher lexical score.\n\n: Lexical tiebreaker promotes wrong chunks.  later stages, there  distractors that require understanding  keyword matching.\n</code></pre>\n<p>Evidence: performance loss when applied BM25 fusion at Stage 2-3. Lexical signals saturate after Stage 1.</p>\n<hr>\n<h3>Why Different Models per Stage?</h3>\n<p><strong>Workload Distribution</strong>:</p>\n<pre><code>\nTask: Bulk filtering (300 -&gt; 50 -&gt; 20)\nModel: Efficient (120B)\nSufficient for filtering obvious chunks\n\n\n\n\nTask: Final precision ranking (20 -&gt; 5)\nModel: Powerful (405B)\nNeed reasoning for nuanced comparison: needles vs. distractors\n\n\n\nResult: Cost-latency-performance trade-off\nSclable efficiency in production\n</code></pre>\n<hr>\n<h3>What about bigger/reasoning models?</h3>\n<p>Thinking models perform better, but they tend to suffer from over-thinking. Increase thinking mode (reasoning=high) has ambiguous effect: higher failure rate and better max performance with retries.</p>\n<p>Real example: <br>\n<code>`\n... Now we need other: maybe chunk 191.\\n\\nOk.\\n\\nNow we need other: maybe chunk 191.\\n\\nStop.\\n\\nOk.\\n\\nThis is stuck.\\n\\nLet\\'s identify other chunks: maybe chunk 191 is the only macro.\\n\\nBut also chunk 191 includes \"General economic inflation.. [seems to recover] ... \\n\\nNow other chunks: maybe chunk 191 ... \n</code><br>\nOh dear 😔</p>\n<p>Increase thinking budget (max_token) with reasoning=medium seems to be more cost-effective for ranking task, but requires careful prompt design.</p>\n<p>Gpt-5 perform better across all experiments (+5-8pp), but higher cost/latency ~5x (did not submit)</p>\n<hr>\n<h3>What about other xyz?</h3>\n<p>Agentic workflow with tool-use (LlamaIndex): high overhead, brittle, structured output not well supported across different reasoning models. Finance domain requires high robustness -&gt; dropped this direction after prototyping.</p>\n<p>Fine-tuning: performance could increase, but quite black-box, hard to debug when things go wrong. Hard to serve models for reproduction.</p>\n<h2>Summary</h2>\n<p>Overall, the challenge is well structured, great dataset on this topic. I've learned some more on agentic design and reasoning models. I hope my solution can provide some useful insights to the community.</p>\n<p>Drop me a message if you want to discuss ideas &amp; real-world application. I'm working in both academic research and industry projects. <a href=\"https://www.linkedin.com/in/hdan0tran/\" target=\"_blank\">LinkedIn</a> </p>",
  "messages": [
    {
      "id": "3305547",
      "postDate": "10/22/2025 23:26:27",
      "content": "<p><strong>Full reproduction</strong>: please refer to the <a href=\"https://github.com/Dan-T24n/Agentic-Retrieval-Challenge-ICAIF25-Public/blob/main/notebooks/fusion_se_databricks.ipynb\" target=\"_blank\">notebook in Github repo</a></p>\n<h2>Notes &amp; disclaimers</h2>\n<p>First, a big thanks to the organizers for this interesting challenge. I've enjoyed working on this project. The baseline notebook provides a solid starting point.</p>\n<p>Disclaimers:</p>\n<ul>\n<li>To maintain fairness, I only used open-source and cheapest models (mainly gpt-oss-120b, llama 3-405b re-ranking) for submission.</li>\n<li>Some design choices are determined by cost/performance trade-offs to replicate real-world usage.</li>\n<li>Total runtime ~2h with Databricks, cost ~$2 USD.</li>\n<li>Using this solution, one can <em>crank up the scores</em> quite a bit (+5-8pp), by forced retries many attempts -&gt;cost 5-10x more (time/$$): not realistic</li>\n</ul>\n<h2>Model Design: Fusion Split-Ensemble</h2>\n<p><strong>Principles:</strong></p>\n<ul>\n<li>maximize robustness and explanability: easy to iterate and reproduce</li>\n<li>modular: swapable models, quickly adjust latency-performance ratio</li>\n</ul>\n<h3>Solution Overview</h3>\n<p><strong>Core ideas</strong>: 3-stage architecture that progressively filters candidates.</p>\n<ul>\n<li>Manage attention allocation, avoid long context, avoid crowded pools</li>\n<li>Stage-specific custom strategy: prompting + model choice + ensemble</li>\n</ul>\n<p><strong>Key mechanism:</strong></p>\n<ul>\n<li>Agentic retry with RRF ensemble, any task/stage -&gt; emulate multiple judges<ul>\n<li>Adaptive retries: based on response quality, remember and fuse multiple partial answers</li>\n<li>Forced retries: redo same query again to get more opinions, then RRF fusion</li></ul></li>\n</ul>\n<p><strong>Tasks allocation:</strong></p>\n<ul>\n<li>Chunk ranking: 3-stage Fusion SE described below. Runtime ~2h</li>\n<li>Document ranking: use same design, only final stage -&gt; run first, not in parallel to avoid QPS throttling. Runtime ~5m</li>\n</ul>\n<p><strong>Results</strong> with  using dev-sets, <em>zero retry</em>: S1: oss-120b /S2: gpt-5-mini  /S3: gpt-5</p>\n<ul>\n<li>average document score: <strong>0.92, MAP@5: 0.905</strong></li>\n<li>average chunk score: <strong>0.74, MAP@5: 0.769</strong></li>\n</ul>\n<p><a href=\"https://drive.google.com/open?id=1xHqayJIs4iXFMGBhib3du-ezPFFQ7B4d&amp;usp=drive_fs\" target=\"_blank\">Overall scores</a></p>\n<p><a href=\"https://drive.google.com/file/d/1VJDaDGirQD73eYomquNiiH_rFLgPhHva/view?usp=sharing\" target=\"_blank\">Breakdown scores at each stage</a></p>\n<p><em>If you want higher scores</em>: just crank up <code>min_retry_attempts</code>, swap bigger models in S2/S3 -&gt; higher cost &amp; latency</p>\n<ul>\n<li>not reproducible in reality</li>\n<li>easy to climb <code>leaderboard</code></li>\n</ul>\n<p>You can check the log and json results in the repo.</p>\n<h2>Data flow</h2>\n<pre><code># using 90-percentile stats\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n</code></pre>\n<h2>Architecture Details</h2>\n<h3>Stage 1: Local Ranking (Recall Focus)</h3>\n<p><strong>Purpose</strong>: Cast wide net to capture all potentially relevant chunks</p>\n<p><strong>Strategy</strong>:</p>\n<ul>\n<li>Split initial pools into equal-chunk parts (&lt;=5 parts per query)</li>\n<li>Manageable parts: with &gt;90%, max 50-70 chunks per part</li>\n<li>BM25 + LLM Fusion with frequencies (consensus) &amp; normalized scores (confidence)</li>\n</ul>\n<p><strong>Key Feature</strong>: Hybrid scoring prevents early loss</p>\n<ul>\n<li>Achieve 100% recall (0% loss) -&gt; validated with n=100 x5 experiments</li>\n<li>If use pure LLM loses 20-30%% of relevant chunks</li>\n</ul>\n<h3>Stage 2: Split Rescore</h3>\n<p><strong>Purpose</strong>: Reduce context overload while maintaining candidate diversity</p>\n<p><strong>Strategy</strong>:</p>\n<ul>\n<li>Split 50 candidates into 2 parts (~25 each)</li>\n<li>Rank each part independently with pure LLM</li>\n<li>Extract top-10 from each part (~20 total)</li>\n</ul>\n<p><strong>Key Feature</strong>: Context reduction improves instruction-following and quality of ranking</p>\n<ul>\n<li>50 chunks in one call: Context Rot &amp; loss-in-the-middle kick in (20k-60k tokens)</li>\n<li>2x25 chunks: better ranking quality, side-effect: less API throttling (TPM rate limits)</li>\n</ul>\n<h3>Stage 3: Final Rescore (Precision Focus)</h3>\n<p><strong>Purpose</strong>: Precise global ranking on reduced candidate pool</p>\n<p><strong>Strategy</strong>:</p>\n<ul>\n<li>Single LLM-call with powerful model on ~20 candidates</li>\n<li>Enable Forced Retry (if time/budget allow): min_attempts &gt; 1 -&gt; independant opinions</li>\n<li>Handle edge-cases: stuck in reasoning loop, correct unfinished answer, one-chunk answer</li>\n</ul>\n<hr>\n<h2>Q&amp;A</h2>\n<h3>Why Lexical Fusion Only at Stage 1?</h3>\n<p>Hypothesis: lexical signals help disambiguate when semantic signals are weak</p>\n<p><strong>Example</strong>:</p>\n<pre><code>Stage : Recall Ranking ( scale)\n\n\nQuery: \"What investor views emerged on geographic expansion?\"\n\nContent Reality:\n  Chunk A: \"State mortgage data: Ohio $2,725, California $2,322...\"\n           (lexically  \"expansion\" context but semantically weak)\n\n  Chunk B: \"Investors expressed optimism about expansion plans...\"\n           (semantically stronger but lower keyword )\n\n chunks  same LLM score (), but Chunk A has higher lexical score.\n\n: Lexical tiebreaker promotes wrong chunks.  later stages, there  distractors that require understanding  keyword matching.\n</code></pre>\n<p>Evidence: performance loss when applied BM25 fusion at Stage 2-3. Lexical signals saturate after Stage 1.</p>\n<hr>\n<h3>Why Different Models per Stage?</h3>\n<p><strong>Workload Distribution</strong>:</p>\n<pre><code>\nTask: Bulk filtering (300 -&gt; 50 -&gt; 20)\nModel: Efficient (120B)\nSufficient for filtering obvious chunks\n\n\n\n\nTask: Final precision ranking (20 -&gt; 5)\nModel: Powerful (405B)\nNeed reasoning for nuanced comparison: needles vs. distractors\n\n\n\nResult: Cost-latency-performance trade-off\nSclable efficiency in production\n</code></pre>\n<hr>\n<h3>What about bigger/reasoning models?</h3>\n<p>Thinking models perform better, but they tend to suffer from over-thinking. Increase thinking mode (reasoning=high) has ambiguous effect: higher failure rate and better max performance with retries.</p>\n<p>Real example: <br>\n<code>`\n... Now we need other: maybe chunk 191.\\n\\nOk.\\n\\nNow we need other: maybe chunk 191.\\n\\nStop.\\n\\nOk.\\n\\nThis is stuck.\\n\\nLet\\'s identify other chunks: maybe chunk 191 is the only macro.\\n\\nBut also chunk 191 includes \"General economic inflation.. [seems to recover] ... \\n\\nNow other chunks: maybe chunk 191 ... \n</code><br>\nOh dear 😔</p>\n<p>Increase thinking budget (max_token) with reasoning=medium seems to be more cost-effective for ranking task, but requires careful prompt design.</p>\n<p>Gpt-5 perform better across all experiments (+5-8pp), but higher cost/latency ~5x (did not submit)</p>\n<hr>\n<h3>What about other xyz?</h3>\n<p>Agentic workflow with tool-use (LlamaIndex): high overhead, brittle, structured output not well supported across different reasoning models. Finance domain requires high robustness -&gt; dropped this direction after prototyping.</p>\n<p>Fine-tuning: performance could increase, but quite black-box, hard to debug when things go wrong. Hard to serve models for reproduction.</p>\n<h2>Summary</h2>\n<p>Overall, the challenge is well structured, great dataset on this topic. I've learned some more on agentic design and reasoning models. I hope my solution can provide some useful insights to the community.</p>\n<p>Drop me a message if you want to discuss ideas &amp; real-world application. I'm working in both academic research and industry projects. <a href=\"https://www.linkedin.com/in/hdan0tran/\" target=\"_blank\">LinkedIn</a> </p>",
      "rawMarkdown": "**Full reproduction**: please refer to the [notebook in Github repo](https://github.com/Dan-T24n/Agentic-Retrieval-Challenge-ICAIF25-Public/blob/main/notebooks/fusion_se_databricks.ipynb)\n\n## Notes & disclaimers\nFirst, a big thanks to the organizers for this interesting challenge. I've enjoyed working on this project. The baseline notebook provides a solid starting point.\n\nDisclaimers:\n- To maintain fairness, I only used open-source and cheapest models (mainly gpt-oss-120b, llama 3-405b re-ranking) for submission.\n- Some design choices are determined by cost/performance trade-offs to replicate real-world usage.\n- Total runtime ~2h with Databricks, cost ~$2 USD.\n- Using this solution, one can *crank up the scores* quite a bit (+5-8pp), by forced retries many attempts ->cost 5-10x more (time/$$): not realistic\n\n\n## Model Design: Fusion Split-Ensemble\n\n**Principles:**\n- maximize robustness and explanability: easy to iterate and reproduce\n- modular: swapable models, quickly adjust latency-performance ratio\n\n### Solution Overview\n\n**Core ideas**: 3-stage architecture that progressively filters candidates.\n- Manage attention allocation, avoid long context, avoid crowded pools\n- Stage-specific custom strategy: prompting + model choice + ensemble\n\n**Key mechanism:**\n- Agentic retry with RRF ensemble, any task/stage -> emulate multiple judges\n    - Adaptive retries: based on response quality, remember and fuse multiple partial answers\n    - Forced retries: redo same query again to get more opinions, then RRF fusion\n\n**Tasks allocation:**\n- Chunk ranking: 3-stage Fusion SE described below. Runtime ~2h\n- Document ranking: use same design, only final stage -> run first, not in parallel to avoid QPS throttling. Runtime ~5m\n\n**Results** with  using dev-sets, *zero retry*: S1: oss-120b /S2: gpt-5-mini  /S3: gpt-5\n- average document score: **0.92, MAP@5: 0.905**\n- average chunk score: **0.74, MAP@5: 0.769**\n\n[Overall scores](https://drive.google.com/open?id=1xHqayJIs4iXFMGBhib3du-ezPFFQ7B4d&usp=drive_fs)\n\n[Breakdown scores at each stage](https://drive.google.com/file/d/1VJDaDGirQD73eYomquNiiH_rFLgPhHva/view?usp=sharing)\n\n*If you want higher scores*: just crank up `min_retry_attempts`, swap bigger models in S2/S3 -> higher cost & latency\n- not reproducible in reality\n- easy to climb `leaderboard`\n\nYou can check the log and json results in the repo.\n\n## Data flow \n\n```\n# using 90-percentile stats\n\nQuery: question + ~300 chunks (~20k-120K tokens)\n------------------------------------------------------\n                    |\n                    v\n\n+---------------------------------------------------+\n| STAGE 1: Local Ranking (Filter: Recall Focus)     |\n+---------------------------------------------------+\n  Input:  300 chunks split into <=5 parts\n          (e.g., Part 1: chunks 0-59,\n                Part 2: chunks 60-119, ...)\n\n  Process: Each part -> Lexical scores + LLM scores\n           - Lexical: Keyword matching (BM25)\n           - RRF fusion: 70% LLM + 30% lexical (normalized scores)\n\n  Extract: Top-k candidates from each part\n                    |\n         ~50 global candidates\n                    |\n                    v\n\n+---------------------------------------------------+\n| STAGE 2: Split Rescore (Balanced focus)           |\n+---------------------------------------------------+\n  Input:  Top-50 from Stage 1\n          Split into 2 parts (~25 each)\n\n  Process: Each part -> Pure LLM\n           - Model: Efficient (120B)\n           - Smart Retry -> emulate multiple judges\n            - n_min_attempts >= 1\n            - disable to retry failed answers only\n            - collect multiple answers -> RFF fusion\n\n  Extract: Top-10 from each part, append (preserve order)\n                    |\n             ~20 candidates\n                    |\n                    v\n\n+---------------------------------------------------+\n| STAGE 3: Final Rescore (Precision Focus)          |\n+---------------------------------------------------+\n  Input:  Top-20 from Stage 2\n          (single global pool)\n\n  Process: Pure LLM ranking top-k\n           - Model: Powerful (405B)\n           - Smart Retry -> same as Stage2\n\n  Extract: Top-5 final ranking\n                    |\n              Top 5 ranked\n                    |\n                    v\n  Final Submission: [idx1, idx2, idx3, idx4, idx5]\n```\n\n\n## Architecture Details\n\n### Stage 1: Local Ranking (Recall Focus)\n\n**Purpose**: Cast wide net to capture all potentially relevant chunks\n\n**Strategy**:\n- Split initial pools into equal-chunk parts (<=5 parts per query)\n- Manageable parts: with >90%, max 50-70 chunks per part\n- BM25 + LLM Fusion with frequencies (consensus) & normalized scores (confidence)\n\n**Key Feature**: Hybrid scoring prevents early loss\n- Achieve 100% recall (0% loss) -> validated with n=100 x5 experiments\n- If use pure LLM loses 20-30%% of relevant chunks\n\n### Stage 2: Split Rescore\n\n**Purpose**: Reduce context overload while maintaining candidate diversity\n\n**Strategy**:\n- Split 50 candidates into 2 parts (~25 each)\n- Rank each part independently with pure LLM\n- Extract top-10 from each part (~20 total)\n\n**Key Feature**: Context reduction improves instruction-following and quality of ranking\n- 50 chunks in one call: Context Rot & loss-in-the-middle kick in (20k-60k tokens)\n- 2x25 chunks: better ranking quality, side-effect: less API throttling (TPM rate limits)\n\n### Stage 3: Final Rescore (Precision Focus)\n\n**Purpose**: Precise global ranking on reduced candidate pool\n\n**Strategy**:\n- Single LLM-call with powerful model on ~20 candidates\n- Enable Forced Retry (if time/budget allow): min_attempts > 1 -> independant opinions\n- Handle edge-cases: stuck in reasoning loop, correct unfinished answer, one-chunk answer\n\n---\n\n## Q&A\n\n### Why Lexical Fusion Only at Stage 1?\nHypothesis: lexical signals help disambiguate when semantic signals are weak\n\n**Example**:\n\n```\nStage 2: Precision/Recall Ranking (0-2 scale)\n---------------------------------------\n\nQuery: \"What investor views emerged on geographic expansion?\"\n\nContent Reality:\n  Chunk A: \"State mortgage data: Ohio $2,725, California $2,322...\"\n           (lexically matches \"expansion\" context but semantically weak)\n\n  Chunk B: \"Investors expressed optimism about expansion plans...\"\n           (semantically stronger but lower keyword match)\n\nBoth chunks get same LLM score (1), but Chunk A has higher lexical score.\n\nResult: Lexical tiebreaker promotes wrong chunks. In later stages, there are distractors that require understanding over keyword matching.\n```\nEvidence: performance loss when applied BM25 fusion at Stage 2-3. Lexical signals saturate after Stage 1.\n\n---\n\n### Why Different Models per Stage?\n\n**Workload Distribution**:\n\n```\nStage 1 + Stage 2 (85% of API calls)\n------------------------------------\nTask: Bulk filtering (300 -> 50 -> 20)\nModel: Efficient (120B)\nSufficient for filtering obvious chunks\n           |\n           v\n\nStage 3 (15% of API calls)\n--------------------------\nTask: Final precision ranking (20 -> 5)\nModel: Powerful (405B)\nNeed reasoning for nuanced comparison: needles vs. distractors\n           |\n           v\n\nResult: Cost-latency-performance trade-off\n- Sclable efficiency in production\n```\n\n---\n\n### What about bigger/reasoning models?\n\nThinking models perform better, but they tend to suffer from over-thinking. Increase thinking mode (reasoning=high) has ambiguous effect: higher failure rate and better max performance with retries.\n\nReal example: \n````\n... Now we need other: maybe chunk 191.\\n\\nOk.\\n\\nNow we need other: maybe chunk 191.\\n\\nStop.\\n\\nOk.\\n\\nThis is stuck.\\n\\nLet\\'s identify other chunks: maybe chunk 191 is the only macro.\\n\\nBut also chunk 191 includes \"General economic inflation.. [seems to recover] ... \\n\\nNow other chunks: maybe chunk 191 ... \n```\nOh dear 😔\n\nIncrease thinking budget (max_token) with reasoning=medium seems to be more cost-effective for ranking task, but requires careful prompt design.\n\nGpt-5 perform better across all experiments (+5-8pp), but higher cost/latency ~5x (did not submit)\n\n---\n\n### What about other xyz?\n\nAgentic workflow with tool-use (LlamaIndex): high overhead, brittle, structured output not well supported across different reasoning models. Finance domain requires high robustness -> dropped this direction after prototyping.\n\nFine-tuning: performance could increase, but quite black-box, hard to debug when things go wrong. Hard to serve models for reproduction.\n\n\n## Summary\n\nOverall, the challenge is well structured, great dataset on this topic. I've learned some more on agentic design and reasoning models. I hope my solution can provide some useful insights to the community.\n\nDrop me a message if you want to discuss ideas & real-world application. I'm working in both academic research and industry projects. [LinkedIn](https://www.linkedin.com/in/hdan0tran/)",
      "votes": null
    },
    {
      "id": "3305822",
      "postDate": "10/23/2025 14:01:44",
      "content": "<p>Great Work! Can you reply to my email?</p>",
      "rawMarkdown": "Great Work! Can you reply to my email?",
      "votes": null
    },
    {
      "id": "3306180",
      "postDate": "10/24/2025 03:52:42",
      "content": "<p>Thank you! I just replied, sorry for the delay</p>",
      "rawMarkdown": "Thank you! I just replied, sorry for the delay",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3305822,
      "author_name": "jihoonkwon",
      "author_url": "",
      "post_date": "10/23/2025 14:01:44",
      "content": "<p>Great Work! Can you reply to my email?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3306180,
          "author_name": "pandalikematcha",
          "author_url": "",
          "post_date": "10/24/2025 03:52:42",
          "content": "<p>Thank you! I just replied, sorry for the delay</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3305547": "**Full reproduction**: please refer to the [notebook in Github repo](https://github.com/Dan-T24n/Agentic-Retrieval-Challenge-ICAIF25-Public/blob/main/notebooks/fusion_se_databricks.ipynb)\n\n## Notes & disclaimers\nFirst, a big thanks to the organizers for this interesting challenge. I've enjoyed working on this project. The baseline notebook provides a solid starting point.\n\nDisclaimers:\n- To maintain fairness, I only used open-source and cheapest models (mainly gpt-oss-120b, llama 3-405b re-ranking) for submission.\n- Some design choices are determined by cost/performance trade-offs to replicate real-world usage.\n- Total runtime ~2h with Databricks, cost ~$2 USD.\n- Using this solution, one can *crank up the scores* quite a bit (+5-8pp), by forced retries many attempts ->cost 5-10x more (time/$$): not realistic\n\n\n## Model Design: Fusion Split-Ensemble\n\n**Principles:**\n- maximize robustness and explanability: easy to iterate and reproduce\n- modular: swapable models, quickly adjust latency-performance ratio\n\n### Solution Overview\n\n**Core ideas**: 3-stage architecture that progressively filters candidates.\n- Manage attention allocation, avoid long context, avoid crowded pools\n- Stage-specific custom strategy: prompting + model choice + ensemble\n\n**Key mechanism:**\n- Agentic retry with RRF ensemble, any task/stage -> emulate multiple judges\n    - Adaptive retries: based on response quality, remember and fuse multiple partial answers\n    - Forced retries: redo same query again to get more opinions, then RRF fusion\n\n**Tasks allocation:**\n- Chunk ranking: 3-stage Fusion SE described below. Runtime ~2h\n- Document ranking: use same design, only final stage -> run first, not in parallel to avoid QPS throttling. Runtime ~5m\n\n**Results** with  using dev-sets, *zero retry*: S1: oss-120b /S2: gpt-5-mini  /S3: gpt-5\n- average document score: **0.92, MAP@5: 0.905**\n- average chunk score: **0.74, MAP@5: 0.769**\n\n[Overall scores](https://drive.google.com/open?id=1xHqayJIs4iXFMGBhib3du-ezPFFQ7B4d&usp=drive_fs)\n\n[Breakdown scores at each stage](https://drive.google.com/file/d/1VJDaDGirQD73eYomquNiiH_rFLgPhHva/view?usp=sharing)\n\n*If you want higher scores*: just crank up `min_retry_attempts`, swap bigger models in S2/S3 -> higher cost & latency\n- not reproducible in reality\n- easy to climb `leaderboard`\n\nYou can check the log and json results in the repo.\n\n## Data flow \n\n```\n# using 90-percentile stats\n\nQuery: question + ~300 chunks (~20k-120K tokens)\n------------------------------------------------------\n                    |\n                    v\n\n+---------------------------------------------------+\n| STAGE 1: Local Ranking (Filter: Recall Focus)     |\n+---------------------------------------------------+\n  Input:  300 chunks split into <=5 parts\n          (e.g., Part 1: chunks 0-59,\n                Part 2: chunks 60-119, ...)\n\n  Process: Each part -> Lexical scores + LLM scores\n           - Lexical: Keyword matching (BM25)\n           - RRF fusion: 70% LLM + 30% lexical (normalized scores)\n\n  Extract: Top-k candidates from each part\n                    |\n         ~50 global candidates\n                    |\n                    v\n\n+---------------------------------------------------+\n| STAGE 2: Split Rescore (Balanced focus)           |\n+---------------------------------------------------+\n  Input:  Top-50 from Stage 1\n          Split into 2 parts (~25 each)\n\n  Process: Each part -> Pure LLM\n           - Model: Efficient (120B)\n           - Smart Retry -> emulate multiple judges\n            - n_min_attempts >= 1\n            - disable to retry failed answers only\n            - collect multiple answers -> RFF fusion\n\n  Extract: Top-10 from each part, append (preserve order)\n                    |\n             ~20 candidates\n                    |\n                    v\n\n+---------------------------------------------------+\n| STAGE 3: Final Rescore (Precision Focus)          |\n+---------------------------------------------------+\n  Input:  Top-20 from Stage 2\n          (single global pool)\n\n  Process: Pure LLM ranking top-k\n           - Model: Powerful (405B)\n           - Smart Retry -> same as Stage2\n\n  Extract: Top-5 final ranking\n                    |\n              Top 5 ranked\n                    |\n                    v\n  Final Submission: [idx1, idx2, idx3, idx4, idx5]\n```\n\n\n## Architecture Details\n\n### Stage 1: Local Ranking (Recall Focus)\n\n**Purpose**: Cast wide net to capture all potentially relevant chunks\n\n**Strategy**:\n- Split initial pools into equal-chunk parts (<=5 parts per query)\n- Manageable parts: with >90%, max 50-70 chunks per part\n- BM25 + LLM Fusion with frequencies (consensus) & normalized scores (confidence)\n\n**Key Feature**: Hybrid scoring prevents early loss\n- Achieve 100% recall (0% loss) -> validated with n=100 x5 experiments\n- If use pure LLM loses 20-30%% of relevant chunks\n\n### Stage 2: Split Rescore\n\n**Purpose**: Reduce context overload while maintaining candidate diversity\n\n**Strategy**:\n- Split 50 candidates into 2 parts (~25 each)\n- Rank each part independently with pure LLM\n- Extract top-10 from each part (~20 total)\n\n**Key Feature**: Context reduction improves instruction-following and quality of ranking\n- 50 chunks in one call: Context Rot & loss-in-the-middle kick in (20k-60k tokens)\n- 2x25 chunks: better ranking quality, side-effect: less API throttling (TPM rate limits)\n\n### Stage 3: Final Rescore (Precision Focus)\n\n**Purpose**: Precise global ranking on reduced candidate pool\n\n**Strategy**:\n- Single LLM-call with powerful model on ~20 candidates\n- Enable Forced Retry (if time/budget allow): min_attempts > 1 -> independant opinions\n- Handle edge-cases: stuck in reasoning loop, correct unfinished answer, one-chunk answer\n\n---\n\n## Q&A\n\n### Why Lexical Fusion Only at Stage 1?\nHypothesis: lexical signals help disambiguate when semantic signals are weak\n\n**Example**:\n\n```\nStage 2: Precision/Recall Ranking (0-2 scale)\n---------------------------------------\n\nQuery: \"What investor views emerged on geographic expansion?\"\n\nContent Reality:\n  Chunk A: \"State mortgage data: Ohio $2,725, California $2,322...\"\n           (lexically matches \"expansion\" context but semantically weak)\n\n  Chunk B: \"Investors expressed optimism about expansion plans...\"\n           (semantically stronger but lower keyword match)\n\nBoth chunks get same LLM score (1), but Chunk A has higher lexical score.\n\nResult: Lexical tiebreaker promotes wrong chunks. In later stages, there are distractors that require understanding over keyword matching.\n```\nEvidence: performance loss when applied BM25 fusion at Stage 2-3. Lexical signals saturate after Stage 1.\n\n---\n\n### Why Different Models per Stage?\n\n**Workload Distribution**:\n\n```\nStage 1 + Stage 2 (85% of API calls)\n------------------------------------\nTask: Bulk filtering (300 -> 50 -> 20)\nModel: Efficient (120B)\nSufficient for filtering obvious chunks\n           |\n           v\n\nStage 3 (15% of API calls)\n--------------------------\nTask: Final precision ranking (20 -> 5)\nModel: Powerful (405B)\nNeed reasoning for nuanced comparison: needles vs. distractors\n           |\n           v\n\nResult: Cost-latency-performance trade-off\n- Sclable efficiency in production\n```\n\n---\n\n### What about bigger/reasoning models?\n\nThinking models perform better, but they tend to suffer from over-thinking. Increase thinking mode (reasoning=high) has ambiguous effect: higher failure rate and better max performance with retries.\n\nReal example: \n````\n... Now we need other: maybe chunk 191.\\n\\nOk.\\n\\nNow we need other: maybe chunk 191.\\n\\nStop.\\n\\nOk.\\n\\nThis is stuck.\\n\\nLet\\'s identify other chunks: maybe chunk 191 is the only macro.\\n\\nBut also chunk 191 includes \"General economic inflation.. [seems to recover] ... \\n\\nNow other chunks: maybe chunk 191 ... \n```\nOh dear 😔\n\nIncrease thinking budget (max_token) with reasoning=medium seems to be more cost-effective for ranking task, but requires careful prompt design.\n\nGpt-5 perform better across all experiments (+5-8pp), but higher cost/latency ~5x (did not submit)\n\n---\n\n### What about other xyz?\n\nAgentic workflow with tool-use (LlamaIndex): high overhead, brittle, structured output not well supported across different reasoning models. Finance domain requires high robustness -> dropped this direction after prototyping.\n\nFine-tuning: performance could increase, but quite black-box, hard to debug when things go wrong. Hard to serve models for reproduction.\n\n\n## Summary\n\nOverall, the challenge is well structured, great dataset on this topic. I've learned some more on agentic design and reasoning models. I hope my solution can provide some useful insights to the community.\n\nDrop me a message if you want to discuss ideas & real-world application. I'm working in both academic research and industry projects. [LinkedIn](https://www.linkedin.com/in/hdan0tran/)",
    "3305822": "Great Work! Can you reply to my email?",
    "3306180": "Thank you! I just replied, sorry for the delay"
  },
  "source": "meta"
}