{
  "id": 612625,
  "title": "Deep Analysis on the Competition",
  "url": "/competitions/acm-icaif-25-ai-agentic-retrieval-grand-challenge/discussion/612625",
  "author_name": "",
  "post_date": "2025-10-21T03:24:46.268481300Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Note: submission deadline in 1h, good luck to everyone. </p>\n<p>I'd like to share with you some insights after ~200h, and $200 of API spending: Databricks/OpenAI/Parasail closing in about 1B tokens spent 😅</p>\n<h1>ACM ICAIF 2025 Competition: Design &amp; Technical Challenges</h1>\n<p><strong>Document Purpose</strong>: Foundation for Kaggle submission justification<br>\n<strong>Date</strong>: 2025-10-21</p>\n<hr>\n<h2>1. Competition Structure</h2>\n<h3>Task Definitions</h3>\n<p><strong>Document Ranking</strong>:</p>\n<ul>\n<li>Input: Financial question + 5 fixed SEC filing types (DEF14A, 10-K, 10-Q, 8-K, Earnings)</li>\n<li>Output: Ranked list of all 5 document indices</li>\n<li>Ground truth: Graded relevance 0-4 (4=most relevant)</li>\n</ul>\n<p><strong>Chunk Ranking</strong>:</p>\n<ul>\n<li>Input: Financial question + 100-300 paragraph-level text chunks</li>\n<li>Output: Top-5 ranked list of most relevant chunk indices</li>\n<li>Ground truth: Graded relevance 0-2 (2=highly relevant), <em>different than required output</em></li>\n</ul>\n<h3>Evaluation Metrics</h3>\n<p>All tasks scored independently using three metrics at top-5 positions:</p>\n<ul>\n<li><strong>MRR@5</strong> (Mean Reciprocal Rank): Rewards first relevant result</li>\n<li><strong>MAP@5</strong> (Mean Average Precision): Measures ranking of all relevant items</li>\n<li><strong>nDCG@5</strong> (Normalized Discounted Cumulative Gain): Accounts for graded relevance with position discounting</li>\n</ul>\n<p>The Leaderboard only shows MAP@5, which is not the full picture (also focus on nDCG@5 mentioned in Discussion topic by organizers)</p>\n<hr>\n<h2>2. Data Characteristics</h2>\n<h3>Document Ranking</h3>\n<p><strong>Query Message Characteristics</strong>:</p>\n<ul>\n<li>Consistent 155 tokens per message (range: 144-173)</li>\n<li>79.1% queries use all 5 scores [0,1,2,3,4]</li>\n<li><strong>20.9% have ties</strong> (should not happen as in competition instructions, but present in data)</li>\n</ul>\n<p><strong>Score Distribution Non-Uniformity</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Score</th>\n<th>Frequency</th>\n<th>Percentage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>6,064</td>\n<td>24.3%</td>\n</tr>\n<tr>\n<td>1</td>\n<td>5,367</td>\n<td>21.5%</td>\n</tr>\n<tr>\n<td>2</td>\n<td>4,951</td>\n<td>19.9%</td>\n</tr>\n<tr>\n<td>3</td>\n<td>4,605</td>\n<td>18.5%</td>\n</tr>\n<tr>\n<td>4</td>\n<td>3,943</td>\n<td>15.8%</td>\n</tr>\n</tbody>\n</table>\n<h3>Chunk Ranking</h3>\n<p><strong>Scale &amp; Context Window</strong>:</p>\n<ul>\n<li>Mean: 190 chunks per query, 65K tokens</li>\n<li>Range: 5-661 chunks, 2.7K-224K tokens</li>\n<li><strong>Critical</strong>: 67.5% queries exceed 32K tokens (eval set)</li>\n<li>Individual chunks: Mean 335 tokens (range 1-71K), median 182 tokens</li>\n</ul>\n<p>On <em>data labelling</em>:</p>\n<ul>\n<li>About &lt;1% of labels are extremely short 1-3 tokens: in data collection, experts see the full document (e.g. 1 relevant chunk with one token \"43\", among ~200 scored as not relevant) vs. this data, we only see chunks.</li>\n<li>There are queries with &gt;10 relevant chunks (max: 142)</li>\n</ul>\n<p><strong>Distribution of Chunk Relevance dev-set</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Relevant chunks (score &gt; 0)</td>\n<td>5.1% of all chunks</td>\n</tr>\n<tr>\n<td>Queries with only 1-2 relevant chunks</td>\n<td>30.67% (needle-in-haystack)</td>\n</tr>\n<tr>\n<td>Queries with &gt;10 highly relevant chunks</td>\n<td>4.25%</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Distribution of ground truth dev-set</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Score</th>\n<th>Frequency</th>\n<th>Percentage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 (irrelevant)</td>\n<td>3,029,827</td>\n<td>94.9%</td>\n</tr>\n<tr>\n<td>1 (relevant)</td>\n<td>118,778</td>\n<td>3.7%</td>\n</tr>\n<tr>\n<td>2 (highly relevant)</td>\n<td>44,847</td>\n<td>1.4%</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Token Distribution</strong> (eval-set):</p>\n<table>\n<thead>\n<tr>\n<th>Context Window</th>\n<th>Queries</th>\n<th>Percentage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>&gt;32K</td>\n<td>135</td>\n<td>67.5%</td>\n</tr>\n<tr>\n<td>&gt;60K</td>\n<td>106</td>\n<td>53.0%</td>\n</tr>\n<tr>\n<td>&gt;128K</td>\n<td>18</td>\n<td>9.0%</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>3. Technical Challenges</h2>\n<h3>Challenge 1: Needle-in-Haystack Problem</h3>\n<p><strong>Definition</strong>: Sparse signal detection - finding 1-2 relevant chunks among 100-300 candidates.</p>\n<p><strong>Quantitative Impact</strong>:</p>\n<ul>\n<li>30.67% of chunk queries have only 1-2 relevant chunks</li>\n<li>5.1% overall chunk relevance rate (94.9% noise)</li>\n<li>12.5% zero-score rate observed in early experiments (fail to find the needles in top-5 = 0x3 scores)</li>\n</ul>\n<p><strong>Concrete Example</strong> (Query q2e28fdef127e):</p>\n<pre><code>Question: \"What investor views emerged on KeyCorp's geographic expansion prospects?\"\nTotal chunks: \nRelevant:  chunk (position / ~ % relative context)\nContent: [] State- loan distribution\n         Washington , | Ohio , |  York  | Colorado ,\n</code></pre>\n<p><strong>Failure Mechanism</strong>:</p>\n<ul>\n<li>Lexical overlap: ZERO (question has \"expansion\", \"prospects\"; chunk has state names, numbers)</li>\n<li>Required inference: \"Loan distribution by state\" → \"Geographic footprint\" → \"Expansion capability\"</li>\n<li>Weak signal buried in noise → Dropped at filtering stage → Final score: 0.0</li>\n</ul>\n<p><strong>Impact</strong>: Even advanced LLMs struggle when relevant chunks lack direct keyword matches and require financial domain inference. Position does matter, needles in early/last relative positions have better chance - \"lost in the middle\".</p>\n<hr>\n<h3>Challenge 2: Context Window Constraints</h3>\n<p><strong>Problem</strong>: 67.5% of chunk queries exceed 32K tokens (refer to Context Rot recent findings), 10% exceed 120k tokens</p>\n<p><strong>Eval Set Statistics</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Percentile</th>\n<th>Message Tokens</th>\n<th>Chunks</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>P50</td>\n<td>63,756</td>\n<td>160</td>\n</tr>\n<tr>\n<td>P90</td>\n<td>120,451</td>\n<td>~300</td>\n</tr>\n<tr>\n<td>Max</td>\n<td>224,000</td>\n<td>661</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Processing Constraint</strong>:</p>\n<ul>\n<li>Most production LLMs: 32K-128K context windows</li>\n<li>Even 200K-1M context models face \"lost in the middle\" phenomenon</li>\n<li>Requires splitting into multiple parts for processing</li>\n</ul>\n<p><strong>Trade-off</strong>: More splits enable good processing early BUT create compound drop problem (Challenge 3).</p>\n<hr>\n<h3>Challenge 3: Multi-Stage Compound Drop Problem</h3>\n<p><strong>Definition</strong>: Local filtering loses globally relevant chunks in multi-stage ranking.</p>\n<p><strong>Dataflow Example</strong> (5-part split):</p>\n<pre><code>Query:  chunks,  relevant\n\nPart :  chunks,  relevant → Rank  →  top → If relevant ranks th → DROPPED\nPart :  chunks,  relevant → Rank  →  top → If relevant ranks th → DROPPED\n...\nPart :  chunks,  relevant →  process\n\n:  part creates  opportunity\n        Dropped chunks NEVER reach  ranking stage\n</code></pre>\n<p><strong>Experimental Evidence</strong> (split size analysis, n=100):</p>\n<table>\n<thead>\n<tr>\n<th>Split Size</th>\n<th>Local Recall</th>\n<th>Zero-Score Rate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3K tokens/part (many splits)</td>\n<td>75.35%</td>\n<td>3.3%</td>\n</tr>\n<tr>\n<td>15K tokens/part (optimal)</td>\n<td>75.35%</td>\n<td>~15-18%</td>\n</tr>\n<tr>\n<td>30K tokens/part (fewer splits)</td>\n<td>64.16%</td>\n<td>28.6%</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Critical Finding</strong>: Splitting enable 'divide and conquer' (insight from the organizers), but too many splits -&gt; global pool dilution (next).</p>\n<hr>\n<h3>Challenge 4: Global Dilution Problem</h3>\n<p><strong>Definition</strong>: Chunks surviving local filtering still face overwhelming competition in global pool.</p>\n<p><strong>Concrete Example</strong> (Query q4054d9fe6d42):</p>\n<pre><code>:  chunks,  relevant, using top-k= (because each split now has few chunks)\n split →  parts \n\n Stage:\n  :  relevant chunks\n  :  chunks\n  :  relevant chunks → Sent to global pool ✓\n\n Pool:\n   candidates:  parts ×  =  candidates\n   density: / = .% relevant\n  : / = .% irrelevant\n\n Re-Ranking (top- selection):\n  :  relevant chunks\n  :  relevant chunks\n\n Drop Rate:  →  = % of surviving chunks lost at global stage\n Score: .\n</code></pre>\n<p><strong>Mechanism</strong>: Surviving chunks compete in global pool. Stronger competition with 'distractors' (seemingly relevant, or weakly relevant).</p>\n<p><strong>Comparison</strong> (same query):</p>\n<pre><code> split →  parts\n\n Stage: ~ chunks survive ( dropped locally)\n Pool:  ×  =  candidates\n density: / = % relevant (vs .% for K)\n Selection: ~ relevant chunks\n Drop Rate:  →  = % (better than K's %)\n Score: .-. (x improvement)\n</code></pre>\n<p><strong>Insight</strong>: <em>Signal-to-noise ratio in later-stage pool</em> determines final performance. There in an inverted-u curve when it comes to splitting (assume fixed top-k, I also tried dynamic k but more issues there).</p>\n<hr>\n<h3>Challenge 5: Semantic Gap</h3>\n<p><strong>Definition</strong>: queries require <em>inference</em> fail more when context is fragmented across splits.</p>\n<p><strong>Example Dataflow</strong> (Query q2e28fdef127e):</p>\n<pre><code> \n\nPart  (chunks ): Contains chunk  (state-level loan table)\n  → ranks in isolation\n  → No  linking  to \n  → Chunk  . (weak relevance)\n  → Dropped from local top\n\nPart  (chunks ): Contains chunks about  (conceptual   → part, never sees chunk  data\n\nPart  (chunks ): Contains chunks about  (meeting transcript)\n  → part, never sees chunk  data\n never synthesizes connection         () Geographic loan data [Part ]\n        () Expansion strategy  [Part ]\n        () Investor view framing [Part ]\n        → Final . (complete failure)\n</code></pre>\n<p><strong>Impact</strong>: Context fragmentation breaks holistic understanding needed for semantic inference in financial domain. </p>\n<p><strong>Hypothesis</strong>:  if chunks are small and uniform (say 50-200 tokens)</p>\n<pre><code>Single-stage ranking (  chunks):\n  → LLM sees chunk  ( loan )\n  → Also sees expansion strategy discussion\n  → Also sees investor question framing\n  → Can infer: \n  → Makes final top-\n</code></pre>\n<p>But this is not the reality. Extreme variance found in chunks' token counts, leptokurtic (1-71K), flood context window.</p>\n<hr>\n<h3>Challenge 6: Context Rot &amp; \"Lost in the middle\"</h3>\n<p><strong>Definition</strong>:  LLMs struggle to find needles when: (1) prompt is longer, and (2) relevant chunk is not in first-last 20% of context window.</p>\n<p><strong>Why Extended Context Fails</strong>:</p>\n<pre><code>\n\n   \n   \n   \n   \n\n\n   \n   \n   \n   \n   \n</code></pre>\n<p><strong>Attention Allocation Comparison</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Strategy</th>\n<th>Chunks/Context</th>\n<th>Attention/Chunk</th>\n<th>Global Pool</th>\n<th>Global Attention</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>30-chunk parts</td>\n<td>30</td>\n<td>3.3%</td>\n<td>50 candidates</td>\n<td>2.0%</td>\n</tr>\n<tr>\n<td>95-chunk parts</td>\n<td>95</td>\n<td>1.1%</td>\n<td>475 candidates</td>\n<td>0.2%</td>\n</tr>\n<tr>\n<td>Extended (full)</td>\n<td>309</td>\n<td>0.3%</td>\n<td>N/A</td>\n<td>0.3%</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Empirical Validation</strong>: Long-context models struggle, perform worse than splitting. Splitting is critical, should NOT fit everything in one query.</p>\n<hr>\n<h2>4. Architecture Trade-offs</h2>\n<h3>Trade-off 1: Split Size Paradox</h3>\n<p><strong>Intuition</strong>: Fewer splits → Fewer local competitions → Less compound drop risk.</p>\n<p><strong>Reality</strong>: Larger splits create saturated ranking tasks that degrade performance.</p>\n<p><strong>Experimental Evidence</strong> (2-split performance comparison):</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>Chunks/Part</th>\n<th>Tokens/Part</th>\n<th>2-Split Local Recall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>15K</td>\n<td>~50-70</td>\n<td>15K</td>\n<td><strong>80.55%</strong> (BEST)</td>\n</tr>\n<tr>\n<td>30K</td>\n<td>~80-100</td>\n<td>30K</td>\n<td>66.07% (-14.5pts)</td>\n</tr>\n<tr>\n<td>45K</td>\n<td>~100-150</td>\n<td>45K</td>\n<td><strong>54.41%</strong> (WORST, -26.1pts)</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Mechanism</strong>:</p>\n<ol>\n<li><strong>Prompt Length Effect</strong>: 15K prompts preserve instruction retention; 45K prompts degrade clarity</li>\n<li><strong>Ranking Capacity</strong>: K/N ratio critical - 10/50 (20%) captures needles better than 10/150 (7%)</li>\n<li><strong>Compound Drops</strong>: More splits create more opportunities to drop, BUT better RECALL per split -&gt; win</li>\n</ol>\n<p><strong>Optimal Point</strong>: ~15K threshold (50-70 chunks/part) balances all three factors.</p>\n<hr>\n<h3>Trade-off 2: Local Pool Paradox</h3>\n<p><strong>Naive Solution</strong>: Increase local top-K from 5 to 10 to capture more relevant chunks.</p>\n<p><strong>Why It Fails</strong>:</p>\n<pre><code>Example: =5  =10, given 10 equal splits\nLocal Ranking (Part-2, 18 chunks, =10):\n  Position 6: Relevant chunk, score 0.73 ← CAPTURED with =10 ✓ (would drop with =5)\n\n\nGlobal Ranking : 50 candidates -&gt; 100 candidates:\n  Relevant chunk STILL has score 0.73 (unchanged)\n  But now competes with 90 instead of 40 chunks.\n  → Relevant chunk ranks ~45th globally\n  → DROPPED  final top-5 ✗\n</code></pre>\n<p><strong>Mathematical Formula</strong>:</p>\n<pre><code>\n\n :\n  \n  \n   :  \n</code></pre>\n<p><strong>Critical Insight</strong>: We cannot improve global outcomes without changing the scores in the 2nd stage -&gt; re-ranking using math formula rarely works. This agrees with the baseline setup from the organizers.</p>\n<p><strong>Solution</strong>: Global re-scoring assigns NEW scores in full context:</p>\n<ul>\n<li>Stage 1 (local): Chunk gets 0.73 in limited context (18 chunks)</li>\n<li>Stage 2 (global): Model sees all 10 x 10 = 100 candidates, assigns NEW score 0.88 </li>\n<li>Result: Chunk becomes competitive, makes final top-5</li>\n</ul>\n<p><strong>Empirical Validation</strong>: Global re-scoring provides +5.8pts improvement over single-stage architecture.</p>\n<hr>\n<h3>Trade-off 3: BM25 Pre-filtering Catastrophe</h3>\n<p><strong>Proposed Approach</strong>: Use BM25/embeddings to filter chunks before LLM ranking.</p>\n<p><strong>Why It Fails</strong>:</p>\n<pre><code>Pipeline: BM25 → :\n: \nChunk : [TABLE] \n\nBM25 :\n  Question : [, , , ]\n  Chunk : [, , , , ]\n  Lexical : MINIMAL ( weak match on geographic terms)\n  BM25 : ~- out of  chunks\n\nFilter to top-:  chunks\n  → Chunk   in top-\n  → DROPPED at Stage \n  → Never reaches LLM\n  → Final : \n</code></pre>\n<p><strong>Fundamental Error</strong>: Financial questions require conceptual understanding:</p>\n<ul>\n<li>\"Geographic expansion\" = loan distribution by state (tables)</li>\n<li>\"Revenue concentration\" = customer segment breakdown (tables)</li>\n<li>\"Margin pressure\" = cost structure details (financial data)</li>\n</ul>\n<p><strong>BM25/embeddings cannot infer these semantic connections</strong> → Drops relevant chunks → Permanent information loss.</p>\n<p><strong>Conclusion</strong>: Lexical filtering BEFORE deep reasoning is more harmful than multi-stage LLM architecture. Only LLM can perform required semantic inference.</p>\n<hr>\n<h3>Trade-off 4: Extended Context Failure</h3>\n<p><strong>Assumption</strong>: Longer context windows (200K) solve splitting problem by processing all chunks together.</p>\n<p><strong>Reality</strong>: Attention degradation worse than multi-stage processing.</p>\n<p><strong>Evidence</strong> (predicted performance for 309-chunk query):</p>\n<pre><code>\n   \n   \n   \n   \n\n\n   \n   \n   \n</code></pre>\n<p><strong>Implication</strong>: More context creates MORE noise, not less. The needle gets buried deep.</p>\n<hr>\n<h2>Summary</h2>\n<h3>Competition Constraints Drive Design</h3>\n<p><strong>Data Characteristics</strong>:</p>\n<ul>\n<li>67.5% queries exceed 32K tokens → Requires multi-part processing</li>\n<li>30.67% queries have 1-2 relevant chunks among ~200-300 candidates → Needles-in-haystack</li>\n<li>5.1% overall relevance rate → Overwhelming noisy</li>\n</ul>\n<p><strong>Technical Challenges</strong>:</p>\n<ol>\n<li>Sparse signal detection in high-noise environments</li>\n<li>Context window constraints requiring splitting</li>\n<li>Multi-stage compound drops (local filtering loses global relevance)</li>\n<li>Global dilution (surviving chunks face 92% drop rate in noisy pools)</li>\n<li>Semantic gaps (inference breaks when context fragments)</li>\n<li>Context rot (attention degradation in extended contexts)</li>\n</ol>\n<h3>Key Trade-offs</h3>\n<p><strong>What Fails</strong>:</p>\n<ul>\n<li>Extended context single-stage: Context rot degrades performance</li>\n<li>BM25 pre-filtering: Drops inference-requiring chunks</li>\n<li>Large split sizes: Saturated ranking (100-150 chunks/split) performs worse than moderate (50-70 chunks/split)</li>\n<li>Pool expansion without re-scoring: Local scores determine global outcomes</li>\n</ul>\n<p><strong>What Works</strong>:</p>\n<ul>\n<li>Optimal split size (15K tokens, 50-70 chunks/part): Balances attention vs drops</li>\n<li>Global re-scoring: NEW scores in full context (+5.8pts improvement)</li>\n<li>Lexical-semantic fusion: <em>BM25 boost (not filter)</em> helps weak signals without dropping chunks</li>\n<li>Multi-stage architecture: Better than extended context despite compound drops</li>\n</ul>\n<p><strong>Critical Design Principles</strong>:  Optimize <em>attention allocation</em> within <em>context rot limit</em> while preserving high <em>local recall</em> at filtering stage.</p>",
  "messages": [
    {
      "id": "3304658",
      "postDate": "10/21/2025 03:24:46",
      "content": "<p>Note: submission deadline in 1h, good luck to everyone. </p>\n<p>I'd like to share with you some insights after ~200h, and $200 of API spending: Databricks/OpenAI/Parasail closing in about 1B tokens spent 😅</p>\n<h1>ACM ICAIF 2025 Competition: Design &amp; Technical Challenges</h1>\n<p><strong>Document Purpose</strong>: Foundation for Kaggle submission justification<br>\n<strong>Date</strong>: 2025-10-21</p>\n<hr>\n<h2>1. Competition Structure</h2>\n<h3>Task Definitions</h3>\n<p><strong>Document Ranking</strong>:</p>\n<ul>\n<li>Input: Financial question + 5 fixed SEC filing types (DEF14A, 10-K, 10-Q, 8-K, Earnings)</li>\n<li>Output: Ranked list of all 5 document indices</li>\n<li>Ground truth: Graded relevance 0-4 (4=most relevant)</li>\n</ul>\n<p><strong>Chunk Ranking</strong>:</p>\n<ul>\n<li>Input: Financial question + 100-300 paragraph-level text chunks</li>\n<li>Output: Top-5 ranked list of most relevant chunk indices</li>\n<li>Ground truth: Graded relevance 0-2 (2=highly relevant), <em>different than required output</em></li>\n</ul>\n<h3>Evaluation Metrics</h3>\n<p>All tasks scored independently using three metrics at top-5 positions:</p>\n<ul>\n<li><strong>MRR@5</strong> (Mean Reciprocal Rank): Rewards first relevant result</li>\n<li><strong>MAP@5</strong> (Mean Average Precision): Measures ranking of all relevant items</li>\n<li><strong>nDCG@5</strong> (Normalized Discounted Cumulative Gain): Accounts for graded relevance with position discounting</li>\n</ul>\n<p>The Leaderboard only shows MAP@5, which is not the full picture (also focus on nDCG@5 mentioned in Discussion topic by organizers)</p>\n<hr>\n<h2>2. Data Characteristics</h2>\n<h3>Document Ranking</h3>\n<p><strong>Query Message Characteristics</strong>:</p>\n<ul>\n<li>Consistent 155 tokens per message (range: 144-173)</li>\n<li>79.1% queries use all 5 scores [0,1,2,3,4]</li>\n<li><strong>20.9% have ties</strong> (should not happen as in competition instructions, but present in data)</li>\n</ul>\n<p><strong>Score Distribution Non-Uniformity</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Score</th>\n<th>Frequency</th>\n<th>Percentage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>6,064</td>\n<td>24.3%</td>\n</tr>\n<tr>\n<td>1</td>\n<td>5,367</td>\n<td>21.5%</td>\n</tr>\n<tr>\n<td>2</td>\n<td>4,951</td>\n<td>19.9%</td>\n</tr>\n<tr>\n<td>3</td>\n<td>4,605</td>\n<td>18.5%</td>\n</tr>\n<tr>\n<td>4</td>\n<td>3,943</td>\n<td>15.8%</td>\n</tr>\n</tbody>\n</table>\n<h3>Chunk Ranking</h3>\n<p><strong>Scale &amp; Context Window</strong>:</p>\n<ul>\n<li>Mean: 190 chunks per query, 65K tokens</li>\n<li>Range: 5-661 chunks, 2.7K-224K tokens</li>\n<li><strong>Critical</strong>: 67.5% queries exceed 32K tokens (eval set)</li>\n<li>Individual chunks: Mean 335 tokens (range 1-71K), median 182 tokens</li>\n</ul>\n<p>On <em>data labelling</em>:</p>\n<ul>\n<li>About &lt;1% of labels are extremely short 1-3 tokens: in data collection, experts see the full document (e.g. 1 relevant chunk with one token \"43\", among ~200 scored as not relevant) vs. this data, we only see chunks.</li>\n<li>There are queries with &gt;10 relevant chunks (max: 142)</li>\n</ul>\n<p><strong>Distribution of Chunk Relevance dev-set</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Relevant chunks (score &gt; 0)</td>\n<td>5.1% of all chunks</td>\n</tr>\n<tr>\n<td>Queries with only 1-2 relevant chunks</td>\n<td>30.67% (needle-in-haystack)</td>\n</tr>\n<tr>\n<td>Queries with &gt;10 highly relevant chunks</td>\n<td>4.25%</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Distribution of ground truth dev-set</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Score</th>\n<th>Frequency</th>\n<th>Percentage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 (irrelevant)</td>\n<td>3,029,827</td>\n<td>94.9%</td>\n</tr>\n<tr>\n<td>1 (relevant)</td>\n<td>118,778</td>\n<td>3.7%</td>\n</tr>\n<tr>\n<td>2 (highly relevant)</td>\n<td>44,847</td>\n<td>1.4%</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Token Distribution</strong> (eval-set):</p>\n<table>\n<thead>\n<tr>\n<th>Context Window</th>\n<th>Queries</th>\n<th>Percentage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>&gt;32K</td>\n<td>135</td>\n<td>67.5%</td>\n</tr>\n<tr>\n<td>&gt;60K</td>\n<td>106</td>\n<td>53.0%</td>\n</tr>\n<tr>\n<td>&gt;128K</td>\n<td>18</td>\n<td>9.0%</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>3. Technical Challenges</h2>\n<h3>Challenge 1: Needle-in-Haystack Problem</h3>\n<p><strong>Definition</strong>: Sparse signal detection - finding 1-2 relevant chunks among 100-300 candidates.</p>\n<p><strong>Quantitative Impact</strong>:</p>\n<ul>\n<li>30.67% of chunk queries have only 1-2 relevant chunks</li>\n<li>5.1% overall chunk relevance rate (94.9% noise)</li>\n<li>12.5% zero-score rate observed in early experiments (fail to find the needles in top-5 = 0x3 scores)</li>\n</ul>\n<p><strong>Concrete Example</strong> (Query q2e28fdef127e):</p>\n<pre><code>Question: \"What investor views emerged on KeyCorp's geographic expansion prospects?\"\nTotal chunks: \nRelevant:  chunk (position / ~ % relative context)\nContent: [] State- loan distribution\n         Washington , | Ohio , |  York  | Colorado ,\n</code></pre>\n<p><strong>Failure Mechanism</strong>:</p>\n<ul>\n<li>Lexical overlap: ZERO (question has \"expansion\", \"prospects\"; chunk has state names, numbers)</li>\n<li>Required inference: \"Loan distribution by state\" → \"Geographic footprint\" → \"Expansion capability\"</li>\n<li>Weak signal buried in noise → Dropped at filtering stage → Final score: 0.0</li>\n</ul>\n<p><strong>Impact</strong>: Even advanced LLMs struggle when relevant chunks lack direct keyword matches and require financial domain inference. Position does matter, needles in early/last relative positions have better chance - \"lost in the middle\".</p>\n<hr>\n<h3>Challenge 2: Context Window Constraints</h3>\n<p><strong>Problem</strong>: 67.5% of chunk queries exceed 32K tokens (refer to Context Rot recent findings), 10% exceed 120k tokens</p>\n<p><strong>Eval Set Statistics</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Percentile</th>\n<th>Message Tokens</th>\n<th>Chunks</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>P50</td>\n<td>63,756</td>\n<td>160</td>\n</tr>\n<tr>\n<td>P90</td>\n<td>120,451</td>\n<td>~300</td>\n</tr>\n<tr>\n<td>Max</td>\n<td>224,000</td>\n<td>661</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Processing Constraint</strong>:</p>\n<ul>\n<li>Most production LLMs: 32K-128K context windows</li>\n<li>Even 200K-1M context models face \"lost in the middle\" phenomenon</li>\n<li>Requires splitting into multiple parts for processing</li>\n</ul>\n<p><strong>Trade-off</strong>: More splits enable good processing early BUT create compound drop problem (Challenge 3).</p>\n<hr>\n<h3>Challenge 3: Multi-Stage Compound Drop Problem</h3>\n<p><strong>Definition</strong>: Local filtering loses globally relevant chunks in multi-stage ranking.</p>\n<p><strong>Dataflow Example</strong> (5-part split):</p>\n<pre><code>Query:  chunks,  relevant\n\nPart :  chunks,  relevant → Rank  →  top → If relevant ranks th → DROPPED\nPart :  chunks,  relevant → Rank  →  top → If relevant ranks th → DROPPED\n...\nPart :  chunks,  relevant →  process\n\n:  part creates  opportunity\n        Dropped chunks NEVER reach  ranking stage\n</code></pre>\n<p><strong>Experimental Evidence</strong> (split size analysis, n=100):</p>\n<table>\n<thead>\n<tr>\n<th>Split Size</th>\n<th>Local Recall</th>\n<th>Zero-Score Rate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3K tokens/part (many splits)</td>\n<td>75.35%</td>\n<td>3.3%</td>\n</tr>\n<tr>\n<td>15K tokens/part (optimal)</td>\n<td>75.35%</td>\n<td>~15-18%</td>\n</tr>\n<tr>\n<td>30K tokens/part (fewer splits)</td>\n<td>64.16%</td>\n<td>28.6%</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Critical Finding</strong>: Splitting enable 'divide and conquer' (insight from the organizers), but too many splits -&gt; global pool dilution (next).</p>\n<hr>\n<h3>Challenge 4: Global Dilution Problem</h3>\n<p><strong>Definition</strong>: Chunks surviving local filtering still face overwhelming competition in global pool.</p>\n<p><strong>Concrete Example</strong> (Query q4054d9fe6d42):</p>\n<pre><code>:  chunks,  relevant, using top-k= (because each split now has few chunks)\n split →  parts \n\n Stage:\n  :  relevant chunks\n  :  chunks\n  :  relevant chunks → Sent to global pool ✓\n\n Pool:\n   candidates:  parts ×  =  candidates\n   density: / = .% relevant\n  : / = .% irrelevant\n\n Re-Ranking (top- selection):\n  :  relevant chunks\n  :  relevant chunks\n\n Drop Rate:  →  = % of surviving chunks lost at global stage\n Score: .\n</code></pre>\n<p><strong>Mechanism</strong>: Surviving chunks compete in global pool. Stronger competition with 'distractors' (seemingly relevant, or weakly relevant).</p>\n<p><strong>Comparison</strong> (same query):</p>\n<pre><code> split →  parts\n\n Stage: ~ chunks survive ( dropped locally)\n Pool:  ×  =  candidates\n density: / = % relevant (vs .% for K)\n Selection: ~ relevant chunks\n Drop Rate:  →  = % (better than K's %)\n Score: .-. (x improvement)\n</code></pre>\n<p><strong>Insight</strong>: <em>Signal-to-noise ratio in later-stage pool</em> determines final performance. There in an inverted-u curve when it comes to splitting (assume fixed top-k, I also tried dynamic k but more issues there).</p>\n<hr>\n<h3>Challenge 5: Semantic Gap</h3>\n<p><strong>Definition</strong>: queries require <em>inference</em> fail more when context is fragmented across splits.</p>\n<p><strong>Example Dataflow</strong> (Query q2e28fdef127e):</p>\n<pre><code> \n\nPart  (chunks ): Contains chunk  (state-level loan table)\n  → ranks in isolation\n  → No  linking  to \n  → Chunk  . (weak relevance)\n  → Dropped from local top\n\nPart  (chunks ): Contains chunks about  (conceptual   → part, never sees chunk  data\n\nPart  (chunks ): Contains chunks about  (meeting transcript)\n  → part, never sees chunk  data\n never synthesizes connection         () Geographic loan data [Part ]\n        () Expansion strategy  [Part ]\n        () Investor view framing [Part ]\n        → Final . (complete failure)\n</code></pre>\n<p><strong>Impact</strong>: Context fragmentation breaks holistic understanding needed for semantic inference in financial domain. </p>\n<p><strong>Hypothesis</strong>:  if chunks are small and uniform (say 50-200 tokens)</p>\n<pre><code>Single-stage ranking (  chunks):\n  → LLM sees chunk  ( loan )\n  → Also sees expansion strategy discussion\n  → Also sees investor question framing\n  → Can infer: \n  → Makes final top-\n</code></pre>\n<p>But this is not the reality. Extreme variance found in chunks' token counts, leptokurtic (1-71K), flood context window.</p>\n<hr>\n<h3>Challenge 6: Context Rot &amp; \"Lost in the middle\"</h3>\n<p><strong>Definition</strong>:  LLMs struggle to find needles when: (1) prompt is longer, and (2) relevant chunk is not in first-last 20% of context window.</p>\n<p><strong>Why Extended Context Fails</strong>:</p>\n<pre><code>\n\n   \n   \n   \n   \n\n\n   \n   \n   \n   \n   \n</code></pre>\n<p><strong>Attention Allocation Comparison</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>Strategy</th>\n<th>Chunks/Context</th>\n<th>Attention/Chunk</th>\n<th>Global Pool</th>\n<th>Global Attention</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>30-chunk parts</td>\n<td>30</td>\n<td>3.3%</td>\n<td>50 candidates</td>\n<td>2.0%</td>\n</tr>\n<tr>\n<td>95-chunk parts</td>\n<td>95</td>\n<td>1.1%</td>\n<td>475 candidates</td>\n<td>0.2%</td>\n</tr>\n<tr>\n<td>Extended (full)</td>\n<td>309</td>\n<td>0.3%</td>\n<td>N/A</td>\n<td>0.3%</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Empirical Validation</strong>: Long-context models struggle, perform worse than splitting. Splitting is critical, should NOT fit everything in one query.</p>\n<hr>\n<h2>4. Architecture Trade-offs</h2>\n<h3>Trade-off 1: Split Size Paradox</h3>\n<p><strong>Intuition</strong>: Fewer splits → Fewer local competitions → Less compound drop risk.</p>\n<p><strong>Reality</strong>: Larger splits create saturated ranking tasks that degrade performance.</p>\n<p><strong>Experimental Evidence</strong> (2-split performance comparison):</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>Chunks/Part</th>\n<th>Tokens/Part</th>\n<th>2-Split Local Recall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>15K</td>\n<td>~50-70</td>\n<td>15K</td>\n<td><strong>80.55%</strong> (BEST)</td>\n</tr>\n<tr>\n<td>30K</td>\n<td>~80-100</td>\n<td>30K</td>\n<td>66.07% (-14.5pts)</td>\n</tr>\n<tr>\n<td>45K</td>\n<td>~100-150</td>\n<td>45K</td>\n<td><strong>54.41%</strong> (WORST, -26.1pts)</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Mechanism</strong>:</p>\n<ol>\n<li><strong>Prompt Length Effect</strong>: 15K prompts preserve instruction retention; 45K prompts degrade clarity</li>\n<li><strong>Ranking Capacity</strong>: K/N ratio critical - 10/50 (20%) captures needles better than 10/150 (7%)</li>\n<li><strong>Compound Drops</strong>: More splits create more opportunities to drop, BUT better RECALL per split -&gt; win</li>\n</ol>\n<p><strong>Optimal Point</strong>: ~15K threshold (50-70 chunks/part) balances all three factors.</p>\n<hr>\n<h3>Trade-off 2: Local Pool Paradox</h3>\n<p><strong>Naive Solution</strong>: Increase local top-K from 5 to 10 to capture more relevant chunks.</p>\n<p><strong>Why It Fails</strong>:</p>\n<pre><code>Example: =5  =10, given 10 equal splits\nLocal Ranking (Part-2, 18 chunks, =10):\n  Position 6: Relevant chunk, score 0.73 ← CAPTURED with =10 ✓ (would drop with =5)\n\n\nGlobal Ranking : 50 candidates -&gt; 100 candidates:\n  Relevant chunk STILL has score 0.73 (unchanged)\n  But now competes with 90 instead of 40 chunks.\n  → Relevant chunk ranks ~45th globally\n  → DROPPED  final top-5 ✗\n</code></pre>\n<p><strong>Mathematical Formula</strong>:</p>\n<pre><code>\n\n :\n  \n  \n   :  \n</code></pre>\n<p><strong>Critical Insight</strong>: We cannot improve global outcomes without changing the scores in the 2nd stage -&gt; re-ranking using math formula rarely works. This agrees with the baseline setup from the organizers.</p>\n<p><strong>Solution</strong>: Global re-scoring assigns NEW scores in full context:</p>\n<ul>\n<li>Stage 1 (local): Chunk gets 0.73 in limited context (18 chunks)</li>\n<li>Stage 2 (global): Model sees all 10 x 10 = 100 candidates, assigns NEW score 0.88 </li>\n<li>Result: Chunk becomes competitive, makes final top-5</li>\n</ul>\n<p><strong>Empirical Validation</strong>: Global re-scoring provides +5.8pts improvement over single-stage architecture.</p>\n<hr>\n<h3>Trade-off 3: BM25 Pre-filtering Catastrophe</h3>\n<p><strong>Proposed Approach</strong>: Use BM25/embeddings to filter chunks before LLM ranking.</p>\n<p><strong>Why It Fails</strong>:</p>\n<pre><code>Pipeline: BM25 → :\n: \nChunk : [TABLE] \n\nBM25 :\n  Question : [, , , ]\n  Chunk : [, , , , ]\n  Lexical : MINIMAL ( weak match on geographic terms)\n  BM25 : ~- out of  chunks\n\nFilter to top-:  chunks\n  → Chunk   in top-\n  → DROPPED at Stage \n  → Never reaches LLM\n  → Final : \n</code></pre>\n<p><strong>Fundamental Error</strong>: Financial questions require conceptual understanding:</p>\n<ul>\n<li>\"Geographic expansion\" = loan distribution by state (tables)</li>\n<li>\"Revenue concentration\" = customer segment breakdown (tables)</li>\n<li>\"Margin pressure\" = cost structure details (financial data)</li>\n</ul>\n<p><strong>BM25/embeddings cannot infer these semantic connections</strong> → Drops relevant chunks → Permanent information loss.</p>\n<p><strong>Conclusion</strong>: Lexical filtering BEFORE deep reasoning is more harmful than multi-stage LLM architecture. Only LLM can perform required semantic inference.</p>\n<hr>\n<h3>Trade-off 4: Extended Context Failure</h3>\n<p><strong>Assumption</strong>: Longer context windows (200K) solve splitting problem by processing all chunks together.</p>\n<p><strong>Reality</strong>: Attention degradation worse than multi-stage processing.</p>\n<p><strong>Evidence</strong> (predicted performance for 309-chunk query):</p>\n<pre><code>\n   \n   \n   \n   \n\n\n   \n   \n   \n</code></pre>\n<p><strong>Implication</strong>: More context creates MORE noise, not less. The needle gets buried deep.</p>\n<hr>\n<h2>Summary</h2>\n<h3>Competition Constraints Drive Design</h3>\n<p><strong>Data Characteristics</strong>:</p>\n<ul>\n<li>67.5% queries exceed 32K tokens → Requires multi-part processing</li>\n<li>30.67% queries have 1-2 relevant chunks among ~200-300 candidates → Needles-in-haystack</li>\n<li>5.1% overall relevance rate → Overwhelming noisy</li>\n</ul>\n<p><strong>Technical Challenges</strong>:</p>\n<ol>\n<li>Sparse signal detection in high-noise environments</li>\n<li>Context window constraints requiring splitting</li>\n<li>Multi-stage compound drops (local filtering loses global relevance)</li>\n<li>Global dilution (surviving chunks face 92% drop rate in noisy pools)</li>\n<li>Semantic gaps (inference breaks when context fragments)</li>\n<li>Context rot (attention degradation in extended contexts)</li>\n</ol>\n<h3>Key Trade-offs</h3>\n<p><strong>What Fails</strong>:</p>\n<ul>\n<li>Extended context single-stage: Context rot degrades performance</li>\n<li>BM25 pre-filtering: Drops inference-requiring chunks</li>\n<li>Large split sizes: Saturated ranking (100-150 chunks/split) performs worse than moderate (50-70 chunks/split)</li>\n<li>Pool expansion without re-scoring: Local scores determine global outcomes</li>\n</ul>\n<p><strong>What Works</strong>:</p>\n<ul>\n<li>Optimal split size (15K tokens, 50-70 chunks/part): Balances attention vs drops</li>\n<li>Global re-scoring: NEW scores in full context (+5.8pts improvement)</li>\n<li>Lexical-semantic fusion: <em>BM25 boost (not filter)</em> helps weak signals without dropping chunks</li>\n<li>Multi-stage architecture: Better than extended context despite compound drops</li>\n</ul>\n<p><strong>Critical Design Principles</strong>:  Optimize <em>attention allocation</em> within <em>context rot limit</em> while preserving high <em>local recall</em> at filtering stage.</p>",
      "rawMarkdown": "Note: submission deadline in 1h, good luck to everyone. \n\nI'd like to share with you some insights after ~200h, and $200 of API spending: Databricks/OpenAI/Parasail closing in about 1B tokens spent 😅\n\n# ACM ICAIF 2025 Competition: Design & Technical Challenges\n\n**Document Purpose**: Foundation for Kaggle submission justification\n**Date**: 2025-10-21\n\n---\n\n## 1. Competition Structure\n\n### Task Definitions\n\n**Document Ranking**:\n- Input: Financial question + 5 fixed SEC filing types (DEF14A, 10-K, 10-Q, 8-K, Earnings)\n- Output: Ranked list of all 5 document indices\n- Ground truth: Graded relevance 0-4 (4=most relevant)\n\n**Chunk Ranking**:\n- Input: Financial question + 100-300 paragraph-level text chunks\n- Output: Top-5 ranked list of most relevant chunk indices\n- Ground truth: Graded relevance 0-2 (2=highly relevant), *different than required output*\n\n\n### Evaluation Metrics\n\nAll tasks scored independently using three metrics at top-5 positions:\n- **MRR@5** (Mean Reciprocal Rank): Rewards first relevant result\n- **MAP@5** (Mean Average Precision): Measures ranking of all relevant items\n- **nDCG@5** (Normalized Discounted Cumulative Gain): Accounts for graded relevance with position discounting\n\nThe Leaderboard only shows MAP@5, which is not the full picture (also focus on nDCG@5 mentioned in Discussion topic by organizers)\n\n---\n\n## 2. Data Characteristics\n\n### Document Ranking\n\n**Query Message Characteristics**:\n- Consistent 155 tokens per message (range: 144-173)\n- 79.1% queries use all 5 scores [0,1,2,3,4]\n- **20.9% have ties** (should not happen as in competition instructions, but present in data)\n\n\n**Score Distribution Non-Uniformity**:\n| Score | Frequency | Percentage |\n|-------|-----------|------------|\n| 0 | 6,064 | 24.3% |\n| 1 | 5,367 | 21.5% |\n| 2 | 4,951 | 19.9% |\n| 3 | 4,605 | 18.5% |\n| 4 | 3,943 | 15.8% |\n\n\n### Chunk Ranking\n\n**Scale & Context Window**:\n- Mean: 190 chunks per query, 65K tokens\n- Range: 5-661 chunks, 2.7K-224K tokens\n- **Critical**: 67.5% queries exceed 32K tokens (eval set)\n- Individual chunks: Mean 335 tokens (range 1-71K), median 182 tokens\n\nOn *data labelling*:\n- About <1% of labels are extremely short 1-3 tokens: in data collection, experts see the full document (e.g. 1 relevant chunk with one token \"43\", among ~200 scored as not relevant) vs. this data, we only see chunks.\n- There are queries with >10 relevant chunks (max: 142)\n\n\n**Distribution of Chunk Relevance dev-set**:\n| Metric | Value |\n|--------|-------|\n| Relevant chunks (score > 0) | 5.1% of all chunks |\n| Queries with only 1-2 relevant chunks | 30.67% (needle-in-haystack) |\n| Queries with >10 highly relevant chunks | 4.25% |\n\n***\n\n**Distribution of ground truth dev-set**:\n\n| Score | Frequency | Percentage |\n|-------|-----------|------------|\n| 0 (irrelevant) | 3,029,827 | 94.9% |\n| 1 (relevant) | 118,778 | 3.7% |\n| 2 (highly relevant) | 44,847 | 1.4%|\n\n***\n\n**Token Distribution** (eval-set):\n| Context Window | Queries | Percentage |\n|----------------|---------|------------|\n| >32K | 135 | 67.5% |\n| >60K | 106 | 53.0% |\n| >128K | 18 | 9.0% |\n\n\n---\n\n## 3. Technical Challenges\n\n### Challenge 1: Needle-in-Haystack Problem\n\n**Definition**: Sparse signal detection - finding 1-2 relevant chunks among 100-300 candidates.\n\n**Quantitative Impact**:\n- 30.67% of chunk queries have only 1-2 relevant chunks\n- 5.1% overall chunk relevance rate (94.9% noise)\n- 12.5% zero-score rate observed in early experiments (fail to find the needles in top-5 = 0x3 scores)\n\n**Concrete Example** (Query q2e28fdef127e):\n```\nQuestion: \"What investor views emerged on KeyCorp's geographic expansion prospects?\"\nTotal chunks: 309\nRelevant: 1 chunk (position 52/309 ~ 17% relative context)\nContent: [TABLE] State-level loan distribution\n         Washington $4,605 | Ohio $2,725 | New York $822 | Colorado $3,027\n```\n\n**Failure Mechanism**:\n- Lexical overlap: ZERO (question has \"expansion\", \"prospects\"; chunk has state names, numbers)\n- Required inference: \"Loan distribution by state\" → \"Geographic footprint\" → \"Expansion capability\"\n- Weak signal buried in noise → Dropped at filtering stage → Final score: 0.0\n\n**Impact**: Even advanced LLMs struggle when relevant chunks lack direct keyword matches and require financial domain inference. Position does matter, needles in early/last relative positions have better chance - \"lost in the middle\".\n\n---\n\n### Challenge 2: Context Window Constraints\n\n**Problem**: 67.5% of chunk queries exceed 32K tokens (refer to Context Rot recent findings), 10% exceed 120k tokens\n\n**Eval Set Statistics**:\n| Percentile | Message Tokens | Chunks |\n|------------|----------------|--------|\n| P50 | 63,756 | 160 |\n| P90 | 120,451 | ~300 |\n| Max | 224,000 | 661 |\n\n\n**Processing Constraint**:\n- Most production LLMs: 32K-128K context windows\n- Even 200K-1M context models face \"lost in the middle\" phenomenon\n- Requires splitting into multiple parts for processing\n\n**Trade-off**: More splits enable good processing early BUT create compound drop problem (Challenge 3).\n\n---\n\n### Challenge 3: Multi-Stage Compound Drop Problem\n\n**Definition**: Local filtering loses globally relevant chunks in multi-stage ranking.\n\n**Dataflow Example** (5-part split):\n```\nQuery: 300 chunks, 15 relevant\n\nPart 1: 60 chunks, 3 relevant → Rank all → Select top-10 → If relevant ranks 11-15th → DROPPED\nPart 2: 60 chunks, 2 relevant → Rank all → Select top-10 → If relevant ranks 11-15th → DROPPED\n...\nPart 5: 60 chunks, 4 relevant → Similar process\n\nResult: Each part creates drop opportunity\n        Dropped chunks NEVER reach global ranking stage\n```\n\n**Experimental Evidence** (split size analysis, n=100):\n\n| Split Size | Local Recall | Zero-Score Rate |\n|-----------|--------------|-----------------|\n| 3K tokens/part (many splits) | 75.35% | 3.3% |\n| 15K tokens/part (optimal) | 75.35% | ~15-18% |\n| 30K tokens/part (fewer splits) | 64.16% | 28.6% |\n\n\n**Critical Finding**: Splitting enable 'divide and conquer' (insight from the organizers), but too many splits -> global pool dilution (next).\n\n---\n\n### Challenge 4: Global Dilution Problem\n\n**Definition**: Chunks surviving local filtering still face overwhelming competition in global pool.\n\n**Concrete Example** (Query q4054d9fe6d42):\n```\nConfiguration: 289 chunks, 28 relevant, using top-k=5 (because each split now has few chunks)\n3K split → 95 parts \n\nLocal Stage:\n  Started: 28 relevant chunks\n  Dropped: 2 chunks\n  Survived: 26 relevant chunks → Sent to global pool ✓\n\nGlobal Pool:\n  Total candidates: 95 parts × 5 = 475 candidates\n  Signal density: 26/475 = 5.5% relevant\n  Noise: 449/475 = 94.5% irrelevant\n\nFinal Re-Ranking (top-5 selection):\n  Selected: 2 relevant chunks\n  Dropped: 24 relevant chunks\n\nGlobal Drop Rate: 26 → 2 = 92% of surviving chunks lost at global stage\nFinal Score: 0.28\n```\n\n**Mechanism**: Surviving chunks compete in global pool. Stronger competition with 'distractors' (seemingly relevant, or weakly relevant).\n\n\n**Comparison** (same query):\n```\n30K split → 10 parts\n\nLocal Stage: ~20 chunks survive (8 dropped locally)\nGlobal Pool: 10 × 5 = 50 candidates\nSignal density: 20/50 = 40% relevant (vs 5.5% for 3K)\nFinal Selection: ~9 relevant chunks\nGlobal Drop Rate: 20 → 9 = 55% (better than 3K's 92%)\nFinal Score: 0.55-0.65 (2x improvement)\n```\n\n**Insight**: *Signal-to-noise ratio in later-stage pool* determines final performance. There in an inverted-u curve when it comes to splitting (assume fixed top-k, I also tried dynamic k but more issues there).\n\n---\n\n### Challenge 5: Semantic Gap\n\n**Definition**: queries require *inference* fail more when context is fragmented across splits.\n\n**Example Dataflow** (Query q2e28fdef127e):\n```\nQuestion: \"What investor views emerged on KeyCorp's geographic expansion prospects?\"\n\nPart 1 (chunks 0-60): Contains chunk 52 (state-level loan table)\n  → LLM ranks in isolation\n  → No context linking \"state distribution\" to \"expansion prospects\"\n  → Chunk 52 score: 0.2 (weak relevance)\n  → Dropped from local top-10\n\nPart 2 (chunks 61-120): Contains chunks about \"expansion strategy\" (conceptual discussion)\n  → Different part, never sees chunk 52 data\n\nPart 3 (chunks 121-180): Contains chunks about \"investor questions\" (meeting transcript)\n  → Different part, never sees chunk 52 data\n\nResult: LLM never synthesizes connection between:\n        (1) Geographic loan data [Part 1]\n        (2) Expansion strategy context [Part 2]\n        (3) Investor view framing [Part 3]\n        → Final score: 0.0 (complete failure)\n```\n\n**Impact**: Context fragmentation breaks holistic understanding needed for semantic inference in financial domain. \n\n**Hypothesis**:  if chunks are small and uniform (say 50-200 tokens)\n```\nSingle-stage ranking (all 309 chunks):\n  → LLM sees chunk 52 (state loan table)\n  → Also sees expansion strategy discussion\n  → Also sees investor question framing\n  → Can infer: \"Table shows geographic presence → Relevant to expansion question\"\n  → Makes final top-5\n```\nBut this is not the reality. Extreme variance found in chunks' token counts, leptokurtic (1-71K), flood context window.\n\n---\n\n### Challenge 6: Context Rot & \"Lost in the middle\" \n\n**Definition**:  LLMs struggle to find needles when: (1) prompt is longer, and (2) relevant chunk is not in first-last 20% of context window.\n\n**Why Extended Context Fails**:\n```\nPrediction for Query q2e28fdef127e (309 chunks, chunk 52 relevant at position 17%):\n\nMulti-stage 30K (10 parts):\n  - Chunk 52 in Part 1 with 30 chunks\n  - Attention per chunk: ~3.3%\n  - Score: 0.3 (low due to semantic gap)\n  - Dropped locally → SCORE: 0.0\n\nExtended context (all 309 chunks):\n  - Chunk 52 buried at position 52/309 (17%)\n  - Attention per chunk: ~0.3% (10x dilution)\n  - \"Lost in middle\" degradation + context rot\n  - Score: 0.1 (even lower)\n  - Dropped in final ranking → SCORE: 0.0 (NO BETTER)\n```\n\n**Attention Allocation Comparison**:\n| Strategy | Chunks/Context | Attention/Chunk | Global Pool | Global Attention |\n|----------|----------------|-----------------|-------------|------------------|\n| 30-chunk parts | 30 | 3.3% | 50 candidates | 2.0% |\n| 95-chunk parts | 95 | 1.1% | 475 candidates | 0.2% |\n| Extended (full) | 309 | 0.3% | N/A | 0.3% |\n\n\n**Empirical Validation**: Long-context models struggle, perform worse than splitting. Splitting is critical, should NOT fit everything in one query.\n\n---\n\n## 4. Architecture Trade-offs\n\n### Trade-off 1: Split Size Paradox\n\n**Intuition**: Fewer splits → Fewer local competitions → Less compound drop risk.\n\n**Reality**: Larger splits create saturated ranking tasks that degrade performance.\n\n**Experimental Evidence** (2-split performance comparison):\n| Threshold | Chunks/Part | Tokens/Part | 2-Split Local Recall |\n|-----------|-------------|-------------|----------------------|\n| 15K | ~50-70 | 15K | **80.55%** (BEST) |\n| 30K | ~80-100 | 30K | 66.07% (-14.5pts) |\n| 45K | ~100-150 | 45K | **54.41%** (WORST, -26.1pts) |\n\n\n\n**Mechanism**:\n1. **Prompt Length Effect**: 15K prompts preserve instruction retention; 45K prompts degrade clarity\n2. **Ranking Capacity**: K/N ratio critical - 10/50 (20%) captures needles better than 10/150 (7%)\n3. **Compound Drops**: More splits create more opportunities to drop, BUT better RECALL per split -> win\n\n**Optimal Point**: ~15K threshold (50-70 chunks/part) balances all three factors.\n\n---\n\n### Trade-off 2: Local Pool Paradox\n\n**Naive Solution**: Increase local top-K from 5 to 10 to capture more relevant chunks.\n\n**Why It Fails**:\n```\nExample: K=5 to K=10, given 10 equal splits\nLocal Ranking (Part-2, 18 chunks, K=10):\n  Position 6: Relevant chunk, score 0.73 ← CAPTURED with K=10 ✓ (would drop with K=5)\n\n\nGlobal Ranking : 50 candidates -> 100 candidates:\n  Relevant chunk STILL has score 0.73 (unchanged)\n  But now competes with 90 instead of 40 chunks.\n  → Relevant chunk ranks ~45th globally\n  → DROPPED from final top-5 ✗\n```\n\n\n**Mathematical Formula**:\n```\nP(final success) = P(local survival) × P(global survival)\n\nIncreasing K:\n  P(local survival) ↑ (more chunks captured)\n  P(global survival) ↓ (same low scores in larger pool)\n  Net effect: Minimal improvement\n```\n\n\n**Critical Insight**: We cannot improve global outcomes without changing the scores in the 2nd stage -> re-ranking using math formula rarely works. This agrees with the baseline setup from the organizers.\n\n**Solution**: Global re-scoring assigns NEW scores in full context:\n- Stage 1 (local): Chunk gets 0.73 in limited context (18 chunks)\n- Stage 2 (global): Model sees all 10 x 10 = 100 candidates, assigns NEW score 0.88 \n- Result: Chunk becomes competitive, makes final top-5\n\n**Empirical Validation**: Global re-scoring provides +5.8pts improvement over single-stage architecture.\n\n---\n\n### Trade-off 3: BM25 Pre-filtering Catastrophe\n\n**Proposed Approach**: Use BM25/embeddings to filter chunks before LLM ranking.\n\n\n**Why It Fails**:\n```\nPipeline: BM25 → Filter to top-30% → LLM ranking\n\nExample (Query q2e28fdef127e):\nQuestion: \"What investor views emerged on KeyCorp's geographic expansion prospects?\"\nChunk 52: [TABLE] \"Washington $4,605 | Ohio $2,725 | New York $822...\"\n\nBM25 Analysis:\n  Question terms: [\"investor\", \"views\", \"expansion\", \"prospects\"]\n  Chunk terms: [\"Washington\", \"Ohio\", \"New York\", \"$4,605\", \"$2,725\"]\n  Lexical overlap: MINIMAL (only weak match on geographic terms)\n  BM25 rank: ~200-250 out of 309 chunks\n\nFilter to top-30%: 93 chunks\n  → Chunk 52 NOT in top-93\n  → DROPPED at Stage 1\n  → Never reaches LLM\n  → Final score: 0.0\n```\n\n**Fundamental Error**: Financial questions require conceptual understanding:\n- \"Geographic expansion\" = loan distribution by state (tables)\n- \"Revenue concentration\" = customer segment breakdown (tables)\n- \"Margin pressure\" = cost structure details (financial data)\n\n**BM25/embeddings cannot infer these semantic connections** → Drops relevant chunks → Permanent information loss.\n\n**Conclusion**: Lexical filtering BEFORE deep reasoning is more harmful than multi-stage LLM architecture. Only LLM can perform required semantic inference.\n\n---\n\n### Trade-off 4: Extended Context Failure\n\n**Assumption**: Longer context windows (200K) solve splitting problem by processing all chunks together.\n\n**Reality**: Attention degradation worse than multi-stage processing.\n\n**Evidence** (predicted performance for 309-chunk query):\n```\nCurrent 30K split (10 parts):\n  Local: 30 chunks/part, 3.3% attention per chunk\n  Global: 50 candidates, 2% attention per chunk\n  Average attention: ~2.5% per relevant chunk\n  Observed score: ~0.4-0.6 (partial success)\n\nExtended context (309 chunks):\n  Single stage: 309 chunks, 0.3% attention per chunk\n  \"Lost in middle\" effect: chunks at 10-90% position get <0.2% attention\n  Predicted score: ~0.3-0.4 (WORSE than multi-stage)\n```\n\n**Implication**: More context creates MORE noise, not less. The needle gets buried deep.\n\n---\n\n## Summary\n\n### Competition Constraints Drive Design\n\n**Data Characteristics**:\n- 67.5% queries exceed 32K tokens → Requires multi-part processing\n- 30.67% queries have 1-2 relevant chunks among ~200-300 candidates → Needles-in-haystack\n- 5.1% overall relevance rate → Overwhelming noisy\n\n**Technical Challenges**:\n1. Sparse signal detection in high-noise environments\n2. Context window constraints requiring splitting\n3. Multi-stage compound drops (local filtering loses global relevance)\n4. Global dilution (surviving chunks face 92% drop rate in noisy pools)\n5. Semantic gaps (inference breaks when context fragments)\n6. Context rot (attention degradation in extended contexts)\n\n### Key Trade-offs\n\n**What Fails**:\n- Extended context single-stage: Context rot degrades performance\n- BM25 pre-filtering: Drops inference-requiring chunks\n- Large split sizes: Saturated ranking (100-150 chunks/split) performs worse than moderate (50-70 chunks/split)\n- Pool expansion without re-scoring: Local scores determine global outcomes\n\n**What Works**:\n- Optimal split size (15K tokens, 50-70 chunks/part): Balances attention vs drops\n- Global re-scoring: NEW scores in full context (+5.8pts improvement)\n- Lexical-semantic fusion: *BM25 boost (not filter)* helps weak signals without dropping chunks\n- Multi-stage architecture: Better than extended context despite compound drops\n\n**Critical Design Principles**:  Optimize *attention allocation* within *context rot limit* while preserving high *local recall* at filtering stage.",
      "votes": null
    },
    {
      "id": "3304678",
      "postDate": "10/21/2025 04:28:02",
      "content": "<p>Wow, $200 for 200 hours — that’s dedication! Thank you for sharing!  I didn’t spend a penny on tokens, just fine-tuned downstream task ,but my GPU is  running so hot I’m getting roasted, and my electricity bill feels it .😂</p>",
      "rawMarkdown": "Wow, $200 for 200 hours — that’s dedication! Thank you for sharing!  I didn’t spend a penny on tokens, just fine-tuned downstream task ,but my GPU is  running so hot I’m getting roasted, and my electricity bill feels it .😂",
      "votes": null
    },
    {
      "id": "3304682",
      "postDate": "10/21/2025 04:41:23",
      "content": "<p>Nice, fine-tune with your GPU is cool. For me it was a chance to test backends also, can brag later that I win -200$ :) Did you notice issues with the labels? Some labels are out-of-context, about &lt;1% relevant chunks have 1-3 tokens</p>",
      "rawMarkdown": "Nice, fine-tune with your GPU is cool. For me it was a chance to test backends also, can brag later that I win -200$ :) Did you notice issues with the labels? Some labels are out-of-context, about <1% relevant chunks have 1-3 tokens",
      "votes": null
    },
    {
      "id": "3305075",
      "postDate": "10/22/2025 00:24:42",
      "content": "<p>Yeah, exactly — because the text is split into chunks, some parts seem irrelevant ,only make sense with more context and domain knowledge.  judging relevance just from one chunk is tricky. I wanted to try adding chunk position and some context info to the model’s encoding, but I didn’t have enough time before the deadline.</p>\n<p>Another challenge is that the content isn’t just about finance — it also includes topics like Apple, healthcare, and many other companies and area , So  finance domain isn’t enough to understand the overall context.</p>",
      "rawMarkdown": "Yeah, exactly — because the text is split into chunks, some parts seem irrelevant ,only make sense with more context and domain knowledge.  judging relevance just from one chunk is tricky. I wanted to try adding chunk position and some context info to the model’s encoding, but I didn’t have enough time before the deadline.\n\nAnother challenge is that the content isn’t just about finance — it also includes topics like Apple, healthcare, and many other companies and area , So  finance domain isn’t enough to understand the overall context.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3304678,
      "author_name": "yuzhenhu",
      "author_url": "",
      "post_date": "10/21/2025 04:28:02",
      "content": "<p>Wow, $200 for 200 hours — that’s dedication! Thank you for sharing!  I didn’t spend a penny on tokens, just fine-tuned downstream task ,but my GPU is  running so hot I’m getting roasted, and my electricity bill feels it .😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 3304682,
          "author_name": "pandalikematcha",
          "author_url": "",
          "post_date": "10/21/2025 04:41:23",
          "content": "<p>Nice, fine-tune with your GPU is cool. For me it was a chance to test backends also, can brag later that I win -200$ :) Did you notice issues with the labels? Some labels are out-of-context, about &lt;1% relevant chunks have 1-3 tokens</p>",
          "votes": null,
          "replies": [
            {
              "id": 3305075,
              "author_name": "yuzhenhu",
              "author_url": "",
              "post_date": "10/22/2025 00:24:42",
              "content": "<p>Yeah, exactly — because the text is split into chunks, some parts seem irrelevant ,only make sense with more context and domain knowledge.  judging relevance just from one chunk is tricky. I wanted to try adding chunk position and some context info to the model’s encoding, but I didn’t have enough time before the deadline.</p>\n<p>Another challenge is that the content isn’t just about finance — it also includes topics like Apple, healthcare, and many other companies and area , So  finance domain isn’t enough to understand the overall context.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3304658": "Note: submission deadline in 1h, good luck to everyone. \n\nI'd like to share with you some insights after ~200h, and $200 of API spending: Databricks/OpenAI/Parasail closing in about 1B tokens spent 😅\n\n# ACM ICAIF 2025 Competition: Design & Technical Challenges\n\n**Document Purpose**: Foundation for Kaggle submission justification\n**Date**: 2025-10-21\n\n---\n\n## 1. Competition Structure\n\n### Task Definitions\n\n**Document Ranking**:\n- Input: Financial question + 5 fixed SEC filing types (DEF14A, 10-K, 10-Q, 8-K, Earnings)\n- Output: Ranked list of all 5 document indices\n- Ground truth: Graded relevance 0-4 (4=most relevant)\n\n**Chunk Ranking**:\n- Input: Financial question + 100-300 paragraph-level text chunks\n- Output: Top-5 ranked list of most relevant chunk indices\n- Ground truth: Graded relevance 0-2 (2=highly relevant), *different than required output*\n\n\n### Evaluation Metrics\n\nAll tasks scored independently using three metrics at top-5 positions:\n- **MRR@5** (Mean Reciprocal Rank): Rewards first relevant result\n- **MAP@5** (Mean Average Precision): Measures ranking of all relevant items\n- **nDCG@5** (Normalized Discounted Cumulative Gain): Accounts for graded relevance with position discounting\n\nThe Leaderboard only shows MAP@5, which is not the full picture (also focus on nDCG@5 mentioned in Discussion topic by organizers)\n\n---\n\n## 2. Data Characteristics\n\n### Document Ranking\n\n**Query Message Characteristics**:\n- Consistent 155 tokens per message (range: 144-173)\n- 79.1% queries use all 5 scores [0,1,2,3,4]\n- **20.9% have ties** (should not happen as in competition instructions, but present in data)\n\n\n**Score Distribution Non-Uniformity**:\n| Score | Frequency | Percentage |\n|-------|-----------|------------|\n| 0 | 6,064 | 24.3% |\n| 1 | 5,367 | 21.5% |\n| 2 | 4,951 | 19.9% |\n| 3 | 4,605 | 18.5% |\n| 4 | 3,943 | 15.8% |\n\n\n### Chunk Ranking\n\n**Scale & Context Window**:\n- Mean: 190 chunks per query, 65K tokens\n- Range: 5-661 chunks, 2.7K-224K tokens\n- **Critical**: 67.5% queries exceed 32K tokens (eval set)\n- Individual chunks: Mean 335 tokens (range 1-71K), median 182 tokens\n\nOn *data labelling*:\n- About <1% of labels are extremely short 1-3 tokens: in data collection, experts see the full document (e.g. 1 relevant chunk with one token \"43\", among ~200 scored as not relevant) vs. this data, we only see chunks.\n- There are queries with >10 relevant chunks (max: 142)\n\n\n**Distribution of Chunk Relevance dev-set**:\n| Metric | Value |\n|--------|-------|\n| Relevant chunks (score > 0) | 5.1% of all chunks |\n| Queries with only 1-2 relevant chunks | 30.67% (needle-in-haystack) |\n| Queries with >10 highly relevant chunks | 4.25% |\n\n***\n\n**Distribution of ground truth dev-set**:\n\n| Score | Frequency | Percentage |\n|-------|-----------|------------|\n| 0 (irrelevant) | 3,029,827 | 94.9% |\n| 1 (relevant) | 118,778 | 3.7% |\n| 2 (highly relevant) | 44,847 | 1.4%|\n\n***\n\n**Token Distribution** (eval-set):\n| Context Window | Queries | Percentage |\n|----------------|---------|------------|\n| >32K | 135 | 67.5% |\n| >60K | 106 | 53.0% |\n| >128K | 18 | 9.0% |\n\n\n---\n\n## 3. Technical Challenges\n\n### Challenge 1: Needle-in-Haystack Problem\n\n**Definition**: Sparse signal detection - finding 1-2 relevant chunks among 100-300 candidates.\n\n**Quantitative Impact**:\n- 30.67% of chunk queries have only 1-2 relevant chunks\n- 5.1% overall chunk relevance rate (94.9% noise)\n- 12.5% zero-score rate observed in early experiments (fail to find the needles in top-5 = 0x3 scores)\n\n**Concrete Example** (Query q2e28fdef127e):\n```\nQuestion: \"What investor views emerged on KeyCorp's geographic expansion prospects?\"\nTotal chunks: 309\nRelevant: 1 chunk (position 52/309 ~ 17% relative context)\nContent: [TABLE] State-level loan distribution\n         Washington $4,605 | Ohio $2,725 | New York $822 | Colorado $3,027\n```\n\n**Failure Mechanism**:\n- Lexical overlap: ZERO (question has \"expansion\", \"prospects\"; chunk has state names, numbers)\n- Required inference: \"Loan distribution by state\" → \"Geographic footprint\" → \"Expansion capability\"\n- Weak signal buried in noise → Dropped at filtering stage → Final score: 0.0\n\n**Impact**: Even advanced LLMs struggle when relevant chunks lack direct keyword matches and require financial domain inference. Position does matter, needles in early/last relative positions have better chance - \"lost in the middle\".\n\n---\n\n### Challenge 2: Context Window Constraints\n\n**Problem**: 67.5% of chunk queries exceed 32K tokens (refer to Context Rot recent findings), 10% exceed 120k tokens\n\n**Eval Set Statistics**:\n| Percentile | Message Tokens | Chunks |\n|------------|----------------|--------|\n| P50 | 63,756 | 160 |\n| P90 | 120,451 | ~300 |\n| Max | 224,000 | 661 |\n\n\n**Processing Constraint**:\n- Most production LLMs: 32K-128K context windows\n- Even 200K-1M context models face \"lost in the middle\" phenomenon\n- Requires splitting into multiple parts for processing\n\n**Trade-off**: More splits enable good processing early BUT create compound drop problem (Challenge 3).\n\n---\n\n### Challenge 3: Multi-Stage Compound Drop Problem\n\n**Definition**: Local filtering loses globally relevant chunks in multi-stage ranking.\n\n**Dataflow Example** (5-part split):\n```\nQuery: 300 chunks, 15 relevant\n\nPart 1: 60 chunks, 3 relevant → Rank all → Select top-10 → If relevant ranks 11-15th → DROPPED\nPart 2: 60 chunks, 2 relevant → Rank all → Select top-10 → If relevant ranks 11-15th → DROPPED\n...\nPart 5: 60 chunks, 4 relevant → Similar process\n\nResult: Each part creates drop opportunity\n        Dropped chunks NEVER reach global ranking stage\n```\n\n**Experimental Evidence** (split size analysis, n=100):\n\n| Split Size | Local Recall | Zero-Score Rate |\n|-----------|--------------|-----------------|\n| 3K tokens/part (many splits) | 75.35% | 3.3% |\n| 15K tokens/part (optimal) | 75.35% | ~15-18% |\n| 30K tokens/part (fewer splits) | 64.16% | 28.6% |\n\n\n**Critical Finding**: Splitting enable 'divide and conquer' (insight from the organizers), but too many splits -> global pool dilution (next).\n\n---\n\n### Challenge 4: Global Dilution Problem\n\n**Definition**: Chunks surviving local filtering still face overwhelming competition in global pool.\n\n**Concrete Example** (Query q4054d9fe6d42):\n```\nConfiguration: 289 chunks, 28 relevant, using top-k=5 (because each split now has few chunks)\n3K split → 95 parts \n\nLocal Stage:\n  Started: 28 relevant chunks\n  Dropped: 2 chunks\n  Survived: 26 relevant chunks → Sent to global pool ✓\n\nGlobal Pool:\n  Total candidates: 95 parts × 5 = 475 candidates\n  Signal density: 26/475 = 5.5% relevant\n  Noise: 449/475 = 94.5% irrelevant\n\nFinal Re-Ranking (top-5 selection):\n  Selected: 2 relevant chunks\n  Dropped: 24 relevant chunks\n\nGlobal Drop Rate: 26 → 2 = 92% of surviving chunks lost at global stage\nFinal Score: 0.28\n```\n\n**Mechanism**: Surviving chunks compete in global pool. Stronger competition with 'distractors' (seemingly relevant, or weakly relevant).\n\n\n**Comparison** (same query):\n```\n30K split → 10 parts\n\nLocal Stage: ~20 chunks survive (8 dropped locally)\nGlobal Pool: 10 × 5 = 50 candidates\nSignal density: 20/50 = 40% relevant (vs 5.5% for 3K)\nFinal Selection: ~9 relevant chunks\nGlobal Drop Rate: 20 → 9 = 55% (better than 3K's 92%)\nFinal Score: 0.55-0.65 (2x improvement)\n```\n\n**Insight**: *Signal-to-noise ratio in later-stage pool* determines final performance. There in an inverted-u curve when it comes to splitting (assume fixed top-k, I also tried dynamic k but more issues there).\n\n---\n\n### Challenge 5: Semantic Gap\n\n**Definition**: queries require *inference* fail more when context is fragmented across splits.\n\n**Example Dataflow** (Query q2e28fdef127e):\n```\nQuestion: \"What investor views emerged on KeyCorp's geographic expansion prospects?\"\n\nPart 1 (chunks 0-60): Contains chunk 52 (state-level loan table)\n  → LLM ranks in isolation\n  → No context linking \"state distribution\" to \"expansion prospects\"\n  → Chunk 52 score: 0.2 (weak relevance)\n  → Dropped from local top-10\n\nPart 2 (chunks 61-120): Contains chunks about \"expansion strategy\" (conceptual discussion)\n  → Different part, never sees chunk 52 data\n\nPart 3 (chunks 121-180): Contains chunks about \"investor questions\" (meeting transcript)\n  → Different part, never sees chunk 52 data\n\nResult: LLM never synthesizes connection between:\n        (1) Geographic loan data [Part 1]\n        (2) Expansion strategy context [Part 2]\n        (3) Investor view framing [Part 3]\n        → Final score: 0.0 (complete failure)\n```\n\n**Impact**: Context fragmentation breaks holistic understanding needed for semantic inference in financial domain. \n\n**Hypothesis**:  if chunks are small and uniform (say 50-200 tokens)\n```\nSingle-stage ranking (all 309 chunks):\n  → LLM sees chunk 52 (state loan table)\n  → Also sees expansion strategy discussion\n  → Also sees investor question framing\n  → Can infer: \"Table shows geographic presence → Relevant to expansion question\"\n  → Makes final top-5\n```\nBut this is not the reality. Extreme variance found in chunks' token counts, leptokurtic (1-71K), flood context window.\n\n---\n\n### Challenge 6: Context Rot & \"Lost in the middle\" \n\n**Definition**:  LLMs struggle to find needles when: (1) prompt is longer, and (2) relevant chunk is not in first-last 20% of context window.\n\n**Why Extended Context Fails**:\n```\nPrediction for Query q2e28fdef127e (309 chunks, chunk 52 relevant at position 17%):\n\nMulti-stage 30K (10 parts):\n  - Chunk 52 in Part 1 with 30 chunks\n  - Attention per chunk: ~3.3%\n  - Score: 0.3 (low due to semantic gap)\n  - Dropped locally → SCORE: 0.0\n\nExtended context (all 309 chunks):\n  - Chunk 52 buried at position 52/309 (17%)\n  - Attention per chunk: ~0.3% (10x dilution)\n  - \"Lost in middle\" degradation + context rot\n  - Score: 0.1 (even lower)\n  - Dropped in final ranking → SCORE: 0.0 (NO BETTER)\n```\n\n**Attention Allocation Comparison**:\n| Strategy | Chunks/Context | Attention/Chunk | Global Pool | Global Attention |\n|----------|----------------|-----------------|-------------|------------------|\n| 30-chunk parts | 30 | 3.3% | 50 candidates | 2.0% |\n| 95-chunk parts | 95 | 1.1% | 475 candidates | 0.2% |\n| Extended (full) | 309 | 0.3% | N/A | 0.3% |\n\n\n**Empirical Validation**: Long-context models struggle, perform worse than splitting. Splitting is critical, should NOT fit everything in one query.\n\n---\n\n## 4. Architecture Trade-offs\n\n### Trade-off 1: Split Size Paradox\n\n**Intuition**: Fewer splits → Fewer local competitions → Less compound drop risk.\n\n**Reality**: Larger splits create saturated ranking tasks that degrade performance.\n\n**Experimental Evidence** (2-split performance comparison):\n| Threshold | Chunks/Part | Tokens/Part | 2-Split Local Recall |\n|-----------|-------------|-------------|----------------------|\n| 15K | ~50-70 | 15K | **80.55%** (BEST) |\n| 30K | ~80-100 | 30K | 66.07% (-14.5pts) |\n| 45K | ~100-150 | 45K | **54.41%** (WORST, -26.1pts) |\n\n\n\n**Mechanism**:\n1. **Prompt Length Effect**: 15K prompts preserve instruction retention; 45K prompts degrade clarity\n2. **Ranking Capacity**: K/N ratio critical - 10/50 (20%) captures needles better than 10/150 (7%)\n3. **Compound Drops**: More splits create more opportunities to drop, BUT better RECALL per split -> win\n\n**Optimal Point**: ~15K threshold (50-70 chunks/part) balances all three factors.\n\n---\n\n### Trade-off 2: Local Pool Paradox\n\n**Naive Solution**: Increase local top-K from 5 to 10 to capture more relevant chunks.\n\n**Why It Fails**:\n```\nExample: K=5 to K=10, given 10 equal splits\nLocal Ranking (Part-2, 18 chunks, K=10):\n  Position 6: Relevant chunk, score 0.73 ← CAPTURED with K=10 ✓ (would drop with K=5)\n\n\nGlobal Ranking : 50 candidates -> 100 candidates:\n  Relevant chunk STILL has score 0.73 (unchanged)\n  But now competes with 90 instead of 40 chunks.\n  → Relevant chunk ranks ~45th globally\n  → DROPPED from final top-5 ✗\n```\n\n\n**Mathematical Formula**:\n```\nP(final success) = P(local survival) × P(global survival)\n\nIncreasing K:\n  P(local survival) ↑ (more chunks captured)\n  P(global survival) ↓ (same low scores in larger pool)\n  Net effect: Minimal improvement\n```\n\n\n**Critical Insight**: We cannot improve global outcomes without changing the scores in the 2nd stage -> re-ranking using math formula rarely works. This agrees with the baseline setup from the organizers.\n\n**Solution**: Global re-scoring assigns NEW scores in full context:\n- Stage 1 (local): Chunk gets 0.73 in limited context (18 chunks)\n- Stage 2 (global): Model sees all 10 x 10 = 100 candidates, assigns NEW score 0.88 \n- Result: Chunk becomes competitive, makes final top-5\n\n**Empirical Validation**: Global re-scoring provides +5.8pts improvement over single-stage architecture.\n\n---\n\n### Trade-off 3: BM25 Pre-filtering Catastrophe\n\n**Proposed Approach**: Use BM25/embeddings to filter chunks before LLM ranking.\n\n\n**Why It Fails**:\n```\nPipeline: BM25 → Filter to top-30% → LLM ranking\n\nExample (Query q2e28fdef127e):\nQuestion: \"What investor views emerged on KeyCorp's geographic expansion prospects?\"\nChunk 52: [TABLE] \"Washington $4,605 | Ohio $2,725 | New York $822...\"\n\nBM25 Analysis:\n  Question terms: [\"investor\", \"views\", \"expansion\", \"prospects\"]\n  Chunk terms: [\"Washington\", \"Ohio\", \"New York\", \"$4,605\", \"$2,725\"]\n  Lexical overlap: MINIMAL (only weak match on geographic terms)\n  BM25 rank: ~200-250 out of 309 chunks\n\nFilter to top-30%: 93 chunks\n  → Chunk 52 NOT in top-93\n  → DROPPED at Stage 1\n  → Never reaches LLM\n  → Final score: 0.0\n```\n\n**Fundamental Error**: Financial questions require conceptual understanding:\n- \"Geographic expansion\" = loan distribution by state (tables)\n- \"Revenue concentration\" = customer segment breakdown (tables)\n- \"Margin pressure\" = cost structure details (financial data)\n\n**BM25/embeddings cannot infer these semantic connections** → Drops relevant chunks → Permanent information loss.\n\n**Conclusion**: Lexical filtering BEFORE deep reasoning is more harmful than multi-stage LLM architecture. Only LLM can perform required semantic inference.\n\n---\n\n### Trade-off 4: Extended Context Failure\n\n**Assumption**: Longer context windows (200K) solve splitting problem by processing all chunks together.\n\n**Reality**: Attention degradation worse than multi-stage processing.\n\n**Evidence** (predicted performance for 309-chunk query):\n```\nCurrent 30K split (10 parts):\n  Local: 30 chunks/part, 3.3% attention per chunk\n  Global: 50 candidates, 2% attention per chunk\n  Average attention: ~2.5% per relevant chunk\n  Observed score: ~0.4-0.6 (partial success)\n\nExtended context (309 chunks):\n  Single stage: 309 chunks, 0.3% attention per chunk\n  \"Lost in middle\" effect: chunks at 10-90% position get <0.2% attention\n  Predicted score: ~0.3-0.4 (WORSE than multi-stage)\n```\n\n**Implication**: More context creates MORE noise, not less. The needle gets buried deep.\n\n---\n\n## Summary\n\n### Competition Constraints Drive Design\n\n**Data Characteristics**:\n- 67.5% queries exceed 32K tokens → Requires multi-part processing\n- 30.67% queries have 1-2 relevant chunks among ~200-300 candidates → Needles-in-haystack\n- 5.1% overall relevance rate → Overwhelming noisy\n\n**Technical Challenges**:\n1. Sparse signal detection in high-noise environments\n2. Context window constraints requiring splitting\n3. Multi-stage compound drops (local filtering loses global relevance)\n4. Global dilution (surviving chunks face 92% drop rate in noisy pools)\n5. Semantic gaps (inference breaks when context fragments)\n6. Context rot (attention degradation in extended contexts)\n\n### Key Trade-offs\n\n**What Fails**:\n- Extended context single-stage: Context rot degrades performance\n- BM25 pre-filtering: Drops inference-requiring chunks\n- Large split sizes: Saturated ranking (100-150 chunks/split) performs worse than moderate (50-70 chunks/split)\n- Pool expansion without re-scoring: Local scores determine global outcomes\n\n**What Works**:\n- Optimal split size (15K tokens, 50-70 chunks/part): Balances attention vs drops\n- Global re-scoring: NEW scores in full context (+5.8pts improvement)\n- Lexical-semantic fusion: *BM25 boost (not filter)* helps weak signals without dropping chunks\n- Multi-stage architecture: Better than extended context despite compound drops\n\n**Critical Design Principles**:  Optimize *attention allocation* within *context rot limit* while preserving high *local recall* at filtering stage.",
    "3304678": "Wow, $200 for 200 hours — that’s dedication! Thank you for sharing!  I didn’t spend a penny on tokens, just fine-tuned downstream task ,but my GPU is  running so hot I’m getting roasted, and my electricity bill feels it .😂",
    "3304682": "Nice, fine-tune with your GPU is cool. For me it was a chance to test backends also, can brag later that I win -200$ :) Did you notice issues with the labels? Some labels are out-of-context, about <1% relevant chunks have 1-3 tokens",
    "3305075": "Yeah, exactly — because the text is split into chunks, some parts seem irrelevant ,only make sense with more context and domain knowledge.  judging relevance just from one chunk is tricky. I wanted to try adding chunk position and some context info to the model’s encoding, but I didn’t have enough time before the deadline.\n\nAnother challenge is that the content isn’t just about finance — it also includes topics like Apple, healthcare, and many other companies and area , So  finance domain isn’t enough to understand the overall context."
  },
  "source": "meta"
}